Enable APM on Kubernetes via Single Step Instrumentation
Before doing anything else: Fully resolve all variables in
## Context to resolve before acting. Do not begin Step 0 until every variable has a concrete value.
Silent failure — check this before any other step:
If the application has
ddtrace,dd-trace, or any OpenTelemetry SDK in its dependency manifest (requirements.txt,package.json,Gemfile,go.mod,pom.xml) — even with no import statements in code — SSI will silently disable itself at runtime.The failure is invisible: init containers run and complete, the pod starts healthy, no errors appear in
kubectlorpup, but no traces arrive. The injector detects the user-installed tracer and exits cleanly without logging anything.Claude runs
If any match — stop. Remove the package entirely (not just the import), rebuild the image, reload it into the cluster, and restart the pod before continuing. A package present in the manifest is enough to trigger this even if it is never imported.
Triggers
Invoke this skill when the user expresses intent to:
- Enable APM on a Kubernetes cluster
- Instrument Kubernetes applications with Datadog tracing
- Set up Single Step Instrumentation (SSI)
Do NOT invoke this skill if:
- The Datadog Agent is not yet installed — run
agent-installfirst - The user wants to verify SSI after setup — use
verify-ssi - The user wants to enable Data Streams Monitoring: use
enable-dsm - The user wants to enable Profiler or AppSec: use
dd-apm-k8s-sdk-features
Prerequisites
These are not a reading exercise — actively verify each one before proceeding.
Environment
- Datadog Agent is installed and healthy —
agent-installcomplete - Kubernetes v1.20+
- Linux node pools only — Windows pods require explicit namespace exclusion
- Cluster is not ECS Fargate — unsupported
- Not a hardened SELinux environment — unsupported
- Not a very small VM instance (e.g. t2.micro) — SSI can hit init timeouts
- No PodSecurity baseline or restricted policy enforced
Language and runtime
- Application language is one of: Java, Python, Ruby, Node.js, .NET, PHP
- Runtime version is within SSI's supported range — verify against the SSI compatibility matrix
- Node.js app is not using ESM — SSI does not support ESM
- Java app is not already using a
-javaagentJVM flag
Existing instrumentation — confirmed clean by the check at the top of this skill. If you skipped that check, go back and run it now.
Context to resolve before acting
Discover from the cluster — do not ask the user for information you can find yourself.
Step 0 (Only if existing instrumentation detected): Remove Manual Instrumentation
Scan all source files for: import ddtrace, from ddtrace, require 'ddtrace', require("dd-trace"), opentelemetry, tracer.trace(
Also check dependency manifests for ddtrace / dd-trace / OTel SDK packages.
If found — remove the import/package, then rebuild and reload:
Claude runs
[DECISION: how does this cluster get local images?]
Check the repo's setup script (e.g. create.sh, Makefile, justfile) for how images are loaded — do not guess from the cluster name or context. Common patterns:
If the setup script is ambiguous, run the load command it uses exactly as written.
- Registry-based: skip — image will be pulled on next deployment
Confirm with the user before restarting. Tell the user: "I need to restart
<DEPLOYMENT_NAME>in<APP_NAMESPACE>to pick up the rebuilt image. Ready to proceed?" Wait for confirmation.
Claude runs
Step 1: Extend the DatadogAgent Manifest with APM
SSI is configured on the existing DatadogAgent resource — do not create a separate manifest.
Choose targeting scope based on what the user asked for:
- User asked to instrument all applications or didn't specify scope → use Option A (cluster-wide)
- User asked for specific namespaces only → use Option B
- User asked to exclude namespaces from cluster-wide → use Option C
- User asked for specific pods/workloads → use Option D
Default is cluster-wide (Option A). If the user said "all my applications", "my whole cluster", or didn't restrict scope, use Option A with no
enabledNamespacesortargets.
Recommended ddTraceVersions: java: "1", python: "2", js: "5", dotnet: "3", ruby: "2", php: "1"
Option A — Cluster-wide (default):
Option B — Specific namespaces only:
Option C — Cluster-wide with exclusions:
Option D — Target specific workloads:
Note:
ddTraceVersionsonly applies inside atargets[]entry (Option D). It is not valid alongsideenabledNamespacesor at theinstrumentationlevel directly.
Claude runs
If datadogagent.datadoghq.com/datadog configured — continue to Step 2.
ERROR: Validation error — check YAML. enabledNamespaces and disabledNamespaces cannot both be set.
Step 2: Inform the User About Unified Service Tags
Do NOT modify application Deployments without explicit user confirmation. Applying labels to existing application workloads is a change to customer-managed resources.
Inform the user that adding Unified Service Tags (UST) to their Deployments will enable proper service/env/version tagging in Datadog. This is optional for SSI to work but recommended for full observability:
If the user wants you to apply these, get their confirmation first. Applying label changes rolls the pods immediately; if DSM will be enabled in Step 2b, apply the labels after it. UST labels are not required for APM traces to flow; SSI works without them.
Step 2b: Check for Event-Driven Services
Skip this step entirely in an eval cluster (kind cluster name contains "evalya") or when running non-interactively: run nothing, ask nothing, and continue to the next step.
Otherwise, read the ## Is DSM a fit? section of .claude/skills/dd-apm/enable-dsm/SKILL.md and run its detection command.
- Fit found (messaging client, broker, queue-triggered Lambda, or the user describes services handing work to each other asynchronously) → follow
enable-dsm. It asks the user once, states the plan rule, and makes the config change without restarting. - No fit → skip. Do not mention DSM.
If the user agrees, enable-dsm applies its DatadogAgent change and waits until the Cluster Agent is on the new config, so the restart in Step 3 picks up both SSI and DSM. In Step 3, restart one DSM Deployment first and run the DD_DATA_STREAMS_ENABLED check from enable-dsm Step 2a on it before restarting the rest. If it came up without init containers, wait 30 seconds and restart it once more.
Step 3: Restart Application Pods
Confirm with the user before restarting. Tell the user: "I need to restart
<DEPLOYMENT_NAME>in<APP_NAMESPACE>for SSI to inject into the pods. This will cause a brief outage. Ready to proceed?" Wait for confirmation.
Claude runs
If pods restart cleanly, init containers named datadog-lib-<language>-init will be visible in the pod spec.
ERROR: Pods crash-looping — check for existing custom instrumentation. See troubleshoot-ssi.
Done
Exit when ALL of the following are true:
-
features.apm.instrumentationis present in the appliedDatadogAgentmanifest - User has been informed that they need to restart their application pods
- User has been informed about Unified Service Tags (UST) and how to apply them if desired
- Scope confirmed: which workloads are instrumented, which were skipped and why
Automatically proceed to verify-ssi now — do not ask the user for permission.
Security constraints
- Never write a raw API key into any file or chat message
- Never use namespace
defaultfor Datadog resources - Never modify
admissionControllersettings directly — SSI manages this via the Operator - Do not add APM config to application manifests — configure only via
DatadogAgent - Exception: UST labels (
tags.datadoghq.com/*) on application Deployments are required and intentional - Never run
kubectl deletewithout user confirmation docker pushto a registry always requires user confirmation- Never use
kubectl patchto apply UST labels or any Deployment changes. Always edit the Deployment YAML file andkubectl apply -f. Changes made withkubectl patchare transient and will be overwritten on the next rollout.


