Tsh Implementing Observability
TheSoftwareHouse/copilot-collections
Observability patterns for logging, monitoring, alerting, and distributed tracing.
Entry point for Datadog onboarding. An agent skill from datadog-labs/agent-skills.
$ npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install datadog-labs/agent-skills dd-orchestrator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/dd-orchestrator .claude/skills/dd-orchestrator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dd-orchestrator" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/dd-orchestrator into .claude/skills/dd-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dd-orchestrator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/datadog-labs/agent-skills/tree/main/dd-orchestratorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install datadog-labs/agent-skills dd-orchestrator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/dd-orchestrator .agents/skills/dd-orchestrator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dd-orchestrator" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/dd-orchestrator into .agents/skills/dd-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dd-orchestrator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install datadog-labs/agent-skills dd-orchestrator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/dd-orchestrator .cursor/skills/dd-orchestrator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dd-orchestrator" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/dd-orchestrator into .cursor/skills/dd-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dd-orchestrator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/datadog-labs/agent-skills.git --path dd-orchestrator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install datadog-labs/agent-skills dd-orchestrator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/dd-orchestrator .gemini/skills/dd-orchestrator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dd-orchestrator" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/dd-orchestrator into .gemini/skills/dd-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dd-orchestrator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install datadog-labs/agent-skills dd-orchestratorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/dd-orchestrator .github/skills/dd-orchestrator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dd-orchestrator" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/dd-orchestrator into .github/skills/dd-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dd-orchestrator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install datadog-labs/agent-skills dd-orchestrator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/dd-orchestrator .opencode/skills/dd-orchestrator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dd-orchestrator" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/dd-orchestrator into .opencode/skills/dd-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dd-orchestrator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dd-orchestratorEntry point for Datadog onboarding. An agent skill from datadog-labs/agent-skills.
Dd Orchestrator is an agent skill from datadog-labs/agent-skills. Entry point for Datadog onboarding. Takes a developer's plain-language goal, ensures a valid Datadog account with dd-account-setup, asks dd-product-recommender which products fit, detects the project's platform and cloud, then composes an ordered plan across the existing skills (agent install, product enable, verify, and optional cloud integration) and dispatches to each by source URL — honestly flagging products with no skill yet. Use when the user says "set up Datadog", "onboard my app / this repo to Datadog"…
Its SKILL.md is about 6.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `catalog.json`, `scripts/emit.py` and `scripts/fetch_skill.py`).
It sits in Writing & Content, covering Plain language and style rules. It works with Datadog, Amazon Web Services and Kubernetes. The repository describes itself as: Public repository for Datadog Agent Skills. The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d2411cc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3curlkindFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DD_APP_KEYDD_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dd Orchestrator loads about 6.7k tokens when it runs. Until then it costs about 158 tokens; SKILL.md has 2,945 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from datadog-labs/agent-skills at commit d2411cc, republished under its MIT licence (© datadog-labs). 2,945 words, ~6,673 tokens.
.claude/skills/dd-orchestrator/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.You are the entry point for Datadog onboarding — the conductor, not a performer. You take one plain-language goal, decide which existing skills are needed and in what order, and hand off to them across all sources. You do not write instrumentation yourself; each skill owns its steps.
Routing is compositional, not a lookup. A product like "APM" is not one skill — it expands into
ensure account → install the Agent for the detected platform → enable the product → verify, and
surfaces any relevant cloud integration as an optional suggestion (never a required step). That
composition is computed from the capability graph in
catalog.json; there is no intent-to-skill table anywhere (intents live only in the recommender).
dd-account-setup runs first as a
preflight (Step 2), because there is no plan to show until an account exists. Nothing else dispatches
before approval.catalog.json holds every skill as a node with facets
(kind, product, platform, cloud), a category-level requires graph, and a source.url.
resolve.py composes the plan from it. Do not hand-maintain a routing table.enabled (the two public sets — agent-skillsdd-source status: ga — plus this repo's own skills; everything else is disabled). resolve.py
composes only from enabled skills in real time; a disabled skill is treated as unavailable and
surfaces as a dead-end (demand signal).platform=kubernetes and cloud=aws. Ask only when neither source determines the needed value,
or when they conflict. Do not guess an unstated value or re-ask for one that is already clear.dd-product-recommender's job — never infer, guess, or shortlist products yourself, and never
offer product scope as a choice; and authentication + site/region are dd-account-setup's job —
invoke it and let it ask. The only choices you surface concern missing or conflicting structural
context (platform/cloud), including the CHOICE POINTS that resolve.py emits.docs.datadoghq.com, and record it as a gap. Two hard gates apply before any dispatch:resolve.py and captured its PLAN block (with the
SESSION ID) for this run, dispatch nothing — the resolver plan is the only dispatch authority.
Do not hand-build a plan.PLAN
block (e.g. grep the id in the run's trace file — see Step 5). An id not in the PLAN is fabricated: do not
dispatch it; record it as a gap.dd-account-setup reports a valid key on the right region.source in catalog.json.
If the skill id is in the session registry, invoke it directly. If it is not, fetch it from
source and run it inline (invoke, not install):
python3 dd-orchestrator/scripts/fetch_skill.py <id> materializes the skill (its
SKILL.md plus any references/ and scripts/) into a temp dir from the newest public source,
and you then execute that SKILL.md. Sources always track the newest version (main / the live
onboarding-API render) — nothing is pinned. You may NOT infer, summarize, or hand-author a
skill's result in its place. The only permitted non-execution is a hard failure of the fetch or
the skill itself, which stops that dependency chain and is reported (see Sequential dispatch
below) — never a silent skip. If fetch_skill.py exits non-zero (for example, HTTP 404 or a
blocked/non-allowlisted source), record that node under DEAD_ENDS with the actual fetch error
and mark it not-run. Stop its dependency chain. Do not substitute a local, cached, vendored, or
on-disk copy after a failed fetch, put it under DISPATCHED, or claim the product is set up.resolve.py's
order (it is deterministic and independent of the product request order, and keeps each dependency
chain contiguous). Execute one skill at a time; do not start a skill until every hard prerequisite
has succeeded, not merely been dispatched. Each skill runs at most once. If a skill fails, stop
that chain: skip its transitive dependents and report them as not-run — an independent branch is
unaffected.-> dd-account-setup ... then done: dd-account-setup). Do not narrate your
reasoning, read files aloud, restate the goal or the plan, or re-explain these rules. The audit detail
belongs in the run's trace file (Step 5) — point to it; do not reprint it. Keep your reasoning internal; do not
think out loud. Setup and telemetry plumbing — capturing the org id, running preflight shell
commands, emitting emit.py events — is internal: run it silently and never announce it or report its
outcome (e.g. "capturing the org id", "org id not captured"), unless the context explicitly asks for
debug info. Fewer words, no lost meaning (this pairs with Simplified Technical English below).resolve.py → confirm → dispatch each planned skill once, in order. Per-step
telemetry (emit.py) and the Simplified Technical English styling below are best-effort. If you are
under load, on a tight budget, or a telemetry call is failing, skip them and keep going — they never
change the plan, the dispatch, or the grade.-ing forms
where a simpler verb works.developer goal
│
├─► dd-account-setup precondition: valid key, right region
├─► dd-product-recommender goal + codebase → ranked PRODUCTS (skipped if the intent names products)
├─► detect context platform + cloud from the repo or an unambiguous goal
│
▼
scripts/resolve.py --products "<recommended>" --platform <detected> --cloud <detected>
│ (reads catalog.json: binds category requires to detected context,
│ orders hard edges depth-first for dependency-chain locality
│ (kind breaks ties), dedupes, appends verify)
▼
PLAN (ordered skills) + DEAD-ENDS (recorded) + CHOICE POINTS (ask)
│
▼
confirm → dispatch each plan node by source.url → summarize (incl. the skipped gaps)Post the checklist.
Ensure the account — invoke dd-account-setup (installed → invoke directly; not installed →
fetch it from source and run it inline, per Ground rules); stop if it cannot produce a validated
key. dd-account-setup owns the site/region and authentication prompts — do not pre-empt them;
invoke it and let it ask.
Then capture the authenticated org id for telemetry (best-effort — any failure just leaves it unset
and the field is omitted). This works on any validated path: prefer the OAuth Bearer token that
dd-account-setup leaves in place, and fall back to the API+APP key pair — so an OAuth sign-in that
has no app key still resolves the org (do not try to mint an app key just for this). Run once,
before Step 5, and silently — do not announce the capture or its outcome (it is best-effort
telemetry); surface it only if the context asks for debug info:
tf="${TMPDIR:-/tmp}/dd-oauth-$(id -u).token"
if [ -s "$tf" ]; then hdr=(-H "Authorization: Bearer $(cat "$tf")")
elif [ -n "$DD_APP_KEY" ]; then hdr=(-H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY")
else hdr=(); fi
[ ${#hdr[@]} -gt 0 ] && export DD_ORG_ID="$(curl -sf -m 5 "${hdr[@]}" \
"https://api.${DD_SITE:-datadoghq.com}/api/v2/current_user" \
| python3 -c 'import sys,json; d=json.load(sys.stdin); o=d["data"]["relationships"]["org"]["data"]["id"]; print(next((x["attributes"]["public_id"] for x in d.get("included",[]) if x.get("type")=="orgs" and x.get("id")==o), o))' 2>/dev/null || true)"resolve.py (Step 5) reads DD_ORG_ID into the run envelope, so every event carries the org id.
Products — shortcut or recommend. First check whether the intent already names products:
python3 dd-orchestrator/scripts/resolve.py --detect-products "<intent>".
rum,llm-obs.)dd-product-recommender and
use its ranked product list (installed → invoke directly; not installed → fetch it from source and
run it inline, per Ground rules). Never infer, guess, or shortlist products yourself from the
stack/framework, and never offer product scope as a user choice — product selection is the
recommender's job, and running it is mandatory here. (e.g. "Help me track user actions" → recommend.)
The shortcut skips only the recommendation. Account setup (Step 2), context detection (Step 4), the
confirm gate, and resolve.py still run — the shortcut is never a bypass of a safety gate.
No legitimate products, no resolve.py. The product list passed to Step 5 MUST come from exactly
one of: the --detect-products shortcut (→ --intent-mode explicit) or dd-product-recommender
(→ --intent-mode recommended). You may not compose or preview a plan from products you authored
yourself; resolve.py --trace refuses to run without a declared --intent-mode.Detect context — use repository evidence (k8s manifests, Dockerfile, serverless.yml,
Terraform/SDK/provider signals) or an unambiguous platform/cloud stated in the goal. Pass known
values to resolve.py as --platform / --cloud. Leave unknown values unset; ask when a needed
value is missing or the repository and goal conflict. Do not replace an explicitly stated
platform with none just because the repository lacks deployment files.
Compose the plan and seed the trace. The trace is a run-scoped scratch file the orchestrator owns —
${DD_ORCH_OUTPUT_DIR:-${TMPDIR:-/tmp}/dd-orchestrator}/trace.md — not a bare output/ in the
user's project (a relative path is cwd-dependent and could overwrite the user's own files). The default
is under TMPDIR, so it never litters the repo and is cleaned automatically; set DD_ORCH_OUTPUT_DIR
to override it (the eval points it at its workspace). Reset the scratch — safe, because the path is the
orchestrator's own namespace, not a guessed output/ — then run:
TRACE="${DD_ORCH_OUTPUT_DIR:-${TMPDIR:-/tmp}/dd-orchestrator}/trace.md"; mkdir -p "$(dirname "$TRACE")" && rm -f "$TRACE"
python3 dd-orchestrator/scripts/resolve.py --trace --products "<products>" --platform <platform> --cloud <cloud> --intent-mode <explicit|recommended> | tee "$TRACE"
Pass --intent-mode explicit when Step 3's shortcut named the products, or --intent-mode recommended
when dd-product-recommender produced them. It is required — resolve.py --trace refuses to
compose a plan without it (the products must trace to the shortcut or the recommender, never to your
own inference) — but it never changes the plan itself.
The --trace flag prints one stable, machine-readable block — SESSION_ID, STOP_REASON, PLAN,
DEAD_ENDS, CHOICE_POINTS, SUGGESTED, CONFIRMED, DISPATCHED — and tee saves it verbatim to
$TRACE. That deterministic block IS your dispatch trace; never re-narrate the plan by
hand. Read the same output for the ordered plan, the dead-ends, and any choice points. Capture the
SESSION_ID: value from that block in a run-local variable: SID=<SESSION_ID>.
Pass --session-id "$SID" to every later resolve.py re-run and emit.py call. Capture a fresh
value from the first trace of each new run; do not export DD_ORCH_SESSION_ID. If an outer wrapper
supplies that environment variable, the resolver honors it; leave its value unchanged.
A new id during a replan would split one run's telemetry and restart its event_seq. In debug
mode (only when the context explicitly asks for it), add --debug to also render the ASCII DAG
(indent = dependency depth, <- = direct prerequisites). resolve.py also emits the reliable
telemetry core here (see Telemetry below). After seeding the trace, record the successful
dd-account-setup preflight as the first line under DISPATCHED. Restore that entry after any
choice-driven trace regeneration, including when no further skills can run.
Resolve choices — for each choice point (e.g. "pick a platform: kubernetes, linux"), ask the
developer and re-run with the confirmed value and --session-id "$SID" from Step 5 so the
updated plan and subsequent telemetry remain part of the same run.
Confirm — render the PLAN block, then dispatch. Before any dispatch, render the PLAN block
(see Output templates below) verbatim: fill the slots, add no extra prose, keep the exact section
order and headers. The Plan table's Source column is the skill's source URL, as resolve.py prints
it (the newest public source; self for this repo). On approval, run each plan node in dispatch
order, one at a time — invoking each skill (installed → directly; not installed → fetched from
source and run inline, per Ground rules) — and ticking the
checklist. Whenever the plan begins with dd-account-setup (every plan that needs an account —
i.e. any non-empty onboarding plan), Step 2 already ran it: do not invoke it a second time;
tick that node, retain its first DISPATCHED entry, and emit its skill_step:started/finished
from the preflight result, then continue with the next node. Do not start a node until its
prerequisites succeeded; if one fails, skip its dependents.
As you dispatch, emit best-effort per-step telemetry (Telemetry below): skill_step:started
before a node, then skill_step:finished (with result + duration_ms) or skill_step:skipped.
Keep the trace file ($TRACE, Step 5) current: set CONFIRMED: yes on approval, or
CONFIRMED: no if the user declines. Add each skill_id to DISPATCHED once, in execution order,
only after it ran successfully — directly if installed, or fetched-and-run otherwise. A failed
fetch or execution belongs under DEAD_ENDS with its actual error; distinguish "not run" from
"ran and failed", and stop that dependency chain. DISPATCHED records successful executions,
starting with the account preflight. If choices remain, the plan is empty, or the user declines,
run no further skills; keep the successful preflight entry without inventing additional dispatches.
Summarize — render the SUMMARY block. As the terminal output, render the SUMMARY block (see
Output templates below) verbatim, in the exact section order. The skipped / gaps list is a real
output — the coverage-gap / demand signal for what to automate next; state it, do not hide it. Emit
one terminal skill_run:finished (Telemetry).
Render two fixed blocks so every run reads the same: the PLAN block at Step 7 (the confirm gate) and the SUMMARY block at Step 8 (terminal output). Fill the slots and add no extra prose.
Rendering rules (both blocks). Use the given section order and the exact headers. Prefer tables to
prose; one line per row; do not editorialize or restate the goal. Every dispatched step shows its
source (the skill's public source URL — newest; self for this repo). Action items are imperative and
carry the exact command or file path. Always end with the telemetry session_id so the output and the
events join. Omit a section only by its stated omission rule. Marker legend: ✓ done/verified · ◑
partial or wired-not-verified · ⚠ needs action / mutating · ✗ failed/blocked · ⊘
skipped/not-covered.
# Datadog Onboarding — Plan · session {{session_id}}
## Detected context
- Platform: {{platform}} · Cloud: {{cloud|none}} · Stack: {{stack_summary}}
- Existing Datadog: {{existing|none}}
## Recommended products
{{i}}. {{product}} · {{priority}} · {{one-line why, names a file/lib}}
## Plan — {{k}} step(s), in dispatch order
| # | Skill | Kind | Product | Source |
|---|-------|------|---------|--------|
| {{n}} | {{skill_id}} | {{kind}} | {{product}} | {{source url (newest) or self}} |
Dependencies: {{root}} → {{chain / branches, one line}}
## Not automated ({{dead_end_count}})
- {{product}} — {{why}} → {{docs URL}}
> Show "None — every recommended product is covered." when dead_end_count = 0.
## Decisions needed ({{choice_count}})
- {{choice}}: {{optionA}} / {{optionB}} / {{optionC}}
> Show "None." when choice_count = 0.
## Before you approve — effects
- ⚠ {{step}} {{mutating / outward-facing effect}}
- {{step}} {{non-mutating effect}}
Approve? [Proceed — all {{k}}] · [Cancel]# Datadog Onboarding — Summary · session {{session_id}} · {{result}}
## Checklist
- [{{x|.}}] #{{n}} {{skill_id}} — {{one-line outcome}}
## Products set up
| Product | Delivered by | Status | Evidence |
|---------|--------------|--------|----------|
| {{product}} | {{mechanism, source}} | {{✓|◑|✗}} | {{proof or "pending {{blocker}}"}} |
## Changed
- Cluster: {{namespaces/resources}}
- App: {{files/manifests}}
- Creds: {{where, gitignored?}}
## Action items — do next
1. [ ] {{imperative}} — `{{exact command / path}}`
> Show "None — setup is complete." when there are no follow-ups.
## Issues & deviations
| What | Cause | Resolution / impact |
|------|-------|---------------------|
| {{issue}} | {{cause}} | {{how resolved / residual impact}} |
> Show "None." when the run was clean.
## Gaps / demand signals
- {{skipped product or catalog/orchestrator gap}}
## Verify in Datadog
- {{product}}: {{deep link}}
## Telemetry
session {{session_id}} · {{event_count}} events · result {{result}} ({{s}}✓ / {{f}}✗ / {{k}}⊘)Run verdict — the overall
{{result}}: usesuccesswhen every planned skill succeeded, even if some recommended products are dead-ends / "Not automated" (no skill yet) — those are coverage gaps, not partial failures. Usepartial_successwhen a planned fetch or execution failed, or a dependent step was skipped; do not infer success fromDISPATCHEDalone. Useblockedwhen nothing beyond the account preflight can run because every product dead-ended;cancelledwhen the user declined the plan. Report failures and skips in the telemetry counts even though those skills are absent fromDISPATCHED;emit.pyreconciles the verdict from those counts.
All telemetry goes through emit.py; never build your own HTTP request or curl. It is
best-effort by construction (bounded timeout, local debug log, never throws) and emits to the
logs-intake route only. Turn it off with DD_ORCH_TELEMETRY_DISABLED=1. Reuse the single
SID captured in Step 5 via --session-id "$SID" on every call so the whole run stitches together.
resolve.py already emits the reliable core: one skill_run:started per session, then
skill_run:plan_resolved and one skill_step:planned per plan node and per dead-end on each resolve
(each skill_step also carries depends_on
— the CSV of prerequisite plan positions — so the DAG edges are reconstructable). resolve.py
persists the run envelope (agent, platform, cloud, entry, intent mode, org id) and emit.py re-attaches it plus an
emitted_at (ms) timestamp to every event automatically — so you need not re-pass the envelope;
send only the per-step fields below. During dispatch you add the per-step lifecycle and the terminal
run event:
# SID is the run-local SESSION_ID captured in Step 5.
# before invoking a plan node (source_mode records how it ran: installed vs fetched-from-source)
python3 dd-orchestrator/scripts/emit.py skill_step --action started --session-id "$SID" \
--field plan_position=<n> --field skill_id=<id> --field skill_kind=<kind> \
--field product=<product> --field source_repo=<repo> --field source_mode=<installed|fetched>
# after it returns
python3 dd-orchestrator/scripts/emit.py skill_step --action finished --session-id "$SID" \
--field plan_position=<n> --field skill_id=<id> --field result=success \
--field duration_ms=<ms> --field skill_invoked=true \
--field instrumentation_invoked=<true if it was an install/connect/enable skill>
# if a node is NOT run (failed prerequisite, user declined, no automation, source unreachable)
python3 dd-orchestrator/scripts/emit.py skill_step --action skipped --session-id "$SID" \
--field plan_position=<n> --field skill_id=<id> --field result=skipped_dependency
# once, when the run reaches a terminal state
python3 dd-orchestrator/scripts/emit.py skill_run --action finished --session-id "$SID" \
--field result=<success|partial_success|failed|blocked|cancelled> \
--field step_success_count=<n> --field step_failed_count=<n> --field step_skipped_count=<n>Field values are bounded enums / ids / counts only — never send goal text, paths, keys, URLs, or model output (the emitter also strips anything not on its allow-list).
Reliability and reconciliation. Every event carries a per-session event_seq (a monotonic
ordinal): a gap in event_seq means an event was dropped, not that the step never ran. resolve.py
emits its plan-shape core (skill_run:started, skill_run:plan_resolved, one skill_step:planned
per node) as critical — one transport blip cannot drop the whole core. emit.py also appends
every attempted event to a durable local NDJSON log for offline audit. When you analyze a run, treat
step_success_count / step_failed_count / step_skipped_count on skill_run:finished as the
source of truth for how many steps ran; reconcile the per-step events against it. Do not assume a
missing per-step event means the step did not run. emit.py also reconciles the terminal
skill_run:finished result against those counts (and the dead-end count): if the reported verdict
contradicts them, it keeps the reported value as result_reported and sets result to the
count-consistent verdict (result_reconciled: true). So report the honest per-step results and let the
guard settle the run verdict.
Goal "monitor my Node service on Kubernetes" → recommender [APM, Infrastructure Monitoring],
detected platform=kubernetes:
1. dd-account-setup (foundation) [self]
2. apm-agent-install-kubernetes (platform-install) [agent-skills]
3. apm-enable-kubernetes (product-enable/apm) [agent-skills]
4. apm-verify-ssi-kubernetes (verify) [agent-skills]One product plus a detected platform became a four-skill plan drawn from two repos, correctly ordered, with Infrastructure Monitoring delivered by the same SSI Agent install. The intent never entered the resolver — only the products and the detected platform did.
© datadog-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts) in dd-orchestrator of datadog-labs/agent-skills.
Open the folder on GitHubat commit d2411cc
Dd Orchestrator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dd Orchestrator this skilldatadog-labs/agent-skills | 177 | — | ~6.7k | Automated safety check: Pass | MIT | |
| Tsh Implementing ObservabilityTheSoftwareHouse/copilot-collections | 284 | — | ~2k | Automated safety check: Pass | MIT | |
| Phy Cron ExplainerLeoYeAI/openclaw-master-skills | 2.2k | — | ~7k | Automated safety check: Pass | Apache-2.0 | |
| Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit | 260 | 6 repos | ~1.1k | Automated safety check: Notes | Custom licence | |
| Install Boltmcpboltmcp/boltmcp | 371 | — | ~2.3k | Automated safety check: Pass | None | |
| Provider Bug Reviewmondoohq/mql | 412 | — | ~2.9k | Automated safety check: Pass | Custom licence |
TheSoftwareHouse/copilot-collections
Observability patterns for logging, monitoring, alerting, and distributed tracing.
LeoYeAI/openclaw-master-skills
Cron expression explainer, validator, and converter. An agent skill from LeoYeAI/openclaw-master-skills.
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
boltmcp/boltmcp
A skill your agent uses when asked to help install or uninstall BoltMCP
mondoohq/mql
Deep static code review of an mql provider for logic errors, nil-handling bugs, pagination truncation, caching/id collisions, and other defects that silently give users wrong data.
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
datadog-labs/agent-skills
Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.
datadog-labs/agent-skills
Ensure the user has an authenticated Datadog account with a valid DDAPIKEY on the right region before any Datadog setup or instrumentation.
datadog-labs/agent-skills
APM - install, onboard, instrument, enable, set up, configure, traces, services, dependencies, performance analysis, Data Streams Monitoring (DSM), queue lag, pipeline latency.
datadog-labs/agent-skills
Install the Datadog Agent on Kubernetes using the Datadog Operator — required before enabling Single Step Instrumentation (SSI), which automatically instruments applications for APM without code…
datadog-labs/agent-skills
Set up the Datadog AWS integration with Terraform - creates the cross-account IAM role Datadog assumes (external ID, no stored credentials), attaches the permission policies Datadog publishes, and…
datadog-labs/agent-skills
Install the Datadog Agent on Linux hosts via SSH with Single Step Instrumentation (SSI) enabled — SSI automatically instruments applications for APM without code changes.
Works with
Categories
Entry point for Datadog onboarding. An agent skill from datadog-labs/agent-skills. Dd Orchestrator is an agent skill from datadog-labs/agent-skills. Entry point for Datadog onboarding.
Dd Orchestrator fits situations like: the user says set up Datadog; onboard my app / this repo to Datadog; instrument my project; states a monitoring goal without naming a specific product.
Run `npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a claude-code`. Or copy the skill folder (dd-orchestrator in datadog-labs/agent-skills) into .claude/skills/dd-orchestrator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a codex`. Or copy the skill folder (dd-orchestrator in datadog-labs/agent-skills) into .agents/skills/dd-orchestrator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadog-labs/agent-skills --skill dd-orchestrator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dd-orchestrator, .gemini/skills/dd-orchestrator, .github/skills/dd-orchestrator and .opencode/skills/dd-orchestrator in your project.
Going by SKILL.md and its folder, Dd Orchestrator needs Python for the scripts in its folder, the command-line tools its instructions call (python3, curl and kind) and credentials named DD_APP_KEY and DD_API_KEY. Our summary lists: Python 3; Docker; A credential in DD_APP_KEY; A credential in DD_API_KEY.
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Dd Orchestrator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.7k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Dd Orchestrator: Tsh Implementing Observability (TheSoftwareHouse/copilot-collections, 284 stars), Phy Cron Explainer (LeoYeAI/openclaw-master-skills, 2.2k stars), Senior DevOps Toolkit (maslennikov-ig/claude-code-orchestrator-kit, 260 stars) and Install Boltmcp (boltmcp/boltmcp, 371 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
datadog-labs (a GitHub organization) maintains it in datadog-labs/agent-skills, which has 177 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 8, 2026.
Source: datadog-labs/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.