Apm Integrations
DataDog/dd-trace-java
Write a new library instrumentation end-to-end. An agent skill from DataDog/dd-trace-java.
Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or…
$ npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-online-experiment --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-observability/agent-observability-online-experiment .claude/skills/agent-observability-online-experiment && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-observability-online-experiment" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-online-experiment into .claude/skills/agent-observability-online-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-online-experiment", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-online-experimentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-online-experiment --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/agent-observability/agent-observability-online-experiment .agents/skills/agent-observability-online-experiment && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-observability-online-experiment" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-online-experiment into .agents/skills/agent-observability-online-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-online-experiment", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-online-experiment --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/agent-observability/agent-observability-online-experiment .cursor/skills/agent-observability-online-experiment && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-observability-online-experiment" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-online-experiment into .cursor/skills/agent-observability-online-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-online-experiment", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/datadog-labs/agent-skills.git --path agent-observability/agent-observability-online-experiment--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-online-experiment --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/agent-observability/agent-observability-online-experiment .gemini/skills/agent-observability-online-experiment && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-observability-online-experiment" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-online-experiment into .gemini/skills/agent-observability-online-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-online-experiment", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install datadog-labs/agent-skills agent-observability-online-experimentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/agent-observability/agent-observability-online-experiment .github/skills/agent-observability-online-experiment && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-observability-online-experiment" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-online-experiment into .github/skills/agent-observability-online-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-online-experiment", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-online-experiment --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/agent-observability/agent-observability-online-experiment .opencode/skills/agent-observability-online-experiment && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-observability-online-experiment" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-online-experiment into .opencode/skills/agent-observability-online-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-online-experiment", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-observability-online-experimentLoad this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or…
Agent Observability Online Experiment is an agent skill from datadog-labs/agent-skills. Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or behavior using a Datadog feature flag, score metrics from evaluations, and cost and token metrics. Covers creating the flag, wiring the app, creating the experiment, attaching metrics, setting the traffic split, and starting it. Triggers: online experiment, A/B test an LLM app, test a prompt change on live traffic…
Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Observability and A/B testing. It works with Datadog. The repository describes itself as: Public repository for Datadog Agent Skills. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d2411cc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Observability Online Experiment loads about 6.1k tokens when it runs. Until then it costs about 150 tokens; SKILL.md has 3,541 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from datadog-labs/agent-skills at commit d2411cc, republished under its MIT licence (© datadog-labs). 3,541 words, ~6,081 tokens.
.claude/skills/agent-observability-online-experiment/SKILL.md (or your agent's skills folder).An online experiment splits live traffic between a control and a treatment version of an LLM application. A Datadog feature flag assigns each subject to a variant, and an Agent Observability evaluation score (plus cost and token usage) measures the outcome. This skill walks through the whole setup, from the flag to starting the experiment.
Online experiments for Agent Observability may require the feature to be enabled for the user's organization. If any step reports the feature is unavailable, tell the user to contact their Datadog representative rather than retrying.
Do not use it for offline experiments over a dataset (use the LLM Observability experiment tooling), for analyzing an experiment that already ran (agent-observability-experiment-analyzer or the Datadog experiments tooling), or for promoting an existing offline experiment.
Do not interrupt the investigation to ask for one input at a time. First perform all available read-only discovery:
llmobs, experiments, and feature-flags capabilities needed by this workflow: environment discovery; flag lookup, creation, enablement, and allocation management; experiment lookup, creation, linking, and start; and experiment-metric lookup, creation, and attachment. The skill's toolset configuration controls its visibility with OR semantics; it does not guarantee that all three toolsets are enabled. If capabilities are missing, report every missing group and the toolset to enable, then stop before any write instead of repeatedly retrying individual operations.ml_app, DD_ENV, and stable subject identifier.Then establish the complete proposal. Infer from context and use the defaults below wherever reasonable:
| Input | Notes |
|---|---|
| Application and its ML app name | The ml_app the application reports to Agent Observability. |
| Control and treatment | What differs between them. Usually the control is the current code path and the treatment is the change under test. |
| Hypothesis | Draft one or two sentences stating the expected mechanism and effect on quality, tokens, or cost. |
| Score metric(s) | Recommend the most relevant existing numeric evaluation score, including its desired-change direction. See Phase 4. |
| Environment | Infer from deployment configuration and the user's live-traffic goal. An experiment links to exactly one environment. |
| Traffic split | Default 50/50 control/treatment. |
| Flag name | Default to a short, descriptive, kebab-case name. |
| Subject | What counts as one subject. See Phase 2. |
When metric-name search is needed, try naming variants before concluding that no metric exists: spaces, hyphens, underscores, and a meaningful substring. Inspect duplicate definitions rather than choosing by name alone. Prefer a definition already used by active experiments only when its source, columns, filters, aggregation, and desired-change direction match the proposal. Do not infer ML-app scope from a metric's name. If the user's requested existing metric lacks an ML-app filter or otherwise differs from the recommended definition, disclose that in the proposal and recommend whether to reuse or replace it.
Present one compact confirmation request containing:
ml_app;End with one question: Confirm this complete setup, or tell me what to change. Use follow-up questions only for a genuinely blocking ambiguity that cannot be resolved from code, server state, or a safe default.
Create a Boolean flag with two variants: control = false and treatment = true. Make control the default variant.
FEATURE_GATE allocation in Phase 5 after linking. Extra rules interfere with the split.Listing environments gives the available ones (typically Production, Staging, Development). Note which are production: some actions on production are not available through the server and must be done in the Datadog UI.
The application must do three things for every subject: evaluate the flag, run the matching behavior, and report the score and spans under the same subject identifier.
Choose the subject identifier. It must be stable across the whole interaction being measured.
Use the same value in three places: the feature flag evaluation context's targetingKey, the subject_identifier tag on the evaluation score (the join key the Datadog online experiments guide requires), and the subject_identifier tag on the root span of each trace, which is what the cost and token metrics read. If these differ, Datadog cannot join exposure to outcome and the experiment shows missing metric data.
Evaluate the flag. Follow the Feature Flags SDK for the application's language. For Python this means a recent ddtrace plus the OpenFeature SDK, registering the Datadog provider once, and evaluating a Boolean flag with targeting_key set to the subject identifier. The SDK needs the Datadog API key, site, and DD_ENV set. Default the evaluation to false (control) and catch failures so a flag outage degrades to control instead of breaking the app.
Feature-flag management guidance may point to a React integration resource regardless of the application's language. Treat that as generic tool guidance, not as an instruction to add React code. Follow the SDK for the application's actual language and prefer an existing working provider/evaluation pattern in the repository.
Choose where the variant is applied. If the behavior is baked into state that persists (for example a system prompt stored in a conversation object), evaluate the flag once when that state is created, not on every call, so a subject cannot change variants mid-interaction. Derive the treatment from the control in code where possible so the two cannot drift.
Report the score. Submit the score evaluation with the Agent Observability SDK and add subject_identifier to its tags. Only evaluations with a numeric score metric type can be experiment metrics; Boolean, categorical, and other types cannot.
Tag the spans and the evaluations. Add the subject_identifier tag (for example subject_identifier:73b2efe1-03e3-40c1-ab6a-bd2a6cfbc865) to the Agent Observability spans and to the evaluation scores in the application's own code. Datadog converts the tag into the event's @usr.id field during ingestion, and that field is what the experiment joins to the flag exposure. Set the tag, not @usr.id: the SDK does not set that field directly, and the conversion happens in Datadog's backend.
Cost and token metrics read each completed root trace's rollup, so the tag must be on the trace's root span: the outermost workflow or agent span. Annotate that span through the SDK's span annotation call. For a span-scoped managed evaluator, also add the same tag to every span that evaluator scores because its outcome derives the subject from the evaluated span's own tags. A trace-scoped managed evaluator can propagate the root span's subject to its outcomes. Verify the evaluator's scope before reusing its metric. Tagging evaluated child spans does not double count cost or tokens because those metrics still measure only root traces.
The code must be deployed for the experiment to collect anything. Tell the user plainly that the flag, allocation, and experiment can exist before the deploy but record nothing until the application evaluates the flag in the linked environment, and that the application's DD_ENV must match that environment.
@usr.id, matching the field produced from Phase 2's subject_identifier tag. Starting requires a subject type, but do not accept an arbitrary server-selected default: subject types mapped to another field cannot join these Agent Observability outcomes. Explicitly select an @usr.id-compatible subject type before proceeding. If the available capabilities cannot inspect or configure the mapping, direct the user to the Datadog UI and stop until it is confirmed.An alternative to creating the flag first is a single operation that creates the flag together with an experiment-linked allocation (the create-experiment-feature-flag tool). It needs the experiment and the environment to exist up front. The step-by-step path above keeps each write small and easy to read back, which is why this skill uses it.
Report the experiment's name and ID. Construct its direct Product Analytics URL as
https://<app-host>/product-analytics/experiments/<experiment-id>, using the ID returned by the server and the application host for the organization's Datadog site. For US1 use app.datadoghq.com; for EU use app.datadoghq.eu; regional sites such as us5.datadoghq.com use that host directly; and datad0g.com uses dd.datad0g.com. Include this link in the final handoff after completing the setup instructions, whether the experiment remains a draft or has been started.
List the existing experiment metrics first. If a metric already measures the user's score for this ML app, reuse it. If any existing metric is built on Agent Observability data, read its stored definition and copy its column names instead of guessing them. The same applies to cost and token metrics.
The primary metric is the score approved in Phase 0. If the user named none, list the evaluations that exist for the ML app and recommend the most relevant numeric score in the single setup proposal. Do not pause here for another choice unless read-back contradicts the approved definition.
METRIC_INCREASES or METRIC_DECREASES. This field is required, and a wrong value inverts how every result reads, so include it explicitly in the Phase 0 proposal.Example definition (the columns are the standard evaluation event fields: label, score value, and ML app):
data source: DATADOG
source type/subtype: LLMOBS / LLMOBS_EVAL_METRICS
aggregation: average of @score_value
source filter: @ml_app = <ml_app> AND @label = <evaluation label>
desired change: METRIC_INCREASES (for a score where higher is better)Add two secondary metrics so the user can see what the change costs. Both measure completed LLM Observability root traces and are averaged per trace:
@trace.total_tokens. Desired change: METRIC_DECREASES.@trace.estimated_total_cost. Desired change: METRIC_DECREASES.Use the Agent Observability span source (LLMOBS / LLMOBS_SPANS) with the average operation and column type int. Filter on @ml_app = <ml_app> through the ordinary event filters so traces from other applications used by the same subject cannot enter the result. Do not add filters for trace completion or root-span status: the platform applies those eligibility rules itself and a metric definition cannot override them. Do not sum over all spans, which counts child spans and unfinished traces. If the organization already has an LLM Observability span metric, copy its column names and type from its stored definition instead of relying on the names here, but reuse it only if its stored filter already scopes it to this @ml_app. The trace rollups are expected to be integers. If the metric shows no data, check the column type first, since the stored type must match what the platform exposes.
Attach one at a time, never concurrently:
Read the experiment back and confirm the primary metric is set. Some experiment reads omit secondary metrics; in that case use the successful attachment response's complete metric set as the confirmation and say plainly that the general experiment read does not expose them.
An experiment links to exactly one environment through one allocation.
Enable the flag in the environment(s) the user asked for. Enabling a flag may be refused for a production environment through the server, which directs the user to the UI. If that happens, give the user the flag's UI link and ask them to enable it there. Changing allocations (step 3) is a separate operation and can still be available for production.
List the flag's complete allocation set in the target environment and in any other environment already linked to this experiment. Setting allocations is a full replacement, so preserve every unrelated allocation and all of its settings in each replacement request.
Add or update exactly one feature-gate allocation linked to the experiment, while retaining unrelated allocations in the complete replacement list. Its variant weights are percentages that total exactly 100 (for example 50 control and 50 treatment). Include an exposure schedule in the same call, or the experiment will not start (see the next paragraph). Setting the allocation without one succeeds, and the start step later rejects it as not ready. The allocation needs an exposure schedule with all of these fields, or the start step reports it as not ready. This holds even though the generic allocation schema says to omit the schedule for feature gates: that advice is for gates that are not linked to an experiment.
selection_interval_ms (required by that strategy);none, so the schedule does not begin by itself and the experiment's own start controls timing;exposure_ratio 1.0 (a 0-1 ratio, not 100), is_pause_record false, and grouped_step_index 0, so the percentages in the variant weights apply to all traffic;control_variant_key, naming the control variant (the allocation can be saved without it, but the start step then reports invalid allocation weights).Do not use a partial rollout ramp here. The 50/50 split is already expressed by the weights.
If the experiment is already linked to an allocation in another environment, read that environment's complete allocation set and remove only the allocation linked to this experiment, preserving every other allocation and its settings, before creating the new allocation. Do not replace an environment with an empty list unless the approved setup explicitly removes all of its allocations. If an allocation key already exists, identify the conflicting allocation: reuse it only when it belongs to the approved setup; otherwise choose a new unique key. Ask again if this recovery differs materially from the approved proposal.
Read the allocations back and confirm the environment, the weights, and the experiment link.
Production allocation changes take effect immediately. If the environment requires approval for flag changes, the change may wait for that approval instead of applying. The approved Phase 0 proposal is the environment and production confirmation; do not ask again unless the existing allocation set differs materially from what that proposal said would be replaced.
Start only when the user explicitly asks. Immediately before starting, independently verify rather than relying only on the start operation's readiness check:
Do not invoke start while the linked environment is known to be disabled. The start readiness check may validate allocation structure without enforcing flag enablement or deployment, so a successful start does not prove traffic can flow.
Then start the experiment (the start-experiment tool). It runs its own readiness check and returns the structural blockers it can detect, each with an action. Apply all of the actions, then retry once. Expect to need more than one pass: it reports the blockers it can currently detect, and fixing one can reveal the next (for example a missing exposure schedule first, then an invalid control variant on that schedule). Repairing an allocation replaces the flag's allocation set, so read the allocations first; ask again only when the required replacement differs from the one approved in Phase 0. If it returns a permissions error, report that the account lacks permission to start experiments, stop, and offer the UI's preview-and-start flow. Do not loop on retries.
After it starts, read the experiment back and report its status.
Once the application is deployed and receiving traffic:
subject_identifier matches the targetingKey exactly; the root span carries the tag; the traces are complete; the metric's column names and types are right.For reading the results once enough data has arrived, use agent-observability-experiment-analyzer or the Datadog experiments tooling available in the active MCP server.
@usr.id.@ml_app filter.control_variant_key, which blocks the start step.desired change set the wrong way round.© datadog-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in agent-observability/agent-observability-online-experiment of datadog-labs/agent-skills.
Open the folder on GitHubat commit d2411cc
Agent Observability Online Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Observability Online Experiment this skilldatadog-labs/agent-skills | 177 | — | ~6.1k | Automated safety check: Pass | MIT | |
| Apm IntegrationsDataDog/dd-trace-java | 736 | — | ~3.7k | Automated safety check: Notes | Apache-2.0 | |
| Redis Observabilityredis/agent-skills | 166 | 2 repos | ~911 | Automated safety check: Pass | MIT | |
| Frontmcp Observabilityagentfront/frontmcp | 146 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Monitoring Observabilityahmedasmar/devops-claude-skills | 203 | — | ~3.9k | Automated safety check: Pass | None | |
| Observability Architecturemajiayu000/litellm-rs | 117 | — | ~1.3k | Automated safety check: Pass | MIT |
DataDog/dd-trace-java
Write a new library instrumentation end-to-end. An agent skill from DataDog/dd-trace-java.
redis/agent-skills
Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO…
agentfront/frontmcp
A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
majiayu000/litellm-rs
LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.
EliasOulkadi/shokunin
Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…
datadog-labs/agent-skills
Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.
datadog-labs/agent-skills
Ensure the user has an authenticated Datadog account with a valid DDAPIKEY on the right region before any Datadog setup or instrumentation.
datadog-labs/agent-skills
Entry point for Datadog onboarding. An agent skill from datadog-labs/agent-skills.
datadog-labs/agent-skills
APM - install, onboard, instrument, enable, set up, configure, traces, services, dependencies, performance analysis, Data Streams Monitoring (DSM), queue lag, pipeline latency.
datadog-labs/agent-skills
Install the Datadog Agent on Kubernetes using the Datadog Operator — required before enabling Single Step Instrumentation (SSI), which automatically instruments applications for APM without code…
datadog-labs/agent-skills
Set up the Datadog AWS integration with Terraform - creates the cross-account IAM role Datadog assumes (external ID, no stored credentials), attaches the permission policies Datadog publishes, and…
Works with
Categories
Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or…. Agent Observability Online Experiment is an agent skill from datadog-labs/agent-skills. Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or behavior using a Datadog feature flag, score metrics from evaluations, and cost and token metrics.
Agent Observability Online Experiment fits situations like: behavior using a Datadog feature flag; score metrics from evaluations; cost and token metrics.
Run `npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a claude-code`. Or copy the skill folder (agent-observability/agent-observability-online-experiment in datadog-labs/agent-skills) into .claude/skills/agent-observability-online-experiment in your project. Claude Code loads it when a task matches its description.
Run `npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a codex`. Or copy the skill folder (agent-observability/agent-observability-online-experiment in datadog-labs/agent-skills) into .agents/skills/agent-observability-online-experiment in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-observability-online-experiment, .gemini/skills/agent-observability-online-experiment, .github/skills/agent-observability-online-experiment and .opencode/skills/agent-observability-online-experiment in your project.
SKILL.md names no scripts, command-line tools or credentials: Agent Observability Online Experiment is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agent Observability Online Experiment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agent Observability Online Experiment: Apm Integrations (DataDog/dd-trace-java, 736 stars), Redis Observability (redis/agent-skills, 166 stars), Frontmcp Observability (agentfront/frontmcp, 146 stars) and Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
datadog-labs (a GitHub organization) maintains it in datadog-labs/agent-skills, which has 177 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 8, 2026.
Source: datadog-labs/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.