Agent skill

Agent Observability Online Experiment

by datadog-labs in datadog-labs/agent-skills

Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or…

MITAuto-check passedDevOps & Cloud

Install Agent Observability Online Experiment

skills CLI
$ npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install datadog-labs/agent-skills agent-observability-online-experiment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-observability/agent-observability-online-experiment .claude/skills/agent-observability-online-experiment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-observability-online-experiment
GitHub stars
177
Token cost
~6.1k tokens
SKILL.md length
3,541 words
Files
1
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or…

  • Works in 8 steps: Discover, recommend, and confirm once → Create the feature flag → Wire the application → …
  • Behavior using a Datadog feature flag
  • SKILL.md covers When to use this skill, Ground rules, Phase 0 — Discover, recommend,… and Phase 1 — Create the feature…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Observability Online Experiment is an agent skill from datadog-labs/agent-skills. Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or behavior using a Datadog feature flag, score metrics from evaluations, and cost and token metrics. Covers creating the flag, wiring the app, creating the experiment, attaching metrics, setting the traffic split, and starting it. Triggers: online experiment, A/B test an LLM app, test a prompt change on live traffic…

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Observability and A/B testing. It works with Datadog. The repository describes itself as: Public repository for Datadog Agent Skills. The licence is MIT.

When your agent uses it

  • Behavior using a Datadog feature flag
  • Score metrics from evaluations
  • Cost and token metrics

Example prompts

  • “/agent-observability-online-experiment”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Discover, recommend, and confirm once
  2. Create the feature flag
  3. Wire the application
  4. Create the experiment and link the flag
  5. Create and attach metrics
  6. Environments and traffic split
  7. Start the experiment
  8. Verify once traffic flows

What it can do on your machine

Read from SKILL.md and the folder at commit d2411cc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Observability Online Experiment loads about 6.1k tokens when it runs. Until then it costs about 150 tokens; SKILL.md has 3,541 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~150
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from datadog-labs/agent-skills at commit d2411cc, republished under its MIT licence (© datadog-labs). 3,541 words, ~6,081 tokens.

Download SKILL.mdSave it as .claude/skills/agent-observability-online-experiment/SKILL.md (or your agent's skills folder).
name
agent-observability-online-experiment
description
Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or behavior using a Datadog feature flag, score metrics from evaluations, and cost and token metrics. Covers creating the flag, wiring the app, creating the experiment, attaching metrics, setting the traffic split, and starting it. Triggers: online experiment, A/B test an LLM app, test a prompt change on live traffic, experiment with feature flag and evaluation score.

Set up an online experiment on an LLM application

An online experiment splits live traffic between a control and a treatment version of an LLM application. A Datadog feature flag assigns each subject to a variant, and an Agent Observability evaluation score (plus cost and token usage) measures the outcome. This skill walks through the whole setup, from the flag to starting the experiment.

Online experiments for Agent Observability may require the feature to be enabled for the user's organization. If any step reports the feature is unavailable, tell the user to contact their Datadog representative rather than retrying.

When to use this skill

  • The user wants to test a change (prompt, model, retrieval strategy, tool set) on real traffic rather than a fixed dataset.
  • The user has a change on a branch and wants to know whether it is better, cheaper, or both in production.
  • The user wants an experiment whose metrics come from Agent Observability evaluations, token counts, or cost.

Do not use it for offline experiments over a dataset (use the LLM Observability experiment tooling), for analyzing an experiment that already ran (agent-observability-experiment-analyzer or the Datadog experiments tooling), or for promoting an existing offline experiment.

Ground rules

  • Use one setup approval checkpoint. Before any write, finish the read-only discovery in Phase 0, make sensible recommendations, and present the complete proposed setup in one confirmation request. Include the flag, variants, metric definitions or reused metric IDs, desired-change directions, environment, subject, traffic split, code change, and every organization write you intend to make. Once the user approves that proposal, do not ask again before each flag, draft, link, metric, enablement, or allocation write that was clearly listed.
  • Make production risk explicit in that same checkpoint. If the proposal enables a production flag or replaces production allocations, say that the change takes effect immediately. A clear approval of that production proposal is sufficient; do not add a second confirmation that repeats the same information. Ask again only if read-back reveals a materially different or destructive replacement that was not in the approved proposal.
  • Starting is separate unless already explicit. Start only if the user explicitly asked to start the experiment. If the original request included starting, include it in the single setup proposal; otherwise finish with a draft and ask no start question until the user requests it.
  • Go one step at a time and read back. After each write, read the object back and report what the server actually stored, not what you sent.
  • Never attach metrics in parallel or alongside other metric edits. The attach operation sends the experiment's complete metric set to the server, so concurrent edits can overwrite each other.
  • Report failures plainly. If a step is refused (permissions, production restrictions), say so and give the user the manual alternative. Do not look for a workaround.
  • Do not guess identifiers. Environment IDs, flag IDs, metric IDs, and variant keys all come from the server.

Phase 0 — Discover, recommend, and confirm once

Do not interrupt the investigation to ask for one input at a time. First perform all available read-only discovery:

  • Inspect the active tool list and verify that the session has the llmobs, experiments, and feature-flags capabilities needed by this workflow: environment discovery; flag lookup, creation, enablement, and allocation management; experiment lookup, creation, linking, and start; and experiment-metric lookup, creation, and attachment. The skill's toolset configuration controls its visibility with OR semantics; it does not guarantee that all three toolsets are enabled. If capabilities are missing, report every missing group and the toolset to enable, then stop before any write instead of repeatedly retrying individual operations.
  • Inspect the application and its deployment configuration. Confirm it reports to Agent Observability, find where it submits evaluations (if it does), find the code path the treatment changes, and infer its ml_app, DD_ENV, and stable subject identifier.
  • List environments, look for an existing flag and experiment for the proposed change, and search existing experiment metrics before proposing new objects.

Then establish the complete proposal. Infer from context and use the defaults below wherever reasonable:

InputNotes
Application and its ML app nameThe ml_app the application reports to Agent Observability.
Control and treatmentWhat differs between them. Usually the control is the current code path and the treatment is the change under test.
HypothesisDraft one or two sentences stating the expected mechanism and effect on quality, tokens, or cost.
Score metric(s)Recommend the most relevant existing numeric evaluation score, including its desired-change direction. See Phase 4.
EnvironmentInfer from deployment configuration and the user's live-traffic goal. An experiment links to exactly one environment.
Traffic splitDefault 50/50 control/treatment.
Flag nameDefault to a short, descriptive, kebab-case name.
SubjectWhat counts as one subject. See Phase 2.

When metric-name search is needed, try naming variants before concluding that no metric exists: spaces, hyphens, underscores, and a meaningful substring. Inspect duplicate definitions rather than choosing by name alone. Prefer a definition already used by active experiments only when its source, columns, filters, aggregation, and desired-change direction match the proposal. Do not infer ML-app scope from a metric's name. If the user's requested existing metric lacks an ML-app filter or otherwise differs from the recommended definition, disclose that in the proposal and recommend whether to reuse or replace it.

Present one compact confirmation request containing:

  • application and ml_app;
  • control, treatment, and hypothesis;
  • proposed flag key and Boolean variants;
  • subject identifier and code location where the flag will be evaluated;
  • target environment and 50/50 or proposed split;
  • primary evaluation metric, stored definition or proposed definition, and desired-change direction;
  • token and cost secondary metrics;
  • whether existing objects will be reused or new ones created;
  • the complete list of writes, including production enablement/allocation when applicable;
  • whether starting is included (only when the user already requested it).

End with one question: Confirm this complete setup, or tell me what to change. Use follow-up questions only for a genuinely blocking ambiguity that cannot be resolved from code, server state, or a safe default.

Phase 1 — Create the feature flag

Create a Boolean flag with two variants: control = false and treatment = true. Make control the default variant.

  • Do not add allocations, targeting rules, or a percentage rollout when creating the flag. Linking it to an experiment does not create an allocation by itself: create or update one experiment-linked FEATURE_GATE allocation in Phase 5 after linking. Extra rules interfere with the split.
  • Check first whether a flag with that key already exists, and reuse it if the user says so.
  • If the organization has no feature flags at all, the server may point to a guided onboarding flow for first-time setup. Direct creation is fine here when the user has asked for a specific flag; mention the onboarding flow only if the user seems new to Datadog Feature Flags.
  • Record the flag's ID and the IDs of its two variants.

Listing environments gives the available ones (typically Production, Staging, Development). Note which are production: some actions on production are not available through the server and must be done in the Datadog UI.

Phase 2 — Wire the application

The application must do three things for every subject: evaluate the flag, run the matching behavior, and report the score and spans under the same subject identifier.

Choose the subject identifier. It must be stable across the whole interaction being measured.

  • Application with signed-in users: a stable user ID.
  • No users, but sessions: the session ID.
  • Autonomous workflows: generate one UUID at the start of a run and reuse it for the whole run.

Use the same value in three places: the feature flag evaluation context's targetingKey, the subject_identifier tag on the evaluation score (the join key the Datadog online experiments guide requires), and the subject_identifier tag on the root span of each trace, which is what the cost and token metrics read. If these differ, Datadog cannot join exposure to outcome and the experiment shows missing metric data.

Evaluate the flag. Follow the Feature Flags SDK for the application's language. For Python this means a recent ddtrace plus the OpenFeature SDK, registering the Datadog provider once, and evaluating a Boolean flag with targeting_key set to the subject identifier. The SDK needs the Datadog API key, site, and DD_ENV set. Default the evaluation to false (control) and catch failures so a flag outage degrades to control instead of breaking the app.

Feature-flag management guidance may point to a React integration resource regardless of the application's language. Treat that as generic tool guidance, not as an instruction to add React code. Follow the SDK for the application's actual language and prefer an existing working provider/evaluation pattern in the repository.

Choose where the variant is applied. If the behavior is baked into state that persists (for example a system prompt stored in a conversation object), evaluate the flag once when that state is created, not on every call, so a subject cannot change variants mid-interaction. Derive the treatment from the control in code where possible so the two cannot drift.

Report the score. Submit the score evaluation with the Agent Observability SDK and add subject_identifier to its tags. Only evaluations with a numeric score metric type can be experiment metrics; Boolean, categorical, and other types cannot.

Tag the spans and the evaluations. Add the subject_identifier tag (for example subject_identifier:73b2efe1-03e3-40c1-ab6a-bd2a6cfbc865) to the Agent Observability spans and to the evaluation scores in the application's own code. Datadog converts the tag into the event's @usr.id field during ingestion, and that field is what the experiment joins to the flag exposure. Set the tag, not @usr.id: the SDK does not set that field directly, and the conversion happens in Datadog's backend.

Cost and token metrics read each completed root trace's rollup, so the tag must be on the trace's root span: the outermost workflow or agent span. Annotate that span through the SDK's span annotation call. For a span-scoped managed evaluator, also add the same tag to every span that evaluator scores because its outcome derives the subject from the evaluated span's own tags. A trace-scoped managed evaluator can propagate the root span's subject to its outcomes. Verify the evaluator's scope before reusing its metric. Tagging evaluated child spans does not double count cost or tokens because those metrics still measure only root traces.

The code must be deployed for the experiment to collect anything. Tell the user plainly that the flag, allocation, and experiment can exist before the deploy but record nothing until the application evaluates the flag in the linked environment, and that the application's DD_ENV must match that environment.

  1. Create a draft experiment with a clear name and the hypothesis. The primary metric can be attached afterwards, so it is fine to create the draft before the metrics exist.
  2. Link the flag from Phase 1 to the experiment.
  3. Read the experiment back. A draft may report an incomplete set-up state until it is started; that alone is not an error. The start step reports concrete blockers if there are any.
  4. Resolve the draft's subject type and verify that its Product Analytics attribute is @usr.id, matching the field produced from Phase 2's subject_identifier tag. Starting requires a subject type, but do not accept an arbitrary server-selected default: subject types mapped to another field cannot join these Agent Observability outcomes. Explicitly select an @usr.id-compatible subject type before proceeding. If the available capabilities cannot inspect or configure the mapping, direct the user to the Datadog UI and stop until it is confirmed.

An alternative to creating the flag first is a single operation that creates the flag together with an experiment-linked allocation (the create-experiment-feature-flag tool). It needs the experiment and the environment to exist up front. The step-by-step path above keeps each write small and easy to read back, which is why this skill uses it.

Report the experiment's name and ID. Construct its direct Product Analytics URL as https://<app-host>/product-analytics/experiments/<experiment-id>, using the ID returned by the server and the application host for the organization's Datadog site. For US1 use app.datadoghq.com; for EU use app.datadoghq.eu; regional sites such as us5.datadoghq.com use that host directly; and datad0g.com uses dd.datad0g.com. Include this link in the final handoff after completing the setup instructions, whether the experiment remains a draft or has been started.

Phase 4 — Create and attach metrics

Reuse before creating

List the existing experiment metrics first. If a metric already measures the user's score for this ML app, reuse it. If any existing metric is built on Agent Observability data, read its stored definition and copy its column names instead of guessing them. The same applies to cost and token metrics.

Show full SKILL.md (1,437 more words)Show less
Primary metric: the score evaluation

The primary metric is the score approved in Phase 0. If the user named none, list the evaluations that exist for the ML app and recommend the most relevant numeric score in the single setup proposal. Do not pause here for another choice unless read-back contradicts the approved definition.

  • Source: Agent Observability evaluation metrics, filtered to the ML app and the evaluation label.
  • Aggregation: average of the score value.
  • Desired change: the direction that counts as an improvement, one of METRIC_INCREASES or METRIC_DECREASES. This field is required, and a wrong value inverts how every result reads, so include it explicitly in the Phase 0 proposal.

Example definition (the columns are the standard evaluation event fields: label, score value, and ML app):

data source:        DATADOG
source type/subtype: LLMOBS / LLMOBS_EVAL_METRICS
aggregation:        average of @score_value
source filter:      @ml_app = <ml_app> AND @label = <evaluation label>
desired change:     METRIC_INCREASES (for a score where higher is better)
Secondary metrics: tokens and cost

Add two secondary metrics so the user can see what the change costs. Both measure completed LLM Observability root traces and are averaged per trace:

  • Average trace tokens: average of @trace.total_tokens. Desired change: METRIC_DECREASES.
  • Average trace cost: average of @trace.estimated_total_cost. Desired change: METRIC_DECREASES.

Use the Agent Observability span source (LLMOBS / LLMOBS_SPANS) with the average operation and column type int. Filter on @ml_app = <ml_app> through the ordinary event filters so traces from other applications used by the same subject cannot enter the result. Do not add filters for trace completion or root-span status: the platform applies those eligibility rules itself and a metric definition cannot override them. Do not sum over all spans, which counts child spans and unfinished traces. If the organization already has an LLM Observability span metric, copy its column names and type from its stored definition instead of relying on the names here, but reuse it only if its stored filter already scopes it to this @ml_app. The trace rollups are expected to be integers. If the metric shows no data, check the column type first, since the stored type must match what the platform exposes.

Attach the metrics

Attach one at a time, never concurrently:

  1. Attach the score metric as the primary.
  2. Attach tokens as a secondary metric.
  3. Attach cost as a secondary metric.

Read the experiment back and confirm the primary metric is set. Some experiment reads omit secondary metrics; in that case use the successful attachment response's complete metric set as the confirmation and say plainly that the general experiment read does not expose them.

Phase 5 — Environments and traffic split

An experiment links to exactly one environment through one allocation.

  1. Enable the flag in the environment(s) the user asked for. Enabling a flag may be refused for a production environment through the server, which directs the user to the UI. If that happens, give the user the flag's UI link and ask them to enable it there. Changing allocations (step 3) is a separate operation and can still be available for production.

  2. List the flag's complete allocation set in the target environment and in any other environment already linked to this experiment. Setting allocations is a full replacement, so preserve every unrelated allocation and all of its settings in each replacement request.

  3. Add or update exactly one feature-gate allocation linked to the experiment, while retaining unrelated allocations in the complete replacement list. Its variant weights are percentages that total exactly 100 (for example 50 control and 50 treatment). Include an exposure schedule in the same call, or the experiment will not start (see the next paragraph). Setting the allocation without one succeeds, and the start step later rejects it as not ready. The allocation needs an exposure schedule with all of these fields, or the start step reports it as not ready. This holds even though the generic allocation schema says to omit the schedule for feature gates: that advice is for gates that are not linked to an experiment.

    • a rollout strategy of uniform intervals, with a selection_interval_ms (required by that strategy);
    • a start of none, so the schedule does not begin by itself and the experiment's own start controls timing;
    • a single rollout step with exposure_ratio 1.0 (a 0-1 ratio, not 100), is_pause_record false, and grouped_step_index 0, so the percentages in the variant weights apply to all traffic;
    • the schedule's control_variant_key, naming the control variant (the allocation can be saved without it, but the start step then reports invalid allocation weights).

    Do not use a partial rollout ramp here. The 50/50 split is already expressed by the weights.

  4. If the experiment is already linked to an allocation in another environment, read that environment's complete allocation set and remove only the allocation linked to this experiment, preserving every other allocation and its settings, before creating the new allocation. Do not replace an environment with an empty list unless the approved setup explicitly removes all of its allocations. If an allocation key already exists, identify the conflicting allocation: reuse it only when it belongs to the approved setup; otherwise choose a new unique key. Ask again if this recovery differs materially from the approved proposal.

  5. Read the allocations back and confirm the environment, the weights, and the experiment link.

Production allocation changes take effect immediately. If the environment requires approval for flag changes, the change may wait for that approval instead of applying. The approved Phase 0 proposal is the environment and production confirmation; do not ask again unless the existing allocation set differs materially from what that proposal said would be replaced.

Phase 6 — Start the experiment

Start only when the user explicitly asks. Immediately before starting, independently verify rather than relying only on the start operation's readiness check:

  • The flag is enabled in the linked environment.
  • The 50/50 (or chosen) allocation is in place and linked.
  • The primary metric is attached.
  • The application change is deployed or the user explicitly accepts starting the clock before deployment.

Do not invoke start while the linked environment is known to be disabled. The start readiness check may validate allocation structure without enforcing flag enablement or deployment, so a successful start does not prove traffic can flow.

Then start the experiment (the start-experiment tool). It runs its own readiness check and returns the structural blockers it can detect, each with an action. Apply all of the actions, then retry once. Expect to need more than one pass: it reports the blockers it can currently detect, and fixing one can reveal the next (for example a missing exposure schedule first, then an invalid control variant on that schedule). Repairing an allocation replaces the flag's allocation set, so read the allocations first; ask again only when the required replacement differs from the one approved in Phase 0. If it returns a permissions error, report that the account lacks permission to start experiments, stop, and offer the UI's preview-and-start flow. Do not loop on retries.

After it starts, read the experiment back and report its status.

Phase 7 — Verify once traffic flows

Once the application is deployed and receiving traffic:

  • Run the experiment diagnostics. They check exposure balance and flag zero-data metrics, unreliable metrics, and sample ratio mismatch.
  • If a metric has no data, check in this order: the flag is being evaluated for the subject (the experiment's code-location lookup shows where the flag is used in source); the evaluation has a numeric score; the subject_identifier matches the targetingKey exactly; the root span carries the tag; the traces are complete; the metric's column names and types are right.
  • Fix a wrong metric definition by updating the metric rather than creating a duplicate.

For reading the results once enough data has arrived, use agent-observability-experiment-analyzer or the Datadog experiments tooling available in the active MCP server.

Pitfalls checklist

  • Subject identifier mismatch between the flag, the evaluation, and the spans.
  • Subject tag missing from the root span, so the cost and token metrics cannot join to exposures.
  • Subject type mapped to an attribute other than @usr.id.
  • Span-scoped managed evaluation missing the subject tag on the span it scores.
  • Cost or token metrics defined as a sum over all spans, or on span attributes other than the trace rollups.
  • Cost or token metric missing its @ml_app filter.
  • Flag enabled in one environment but the application running in another.
  • Experiment linked to the wrong environment's allocation (it can only link one).
  • Allocation replacement that drops unrelated allocations or their exposure schedules.
  • Allocation weights not totaling 100.
  • Allocation set without an exposure schedule, or with a schedule that has no control_variant_key, which blocks the start step.
  • Evaluation is Boolean or categorical, not a score.
  • Metric desired change set the wrong way round.
  • Attaching metrics in parallel.
  • Starting before the user has asked, or before the flag is enabled where the app runs.

© datadog-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in agent-observability/agent-observability-online-experiment of datadog-labs/agent-skills.

Open the folder on GitHubat commit d2411cc

Compare with similar skills

Agent Observability Online Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Observability Online Experiment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Observability Online Experiment this skilldatadog-labs/agent-skills177—~6.1kAutomated safety check: PassMIT
Apm IntegrationsDataDog/dd-trace-java736—~3.7kAutomated safety check: NotesApache-2.0
Redis Observabilityredis/agent-skills1662 repos~911Automated safety check: PassMIT
Frontmcp Observabilityagentfront/frontmcp146—~4.6kAutomated safety check: PassApache-2.0
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Observability Architecturemajiayu000/litellm-rs117—~1.3kAutomated safety check: PassMIT

Similar skills

  • Apm Integrations

    DataDog/dd-trace-java

    Official

    Write a new library instrumentation end-to-end. An agent skill from DataDog/dd-trace-java.

    736 GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Redis Observability

    redis/agent-skills

    Official

    Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO…

    166 GitHub starsUsed in 2 repos~911 tokens
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • Observability Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

    117 GitHub stars~1.3k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Error Handler

    EliasOulkadi/shokunin

    Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

    114 GitHub stars~3.6k tokensUpdated 5 days ago
    DevOps & CloudAuto-check: notes

More from datadog-labs/agent-skills

All 39 skills in this repo
  • Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.

    177 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Dd Account Setup

    datadog-labs/agent-skills

    Ensure the user has an authenticated Datadog account with a valid DDAPIKEY on the right region before any Datadog setup or instrumentation.

    177 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check: notes
  • Dd Orchestrator

    datadog-labs/agent-skills

    Entry point for Datadog onboarding. An agent skill from datadog-labs/agent-skills.

    177 GitHub stars~6.7k tokensUpdated yesterday
    Auto-check passed
  • Dd Apm

    datadog-labs/agent-skills

    APM - install, onboard, instrument, enable, set up, configure, traces, services, dependencies, performance analysis, Data Streams Monitoring (DSM), queue lag, pipeline latency.

    177 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Agent Install

    datadog-labs/agent-skills

    Install the Datadog Agent on Kubernetes using the Datadog Operator — required before enabling Single Step Instrumentation (SSI), which automatically instruments applications for APM without code…

    177 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check: warnings
  • Dd AWS Integration

    datadog-labs/agent-skills

    Set up the Datadog AWS integration with Terraform - creates the cross-account IAM role Datadog assumes (external ID, no stored credentials), attaches the permission policies Datadog publishes, and…

    177 GitHub stars~6.8k tokensUpdated yesterday
    Auto-check: notes

Works with

Categories

Questions about Agent Observability Online Experiment

What does Agent Observability Online Experiment do?

Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or…. Agent Observability Online Experiment is an agent skill from datadog-labs/agent-skills. Load this skill when the user wants to run an online experiment (live-traffic A/B test) on an LLM application instrumented with Agent Observability: compare two versions of a prompt, model, or behavior using a Datadog feature flag, score metrics from evaluations, and cost and token metrics.

When should I use Agent Observability Online Experiment?

Agent Observability Online Experiment fits situations like: behavior using a Datadog feature flag; score metrics from evaluations; cost and token metrics.

How do I install Agent Observability Online Experiment in Claude Code?

Run `npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a claude-code`. Or copy the skill folder (agent-observability/agent-observability-online-experiment in datadog-labs/agent-skills) into .claude/skills/agent-observability-online-experiment in your project. Claude Code loads it when a task matches its description.

How do I install Agent Observability Online Experiment in Codex?

Run `npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a codex`. Or copy the skill folder (agent-observability/agent-observability-online-experiment in datadog-labs/agent-skills) into .agents/skills/agent-observability-online-experiment in your project. Codex loads it when a task matches its description.

Can I use Agent Observability Online Experiment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadog-labs/agent-skills --skill agent-observability-online-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-observability-online-experiment, .gemini/skills/agent-observability-online-experiment, .github/skills/agent-observability-online-experiment and .opencode/skills/agent-observability-online-experiment in your project.

What does Agent Observability Online Experiment need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Observability Online Experiment is instructions for the agent only. Our summary lists: Python 3.

Does Agent Observability Online Experiment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Observability Online Experiment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Observability Online Experiment use?

Agent Observability Online Experiment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Observability Online Experiment use?

About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Observability Online Experiment?

Skills that share tags, products or a category with Agent Observability Online Experiment: Apm Integrations (DataDog/dd-trace-java, 736 stars), Redis Observability (redis/agent-skills, 166 stars), Frontmcp Observability (agentfront/frontmcp, 146 stars) and Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Observability Online Experiment?

datadog-labs (a GitHub organization) maintains it in datadog-labs/agent-skills, which has 177 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 8, 2026.

Source: datadog-labs/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.