Failproof AI SDK Integration
FailproofAI/failproofai
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-experiment-bootstrap --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-observability/agent-observability-experiment-bootstrap .claude/skills/agent-observability-experiment-bootstrap && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-observability-experiment-bootstrap" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-experiment-bootstrap into .claude/skills/agent-observability-experiment-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-experiment-bootstrap", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-experiment-bootstrapType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-experiment-bootstrap --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/agent-observability/agent-observability-experiment-bootstrap .agents/skills/agent-observability-experiment-bootstrap && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-observability-experiment-bootstrap" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-experiment-bootstrap into .agents/skills/agent-observability-experiment-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-experiment-bootstrap", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-experiment-bootstrap --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/agent-observability/agent-observability-experiment-bootstrap .cursor/skills/agent-observability-experiment-bootstrap && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-observability-experiment-bootstrap" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-experiment-bootstrap into .cursor/skills/agent-observability-experiment-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-experiment-bootstrap", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/datadog-labs/agent-skills.git --path agent-observability/agent-observability-experiment-bootstrap--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-experiment-bootstrap --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/agent-observability/agent-observability-experiment-bootstrap .gemini/skills/agent-observability-experiment-bootstrap && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-observability-experiment-bootstrap" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-experiment-bootstrap into .gemini/skills/agent-observability-experiment-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-experiment-bootstrap", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install datadog-labs/agent-skills agent-observability-experiment-bootstrapInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/agent-observability/agent-observability-experiment-bootstrap .github/skills/agent-observability-experiment-bootstrap && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-observability-experiment-bootstrap" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-experiment-bootstrap into .github/skills/agent-observability-experiment-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-experiment-bootstrap", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install datadog-labs/agent-skills agent-observability-experiment-bootstrap --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadog-labs/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/agent-observability/agent-observability-experiment-bootstrap .opencode/skills/agent-observability-experiment-bootstrap && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-observability-experiment-bootstrap" agent skill from https://github.com/datadog-labs/agent-skills/tree/main/agent-observability/agent-observability-experiment-bootstrap into .opencode/skills/agent-observability-experiment-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-observability-experiment-bootstrap", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-observability-experiment-bootstrapBootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.
Agent Observability Experiment Bootstrap is an agent skill from datadog-labs/agent-skills. Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK. Use for experiment, dataset, evaluator, benchmark, regression, or LLM-as-a-judge scaffolding. The legacy Python invocation remains supported.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including reference files (for example `references/node/nodejs.md`, `references/python/env_setup_template.py` and `references/python/evaluator-styles/class.md`).
It sits in DevOps & Cloud, covering Observability, LLM observability and Project scaffolding. It works with Python. The repository describes itself as: Public repository for Datadog Agent Skills. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d2411cc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Observability Experiment Bootstrap loads about 2.3k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 73 tokens; SKILL.md has 1,001 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from datadog-labs/agent-skills at commit d2411cc, republished under its MIT licence (© datadog-labs). 1,001 words, ~2,260 tokens.
.claude/skills/agent-observability-experiment-bootstrap/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.Generate one reproducible experiment artifact. The artifact evaluates a task over a versioned dataset, records outputs and evaluator metrics, carries configuration and provenance, and prints a result link or identifiers when possible.
This skill is adapter-independent. Each adapter owns a language-specific directory under references/; load only the selected adapter contract. The directories are intentionally symmetric even when one adapter currently has fewer supporting references.
The installed directory and legacy invocation remain valid:
/agent-observability-experiment-bootstrap [--purpose TEXT] [--format py|ipynb|mjs]
[--dataset PATH | --dataset-name NAME] [--dataset-version N]
[--project-name NAME] [--evaluator-style function|class|remote]
[--jobs N] [--output PATH] [--task-source module:function]
[--placeholder-task] [--app-root PATH] [--env-file PATH]General options:
--adapter python|node # default: python
--format py|ipynb|mjs # Python: py/ipynb; Node: mjs
--site SITE # otherwise DD_SITE or datadoghq.comDo not prompt for optional defaults. Resolve a non-empty purpose from --purpose, the request, or a focused question. Keep the purpose as reasoning context, not a fixed taxonomy.
Load context in this order:
references/python/python.mdreferences/node/nodejs.mdreferences/python/providers/.references/python/evaluator-styles/.Do not load all provider, evaluator, Python, and Node references “for completeness.” The selected reference is the source of truth for syntax and API behavior.
Use Python when the application or requested artifact is Python, or when no adapter is specified. Use Node when the application is JavaScript/TypeScript and the local dd-trace package exposes tracer.llmobs.experiments.
Never mix the Python and Node SDKs in one generated artifact. Do not use private SDK modules or invent a missing symbol. If local source and an installed package disagree, report the discrepancy and generate against the selected version.
Every adapter must represent the following concepts:
expected_output is optional and must not be synthesized from an observed production output without explicit validation. Distinguish a missing value from an intentionally empty object. Dataset tags must use the backend’s validated key:value form where the selected reference requires it.
Derive the purpose and project without guessing across product boundaries. A project is not automatically the same as an ml_app, service, dataset, or repository name. Record how each value was resolved.
Support:
For local JSON, require a top-level array, validate the selected adapter’s record shape, scrub obvious PII and credential-like values, and report affected record indices. Do not invent canonical or remote record IDs.
For CSV, preserve the runtime path and document the dependency. Use the Python CSV column contract from references/python/python.md; Node generation must not pretend that a Python-only CSV helper exists.
Use --task-source when provided. Otherwise use the selected language’s bounded application discovery rules:
Never claim that an invented import is wired. Preserve side-effect warnings for network, database, filesystem, environment, or tool calls.
Select two or three evaluators based on purpose and available signals. Keep labels unique and stable.
Evaluator failures must not become passing values. Summary evaluators must remain distinct from row evaluators.
Use the selected adapter reference for the exact generated code. Include:
Preserve the historical Python section ordering and evaluator/provider reference behavior when using the Python adapter.
Before presenting the artifact:
.py: python -m py_compile <path>..ipynb: parse JSON and require code/markdown cells..mjs: node --check <path>.For every adapter, check for private imports, literal credentials, malformed tags, missing provenance, mismatched dataset versions, fabricated IDs, and task/evaluator errors that were collapsed into false or pass.
Use this compact structure:
Generated LLM Observability experiment: <adapter>/<format>
Path: <path>
Purpose: "<purpose>"
Project: <project>
Dataset: <local path | name>, version=<version or latest>
Task: <wired source | placeholder>
Evaluators: <labels>
Provenance: generated_by=claude-code, adapter=<adapter>, skill=agent-observability-experiment-bootstrap
Validation: <commands and pass/fail>
Result link: <URL or pending until run>
Next steps:
1. Verify the task source and evaluator semantics.
2. Set the credentials required by the selected SDK.
3. Install the selected SDK and run the generated artifact.
4. Review per-row errors before treating metrics as a successful run.Each adapter reference must identify the public source links and branch used to verify it. Re-check the reference when the SDK version changes. The Python reference uses the public dd-trace-py main branch; the Node reference uses the public dd-trace-js master branch.
Keep shared workflow guidance here and language-specific syntax in the references. If a detail is only true for one SDK, do not duplicate it in this file.
references/python/ — Python ddtrace.llmobs API, providers, evaluator styles, environment template, and legacy compatibility.references/node/ — Node tracer.llmobs.experiments API and future Node-specific references.Do not modify dd-trace-py or dd-trace-js while updating this skill.
© datadog-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (references) in agent-observability/agent-observability-experiment-bootstrap of datadog-labs/agent-skills.
Open the folder on GitHubat commit d2411cc
Agent Observability Experiment Bootstrap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Observability Experiment Bootstrap this skilldatadog-labs/agent-skills | 177 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Failproof AI SDK IntegrationFailproofAI/failproofai | 5.3k | — | ~6k | Automated safety check: Pass | Custom licence | |
| Phoenix Tracinggithub/awesome-copilot | 40k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Caveman Gateway SetupJuliusBrussee/caveman | 111k | 1 repos | ~2.6k | Automated safety check: Warn | Apache-2.0 | |
| Agent Kill Switchvivekchand/clawmetry | 426 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Clawmetry Selfcheckvivekchand/clawmetry | 426 | — | ~515 | Automated safety check: Pass | MIT |
FailproofAI/failproofai
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
github/awesome-copilot
OpenInference semantic conventions and instrumentation for Phoenix AI observability.
JuliusBrussee/caveman
Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.
vivekchand/clawmetry
Give the human an off switch and a cost meter for the coding agents on this machine, using ClawMetry.
vivekchand/clawmetry
Read your own agent telemetry from ClawMetry (waste, progress, cost) and act on it before finishing a task.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
datadog-labs/agent-skills
Ensure the user has an authenticated Datadog account with a valid DDAPIKEY on the right region before any Datadog setup or instrumentation.
datadog-labs/agent-skills
Entry point for Datadog onboarding. An agent skill from datadog-labs/agent-skills.
datadog-labs/agent-skills
APM - install, onboard, instrument, enable, set up, configure, traces, services, dependencies, performance analysis, Data Streams Monitoring (DSM), queue lag, pipeline latency.
datadog-labs/agent-skills
Install the Datadog Agent on Kubernetes using the Datadog Operator — required before enabling Single Step Instrumentation (SSI), which automatically instruments applications for APM without code…
datadog-labs/agent-skills
Set up the Datadog AWS integration with Terraform - creates the cross-account IAM role Datadog assumes (external ID, no stored credentials), attaches the permission policies Datadog publishes, and…
datadog-labs/agent-skills
Install the Datadog Agent on Linux hosts via SSH with Single Step Instrumentation (SSI) enabled — SSI automatically instruments applications for APM without code changes.
Works with
Categories
Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK. Agent Observability Experiment Bootstrap is an agent skill from datadog-labs/agent-skills. Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.
Agent Observability Experiment Bootstrap fits situations like: LLM-as-a-judge scaffolding; tasks that involve Observability; tasks that involve LLM observability.
Run `npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a claude-code`. Or copy the skill folder (agent-observability/agent-observability-experiment-bootstrap in datadog-labs/agent-skills) into .claude/skills/agent-observability-experiment-bootstrap in your project. Claude Code loads it when a task matches its description.
Run `npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a codex`. Or copy the skill folder (agent-observability/agent-observability-experiment-bootstrap in datadog-labs/agent-skills) into .agents/skills/agent-observability-experiment-bootstrap in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadog-labs/agent-skills --skill agent-observability-experiment-bootstrap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-observability-experiment-bootstrap, .gemini/skills/agent-observability-experiment-bootstrap, .github/skills/agent-observability-experiment-bootstrap and .opencode/skills/agent-observability-experiment-bootstrap in your project.
Going by SKILL.md and its folder, Agent Observability Experiment Bootstrap needs Python for the scripts in its folder and the command-line tools its instructions call (python and node). Our summary lists: Python 3; Node.js.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agent Observability Experiment Bootstrap is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agent Observability Experiment Bootstrap: Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), Phoenix Tracing (github/awesome-copilot, 40k stars), Caveman Gateway Setup (JuliusBrussee/caveman, 111k stars) and Agent Kill Switch (vivekchand/clawmetry, 426 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
datadog-labs (a GitHub organization) maintains it in datadog-labs/agent-skills, which has 177 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 8, 2026.
Source: datadog-labs/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.