Flaky Test Fixer
DataDog/dd-trace-js
A skill your agent uses when classifying, investigating, or fixing a suspected flaky test, intermittent test result, nondeterministic CI test failure, timing race, hang, or test-order dependency in…
Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.
$ npx skills add DataDog/datadog-agent --skill triage-ci-failure -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install DataDog/datadog-agent triage-ci-failure --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/triage-ci-failure .claude/skills/triage-ci-failure && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "triage-ci-failure" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/triage-ci-failure into .claude/skills/triage-ci-failure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-ci-failure", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/triage-ci-failureType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add DataDog/datadog-agent --skill triage-ci-failure -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install DataDog/datadog-agent triage-ci-failure --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/triage-ci-failure .agents/skills/triage-ci-failure && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "triage-ci-failure" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/triage-ci-failure into .agents/skills/triage-ci-failure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-ci-failure", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add DataDog/datadog-agent --skill triage-ci-failure -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install DataDog/datadog-agent triage-ci-failure --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/triage-ci-failure .cursor/skills/triage-ci-failure && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "triage-ci-failure" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/triage-ci-failure into .cursor/skills/triage-ci-failure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-ci-failure", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/DataDog/datadog-agent.git --path .agents/skills/triage-ci-failure--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add DataDog/datadog-agent --skill triage-ci-failure -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install DataDog/datadog-agent triage-ci-failure --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/triage-ci-failure .gemini/skills/triage-ci-failure && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "triage-ci-failure" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/triage-ci-failure into .gemini/skills/triage-ci-failure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-ci-failure", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install DataDog/datadog-agent triage-ci-failureInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add DataDog/datadog-agent --skill triage-ci-failure -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/triage-ci-failure .github/skills/triage-ci-failure && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "triage-ci-failure" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/triage-ci-failure into .github/skills/triage-ci-failure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-ci-failure", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add DataDog/datadog-agent --skill triage-ci-failure -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install DataDog/datadog-agent triage-ci-failure --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/triage-ci-failure .opencode/skills/triage-ci-failure && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "triage-ci-failure" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/triage-ci-failure into .opencode/skills/triage-ci-failure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-ci-failure", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
triage-ci-failureClassify a failed CI as either caused by an active incident, flakiness, or a true code regression.
Triage CI Failure is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Classify a failed CI as either caused by an active incident, flakiness, or a true code regression. Use when a PR's pipeline is red and it isn't obvious whether the PR's own changes are at fault. Trigger phrases include: - "investigate this CI failure" - "please fix CI" - "why did this job fail" - "is there an incident affecting CI" - "should I retry this" This should also be invoked whenever the user asks you to investigate or fix a failing CI, to ensure we don't spend hours trying to fix something broken…
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/evidence.md`, `references/signals.md` and `scripts/incidents.py`).
It sits in Testing & QA, covering Failing and flaky tests. It works with Datadog. The repository describes itself as: Main repository for Datadog Agent. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a706f1a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Triage CI Failure loads about 2.3k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 155 tokens; SKILL.md has 975 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from DataDog/datadog-agent at commit a706f1a, republished under its Apache-2.0 licence (© DataDog). 975 words, ~2,251 tokens.
.claude/skills/triage-ci-failure/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Answer one question: is this failure caused by this PR's own changes ? with hard evidence. Every verdict below must cite the evidence that produced it — a bare "looks flaky, retry" or "looks broken, fix it" is not an acceptable output.
This skill only diagnoses. Never take action (writing a fix, retrying a job) on your own: only present your investigation results to the user.
Owning team: @DataDog/agent-devx
Both ddgl and pup are required, and both live in the same places: locally, or inside a dda env dev.
which ddgl pupIf pup is present but not authenticated, either run pup auth login or use the dd-auth skill.
If pup can't be made to work, say so and continue with Steps 1 and 4 only:
Steps 2 and 3 are unavailable, and the verdict should state that limitation rather than silently producing a weaker one.
ddgl jobs list --failed --json --no-pager [--ref <ref> | --pipeline <id>]An empty [] means there's nothing to triage — stop here.
Also fetch pipeline state:
ddgl pipelines get --json [--ref <ref> | --pipeline <id>]If the pipeline is still running, a job you're about to triage may yet be auto-retried into success. Note that in the verdict rather than treating the failure as final.
For each failed job, look at its failure_reason, i.e. the failure reason as determined by gitlab.
Treat it as an aditionnal data point, not the be-all-end-all. For example, a runner_system_failure can be caused by a change of this PR (e.g. a malformed image:).
See @references/signals.md for more details.
Check if this job is failing everywhere (i.e. on main), or if it is often flaky using CI Visibility.
For each failed job, ask how it behaves elsewhere:
pup cicd events aggregate \
--query='ci_level:job @ci.pipeline.name:DataDog/datadog-agent @git.branch:main @ci.job.name:"<exact job name>"' \
--compute=count --group-by='@ci.status' --from='2d'Check @references/signals.md for additional queries that can help if this first one is inconclusive.
Come out of this step with a working hypothesis (upstream, flake, or pr-code) for later steps to confirm or overturn — not a final verdict.
If Step 2 pointed clearly at pr-code, skip to Step 4.
Use the helper script to search for an active CI incident matching the failing job:
.agents/skills/triage-ci-failure/scripts/incidents.py search \
--at <job-failure-ISO8601-timestamp> \
--job '<exact failing job name>' [--job '<another one>' ...]Read the match tier in the output (exact, base, prefix, token, none) — anything but none is worth reading the timeline for:
.agents/skills/triage-ci-failure/scripts/incidents.py timeline <IR-nnnnn>This is where you find out how far along the fix is — not just whether one exists.
stable usually means a rollback or workaround has already landed and the affected job(s) should pass again on a rebase.resolved (or completed) is the stronger signal: the incident is fully closed out.Look for a rollback, a merged fix PR, or an explicit state transition to tell which.
If nothing matches, widen deliberately rather than re-running the same call — escalate through the tier ladder in references/signals.md:
services:datadog-agent-ci, default window.main long before you rebased onto them.Always do a quick sanity check here, even when Step 3 was conclusive — a time-and-name correlation is strong evidence but not proof.
Skim the job's diff against main and the last ~50 lines of its log, and confirm the failure signature actually looks like what the incident describes.
You can obtain the job's log via ddgl:
ddgl logs --job <ID> [--output <some_file>]If it lines up, you're done — the full cookbook below is skippable. If it doesn't, or Step 3 didn't produce a confident match at all, work through @references/evidence.md's cookbook.
You're looking for two things:
State your verdict among the below options, as well as a recommended course of action and the linked incident if any.
| blame | incident status | Suggested action |
|---|---|---|
pr-code | — | Propose the smallest concrete fix. Don't apply it. |
upstream | active, still breaking | Don't suggest rebasing yet. Report the incident. |
upstream | stable | Suggest a rebase and retry — stable usually means a rollback or workaround already landed — but say plainly that this is a weaker signal than resolved: the underlying fix may still be in progress. |
upstream | resolved | Rebase onto latest main and re-run with confidence. Name the fixing commit/PR if the timeline gave you one. |
upstream | none declared | Say CI looks broken on main with nothing declared for it — worth surfacing loudly. |
infra | any | Suggest a retry. Note whether the job already burned its one automatic retry (references/signals.md). |
flake | any | Suggest a retry, citing the measured cross-branch failure rate from Step 2 as the reason — not just a feeling. |
inconclusive | any | Present the evidence and the two most likely readings. Don't guess past what you found. |
End with one block per failed job, exactly in this shape, so a caller like /follow-pr or /handle-pr-ci-failure can act without re-deriving your reasoning or guessing which job a verdict belongs to:
CI triage result
Job: <exact GitLab job name>
Pipeline SHA: <full SHA of the pipeline you inspected>
Blame: pr-code | upstream | infra | flake | inconclusive
Failure signature: <stable failing command/test/error, e.g. "TestFoo/bar: assert.Equal want=1 got=2">
Evidence: <one-line summary of the hard evidence from Steps 1-4>
Proposed fix: <smallest concrete fix, or none>
Incident: <one of the forms below, or none>
End CI triage resultConcrete Incident forms:
IR-59848 (active, still breaking) — https://app.datadoghq.com/incidents/59848IR-59848 (stable, probably safe to retry) — https://app.datadoghq.com/incidents/59848IR-59848 (resolved) — https://app.datadoghq.com/incidents/59848noneProposed fix is none for every Blame except pr-code — only a PR-caused failure gets a concrete fix proposed. For example:
CI triage result
Job: lint_go_linux-x64
Pipeline SHA: 8f2c1e9a4b1d7e3f0a9c6b5d4e3f2a1b0c9d8e7f
Blame: pr-code
Failure signature: pkg/foo/bar.go: ineffectual assignment to err (ineffassign)
Evidence: introduced in this PR's commit a1b2c3d; the same job passes on main at the same base commit
Proposed fix: remove the unused `err :=` reassignment on line 42
Incident: none
End CI triage resultA second example, for a failure that turned out inconclusive rather than PR-caused:
CI triage result
Job: new-e2e-container-images
Pipeline SHA: 3c7b1a0f9e8d6c5b4a3f2e1d0c9b8a7f6e5d4c3b
Blame: inconclusive
Failure signature: TestContainerImages/pull_public_image: context deadline exceeded
Evidence: fails intermittently on main too (3/40 runs over the last 2 days); no matching incident found; the job's own diff and log show no clear infra or PR-code signal
Proposed fix: none
Incident: none
End CI triage resultFailure signature must stay stable across replacement pipelines for the same underlying defect: strip timestamps, job/pipeline IDs, temp paths, and line numbers, keeping the failing command/test name and the error itself. A caller diffs signatures across pipelines to tell "same bug, still broken" from "new bug" — don't let cosmetic noise make two identical failures look different.
Never collapse multiple failed jobs into one block, and never let a none incident stand in for a pr-code verdict — a caller must branch on Blame alone, not on the absence of an incident.
© DataDog, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts, references) in .agents/skills/triage-ci-failure of DataDog/datadog-agent.
Open the folder on GitHubat commit a706f1a
Triage CI Failure next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Triage CI Failure this skillDataDog/datadog-agent | 3.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Flaky Test FixerDataDog/dd-trace-js | 837 | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Dd Unblock PRDataDog/pup | 1k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Analyze Azdo BuildDataDog/dd-trace-dotnet | 573 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Resolve Muzzle CIDataDog/dd-trace-java | 736 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Dd Triage Flaky TestDataDog/pup | 1k | 1 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 |
DataDog/dd-trace-js
A skill your agent uses when classifying, investigating, or fixing a suspected flaky test, intermittent test result, nondeterministic CI test failure, timing race, hang, or test-order dependency in…
DataDog/pup
Load when investigating a failing PR CI pipeline or checking PR health.
DataDog/dd-trace-dotnet
Analyze Azure DevOps CI build failures in dd-trace-dotnet pipeline.
DataDog/dd-trace-java
Diagnose and resolve dd-trace-java CI failures from a module's muzzle task or the runMuzzle aggregate.
DataDog/pup
Load when investigating a specific flaky test. An agent skill from DataDog/pup.
DataDog/terraform-provider-datadog
Runs an end-to-end workflow to diagnose, reproduce, fix and validate a failing integration test in the Datadog Terraform provider, ending with a draft PR.
DataDog/datadog-agent
Run a structured discovery session to build an Allium specification through conversation.
DataDog/datadog-agent
Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure.
DataDog/datadog-agent
A skill your agent uses when an engineer or manager asks to recap, summarize, or post an update on a Jira Epic — a progress update for an in-progress Epic (how far along it is, what's shipped so…
DataDog/datadog-agent
Explains a lading.yaml config file from the regression test suite, using the lading Rust source as ground truth for field meanings and defaults.
DataDog/datadog-agent
Extract an Allium specification from an existing codebase. An agent skill from DataDog/datadog-agent.
DataDog/datadog-agent
Run one already-written new-e2e test locally and triage the setup failures that stop it — "run the containers e2e tests", "my e2e run fails before any test starts".
Works with
Categories
Classify a failed CI as either caused by an active incident, flakiness, or a true code regression. Triage CI Failure is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.
Triage CI Failure fits situations like: A PRs pipeline is red and it isnt obvious whether the PRs own changes are at fault; fix a failing CI; ensure we dont spend hours trying to fix something broken upstream.
Run `npx skills add DataDog/datadog-agent --skill triage-ci-failure -a claude-code`. Or copy the skill folder (.agents/skills/triage-ci-failure in DataDog/datadog-agent) into .claude/skills/triage-ci-failure in your project. Claude Code loads it when a task matches its description.
Run `npx skills add DataDog/datadog-agent --skill triage-ci-failure -a codex`. Or copy the skill folder (.agents/skills/triage-ci-failure in DataDog/datadog-agent) into .agents/skills/triage-ci-failure in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DataDog/datadog-agent --skill triage-ci-failure -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage-ci-failure, .gemini/skills/triage-ci-failure, .github/skills/triage-ci-failure and .opencode/skills/triage-ci-failure in your project.
Going by SKILL.md and its folder, Triage CI Failure needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Triage CI Failure is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Triage CI Failure: Flaky Test Fixer (DataDog/dd-trace-js, 837 stars), Dd Unblock PR (DataDog/pup, 1k stars), Analyze Azdo Build (DataDog/dd-trace-dotnet, 573 stars) and Resolve Muzzle CI (DataDog/dd-trace-java, 736 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
DataDog (a GitHub organization, an official publisher) maintains it in DataDog/datadog-agent, which has 3,759 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 9, 2026.
Source: DataDog/datadog-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.