Official agent skill

Triage CI Failure

by DataDog in DataDog/datadog-agent

Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

OfficialApache-2.0Auto-check passedTesting & QA

Install Triage CI Failure

skills CLI
$ npx skills add DataDog/datadog-agent --skill triage-ci-failure -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DataDog/datadog-agent triage-ci-failure --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/triage-ci-failure .claude/skills/triage-ci-failure && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triage-ci-failure
GitHub stars
3.8k
Token cost
~2.3k tokens
SKILL.md length
975 words
Files
5 (incl. scripts, references)
Skills in repo
35
Repo updated
First seen
Licence
Apache-2.0

At a glance

Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

  • Works in 6 steps: Preflight → Collect the failures → CI Visibility baseline → …
  • A PRs pipeline is red and it isnt obvious whether the PRs own changes are at fault
  • SKILL.md covers Goal, Step 0 — Preflight, Step 1 — Collect the failures and Step 2 — CI Visibility baseline, plus 3 more sections
  • Runs Python scripts from its folder

What it does

Triage CI Failure is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Classify a failed CI as either caused by an active incident, flakiness, or a true code regression. Use when a PR's pipeline is red and it isn't obvious whether the PR's own changes are at fault. Trigger phrases include: - "investigate this CI failure" - "please fix CI" - "why did this job fail" - "is there an incident affecting CI" - "should I retry this" This should also be invoked whenever the user asks you to investigate or fix a failing CI, to ensure we don't spend hours trying to fix something broken…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/evidence.md`, `references/signals.md` and `scripts/incidents.py`).

It sits in Testing & QA, covering Failing and flaky tests. It works with Datadog. The repository describes itself as: Main repository for Datadog Agent. The licence is Apache-2.0.

When your agent uses it

  • A PRs pipeline is red and it isnt obvious whether the PRs own changes are at fault
  • Fix a failing CI
  • Ensure we dont spend hours trying to fix something broken upstream

Example prompts

  • “s pipeline is red and it isn”
  • “s own changes are at fault. Trigger phrases include: -”
  • “please fix CI”
  • “/triage-ci-failure”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Preflight
  2. Collect the failures
  3. CI Visibility baseline
  4. Incident correlation
  5. Read the log
  6. Verdict

What it can do on your machine

Read from SKILL.md and the folder at commit a706f1a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triage CI Failure loads about 2.3k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 155 tokens; SKILL.md has 975 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~155
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from DataDog/datadog-agent at commit a706f1a, republished under its Apache-2.0 licence (© DataDog). 975 words, ~2,251 tokens.

Download SKILL.mdSave it as .claude/skills/triage-ci-failure/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
triage-ci-failure
description
Classify a failed CI as either caused by an active incident, flakiness, or a true code regression. Use when a PR's pipeline is red and it isn't obvious whether the PR's own changes are at fault. Trigger phrases include: - "investigate this CI failure" - "please fix CI" - "why did this job fail" - "is there an incident affecting CI" - "should I retry this" This should also be invoked whenever the user asks you to investigate _or fix_ a failing CI, to ensure we don't spend hours trying to fix something broken upstream. Diagnosis only — pair with handle-pr-ci-failure to act on a pr-code verdict.
model
sonnet

Triage CI failure

Goal

Answer one question: is this failure caused by this PR's own changes ? with hard evidence. Every verdict below must cite the evidence that produced it — a bare "looks flaky, retry" or "looks broken, fix it" is not an acceptable output.

This skill only diagnoses. Never take action (writing a fix, retrying a job) on your own: only present your investigation results to the user.

Owning team: @DataDog/agent-devx

Step 0 — Preflight

Both ddgl and pup are required, and both live in the same places: locally, or inside a dda env dev.

bash
which ddgl pup

If pup is present but not authenticated, either run pup auth login or use the dd-auth skill. If pup can't be made to work, say so and continue with Steps 1 and 4 only: Steps 2 and 3 are unavailable, and the verdict should state that limitation rather than silently producing a weaker one.

Step 1 — Collect the failures

bash
ddgl jobs list --failed --json --no-pager [--ref <ref> | --pipeline <id>]

An empty [] means there's nothing to triage — stop here.

Also fetch pipeline state:

bash
ddgl pipelines get --json [--ref <ref> | --pipeline <id>]

If the pipeline is still running, a job you're about to triage may yet be auto-retried into success. Note that in the verdict rather than treating the failure as final.

For each failed job, look at its failure_reason, i.e. the failure reason as determined by gitlab. Treat it as an aditionnal data point, not the be-all-end-all. For example, a runner_system_failure can be caused by a change of this PR (e.g. a malformed image:). See @references/signals.md for more details.

Step 2 — CI Visibility baseline

Check if this job is failing everywhere (i.e. on main), or if it is often flaky using CI Visibility.

For each failed job, ask how it behaves elsewhere:

bash
pup cicd events aggregate \
  --query='ci_level:job @ci.pipeline.name:DataDog/datadog-agent @git.branch:main @ci.job.name:"<exact job name>"' \
  --compute=count --group-by='@ci.status' --from='2d'

Check @references/signals.md for additional queries that can help if this first one is inconclusive. Come out of this step with a working hypothesis (upstream, flake, or pr-code) for later steps to confirm or overturn — not a final verdict.

Step 3 — Incident correlation

If Step 2 pointed clearly at pr-code, skip to Step 4.

Use the helper script to search for an active CI incident matching the failing job:

bash
.agents/skills/triage-ci-failure/scripts/incidents.py search \
  --at <job-failure-ISO8601-timestamp> \
  --job '<exact failing job name>' [--job '<another one>' ...]

Read the match tier in the output (exact, base, prefix, token, none) — anything but none is worth reading the timeline for:

bash
.agents/skills/triage-ci-failure/scripts/incidents.py timeline <IR-nnnnn>

This is where you find out how far along the fix is — not just whether one exists.

  • stable usually means a rollback or workaround has already landed and the affected job(s) should pass again on a rebase.
  • resolved (or completed) is the stronger signal: the incident is fully closed out.

Look for a rollback, a merged fix PR, or an explicit state transition to tell which.

If nothing matches, widen deliberately rather than re-running the same call — escalate through the tier ladder in references/signals.md:

  1. (default, above) services:datadog-agent-ci, default window.
  2. Same, a much wider window — for old branches whose failure was fixed on main long before you rebased onto them.
  3. Drop the service filter, search free text instead, using a keyword pulled from the job log in Step 4 (an image reference, a host, an endpoint, a bucket name).

Step 4 — Read the log

Always do a quick sanity check here, even when Step 3 was conclusive — a time-and-name correlation is strong evidence but not proof. Skim the job's diff against main and the last ~50 lines of its log, and confirm the failure signature actually looks like what the incident describes.

You can obtain the job's log via ddgl:

bash
ddgl logs --job <ID> [--output <some_file>]

If it lines up, you're done — the full cookbook below is skippable. If it doesn't, or Step 3 didn't produce a confident match at all, work through @references/evidence.md's cookbook.

You're looking for two things:

  1. the command that actually failed and its exit status
  2. whether the failure happened in the job's own work or in its setup/teardown.
Show full SKILL.md (349 more words)Show less

Step 5 — Verdict

State your verdict among the below options, as well as a recommended course of action and the linked incident if any.

blameincident statusSuggested action
pr-code—Propose the smallest concrete fix. Don't apply it.
upstreamactive, still breakingDon't suggest rebasing yet. Report the incident.
upstreamstableSuggest a rebase and retry — stable usually means a rollback or workaround already landed — but say plainly that this is a weaker signal than resolved: the underlying fix may still be in progress.
upstreamresolvedRebase onto latest main and re-run with confidence. Name the fixing commit/PR if the timeline gave you one.
upstreamnone declaredSay CI looks broken on main with nothing declared for it — worth surfacing loudly.
infraanySuggest a retry. Note whether the job already burned its one automatic retry (references/signals.md).
flakeanySuggest a retry, citing the measured cross-branch failure rate from Step 2 as the reason — not just a feeling.
inconclusiveanyPresent the evidence and the two most likely readings. Don't guess past what you found.

End with one block per failed job, exactly in this shape, so a caller like /follow-pr or /handle-pr-ci-failure can act without re-deriving your reasoning or guessing which job a verdict belongs to:

CI triage result
Job: <exact GitLab job name>
Pipeline SHA: <full SHA of the pipeline you inspected>
Blame: pr-code | upstream | infra | flake | inconclusive
Failure signature: <stable failing command/test/error, e.g. "TestFoo/bar: assert.Equal want=1 got=2">
Evidence: <one-line summary of the hard evidence from Steps 1-4>
Proposed fix: <smallest concrete fix, or none>
Incident: <one of the forms below, or none>
End CI triage result

Concrete Incident forms:

  • IR-59848 (active, still breaking) — https://app.datadoghq.com/incidents/59848
  • IR-59848 (stable, probably safe to retry) — https://app.datadoghq.com/incidents/59848
  • IR-59848 (resolved) — https://app.datadoghq.com/incidents/59848
  • none

Proposed fix is none for every Blame except pr-code — only a PR-caused failure gets a concrete fix proposed. For example:

CI triage result
Job: lint_go_linux-x64
Pipeline SHA: 8f2c1e9a4b1d7e3f0a9c6b5d4e3f2a1b0c9d8e7f
Blame: pr-code
Failure signature: pkg/foo/bar.go: ineffectual assignment to err (ineffassign)
Evidence: introduced in this PR's commit a1b2c3d; the same job passes on main at the same base commit
Proposed fix: remove the unused `err :=` reassignment on line 42
Incident: none
End CI triage result

A second example, for a failure that turned out inconclusive rather than PR-caused:

CI triage result
Job: new-e2e-container-images
Pipeline SHA: 3c7b1a0f9e8d6c5b4a3f2e1d0c9b8a7f6e5d4c3b
Blame: inconclusive
Failure signature: TestContainerImages/pull_public_image: context deadline exceeded
Evidence: fails intermittently on main too (3/40 runs over the last 2 days); no matching incident found; the job's own diff and log show no clear infra or PR-code signal
Proposed fix: none
Incident: none
End CI triage result

Failure signature must stay stable across replacement pipelines for the same underlying defect: strip timestamps, job/pipeline IDs, temp paths, and line numbers, keeping the failing command/test name and the error itself. A caller diffs signatures across pipelines to tell "same bug, still broken" from "new bug" — don't let cosmetic noise make two identical failures look different.

Never collapse multiple failed jobs into one block, and never let a none incident stand in for a pr-code verdict — a caller must branch on Blame alone, not on the absence of an incident.

© DataDog, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in .agents/skills/triage-ci-failure of DataDog/datadog-agent.

  • SKILL.md
  • references/evidence.md
  • references/signals.md
  • scripts/incidents.py
  • scripts/test_incidents.py

Open the folder on GitHubat commit a706f1a

Compare with similar skills

Triage CI Failure next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triage CI Failure compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triage CI Failure this skillDataDog/datadog-agent3.8k—~2.3kAutomated safety check: PassApache-2.0
Flaky Test FixerDataDog/dd-trace-js837—~1.6kAutomated safety check: PassCustom licence
Dd Unblock PRDataDog/pup1k—~1.7kAutomated safety check: PassApache-2.0
Analyze Azdo BuildDataDog/dd-trace-dotnet573—~3.5kAutomated safety check: PassApache-2.0
Resolve Muzzle CIDataDog/dd-trace-java736—~3.2kAutomated safety check: PassApache-2.0
Dd Triage Flaky TestDataDog/pup1k1 repos~2.3kAutomated safety check: PassApache-2.0

Similar skills

  • Flaky Test Fixer

    DataDog/dd-trace-js

    Official

    A skill your agent uses when classifying, investigating, or fixing a suspected flaky test, intermittent test result, nondeterministic CI test failure, timing race, hang, or test-order dependency in…

    837 GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Dd Unblock PR

    DataDog/pup

    Official

    Load when investigating a failing PR CI pipeline or checking PR health.

    1k GitHub stars~1.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Analyze Azdo Build

    DataDog/dd-trace-dotnet

    Official

    Analyze Azure DevOps CI build failures in dd-trace-dotnet pipeline.

    573 GitHub stars~3.5k tokensUpdated today
    Testing & QAAuto-check passed
  • Resolve Muzzle CI

    DataDog/dd-trace-java

    Official

    Diagnose and resolve dd-trace-java CI failures from a module's muzzle task or the runMuzzle aggregate.

    736 GitHub stars~3.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Official

    Load when investigating a specific flaky test. An agent skill from DataDog/pup.

    1k GitHub starsUsed in 1 repo~2.3k tokens
    Testing & QAAuto-check passed
  • Fix Broken Datadog Provider Tests

    DataDog/terraform-provider-datadog

    Official

    Runs an end-to-end workflow to diagnose, reproduce, fix and validate a failing integration test in the Datadog Terraform provider, ending with a draft PR.

    468 GitHub stars~2.4k tokensUpdated yesterday
    Testing & QAAuto-check: notes

More from DataDog/datadog-agent

All 35 skills in this repo
  • Elicit

    DataDog/datadog-agent

    Official

    Run a structured discovery session to build an Allium specification through conversation.

    3.8k GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Follow PR

    DataDog/datadog-agent

    Official

    Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure.

    3.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Create Epic Recap

    DataDog/datadog-agent

    Official

    A skill your agent uses when an engineer or manager asks to recap, summarize, or post an update on a Jira Epic — a progress update for an in-progress Epic (how far along it is, what's shipped so…

    3.8k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Explain Lading Config

    DataDog/datadog-agent

    Official

    Explains a lading.yaml config file from the regression test suite, using the lading Rust source as ground truth for field meanings and defaults.

    3.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Distill

    DataDog/datadog-agent

    Official

    Extract an Allium specification from an existing codebase. An agent skill from DataDog/datadog-agent.

    3.8k GitHub starsUsed in 1 repo~7k tokens
    Auto-check passed
  • Run E2E

    DataDog/datadog-agent

    Official

    Run one already-written new-e2e test locally and triage the setup failures that stop it — "run the containers e2e tests", "my e2e run fails before any test starts".

    3.8k GitHub stars~2.2k tokensUpdated today
    Auto-check: notes

Works with

Categories

Questions about Triage CI Failure

What does Triage CI Failure do?

Classify a failed CI as either caused by an active incident, flakiness, or a true code regression. Triage CI Failure is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

When should I use Triage CI Failure?

Triage CI Failure fits situations like: A PRs pipeline is red and it isnt obvious whether the PRs own changes are at fault; fix a failing CI; ensure we dont spend hours trying to fix something broken upstream.

How do I install Triage CI Failure in Claude Code?

Run `npx skills add DataDog/datadog-agent --skill triage-ci-failure -a claude-code`. Or copy the skill folder (.agents/skills/triage-ci-failure in DataDog/datadog-agent) into .claude/skills/triage-ci-failure in your project. Claude Code loads it when a task matches its description.

How do I install Triage CI Failure in Codex?

Run `npx skills add DataDog/datadog-agent --skill triage-ci-failure -a codex`. Or copy the skill folder (.agents/skills/triage-ci-failure in DataDog/datadog-agent) into .agents/skills/triage-ci-failure in your project. Codex loads it when a task matches its description.

Can I use Triage CI Failure in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DataDog/datadog-agent --skill triage-ci-failure -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage-ci-failure, .gemini/skills/triage-ci-failure, .github/skills/triage-ci-failure and .opencode/skills/triage-ci-failure in your project.

What does Triage CI Failure need to run?

Going by SKILL.md and its folder, Triage CI Failure needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Triage CI Failure access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triage CI Failure safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Triage CI Failure use?

Triage CI Failure is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triage CI Failure use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Triage CI Failure?

Skills that share tags, products or a category with Triage CI Failure: Flaky Test Fixer (DataDog/dd-trace-js, 837 stars), Dd Unblock PR (DataDog/pup, 1k stars), Analyze Azdo Build (DataDog/dd-trace-dotnet, 573 stars) and Resolve Muzzle CI (DataDog/dd-trace-java, 736 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triage CI Failure?

DataDog (a GitHub organization, an official publisher) maintains it in DataDog/datadog-agent, which has 3,759 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 9, 2026.

Source: DataDog/datadog-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.