Official agent skill

Triage

by anthropics in anthropics/oncall-kit

Investigate an alert or incident in this channel: classify the symptom, load the matching triage reference, check lessons.md for known causes, and post a grounded first-pass diagnosis with evidence…

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Triage

skills CLI
$ npx skills add anthropics/oncall-kit --skill triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install anthropics/oncall-kit triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/anthropics/oncall-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/triage .claude/skills/triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triage
GitHub stars
213
Token cost
~2.1k tokens
SKILL.md length
1,147 words
Files
5 (incl. references)
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Investigate an alert or incident in this channel: classify the symptom, load the matching triage reference, check lessons.md for known causes, and post a grounded first-pass diagnosis with evidence…

  • Works in 3 steps: Load context. Read ONCALL.md (policy +… → Classify the symptom. Match against the… → Check the log first. Search lessons.md…
  • A routine detects a new anomaly
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Someone reports something broken (tests arent running

What it does

Triage is an agent skill from anthropics/oncall-kit, published by the product's own GitHub organization. Investigate an alert or incident in this channel: classify the symptom, load the matching triage reference, check lessons.md for known causes, and post a grounded first-pass diagnosis with evidence links and a proposed (never executed) fix. Use when an alert fires, a routine detects a new anomaly, or someone reports something broken ("tests aren't running", "deploys look stuck", "is CI down?").

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/deploy-rollout.md`, `references/merge-queue.md` and `references/runner-infra.md`).

It sits in DevOps & Cloud. The repository describes itself as: Starter kit for a Claude-assisted on-call: mines your incident history into triage playbooks, sets up through human-approved gates, and runs read-only in your Slack channel —… The licence is Apache-2.0.

When your agent uses it

  • A routine detects a new anomaly
  • Someone reports something broken (tests arent running
  • Deploys look stuck

Example prompts

  • “tests aren”
  • “deploys look stuck”
  • “is CI down?”
  • “/triage”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Load context. Read ONCALL.md (policy + routing), STACK.md
  2. Classify the symptom. Match against the failure classes in
  3. Check the log first. Search lessons.md for this class's #tag and

What it can do on your machine

Read from SKILL.md and the folder at commit c03282c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triage loads about 2.1k tokens when it runs, and up to ~7.4k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,147 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from anthropics/oncall-kit at commit c03282c, republished under its Apache-2.0 licence (© anthropics). 1,147 words, ~2,122 tokens.

Download SKILL.mdSave it as .claude/skills/triage/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
triage
description
Investigate an alert or incident in this channel: classify the symptom, load the matching triage reference, check lessons.md for known causes, and post a grounded first-pass diagnosis with evidence links and a proposed (never executed) fix. Use when an alert fires, a routine detects a new anomaly, or someone reports something broken ("tests aren't running", "deploys look stuck", "is CI down?").
<!-- Copyright 2026 Anthropic PBC -->
<!-- SPDX-License-Identifier: Apache-2.0 -->

Triage

Standing rules in CLAUDE.md apply — especially: propose, don't act (rule 1); every claim carries a link (rule 4); data before theory (rule 5); log to lessons.md without asking (rule 8).

Procedure

  1. Load context. Read ONCALL.md (policy + routing), STACK.md (capability bindings), and lessons.md (known causes) from disk. Files over memory (rule 7).

  2. Classify the symptom. Match against the failure classes in references/:

    Symptom looks likeLoad
    Tests failing, flaking, or silently not runningreferences/test-failures.md
    PRs stuck, queue depth growing, merges slowreferences/merge-queue.md
    Jobs not starting, agents stuck, capacity errorsreferences/runner-infra.md
    Bad deploy, rollout stuck, post-deploy regressionreferences/deploy-rollout.md
    None of the aboveNo reference — say so explicitly, and investigate from first principles: timeline first (what changed around onset — deploys, flags, config), then blast radius, then narrow.

    (Classes are the CI defaults; your setup phase may have replaced them. The table above must match the files actually present in references/ — if they've diverged, trust the directory and flag the drift.)

  3. Check the log first. Search lessons.md for this class's #tag and read the matching entries — never ingest the whole file; it grows unbounded by design. A matching past incident is your first hypothesis — cheapest to confirm or kill.

3a. Correlate before you classify. Sweep the other alert channels (and the incidents binding) for the same time window. Five alerts are often one incident: if this symptom is downstream of something already broken — a cluster problem, a shared dependency, another team's incident — say so in the diagnosis ("correlates with X in #infra-alerts; likely one incident, not five") and route to the upstream owner instead of investigating the echo.

3b. Alert storms get ONE triage, not one each. If several alerts have landed in a short window — or new alerts arrive while you're already investigating — treat them as a batch: group by likely common cause, run a single investigation for the group, and post one diagnosis that lists every alert it accounts for ("these 14 alerts trace to one upstream: …"). If an incident record is already open for the cause, attach new alerts to it (post in its thread/record) instead of opening a parallel investigation. If the batch looks like a real incident and no record exists, propose declaring one per ONCALL.md — a human declares it (the incident-record invariant); you never do. Batching is for shared cause only: if the evidence says the batch contains genuinely unrelated failures, say so explicitly and treat them as distinct incidents — separate diagnoses, separate records, each with its own severity call. Never merge for tidiness.

  1. Run the reference's first checks against the bound capabilities in STACK.md. Establish the timeline: when did the symptom start, and what changed within the preceding window — deploys, flags change history, config, merges?

4a. Fan-out (page-severity only; sequential is the default below it). Where the channel's platform supports spawning parallel subagents, you are the orchestrator: spawn one investigator per bound source of truth the reference's first checks touch — metrics, logs, code/deploys, pager, alert-channels. Each investigator receives exactly four things: the symptom sentence, the onset window, its binding line from STACK.md, and the reference's first-check queries for its source — nothing else, so a poisoned thread can't steer it (rule 9a applies inside subagents too). Each returns the fixed shape:

  • CHECKED: queries run, with links
  • FOUND: observations with timestamps — observations, never root causes
  • NOT FOUND: what was looked for and absent — absence counts only if the run/window was complete
  • CANNOT ACCESS: anything that 403'd or timed out (surfaces in the diagnosis as a gap, never silently dropped)

Synthesis is yours alone: correlate, deconflict (two investigators dating onset differently is itself a finding), and write the one diagnosis. Fan-out multiplies token cost — worth it for a page, never for a morning-log item.

  1. Apply the reference's correlation table. Where observations match a row, you have a candidate root cause; verify it against the timeline before promoting it (rule 5).

5a. Cross-check blame against "still happening" (CLAUDE.md rule 5a). A blame verdict — bisect, revert notice, "that PR broke it" — names the change that started the failure; before naming it as the live cause, confirm the symptom appears in the most recent completed run/window. Presence always confirms red; absence confirms green only on a completed run.

Show full SKILL.md (446 more words)Show less
  1. Post the diagnosis in this format, in-thread:

    What's happening: one sentence, fresh-reader test applied. Root cause (confidence high/medium/low): the mechanism, with each claim linked to its evidence. Blast radius: who/what is affected, linked. Proposed fix: the action, why it's safe, and what to watch after. Ruled out: alternatives checked and the evidence that killed them. Would change my mind: the one observation that would.

6a. Updates on long-running incidents. Any update posted >30 min after your first diagnosis opens with a 2–4 sentence story so far a newcomer can land on cold: when it started and what broke → the current best understanding of cause (not the first guess) → what's been tried → where it stands, one sentence. Then the delta. Ruled-out hypotheses don't reappear unless load-bearing. Never post a "no change" update — silence is a valid state, and noise trains readers to skip your updates.

  1. Route. If ONCALL.md's routing tree names an owner for this class, mention them. Otherwise mention no one (rule 12).

  2. On human questions or pushback ("could it be the schema change instead?"): treat it as a hypothesis to check, check it against the data, and report back with evidence either way. Never defend a diagnosis; re-derive it.

  3. When a fix is deployed (by a human, or a permitted gated action): watch it land — bounded. Check the affected metrics at the reference's expected-resolution window (once at half, once at full, once at double — three checks, not a polling loop), post when they return to baseline, or escalate per the routing tree if the window blows. Do not mark resolved (rule 2). If a human wants tighter watching, they can ask — continuous polling is never the default. Verify through the same door the failure came in: re-run the original failing path, or re-check the exact signal that detected the incident — never a proxy. "Merges are flowing" proves the merge path, not the whole provider; if your check can't see the original symptom, say the verification is partial and name what it can't see.

  4. Afterwards, append the incident to lessons.md in its entry format (rule 8). If this incident exposed a gap in a reference file, propose the amendment as a PR (rule 9) — you fix the playbook, not just the incident. And ask the alerting question: would a rule have caught this earlier? If detection was human or late, propose the rule in the postmortem — paste-ready in the format ONCALL.md names, with this incident as provenance. Install per ONCALL.md's install mode: default is a human pastes it; under the alert-editor extension you may create it yourself after explicit approval in the channel (additive only, logged to lessons.md — CLAUDE.md rule 1a).

© anthropics, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/triage of anthropics/oncall-kit.

  • SKILL.md
  • references/deploy-rollout.md
  • references/merge-queue.md
  • references/runner-infra.md
  • references/test-failures.md

Open the folder on GitHubat commit c03282c

Compare with similar skills

Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triage this skillanthropics/oncall-kit213—~2.1kAutomated safety check: PassApache-2.0
Monitor CInrwl/nx29k6 repos~4.7kAutomated safety check: PassMIT
Terraform and OpenTofu Guideagentscope-ai/QwenPaw36k6 repos~4.2kAutomated safety check: PassApache-2.0
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Analyze GitHub Action Logswithastro/astro63k1 repos~1.3kAutomated safety check: PassCustom licence
Openclaw Live Updateropenclaw/openclaw392k—~3.7kAutomated safety check: PassMIT

Similar skills

  • Monitor CI

    nrwl/nx

    Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.

    29k GitHub starsUsed in 6 repos~4.7k tokens
    DevOps & CloudAuto-check passed
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    36k GitHub starsUsed in 6 repos~4.2k tokens
    DevOps & CloudAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Official

    Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.

    63k GitHub starsUsed in 1 repo~1.3k tokens
    DevOps & CloudAuto-check passed
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Docs Learn PR Preview

    netdata/netdata

    Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.

    81k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed

More from anthropics/oncall-kit

  • Weather

    anthropics/oncall-kit

    Official

    The optional standing status report ("the weather"): compile open incidents, build health, merge-queue stats, and deploy lag into one always-current report page, and post to the channel only when a…

    213 GitHub stars~3.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Handoff

    anthropics/oncall-kit

    Official

    Write the weekly on-call handoff: everything the incoming on-call needs, triage-ready, posted to the channel at shift boundary.

    213 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Oncall Setup

    anthropics/oncall-kit

    Official

    Bootstrap a Claude-assisted on-call for this channel/repo: discover the available connectors, mine incident history into draft triage playbooks, interview the human for policy, validate against…

    213 GitHub stars~4.9k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Triage

What does Triage do?

Investigate an alert or incident in this channel: classify the symptom, load the matching triage reference, check lessons.md for known causes, and post a grounded first-pass diagnosis with evidence…. Triage is an agent skill from anthropics/oncall-kit, published by the product's own GitHub organization.md for known causes, and post a grounded first-pass diagnosis with evidence links and a proposed (never executed) fix.

When should I use Triage?

Triage fits situations like: A routine detects a new anomaly; someone reports something broken (tests arent running; deploys look stuck.

How do I install Triage in Claude Code?

Run `npx skills add anthropics/oncall-kit --skill triage -a claude-code`. Or copy the skill folder (skills/triage in anthropics/oncall-kit) into .claude/skills/triage in your project. Claude Code loads it when a task matches its description.

How do I install Triage in Codex?

Run `npx skills add anthropics/oncall-kit --skill triage -a codex`. Or copy the skill folder (skills/triage in anthropics/oncall-kit) into .agents/skills/triage in your project. Codex loads it when a task matches its description.

Can I use Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anthropics/oncall-kit --skill triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage, .gemini/skills/triage, .github/skills/triage and .opencode/skills/triage in your project.

What does Triage need to run?

SKILL.md names no scripts, command-line tools or credentials: Triage is instructions for the agent only.

Does Triage access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triage use?

Triage is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triage use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.

What are the alternatives to Triage?

Skills that share tags, products or a category with Triage: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 36k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triage?

anthropics (a GitHub organization, an official publisher) maintains it in anthropics/oncall-kit, which has 213 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on August 6, 2026.

Source: anthropics/oncall-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.