Agent skill

Incident Management

by cbrock84 in cbrock84/headcount

Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it…

MITAuto-check passedDevOps & Cloud

Install Incident Management

skills CLI
$ npx skills add cbrock84/headcount --skill incident-management -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cbrock84/headcount incident-management --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/operations/skills/incident-management .claude/skills/incident-management && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
incident-management
GitHub stars
2k
Token cost
~1.3k tokens
SKILL.md length
733 words
Files
2 (incl. references)
Skills in repo
178
Repo updated
First seen
Licence
MIT

At a glance

Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it…

  • Tasks that involve Incident response
  • SKILL.md covers Declare it, and say so out loud, Name a commander who does not…, Restore first, understand later and Keep a timeline while it is…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Incident Management is an agent skill from cbrock84/headcount. Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it happens, deciding when it is over, and running a review that produces a small number of changes someone actually completes. Use this to set up an incident process, run one, work out why the same failure keeps recurring, or fix a review practice that generates findings nobody closes.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/sources.md`).

It sits in DevOps & Cloud, covering Incident response. The repository describes itself as: An agent organization structured as a company — 15+ departments, 125+ skills, each independently installable, citing the standards and regulators that settle the question. Runs… The licence is MIT.

When your agent uses it

  • Tasks that involve Incident response

Example prompts

  • “/incident-management”

What it can do on your machine

Read from SKILL.md and the folder at commit 98d1c17. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Incident Management loads about 1.3k tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 125 tokens; SKILL.md has 733 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cbrock84/headcount at commit 98d1c17, republished under its MIT licence (© cbrock84). 733 words, ~1,273 tokens.

Download SKILL.mdSave it as .claude/skills/incident-management/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
incident-management
description
Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it happens, deciding when it is over, and running a review that produces a small number of changes someone actually completes. Use this to set up an incident process, run one, work out why the same failure keeps recurring, or fix a review practice that generates findings nobody closes.

Incident management

An incident is any unplanned disruption significant enough that normal work stops until it is resolved. The discipline exists because the instincts that serve people well in ordinary work — investigate thoroughly, decide carefully, keep everyone informed — all fail under time pressure unless someone has structured them in advance.

This covers operational incidents generally. For security incidents, where evidence preservation and disclosure obligations change the order of operations, see security:incident-response.

Declare it, and say so out loud

The most expensive minutes are the ones spent deciding whether this is an incident. Set a low threshold for declaring and accept that some declarations will be withdrawn — an incident stood down after twenty minutes costs far less than one that ran for two hours as a conversation between three people who each assumed someone else had it.

Severity should be defined in advance, in terms of customer impact rather than internal inconvenience, with each level carrying a stated response: who is notified, how fast, and who can be woken.

Name a commander who does not fix anything

One person owns the incident: they decide, they sequence, they assign. They should not be the person with their hands in the system — the moment the commander starts debugging, nobody is running the incident and the timeline stops being kept.

Separate three roles even in a small response: the commander, the people restoring service, and one person handling communication. Combining the first and third is survivable; combining the first and second is how incidents run long without anyone noticing they have.

Restore first, understand later

The goal during the incident is service restored, not cause understood. Roll back, fail over, disable the feature, add capacity — whatever returns the customer to working. Diagnosis is tomorrow's work, and pursuing it while people are affected is the most common way a thirty-minute outage becomes a four-hour one.

Preserve what you will need to diagnose before you destroy it. Capture logs, a snapshot, the current state — then restore. This is the one step where a minute spent now saves the entire review.

Keep a timeline while it is happening

Written as it happens, not reconstructed afterward. What was observed, what was changed, at what time, by whom. Memory of an incident is unreliable within hours and the timeline is what makes the review worth anything.

One channel, and everything in it. Side conversations produce a response where two people are acting on different information.

Show full SKILL.md (327 more words)Show less

Communicate on a rhythm, including when there is nothing new

Say what is happening, what you are doing, what the impact is, and when you will next update — then send that next update on time even if it says nothing has changed. Silence is read as absence, and the update cost is trivial compared to the calls it prevents.

Do not speculate on cause while the incident is open. An early theory shared externally and later withdrawn does more damage than the outage.

Declare it over deliberately

Service restored is not the same as incident closed. Confirm recovery is holding, confirm the backlog it created has been worked through, and say explicitly that it is over so people can stop.

Run the review on the system, not the person

Within a few days, while memory is fresh. The question is what about the system made this failure possible and made it take this long to resolve — not who typed the command. A review that produces blame produces less information every subsequent time, because people stop volunteering what actually happened.

Produce one or two changes, not fifteen. A short list that gets done beats a thorough list that does not, and an unclosed action from a previous incident is the most common finding in the next one. Track them to completion somewhere visible, with owners and dates.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Never

  • Debug while nobody is commanding.
  • Destroy the state you will need to diagnose in order to restore a minute sooner.
  • Skip a scheduled update because there is nothing new to say.
  • Close a review with more actions than the organization will actually complete.

© cbrock84, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/operations/skills/incident-management of cbrock84/headcount.

  • SKILL.md
  • references/sources.md

Open the folder on GitHubat commit 98d1c17

Compare with similar skills

Incident Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Incident Management compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Incident Management this skillcbrock84/headcount2k—~1.3kAutomated safety check: PassMIT
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0
Nix Config Debugryan4yin/nix-config2.1k—~1.2kAutomated safety check: PassMIT
UModel Root Cause Analysisalibaba/UnifiedModel415—~1.9kAutomated safety check: PassCustom licence
Learningskortix-ai/suna20k—~1.1kAutomated safety check: PassCustom licence
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0

Similar skills

  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Nix Config Debug

    ryan4yin/nix-config

    A skill your agent uses when something here is broken or stops working: an eval or build error, a failed activation, a dead or restarting unit, a mihomo or DNS outage, an unreachable host or MicroVM…

    2.1k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    415 GitHub stars~1.9k tokensUpdated 16 days ago
    DevOps & CloudAuto-check passed
  • Learnings

    kortix-ai/suna

    The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.

    20k GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Loop Triage Report

    cobusgreyling/loop-engineering

    Turns CI failures, open issues, recent commits and chat threads into a prioritized markdown report that an automation loop can act on without inventing architecture work.

    11k GitHub stars~500 tokensUpdated today
    DevOps & CloudAuto-check passed

More from cbrock84/headcount

All 178 skills in this repo
  • Agent Hierarchy

    cbrock84/headcount

    Designs orchestrator-and-subagent hierarchies for a repository — splitting agents by exclusive write surface, pairing every producer with an independent auditor, and enforcing the split with a…

    2k GitHub stars~1.2k tokensUpdated 23 days ago
    Auto-check passed
  • Access And Identity

    cbrock84/headcount

    Designs and audits who can reach what — authentication, authorization models, privileged access, service credentials, and joiner-mover-leaver process.

    2k GitHub stars~1.1k tokensUpdated 23 days ago
    Auto-check passed
  • Account Based Marketing

    cbrock84/headcount

    Concentrates marketing and sales effort on a named set of accounts rather than on volume — qualifying whether the model fits your economics at all, building the account list and the buying group…

    2k GitHub stars~1.2k tokensUpdated 23 days ago
    Auto-check passed
  • Activation

    cbrock84/headcount

    Gets new users from signup to first real value — signup flow, onboarding, time-to-value, and the early experience that determines whether someone becomes a user or a lapsed account.

    2k GitHub stars~865 tokensUpdated 23 days ago
    Auto-check passed
  • AI ML Governance

    cbrock84/headcount

    Governs models and AI systems in production — intended use, evaluation, monitoring, human oversight, documentation, and the decision to deploy or retire.

    2k GitHub stars~1k tokensUpdated 23 days ago
    Auto-check passed
  • AI Research Analyst

    cbrock84/headcount

    Produces executive-level research — market sizing, competitor mapping, trend analysis, and strategic intelligence — grounded in cited sources with the confidence in each claim made explicit.

    2k GitHub stars~916 tokensUpdated 23 days ago
    Auto-check passed

Categories

Questions about Incident Management

What does Incident Management do?

Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it…. Incident Management is an agent skill from cbrock84/headcount. Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it happens, deciding when it is over, and running a review that produces a small number of changes someone actually completes.

When should I use Incident Management?

Incident Management fits situations like: tasks that involve Incident response.

How do I install Incident Management in Claude Code?

Run `npx skills add cbrock84/headcount --skill incident-management -a claude-code`. Or copy the skill folder (plugins/operations/skills/incident-management in cbrock84/headcount) into .claude/skills/incident-management in your project. Claude Code loads it when a task matches its description.

How do I install Incident Management in Codex?

Run `npx skills add cbrock84/headcount --skill incident-management -a codex`. Or copy the skill folder (plugins/operations/skills/incident-management in cbrock84/headcount) into .agents/skills/incident-management in your project. Codex loads it when a task matches its description.

Can I use Incident Management in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cbrock84/headcount --skill incident-management -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-management, .gemini/skills/incident-management, .github/skills/incident-management and .opencode/skills/incident-management in your project.

What does Incident Management need to run?

SKILL.md names no scripts, command-line tools or credentials: Incident Management is instructions for the agent only.

Does Incident Management access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Incident Management safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Incident Management use?

Incident Management is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Incident Management use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 448 tokens, read only when the agent opens those files.

What are the alternatives to Incident Management?

Skills that share tags, products or a category with Incident Management: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), Nix Config Debug (ryan4yin/nix-config, 2.1k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 415 stars) and Learnings (kortix-ai/suna, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Incident Management?

cbrock84 (a GitHub user) maintains it in cbrock84/headcount, which has 2,022 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on September 17, 2026.

Source: cbrock84/headcount on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.