Kubernetes Network Root Cause Analysis
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it…
$ npx skills add cbrock84/headcount --skill incident-management -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cbrock84/headcount incident-management --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/operations/skills/incident-management .claude/skills/incident-management && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "incident-management" agent skill from https://github.com/cbrock84/headcount/tree/main/plugins/operations/skills/incident-management into .claude/skills/incident-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-management", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cbrock84/headcount/tree/main/plugins/operations/skills/incident-managementType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cbrock84/headcount --skill incident-management -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cbrock84/headcount incident-management --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/operations/skills/incident-management .agents/skills/incident-management && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "incident-management" agent skill from https://github.com/cbrock84/headcount/tree/main/plugins/operations/skills/incident-management into .agents/skills/incident-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-management", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cbrock84/headcount --skill incident-management -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cbrock84/headcount incident-management --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/operations/skills/incident-management .cursor/skills/incident-management && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "incident-management" agent skill from https://github.com/cbrock84/headcount/tree/main/plugins/operations/skills/incident-management into .cursor/skills/incident-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-management", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cbrock84/headcount.git --path plugins/operations/skills/incident-management--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cbrock84/headcount --skill incident-management -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cbrock84/headcount incident-management --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/operations/skills/incident-management .gemini/skills/incident-management && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "incident-management" agent skill from https://github.com/cbrock84/headcount/tree/main/plugins/operations/skills/incident-management into .gemini/skills/incident-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-management", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cbrock84/headcount incident-managementInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cbrock84/headcount --skill incident-management -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/operations/skills/incident-management .github/skills/incident-management && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "incident-management" agent skill from https://github.com/cbrock84/headcount/tree/main/plugins/operations/skills/incident-management into .github/skills/incident-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-management", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cbrock84/headcount --skill incident-management -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cbrock84/headcount incident-management --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/operations/skills/incident-management .opencode/skills/incident-management && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "incident-management" agent skill from https://github.com/cbrock84/headcount/tree/main/plugins/operations/skills/incident-management into .opencode/skills/incident-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-management", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
incident-managementRuns an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it…
Incident Management is an agent skill from cbrock84/headcount. Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it happens, deciding when it is over, and running a review that produces a small number of changes someone actually completes. Use this to set up an incident process, run one, work out why the same failure keeps recurring, or fix a review practice that generates findings nobody closes.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/sources.md`).
It sits in DevOps & Cloud, covering Incident response. The repository describes itself as: An agent organization structured as a company — 15+ departments, 125+ skills, each independently installable, citing the standards and regulators that settle the question. Runs… The licence is MIT.
Read from SKILL.md and the folder at commit 98d1c17. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Incident Management loads about 1.3k tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 125 tokens; SKILL.md has 733 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from cbrock84/headcount at commit 98d1c17, republished under its MIT licence (© cbrock84). 733 words, ~1,273 tokens.
.claude/skills/incident-management/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.An incident is any unplanned disruption significant enough that normal work stops until it is resolved. The discipline exists because the instincts that serve people well in ordinary work — investigate thoroughly, decide carefully, keep everyone informed — all fail under time pressure unless someone has structured them in advance.
This covers operational incidents generally. For security incidents, where evidence preservation
and disclosure obligations change the order of operations, see security:incident-response.
The most expensive minutes are the ones spent deciding whether this is an incident. Set a low threshold for declaring and accept that some declarations will be withdrawn — an incident stood down after twenty minutes costs far less than one that ran for two hours as a conversation between three people who each assumed someone else had it.
Severity should be defined in advance, in terms of customer impact rather than internal inconvenience, with each level carrying a stated response: who is notified, how fast, and who can be woken.
One person owns the incident: they decide, they sequence, they assign. They should not be the person with their hands in the system — the moment the commander starts debugging, nobody is running the incident and the timeline stops being kept.
Separate three roles even in a small response: the commander, the people restoring service, and one person handling communication. Combining the first and third is survivable; combining the first and second is how incidents run long without anyone noticing they have.
The goal during the incident is service restored, not cause understood. Roll back, fail over, disable the feature, add capacity — whatever returns the customer to working. Diagnosis is tomorrow's work, and pursuing it while people are affected is the most common way a thirty-minute outage becomes a four-hour one.
Preserve what you will need to diagnose before you destroy it. Capture logs, a snapshot, the current state — then restore. This is the one step where a minute spent now saves the entire review.
Written as it happens, not reconstructed afterward. What was observed, what was changed, at what time, by whom. Memory of an incident is unreliable within hours and the timeline is what makes the review worth anything.
One channel, and everything in it. Side conversations produce a response where two people are acting on different information.
Say what is happening, what you are doing, what the impact is, and when you will next update — then send that next update on time even if it says nothing has changed. Silence is read as absence, and the update cost is trivial compared to the calls it prevents.
Do not speculate on cause while the incident is open. An early theory shared externally and later withdrawn does more damage than the outage.
Service restored is not the same as incident closed. Confirm recovery is holding, confirm the backlog it created has been worked through, and say explicitly that it is over so people can stop.
Within a few days, while memory is fresh. The question is what about the system made this failure possible and made it take this long to resolve — not who typed the command. A review that produces blame produces less information every subsequent time, because people stop volunteering what actually happened.
Produce one or two changes, not fifteen. A short list that gets done beats a thorough list that does not, and an unclosed action from a previous incident is the most common finding in the next one. Track them to completion somewhere visible, with owners and dates.
references/sources.md in this skill lists the outside authorities that settle the questions
here — what each one is authoritative for, and what you may do with it. Check them before
answering on anything they cover, and cite what you used. Most are free to read and not free
to reproduce; the use note on each is binding.
© cbrock84, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/operations/skills/incident-management of cbrock84/headcount.
Open the folder on GitHubat commit 98d1c17
Incident Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Incident Management this skillcbrock84/headcount | 2k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Kubernetes Network Root Cause Analysiskubeshark/kubeshark | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | |
| Nix Config Debugryan4yin/nix-config | 2.1k | — | ~1.2k | Automated safety check: Pass | MIT | |
| UModel Root Cause Analysisalibaba/UnifiedModel | 415 | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Learningskortix-ai/suna | 20k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Oncallpigweed-project/pigweed | 548 | — | ~963 | Automated safety check: Pass | Apache-2.0 |
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
ryan4yin/nix-config
A skill your agent uses when something here is broken or stops working: an eval or build error, a failed activation, a dead or restarting unit, a mihomo or DNS outage, an unreachable host or MicroVM…
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
kortix-ai/suna
The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
cobusgreyling/loop-engineering
Turns CI failures, open issues, recent commits and chat threads into a prioritized markdown report that an automation loop can act on without inventing architecture work.
cbrock84/headcount
Designs orchestrator-and-subagent hierarchies for a repository — splitting agents by exclusive write surface, pairing every producer with an independent auditor, and enforcing the split with a…
cbrock84/headcount
Designs and audits who can reach what — authentication, authorization models, privileged access, service credentials, and joiner-mover-leaver process.
cbrock84/headcount
Concentrates marketing and sales effort on a named set of accounts rather than on volume — qualifying whether the model fits your economics at all, building the account list and the buying group…
cbrock84/headcount
Gets new users from signup to first real value — signup flow, onboarding, time-to-value, and the early experience that determines whether someone becomes a user or a lapsed account.
cbrock84/headcount
Governs models and AI systems in production — intended use, evaluation, monitoring, human oversight, documentation, and the decision to deploy or retire.
cbrock84/headcount
Produces executive-level research — market sizing, competitor mapping, trend analysis, and strategic intelligence — grounded in cited sources with the confidence in each claim made explicit.
Categories
Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it…. Incident Management is an agent skill from cbrock84/headcount. Runs an operational incident from detection to closed action — declaring it and naming a commander, separating the people restoring service from the people communicating, keeping a timeline as it happens, deciding when it is over, and running a review that produces a small number of changes someone actually completes.
Incident Management fits situations like: tasks that involve Incident response.
Run `npx skills add cbrock84/headcount --skill incident-management -a claude-code`. Or copy the skill folder (plugins/operations/skills/incident-management in cbrock84/headcount) into .claude/skills/incident-management in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cbrock84/headcount --skill incident-management -a codex`. Or copy the skill folder (plugins/operations/skills/incident-management in cbrock84/headcount) into .agents/skills/incident-management in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cbrock84/headcount --skill incident-management -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-management, .gemini/skills/incident-management, .github/skills/incident-management and .opencode/skills/incident-management in your project.
SKILL.md names no scripts, command-line tools or credentials: Incident Management is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Incident Management is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 448 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Incident Management: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), Nix Config Debug (ryan4yin/nix-config, 2.1k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 415 stars) and Learnings (kortix-ai/suna, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cbrock84 (a GitHub user) maintains it in cbrock84/headcount, which has 2,022 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on September 17, 2026.
Source: cbrock84/headcount on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.