Kubernetes Network Root Cause Analysis
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making.
$ npx skills add rampstackco/claude-skills --skill incident-response -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install rampstackco/claude-skills incident-response --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/incident-response .claude/skills/incident-response && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "incident-response" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/incident-response into .claude/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/rampstackco/claude-skills/tree/main/skills/incident-responseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add rampstackco/claude-skills --skill incident-response -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install rampstackco/claude-skills incident-response --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/incident-response .agents/skills/incident-response && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "incident-response" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/incident-response into .agents/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rampstackco/claude-skills --skill incident-response -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install rampstackco/claude-skills incident-response --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/incident-response .cursor/skills/incident-response && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "incident-response" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/incident-response into .cursor/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/rampstackco/claude-skills.git --path skills/incident-response--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add rampstackco/claude-skills --skill incident-response -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install rampstackco/claude-skills incident-response --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/incident-response .gemini/skills/incident-response && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/incident-response into .gemini/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install rampstackco/claude-skills incident-responseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add rampstackco/claude-skills --skill incident-response -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/incident-response .github/skills/incident-response && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/incident-response into .github/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rampstackco/claude-skills --skill incident-response -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install rampstackco/claude-skills incident-response --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/incident-response .opencode/skills/incident-response && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/incident-response into .opencode/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
incident-responseManage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making.
Incident Response is an agent skill from rampstackco/claude-skills. Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making. Use this skill whenever the user has an active incident, a production issue, a service outage, a security incident, or needs to plan incident response procedures. Triggers on incident response, production incident, outage, service down, site down, P0, P1, severity, downtime, on-call, incident commander, status page, postmortem prep. Also triggers when something is…
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `README.md` and `references/incident-playbook.md`).
It sits in DevOps & Cloud, covering Incident response. The repository describes itself as: Stack-agnostic Claude Skills covering the full website lifecycle: brand, design, content, SEO, dev, ops, growth, and research. Build, ship, audit, optimize. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 482c9bf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Incident Response loads about 2.5k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 151 tokens; SKILL.md has 1,137 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from rampstackco/claude-skills at commit 482c9bf, republished under its MIT licence (© rampstackco). 1,137 words, ~2,490 tokens.
.claude/skills/incident-response/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Manage active production incidents from detection to resolution. Stack-agnostic. Tool-agnostic.
This skill is for active incidents and incident process. For after-the-fact analysis, use after-action-report. For planned launches, use launch-runbook.
after-action-report)launch-runbook)qa-testing)How the incident becomes known.
Detection sources:
On detection:
Establish severity and impact.
Severity rubric:
| Severity | Definition | Response |
|---|---|---|
| SEV-1 (Critical) | Major customer-facing functionality broken. Data integrity at risk. Security breach. | All-hands. Incident commander. Active war room. Public communication required. |
| SEV-2 (Major) | Significant degradation. Some customers affected. Revenue impact. | Incident commander assigned. Active response. Internal communication. May or may not need public communication. |
| SEV-3 (Minor) | Limited impact. Workaround available. Affecting a small group of users. | Standard on-call response. Single owner. |
| SEV-4 (Low) | Cosmetic, edge-case, or low-frequency. No urgent action needed. | Tracked as bug. Addressed in normal queue. |
Severity can change. Re-evaluate as more info emerges.
Stop the bleeding before fixing the cause.
Mitigation patterns (faster than full fix):
Mitigation principle: Stop user impact first. Cause analysis second.
Three audiences during an incident:
Internal team:
Internal stakeholders:
External / customers:
Communication principles:
Verified fix, customers restored, incident closed.
Resolution criteria:
After closure:
| Role | Responsibility |
|---|---|
| Incident commander (IC) | Owns the response. Calls decisions. Assigns work. Not necessarily the most technical person; needs to coordinate. |
| Communications lead | Owns internal and external messaging. Reduces IC's communication burden. |
| Operations lead | Drives the technical investigation and mitigation. Often the most senior on-call engineer. |
| Scribe | Captures the timeline as the incident unfolds. Critical for AAR. |
| Subject matter experts | Pulled in as needed. Service owners, database experts, security experts. |
For small teams or low-severity incidents, one person can hold multiple roles. Each role's responsibilities should still be explicit.
The IC's authority:
Non-decisions to avoid:
When in doubt: act. A wrong action that can be rolled back beats inaction while users suffer.
Initial:
"We are investigating reports of [issue]. Updates to follow."
Identified:
"We have identified the issue affecting [scope]. Engineers are working on a fix. Next update by [time]."
Monitoring:
"A fix has been applied. We are monitoring to confirm resolution. Next update by [time]."
Resolved:
"This incident has been resolved. Service has been restored. A full incident report will be posted within [timeframe]."
Patterns to avoid:
During an active incident: incident channel updates and status page updates as per the framework above.
After incident close: a brief incident summary feeding into the AAR.
# Incident: [Brief title]
**Date:** [YYYY-MM-DD]
**Severity:** [SEV-1 / 2 / 3 / 4]
**Duration:** [Detection to resolution]
**Customer impact:** [Who, how many, how, or state the gap per the data-availability rule]
## Summary
[1 to 2 paragraphs]
## Timeline
[Timestamped events]
## Mitigation
[What was done]
## Action items
[Follow-ups, with owners]
## AAR scheduled for
[Date]This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
references/incident-playbook.md - Severity definitions, roles, status page templates, decision rubrics.© rampstackco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/incident-response of rampstackco/claude-skills.
Open the folder on GitHubat commit 482c9bf
Incident Response next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Incident Response this skillrampstackco/claude-skills | 940 | — | ~2.5k | Automated safety check: Pass | MIT | |
| Kubernetes Network Root Cause Analysiskubeshark/kubeshark | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | |
| UModel Root Cause Analysisalibaba/UnifiedModel | 412 | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Learningskortix-ai/suna | 20k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Oncallpigweed-project/pigweed | 548 | — | ~992 | Automated safety check: Pass | Apache-2.0 | |
| Loop Triage Reportcobusgreyling/loop-engineering | 11k | — | ~500 | Automated safety check: Pass | MIT |
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
kortix-ai/suna
The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
cobusgreyling/loop-engineering
Turns CI failures, open issues, recent commits and chat threads into a prioritized markdown report that an automation loop can act on without inventing architecture work.
openclaw/clawhub
Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.
rampstackco/claude-skills
Run a structured after-action review (postmortem, retrospective) on a launch, incident, or completed project to capture timeline, root cause analysis, contributing factors, and actionable lessons.
rampstackco/claude-skills
Design measurement frameworks including event taxonomy, KPI hierarchy, dashboard architecture, attribution models, and analytics implementation strategy.
rampstackco/claude-skills
Build or audit a comprehensive brand style guide that documents the full brand system including story, logo system, color, typography, imagery, voice, applications, and dos/don'ts.
rampstackco/claude-skills
Develop or document a complete brand voice and tone system covering voice attributes, tone shifts by context, vocabulary preferences, grammar rules, and copy examples.
rampstackco/claude-skills
Write or edit website copy, blog content, and editorial pieces with attention to voice, structure, and goal.
rampstackco/claude-skills
Develop a content strategy covering editorial positioning, content pillars, formats, calendar, governance, and topical authority planning.
Categories
Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making. Incident Response is an agent skill from rampstackco/claude-skills. Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making.
Incident Response fits situations like: the user has an active incident; A production issue; A service outage; A security incident.
Run `npx skills add rampstackco/claude-skills --skill incident-response -a claude-code`. Or copy the skill folder (skills/incident-response in rampstackco/claude-skills) into .claude/skills/incident-response in your project. Claude Code loads it when a task matches its description.
Run `npx skills add rampstackco/claude-skills --skill incident-response -a codex`. Or copy the skill folder (skills/incident-response in rampstackco/claude-skills) into .agents/skills/incident-response in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rampstackco/claude-skills --skill incident-response -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-response, .gemini/skills/incident-response, .github/skills/incident-response and .opencode/skills/incident-response in your project.
SKILL.md names no scripts, command-line tools or credentials: Incident Response is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Incident Response is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Incident Response: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 412 stars), Learnings (kortix-ai/suna, 20k stars) and Oncall (pigweed-project/pigweed, 548 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
rampstackco (a GitHub organization) maintains it in rampstackco/claude-skills, which has 940 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 7, 2026.
Source: rampstackco/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.