Agent skill

Incident Response

by rampstackco in rampstackco/claude-skills

Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making.

MITAuto-check passedDevOps & Cloud

Install Incident Response

skills CLI
$ npx skills add rampstackco/claude-skills --skill incident-response -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rampstackco/claude-skills incident-response --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/incident-response .claude/skills/incident-response && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
incident-response
GitHub stars
940
Token cost
~2.5k tokens
SKILL.md length
1,137 words
Files
3 (incl. references)
Skills in repo
103
Repo updated
First seen
Licence
MIT

At a glance

Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making.

  • Works in 5 steps: Detection → Triage → Mitigation → …
  • The user has an active incident
  • SKILL.md covers When to use, When NOT to use, Required inputs and The framework: 5 phases, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Incident Response is an agent skill from rampstackco/claude-skills. Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making. Use this skill whenever the user has an active incident, a production issue, a service outage, a security incident, or needs to plan incident response procedures. Triggers on incident response, production incident, outage, service down, site down, P0, P1, severity, downtime, on-call, incident commander, status page, postmortem prep. Also triggers when something is…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `README.md` and `references/incident-playbook.md`).

It sits in DevOps & Cloud, covering Incident response. The repository describes itself as: Stack-agnostic Claude Skills covering the full website lifecycle: brand, design, content, SEO, dev, ops, growth, and research. Build, ship, audit, optimize. The licence is MIT.

When your agent uses it

  • The user has an active incident
  • A production issue
  • A service outage
  • A security incident

Example prompts

  • “/incident-response”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Detection
  2. Triage
  3. Mitigation
  4. Communication
  5. Resolution

What it can do on your machine

Read from SKILL.md and the folder at commit 482c9bf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Incident Response loads about 2.5k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 151 tokens; SKILL.md has 1,137 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~151
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rampstackco/claude-skills at commit 482c9bf, republished under its MIT licence (© rampstackco). 1,137 words, ~2,490 tokens.

Download SKILL.mdSave it as .claude/skills/incident-response/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
incident-response
description
Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making. Use this skill whenever the user has an active incident, a production issue, a service outage, a security incident, or needs to plan incident response procedures. Triggers on incident response, production incident, outage, service down, site down, P0, P1, severity, downtime, on-call, incident commander, status page, postmortem prep. Also triggers when something is actively broken in production and the user is figuring out what to do.
category
operations
catalog_summary
Incident triage, comms, mitigation, escalation
display_order
2

Incident Response

Manage active production incidents from detection to resolution. Stack-agnostic. Tool-agnostic.

This skill is for active incidents and incident process. For after-the-fact analysis, use after-action-report. For planned launches, use launch-runbook.


When to use

  • An active incident is happening
  • Building incident response procedures
  • Defining severity levels
  • Setting up on-call rotations
  • Training a team on incident response

When NOT to use

  • Post-incident retrospective (use after-action-report)
  • Planned launches (use launch-runbook)
  • Pre-launch issue triage (use qa-testing)

Required inputs

  • Awareness of the incident (alert, customer report, internal observation)
  • Access to production systems and monitoring
  • Roles and authorities clearly defined
  • Communication channels operational

The framework: 5 phases

1. Detection

How the incident becomes known.

Detection sources:

  • Automated alerts (monitoring, SLO violations, error rate spikes)
  • Customer reports (support tickets, social media, status page subscribers)
  • Internal observation (engineer notices something off)
  • Third-party (security researchers, partners)

On detection:

  • Acknowledge within target time (typically 5 to 15 minutes for critical)
  • Assess severity (see severity rubric below)
  • Page the on-call if not already paged
  • Open the incident channel
2. Triage

Establish severity and impact.

Severity rubric:

SeverityDefinitionResponse
SEV-1 (Critical)Major customer-facing functionality broken. Data integrity at risk. Security breach.All-hands. Incident commander. Active war room. Public communication required.
SEV-2 (Major)Significant degradation. Some customers affected. Revenue impact.Incident commander assigned. Active response. Internal communication. May or may not need public communication.
SEV-3 (Minor)Limited impact. Workaround available. Affecting a small group of users.Standard on-call response. Single owner.
SEV-4 (Low)Cosmetic, edge-case, or low-frequency. No urgent action needed.Tracked as bug. Addressed in normal queue.

Severity can change. Re-evaluate as more info emerges.

3. Mitigation

Stop the bleeding before fixing the cause.

Mitigation patterns (faster than full fix):

  • Rollback (revert recent deploy)
  • Feature flag off (disable the broken feature without deploy)
  • Failover (route to healthy replica or region)
  • Scale up (more capacity to absorb the load)
  • Throttle (reject some traffic to protect the rest)
  • Graceful degradation (turn off non-essential features to keep core functional)
  • Maintenance mode (last resort, blocks all users)

Mitigation principle: Stop user impact first. Cause analysis second.

4. Communication

Three audiences during an incident:

Internal team:

  • Real-time updates in incident channel
  • Cadence: every 15 minutes minimum during active incident
  • Format: timestamped status updates with what we know, what we're doing, ETA

Internal stakeholders:

  • Higher-level updates to broader org
  • Cadence: every 30 to 60 minutes
  • Format: business-impact framing, not technical detail

External / customers:

  • Status page updates
  • Cadence: every 30 minutes minimum during active incident
  • Format: plain language, no blame, what users are experiencing, what to expect

Communication principles:

  • Acknowledge before you have answers ("We're aware and investigating")
  • Update on schedule even if no progress ("Still investigating, no new information")
  • Never speculate publicly about cause
  • Confirm resolution explicitly when restored
5. Resolution

Verified fix, customers restored, incident closed.

Resolution criteria:

  • Mitigation in place and verified
  • Root cause identified (or explicitly deferred to AAR)
  • All affected systems back to normal
  • Customers can resume normal use
  • Final status update posted (internal and external)
  • Incident channel can be closed (or archived for AAR)

After closure:

  • Schedule AAR within 1 to 2 weeks
  • Capture initial timeline while memories are fresh
  • Track follow-up action items

Roles during an incident

RoleResponsibility
Incident commander (IC)Owns the response. Calls decisions. Assigns work. Not necessarily the most technical person; needs to coordinate.
Communications leadOwns internal and external messaging. Reduces IC's communication burden.
Operations leadDrives the technical investigation and mitigation. Often the most senior on-call engineer.
ScribeCaptures the timeline as the incident unfolds. Critical for AAR.
Subject matter expertsPulled in as needed. Service owners, database experts, security experts.

For small teams or low-severity incidents, one person can hold multiple roles. Each role's responsibilities should still be explicit.


Decision-making during an incident

The IC's authority:

  • Call rollback or other mitigations
  • Pull additional people in
  • Escalate severity
  • Make the call when unclear options exist

Non-decisions to avoid:

  • "Let's wait and see" when mitigations are available and impact is occurring
  • Discussing root cause while users are actively impacted (mitigate first)
  • Premature resolution announcements before verification
  • Death-by-committee (pull in lots of people, no one decides)

When in doubt: act. A wrong action that can be rolled back beats inaction while users suffer.


Show full SKILL.md (440 more words)Show less

Status page communication patterns

Initial:

"We are investigating reports of [issue]. Updates to follow."

Identified:

"We have identified the issue affecting [scope]. Engineers are working on a fix. Next update by [time]."

Monitoring:

"A fix has been applied. We are monitoring to confirm resolution. Next update by [time]."

Resolved:

"This incident has been resolved. Service has been restored. A full incident report will be posted within [timeframe]."

Patterns to avoid:

  • Vague language ("experiencing some issues" - what kind?)
  • Missing affected scope ("login is down" - everywhere or just one region?)
  • Missing time commitments
  • "Should be resolved soon" without verification
  • Using "back up" before verification

Workflow

  1. Acknowledge. First responder acknowledges within target time.
  2. Assess severity. Use the rubric. Open the appropriate response channel.
  3. Assign roles. IC, comms, ops at minimum.
  4. Communicate. Initial status update. Internal channel active.
  5. Investigate. Logs, metrics, recent changes. The four most common causes: a recent deploy, a configuration change, a third-party dependency change, a load spike.
  6. Mitigate. Stop the bleeding. Don't wait for full root cause.
  7. Verify mitigation. Don't trust dashboards alone; test the user flow.
  8. Communicate resolution. Internal and external.
  9. Close incident. Final timeline noted. Action items tracked.
  10. Schedule AAR. Within 1 to 2 weeks.

Failure patterns

  • No clear IC. Multiple people debugging in parallel, no coordination. Slower to mitigate, easier to make conflicting changes.
  • Skipping mitigation, going straight to root cause. Users keep suffering while engineers debug.
  • Premature "all clear." Announcing resolution before verification.
  • Communication silence. Users don't know if anyone is working on it.
  • Status updates too vague. "We're working on it" with no detail.
  • Speculating publicly about cause. Often wrong, always damaging trust.
  • Pulling in too many people. Coordination overhead exceeds value.
  • No scribe. The timeline gets lost. AAR has to reconstruct from chat logs.
  • Skipping AAR for "minor" incidents. Patterns get missed. Lessons get re-learned.
  • Blame culture. People hide mistakes, incidents take longer.

Output format

During an active incident: incident channel updates and status page updates as per the framework above.

After incident close: a brief incident summary feeding into the AAR.

markdown
# Incident: [Brief title]

**Date:** [YYYY-MM-DD]
**Severity:** [SEV-1 / 2 / 3 / 4]
**Duration:** [Detection to resolution]
**Customer impact:** [Who, how many, how, or state the gap per the data-availability rule]

## Summary
[1 to 2 paragraphs]

## Timeline
[Timestamped events]

## Mitigation
[What was done]

## Action items
[Follow-ups, with owners]

## AAR scheduled for
[Date]

If required data is unavailable

This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.


Reference files

© rampstackco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/incident-response of rampstackco/claude-skills.

  • SKILL.md
  • README.md
  • references/incident-playbook.md

Open the folder on GitHubat commit 482c9bf

Compare with similar skills

Incident Response next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Incident Response compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Incident Response this skillrampstackco/claude-skills940—~2.5kAutomated safety check: PassMIT
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0
UModel Root Cause Analysisalibaba/UnifiedModel412—~1.9kAutomated safety check: PassCustom licence
Learningskortix-ai/suna20k—~1.1kAutomated safety check: PassCustom licence
Oncallpigweed-project/pigweed548—~992Automated safety check: PassApache-2.0
Loop Triage Reportcobusgreyling/loop-engineering11k—~500Automated safety check: PassMIT

Similar skills

  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    412 GitHub stars~1.9k tokensUpdated 14 days ago
    DevOps & CloudAuto-check passed
  • Learnings

    kortix-ai/suna

    The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.

    20k GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~992 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Loop Triage Report

    cobusgreyling/loop-engineering

    Turns CI failures, open issues, recent commits and chat threads into a prioritized markdown report that an automation loop can act on without inventing architecture work.

    11k GitHub stars~500 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated today
    DevOps & CloudAuto-check passed

More from rampstackco/claude-skills

All 103 skills in this repo
  • After Action Report

    rampstackco/claude-skills

    Run a structured after-action review (postmortem, retrospective) on a launch, incident, or completed project to capture timeline, root cause analysis, contributing factors, and actionable lessons.

    940 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Analytics Strategy

    rampstackco/claude-skills

    Design measurement frameworks including event taxonomy, KPI hierarchy, dashboard architecture, attribution models, and analytics implementation strategy.

    940 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Brand Style Guide

    rampstackco/claude-skills

    Build or audit a comprehensive brand style guide that documents the full brand system including story, logo system, color, typography, imagery, voice, applications, and dos/don'ts.

    940 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Brand Voice

    rampstackco/claude-skills

    Develop or document a complete brand voice and tone system covering voice attributes, tone shifts by context, vocabulary preferences, grammar rules, and copy examples.

    940 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Content And Copy

    rampstackco/claude-skills

    Write or edit website copy, blog content, and editorial pieces with attention to voice, structure, and goal.

    940 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Content Strategy

    rampstackco/claude-skills

    Develop a content strategy covering editorial positioning, content pillars, formats, calendar, governance, and topical authority planning.

    940 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Incident Response

What does Incident Response do?

Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making. Incident Response is an agent skill from rampstackco/claude-skills. Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making.

When should I use Incident Response?

Incident Response fits situations like: the user has an active incident; A production issue; A service outage; A security incident.

How do I install Incident Response in Claude Code?

Run `npx skills add rampstackco/claude-skills --skill incident-response -a claude-code`. Or copy the skill folder (skills/incident-response in rampstackco/claude-skills) into .claude/skills/incident-response in your project. Claude Code loads it when a task matches its description.

How do I install Incident Response in Codex?

Run `npx skills add rampstackco/claude-skills --skill incident-response -a codex`. Or copy the skill folder (skills/incident-response in rampstackco/claude-skills) into .agents/skills/incident-response in your project. Codex loads it when a task matches its description.

Can I use Incident Response in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rampstackco/claude-skills --skill incident-response -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-response, .gemini/skills/incident-response, .github/skills/incident-response and .opencode/skills/incident-response in your project.

What does Incident Response need to run?

SKILL.md names no scripts, command-line tools or credentials: Incident Response is instructions for the agent only.

Does Incident Response access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Incident Response safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Incident Response use?

Incident Response is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Incident Response use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Incident Response?

Skills that share tags, products or a category with Incident Response: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 412 stars), Learnings (kortix-ai/suna, 20k stars) and Oncall (pigweed-project/pigweed, 548 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Incident Response?

rampstackco (a GitHub organization) maintains it in rampstackco/claude-skills, which has 940 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 7, 2026.

Source: rampstackco/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.