Agent skill

Incident Response

by kid-sid in kid-sid/claude-spellbook

A skill your agent uses when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.

MITAuto-check passedDevOps & Cloud

Install Incident Response

skills CLI
$ npx skills add kid-sid/claude-spellbook --skill incident-response -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kid-sid/claude-spellbook incident-response --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/incident-response .claude/skills/incident-response && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
incident-response
GitHub stars
190
Token cost
~3k tokens
SKILL.md length
866 words
Files
1
Skills in repo
55
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.

  • Works in 2 steps: Check DB connection pool → Check for slow queries
  • Triaging a production alert
  • SKILL.md covers When to Activate, Severity Classification, Incident Lifecycle and Communication Templates, plus 8 more sections
  • Calls kubectl

What it does

Incident Response is an agent skill from kid-sid/claude-spellbook. Use when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Incident response and Runbooks and postmortems. The repository describes itself as: A curated collection of skills, prompts, and workflows that extend Claude's capabilities — your personal grimoire for AI-powered development. The licence is MIT.

When your agent uses it

  • Triaging a production alert
  • Writing a postmortem
  • Updating a runbook
  • Classifying incident severity

Example prompts

  • “/incident-response”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Check DB connection pool
  2. Check for slow queries

What it can do on your machine

Read from SKILL.md and the folder at commit a7c2ac9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Incident Response loads about 3k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 866 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kid-sid/claude-spellbook at commit a7c2ac9, republished under its MIT licence (© kid-sid). 866 words, ~2,953 tokens.

Download SKILL.mdSave it as .claude/skills/incident-response/SKILL.md (or your agent's skills folder).
name
incident-response
description
Use when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.

Incident Response

Incident response is the structured process of detecting, mitigating, communicating, and learning from production failures to minimise user impact and prevent recurrence.

When to Activate

  • Triaging a production alert or on-call page
  • Writing a postmortem after an incident
  • Creating or updating a runbook for a service
  • Defining severity levels and escalation paths for a team
  • Setting up an on-call rotation
  • Running an incident response drill or game day

Severity Classification

SeverityDefinitionResponse SLAComms cadenceExample
P0Total outage or data loss — all users affectedPage immediately, < 5 minEvery 15 minPayment service down, DB unreachable
P1Major feature broken — most users affected< 15 min acknowledgementEvery 30 minLogin failing for 50%+ of users
P2Significant degradation — subset of users affected< 1 hourEvery 2 hoursSearch slow for US region
P3Minor issue — small impact, workaround availableNext business dayOnce resolvedNon-critical dashboard shows stale data
P4Cosmetic / no user impactSprint backlogN/ALog noise, minor UI misalignment

Escalation path:

  • P0/P1: page on-call engineer → page on-call lead if not ack'd in 5 min → escalate to eng manager
  • P2: page on-call engineer
  • P3/P4: create ticket, no page

Incident Lifecycle

Detection → Triage → Mitigate → Communicate → Resolve → Review (Postmortem)
First 5 Minutes — Triage Checklist
  • Acknowledge the alert and claim the incident in your incident tool (PagerDuty / Opsgenie)
  • Identify: what is broken, who is affected, since when?
  • Check the deployment timeline: was anything deployed in the last 2 hours?
  • Check the dashboards: error rate, latency, saturation — which service is the origin?
  • Open an incident channel: #inc-YYYY-MM-DD-short-description
  • Post initial acknowledgement message (see template below)
  • Assign roles: Incident Commander (IC), Communicator, Subject Matter Expert (SME)

Communication Templates

Initial Acknowledgement
🔴 [P0/P1 INCIDENT] Payment service degradation

Status: Investigating
Impact: ~30% of payment requests failing with 500 errors since 14:23 UTC
Affected: All users attempting checkout

IC: @alice
SME: @bob
Next update: 14:45 UTC

Tracking: https://incident.example.com/inc-2024-0042
Status Update (every 15–30 min for P0/P1)
🟡 [P1 UPDATE] Payment service — 14:45 UTC

Status: Mitigating
Root cause identified: Connection pool exhaustion after deploy at 14:15
Action taken: Rolled back to v2.3.1, monitoring error rate
Current error rate: 2% (down from 30%)

Next update: 15:00 UTC
Resolution
✅ [P1 RESOLVED] Payment service — 15:02 UTC

Status: Resolved
Duration: 39 minutes (14:23 – 15:02 UTC)
Root cause: Deploy v2.4.0 introduced a connection leak; pool exhausted under load
Resolution: Rolled back to v2.3.1; error rate returned to baseline at 15:00

Users impacted: ~15,000 failed checkout attempts
Follow-up: Postmortem scheduled for 2024-01-16 15:00 UTC
Incident report: https://incident.example.com/inc-2024-0042

Mitigation Decision Tree

Error rate > SLO threshold?
├── Yes
│   ├── Was something deployed in the last 2 hours?
│   │   ├── Yes → ROLLBACK first, investigate after
│   │   └── No  → Check: DB, cache, upstream dependency, config change
│   ├── Can we isolate the impact with a feature flag kill?
│   │   └── Yes → Kill the flag immediately
│   └── Is this a traffic spike?
│       └── Yes → Scale up horizontally, enable circuit breaker
└── No — latency degraded only?
    ├── Check DB: slow queries, lock contention, pool saturation
    ├── Check cache hit rate: has cache been evicted?
    └── Check upstream service latency

When NOT to roll back immediately:

  • The new version fixes a critical security issue (rolling back re-introduces the vulnerability)
  • Rollback would itself cause data migration issues
  • The issue is cosmetic (P3/P4) and the fix is already in progress

Runbook Structure

Runbooks must be written for the 3am engineer who has never seen this service.

markdown
# Runbook: [Service Name] — [Alert Name]

## Service Overview
[2–3 sentences: what does this service do, what does it depend on?]

## Alert: [Alert Name]
**Trigger condition:** [e.g., error rate > 1% for 5 minutes]
**Severity:** P1
**Dashboard:** [link]
**Logs:** [link to log query]

## Diagnostic Steps
1. Check the error rate panel on the [service dashboard](link)
   - Expected: < 0.1%
   - If > 1%: proceed to step 2
2. Check recent deployments:
   ```bash
   kubectl rollout history deployment/payment-service -n production
  1. Check DB connection pool:
    bash
    kubectl exec -it $(kubectl get pod -l app=payment-service -o name | head -1) \
      -- curl -s localhost:8080/metrics | grep db_pool
    • If db_pool_wait_duration_seconds > 1s: pool is exhausted, proceed to step 4
  2. Check for slow queries:
    sql
    SELECT query, mean_exec_time, calls
    FROM pg_stat_statements
    ORDER BY mean_exec_time DESC
    LIMIT 10;

Mitigation Steps

  • If recent deployment: kubectl rollout undo deployment/payment-service -n production
  • If DB pool exhausted: Scale up replicas: kubectl scale deployment/payment-service --replicas=6
  • If upstream dependency: Enable circuit breaker feature flag: [link to flag]

Escalation

  • If not resolved in 30 minutes: page @payment-team-lead
  • DB issues: page @dba-on-call
  • Infrastructure: page @infra-on-call

**Runbook quality checks:**
- Every step has an expected output — the engineer knows what "normal" looks like
- Commands are copy-paste ready (no placeholders that need substitution)
- Decision points have explicit branches ("if X, do Y; if Z, do W")
- Links to dashboards, log queries, and escalation contacts are current

## Blameless Postmortem

Write the postmortem within 48 hours while details are fresh. **Blameless = focus on systems and processes, not individuals.**

```markdown
# Postmortem: [Service] [Brief Description] — [Date]

## Summary
[2–3 sentences: what happened, impact, how it was resolved]

**Impact:** [number of users affected, % error rate, duration]
**Detection time:** [how long from start to detection]
**Resolution time:** [how long from detection to resolution]

## Timeline (UTC)
| Time  | Event |
|-------|-------|
| 14:15 | Deploy v2.4.0 rolled out to 100% |
| 14:23 | Alert fired: error rate > 1% |
| 14:28 | On-call acknowledged, started investigation |
| 14:38 | Root cause identified: connection pool exhausted |
| 14:45 | Rollback initiated |
| 15:00 | Error rate returned to baseline |
| 15:02 | Incident declared resolved |

## Root Cause Analysis (5 Whys)
1. **Why** did payment requests fail?
   → DB connection pool was exhausted
2. **Why** was the pool exhausted?
   → v2.4.0 introduced a connection leak in the retry handler
3. **Why** did the retry handler leak connections?
   → The `defer conn.Close()` was placed inside the retry loop, closing on each attempt but not releasing the acquired connection back to the pool
4. **Why** wasn't this caught in testing?
   → Integration tests used a single-connection test DB; pool exhaustion only manifests at scale
5. **Why** wasn't this caught by the integration test DB pool?
   → Test pool size was set to 100 (no practical limit); prod pool size is 20

## Contributing Factors
- No load test run before this deploy
- No DB pool exhaustion alert existed
- Code review missed the subtle connection lifecycle issue

## What Went Well
- Alert fired within 8 minutes of degradation starting
- On-call was paged and acknowledged quickly
- Rollback decision was made in < 10 minutes

## Action Items

| Action | Owner | Due | Category |
|--------|-------|-----|----------|
| Add DB pool wait time alert (threshold: > 1s for 5 min) | @alice | 2024-01-19 | Detection |
| Add integration test that simulates pool exhaustion under concurrent load | @bob | 2024-01-26 | Prevention |
| Add `db_pool_size` check to pre-deploy checklist | @alice | 2024-01-19 | Prevention |
| Run k6 load test before all deploys touching DB connection code | @bob | 2024-01-26 | Prevention |
Action Item Categories
  • Prevention: stops this class of failure from happening
  • Detection: reduces time-to-detection (MTTD)
  • Response: reduces time-to-resolution (MTTR)

Metrics to Track

MetricDefinitionTarget
MTTDMean Time To Detect — start of incident to first alert firing< 5 min
MTTAMean Time To Acknowledge — alert fires to on-call acks< 5 min
MTTRMean Time To Resolve — detection to resolution< 30 min for P0/P1
Incident frequencyNumber of P0/P1 incidents per month per serviceTrack trend; goal: decreasing
Repeat incidentsIncidents with the same root cause as a prior incidentGoal: 0

Review these monthly per service. Rising MTTR = runbooks need updating. Repeat incidents = action items not implemented.

See also: observability, deployment-strategies

Show full SKILL.md (340 more words)Show less

Red Flags

  • Postmortem that names individuals as root cause — "Alice deployed bad code" stops at the human rather than the system that allowed the bad code to reach production; blameless postmortems ask why the system made it possible
  • Action items with no owner or no due date — "Improve monitoring" as an action item is never done; every item must have a named owner and a specific due date to be tracked and closed
  • Runbook that assumes the on-call engineer knows the service — runbooks must include what "normal" looks like and copy-paste commands; a 3am engineer touching an unfamiliar service cannot safely improvise
  • Rolling back immediately without checking if the rollback itself causes data loss — rolling back a deploy that ran a destructive migration may orphan or corrupt rows that were written against the new schema
  • Posting a P0 incident only in an engineering Slack channel — stakeholders (product, support, leadership) need timely updates via their own channels; the Communicator role exists specifically to bridge this gap
  • Severity P0 declared for every outage regardless of blast radius — "P0" becomes meaningless if used for single-user bugs; a calibrated P0 ensures the right resources are mobilized and avoids on-call fatigue
  • MTTD and MTTR tracked per-incident but never aggregated — individual numbers without a monthly trend hide whether the team is improving; review rolling averages per service each month
  • Closing an incident before a postmortem is scheduled — if the postmortem is not scheduled at resolution time it rarely happens; require a postmortem date as a condition of closing any P0 or P1

Checklist

  • Incident acknowledged within SLA (P0: 5 min, P1: 15 min)
  • Incident channel opened and IC/SME roles assigned
  • Initial acknowledgement posted to stakeholder channel
  • Status updates sent on cadence (every 15 min for P0, 30 min for P1)
  • Resolution announcement sent with impact summary
  • Postmortem written within 48 hours of resolution
  • 5 Whys root cause analysis complete (not just "human error")
  • Action items are SMART: owner, due date, and category (prevention/detection/response)
  • Runbook updated based on lessons learned
  • MTTD, MTTA, MTTR recorded for this incident

© kid-sid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/incident-response of kid-sid/claude-spellbook.

Open the folder on GitHubat commit a7c2ac9

Compare with similar skills

Incident Response next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Incident Response compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Incident Response this skillkid-sid/claude-spellbook190—~3kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~992Automated safety check: PassApache-2.0
Activation Governance Chaos RolloutAli-Marandi/DataSense107—~1.9kAutomated safety check: PassMIT
Incident Response686f6c61/alfred-dev117—~1.1kAutomated safety check: PassMIT
Superset Incident Triagesuperset-sh/superset15k—~1kAutomated safety check: PassCustom licence
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT

Similar skills

  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~992 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Incident Response

    686f6c61/alfred-dev

    Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.

    117 GitHub stars~1.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed

More from kid-sid/claude-spellbook

All 55 skills in this repo
  • Accessibility

    kid-sid/claude-spellbook

    A skill your agent uses when building or reviewing UI components for keyboard and screen reader compatibility, adding ARIA to custom widgets, auditing a page for WCAG AA conformance, or preparing…

    190 GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Agentex

    kid-sid/claude-spellbook

    A skill your agent uses when building, wiring, or debugging an Agentex agent — choosing agent type, configuring acp.py and manifest.yaml, using adk.messages or adk.state, or resolving…

    190 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • AI Engineer

    kid-sid/claude-spellbook

    A skill your agent uses when building production LLM applications — designing RAG pipelines, choosing vector databases, implementing agent orchestration, optimizing cost, or adding AI safety…

    190 GitHub stars~3.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Angular

    kid-sid/claude-spellbook

    A skill your agent uses when building or refactoring Angular applications — choosing between signals, RxJS, and NgRx for state, configuring routing with guards and lazy loading, optimizing change…

    190 GitHub stars~5k tokensUpdated 2 mo ago
    Auto-check passed
  • API Design

    kid-sid/claude-spellbook

    A skill your agent uses when designing new REST endpoints, reviewing an existing API contract, adding pagination or filtering, planning a versioning strategy, or building a public or partner-facing…

    190 GitHub stars~3.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Auth

    kid-sid/claude-spellbook

    A skill your agent uses when implementing login flows, issuing or validating JWTs, setting up OAuth2/OIDC with a provider, designing role-based or attribute-based access control, securing API…

    190 GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Incident Response

What does Incident Response do?

A skill your agent uses when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths. Incident Response is an agent skill from kid-sid/claude-spellbook. Use when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.

When should I use Incident Response?

Incident Response fits situations like: triaging a production alert; writing a postmortem; updating a runbook; classifying incident severity.

How do I install Incident Response in Claude Code?

Run `npx skills add kid-sid/claude-spellbook --skill incident-response -a claude-code`. Or copy the skill folder (skills/incident-response in kid-sid/claude-spellbook) into .claude/skills/incident-response in your project. Claude Code loads it when a task matches its description.

How do I install Incident Response in Codex?

Run `npx skills add kid-sid/claude-spellbook --skill incident-response -a codex`. Or copy the skill folder (skills/incident-response in kid-sid/claude-spellbook) into .agents/skills/incident-response in your project. Codex loads it when a task matches its description.

Can I use Incident Response in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kid-sid/claude-spellbook --skill incident-response -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-response, .gemini/skills/incident-response, .github/skills/incident-response and .opencode/skills/incident-response in your project.

What does Incident Response need to run?

Going by SKILL.md and its folder, Incident Response needs the command-line tools its instructions call (kubectl).

Does Incident Response access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Incident Response safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Incident Response use?

Incident Response is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Incident Response use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Incident Response?

Skills that share tags, products or a category with Incident Response: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Incident Response?

kid-sid (a GitHub user) maintains it in kid-sid/claude-spellbook, which has 190 GitHub stars. The repository holds 55 skills in this directory. The repository was last updated on August 5, 2026.

Source: kid-sid/claude-spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.