Agent skill

Incident Runbook Templates

by wshobson in wshobson/agents

Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions.

MITAuto-check passedDevOps & Cloud

Install Incident Runbook Templates

skills CLI
$ npx skills add wshobson/agents --skill incident-runbook-templates -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents incident-runbook-templates --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/incident-response/skills/incident-runbook-templates .claude/skills/incident-runbook-templates && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
incident-runbook-templates
GitHub stars
40k
Token cost
~1.4k tokens
SKILL.md length
401 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions.

  • Works in 2 steps: Incident Severity Levels → Runbook Structure
  • Building a service outage runbook for a payment processing system
  • SKILL.md covers When to Use This Skill, Core Concepts, Detailed patterns and worked… and Best Practices, plus 2 more sections
  • Calls kubectl

What it does

Incident Runbook Templates is an agent skill from wshobson/agents. Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions. Use this skill when building a service outage runbook for a payment processing system; creating database incident procedures covering connection pool exhaustion, replication lag, and disk space alerts; onboarding new on-call engineers who need step-by-step recovery guides written for a 3 AM brain; or standardizing escalation matrices across multiple engineering teams.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/details.md`).

It sits in DevOps & Cloud, covering Runbooks and postmortems and Incident response. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Building a service outage runbook for a payment processing system
  • Creating database incident procedures covering connection pool exhaustion
  • Replication lag
  • Disk space alerts

Example prompts

  • “/incident-runbook-templates”

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Incident Severity Levels
  2. Runbook Structure

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Incident Runbook Templates loads about 1.4k tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 401 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 401 words, ~1,366 tokens.

Download SKILL.mdSave it as .claude/skills/incident-runbook-templates/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
incident-runbook-templates
description
Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions. Use this skill when building a service outage runbook for a payment processing system; creating database incident procedures covering connection pool exhaustion, replication lag, and disk space alerts; onboarding new on-call engineers who need step-by-step recovery guides written for a 3 AM brain; or standardizing escalation matrices across multiple engineering teams.

Incident Runbook Templates

Production-ready templates for incident response runbooks covering detection, triage, mitigation, resolution, and communication.

When to Use This Skill

  • Creating incident response procedures
  • Building service-specific runbooks
  • Establishing escalation paths
  • Documenting recovery procedures
  • Responding to active incidents
  • Onboarding on-call engineers

Core Concepts

1. Incident Severity Levels
SeverityImpactResponse TimeExample
SEV1Complete outage, data loss15 minProduction down
SEV2Major degradation30 minCritical feature broken
SEV3Minor impact2 hoursNon-critical bug
SEV4Minimal impactNext business dayCosmetic issue
2. Runbook Structure
1. Overview & Impact
2. Detection & Alerts
3. Initial Triage
4. Mitigation Steps
5. Root Cause Investigation
6. Resolution Procedures
7. Verification & Rollback
8. Communication Templates
9. Escalation Matrix

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Best Practices

Do's
  • Keep runbooks updated - Review after every incident
  • Test runbooks regularly - Game days, chaos engineering
  • Include rollback steps - Always have an escape hatch
  • Document assumptions - What must be true for steps to work
  • Link to dashboards - Quick access during stress
Don'ts
  • Don't assume knowledge - Write for 3 AM brain
  • Don't skip verification - Confirm each step worked
  • Don't forget communication - Keep stakeholders informed
  • Don't work alone - Escalate early
  • Don't skip postmortems - Learn from every incident

Troubleshooting

Runbook steps work in staging but fail during a real incident

Steps often assume preconditions that are true in a healthy environment but not during an outage. For each command in your runbook, add a prerequisite check and a "what to do if this command fails" note:

bash
# Step: Check pod status
kubectl get pods -n payments

# Prerequisites: kubectl configured, kubeconfig points to correct cluster
# If this fails: run `aws eks update-kubeconfig --name prod-cluster --region us-east-1`
# Expected output: pods in Running state
Show full SKILL.md (169 more words)Show less
On-call engineer panics and skips steps out of order

Add a numbered checklist at the top of the runbook that mirrors the section numbers, so responders can track progress under stress without reading the full document:

markdown
## Quick Checklist
- [ ] 1. Declare incident severity and open war room
- [ ] 2. Check service health (Section 4.1)
- [ ] 3. Check recent deployments (Section 4.1)
- [ ] 4. Roll back if deploy is suspect (Section 4.1)
- [ ] 5. Post initial notification to #payments-incidents
- [ ] 6. Escalate if > 15 min unresolved
Runbook is outdated — commands reference old cluster names or endpoints

Runbooks rot because they're updated manually. Include a "Last Verified" date and owner at the top, and add a CI check that validates all curl endpoints and kubectl context names are still valid:

markdown
## Runbook Metadata
| Field | Value |
|---|---|
| Last verified | 2024-11-15 |
| Owner | @platform-team |
| Review cadence | After every SEV1/SEV2 |
Stakeholder communication is delayed while engineers are heads-down

Assign a dedicated incident communicator role (separate from the incident commander) whose only job is to post status updates. Add a standing agenda in the communication template:

Update every 15 minutes (even if no new information):
- Current status (Investigating / Mitigating / Monitoring)
- Impact (what is broken, who is affected, % of traffic)
- What we are doing right now
- Next update in: 15 minutes
Database runbook commands cause additional downtime when run incorrectly

Add explicit warnings before destructive SQL commands and require a dry-run output check before executing:

sql
-- WARNING: This terminates active connections. Verify count first.
-- DRY RUN (check count before terminating):
SELECT count(*) FROM pg_stat_activity WHERE state = 'idle' AND query_start < now() - interval '10 minutes';

-- EXECUTE only after verifying count is reasonable (< 50):
SELECT pg_terminate_backend(pid) FROM pg_stat_activity
WHERE state = 'idle' AND query_start < now() - interval '10 minutes';
  • postmortem-writing - After resolving an incident, use postmortem templates to capture root cause and preventive actions
  • on-call-handoff-patterns - Structure shift handoffs so the incoming responder has full context on active incidents

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/incident-response/skills/incident-runbook-templates of wshobson/agents.

  • SKILL.md
  • references/details.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Incident Runbook Templates next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Incident Runbook Templates compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Incident Runbook Templates this skillwshobson/agents40k—~1.4kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0
Activation Governance Chaos RolloutAli-Marandi/DataSense107—~1.9kAutomated safety check: PassMIT
Incident Response686f6c61/alfred-dev117—~1.1kAutomated safety check: PassMIT
Superset Incident Triagesuperset-sh/superset15k—~1kAutomated safety check: PassCustom licence
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT

Similar skills

  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Incident Response

    686f6c61/alfred-dev

    Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.

    117 GitHub stars~1.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 6 days ago
    Auto-check passed

Categories

Questions about Incident Runbook Templates

What does Incident Runbook Templates do?

Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions. Incident Runbook Templates is an agent skill from wshobson/agents. Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions.

When should I use Incident Runbook Templates?

Incident Runbook Templates fits situations like: building a service outage runbook for a payment processing system; creating database incident procedures covering connection pool exhaustion; replication lag; disk space alerts.

How do I install Incident Runbook Templates in Claude Code?

Run `npx skills add wshobson/agents --skill incident-runbook-templates -a claude-code`. Or copy the skill folder (plugins/incident-response/skills/incident-runbook-templates in wshobson/agents) into .claude/skills/incident-runbook-templates in your project. Claude Code loads it when a task matches its description.

How do I install Incident Runbook Templates in Codex?

Run `npx skills add wshobson/agents --skill incident-runbook-templates -a codex`. Or copy the skill folder (plugins/incident-response/skills/incident-runbook-templates in wshobson/agents) into .agents/skills/incident-runbook-templates in your project. Codex loads it when a task matches its description.

Can I use Incident Runbook Templates in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill incident-runbook-templates -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-runbook-templates, .gemini/skills/incident-runbook-templates, .github/skills/incident-runbook-templates and .opencode/skills/incident-runbook-templates in your project.

What does Incident Runbook Templates need to run?

Going by SKILL.md and its folder, Incident Runbook Templates needs the command-line tools its instructions call (kubectl).

Does Incident Runbook Templates access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Incident Runbook Templates safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Incident Runbook Templates use?

Incident Runbook Templates is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Incident Runbook Templates use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Incident Runbook Templates?

Skills that share tags, products or a category with Incident Runbook Templates: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Incident Runbook Templates?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.