Oncall
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
A skill your agent uses when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.
$ npx skills add kid-sid/claude-spellbook --skill incident-response -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install kid-sid/claude-spellbook incident-response --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/incident-response .claude/skills/incident-response && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "incident-response" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/incident-response into .claude/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/kid-sid/claude-spellbook/tree/main/skills/incident-responseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add kid-sid/claude-spellbook --skill incident-response -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install kid-sid/claude-spellbook incident-response --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/incident-response .agents/skills/incident-response && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "incident-response" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/incident-response into .agents/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kid-sid/claude-spellbook --skill incident-response -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install kid-sid/claude-spellbook incident-response --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/incident-response .cursor/skills/incident-response && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "incident-response" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/incident-response into .cursor/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/kid-sid/claude-spellbook.git --path skills/incident-response--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add kid-sid/claude-spellbook --skill incident-response -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install kid-sid/claude-spellbook incident-response --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/incident-response .gemini/skills/incident-response && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/incident-response into .gemini/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install kid-sid/claude-spellbook incident-responseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add kid-sid/claude-spellbook --skill incident-response -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/incident-response .github/skills/incident-response && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/incident-response into .github/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kid-sid/claude-spellbook --skill incident-response -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install kid-sid/claude-spellbook incident-response --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/incident-response .opencode/skills/incident-response && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/incident-response into .opencode/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
incident-responseA skill your agent uses when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.
Incident Response is an agent skill from kid-sid/claude-spellbook. Use when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Incident response and Runbooks and postmortems. The repository describes itself as: A curated collection of skills, prompts, and workflows that extend Claude's capabilities — your personal grimoire for AI-powered development. The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit a7c2ac9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Incident Response loads about 3k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 866 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from kid-sid/claude-spellbook at commit a7c2ac9, republished under its MIT licence (© kid-sid). 866 words, ~2,953 tokens.
.claude/skills/incident-response/SKILL.md (or your agent's skills folder).Incident response is the structured process of detecting, mitigating, communicating, and learning from production failures to minimise user impact and prevent recurrence.
| Severity | Definition | Response SLA | Comms cadence | Example |
|---|---|---|---|---|
| P0 | Total outage or data loss — all users affected | Page immediately, < 5 min | Every 15 min | Payment service down, DB unreachable |
| P1 | Major feature broken — most users affected | < 15 min acknowledgement | Every 30 min | Login failing for 50%+ of users |
| P2 | Significant degradation — subset of users affected | < 1 hour | Every 2 hours | Search slow for US region |
| P3 | Minor issue — small impact, workaround available | Next business day | Once resolved | Non-critical dashboard shows stale data |
| P4 | Cosmetic / no user impact | Sprint backlog | N/A | Log noise, minor UI misalignment |
Escalation path:
Detection → Triage → Mitigate → Communicate → Resolve → Review (Postmortem)#inc-YYYY-MM-DD-short-description🔴 [P0/P1 INCIDENT] Payment service degradation
Status: Investigating
Impact: ~30% of payment requests failing with 500 errors since 14:23 UTC
Affected: All users attempting checkout
IC: @alice
SME: @bob
Next update: 14:45 UTC
Tracking: https://incident.example.com/inc-2024-0042🟡 [P1 UPDATE] Payment service — 14:45 UTC
Status: Mitigating
Root cause identified: Connection pool exhaustion after deploy at 14:15
Action taken: Rolled back to v2.3.1, monitoring error rate
Current error rate: 2% (down from 30%)
Next update: 15:00 UTC✅ [P1 RESOLVED] Payment service — 15:02 UTC
Status: Resolved
Duration: 39 minutes (14:23 – 15:02 UTC)
Root cause: Deploy v2.4.0 introduced a connection leak; pool exhausted under load
Resolution: Rolled back to v2.3.1; error rate returned to baseline at 15:00
Users impacted: ~15,000 failed checkout attempts
Follow-up: Postmortem scheduled for 2024-01-16 15:00 UTC
Incident report: https://incident.example.com/inc-2024-0042Error rate > SLO threshold?
├── Yes
│ ├── Was something deployed in the last 2 hours?
│ │ ├── Yes → ROLLBACK first, investigate after
│ │ └── No → Check: DB, cache, upstream dependency, config change
│ ├── Can we isolate the impact with a feature flag kill?
│ │ └── Yes → Kill the flag immediately
│ └── Is this a traffic spike?
│ └── Yes → Scale up horizontally, enable circuit breaker
└── No — latency degraded only?
├── Check DB: slow queries, lock contention, pool saturation
├── Check cache hit rate: has cache been evicted?
└── Check upstream service latencyWhen NOT to roll back immediately:
Runbooks must be written for the 3am engineer who has never seen this service.
# Runbook: [Service Name] — [Alert Name]
## Service Overview
[2–3 sentences: what does this service do, what does it depend on?]
## Alert: [Alert Name]
**Trigger condition:** [e.g., error rate > 1% for 5 minutes]
**Severity:** P1
**Dashboard:** [link]
**Logs:** [link to log query]
## Diagnostic Steps
1. Check the error rate panel on the [service dashboard](link)
- Expected: < 0.1%
- If > 1%: proceed to step 2
2. Check recent deployments:
```bash
kubectl rollout history deployment/payment-service -n productionkubectl exec -it $(kubectl get pod -l app=payment-service -o name | head -1) \
-- curl -s localhost:8080/metrics | grep db_pooldb_pool_wait_duration_seconds > 1s: pool is exhausted, proceed to step 4SELECT query, mean_exec_time, calls
FROM pg_stat_statements
ORDER BY mean_exec_time DESC
LIMIT 10;kubectl rollout undo deployment/payment-service -n productionkubectl scale deployment/payment-service --replicas=6[link to flag]
**Runbook quality checks:**
- Every step has an expected output — the engineer knows what "normal" looks like
- Commands are copy-paste ready (no placeholders that need substitution)
- Decision points have explicit branches ("if X, do Y; if Z, do W")
- Links to dashboards, log queries, and escalation contacts are current
## Blameless Postmortem
Write the postmortem within 48 hours while details are fresh. **Blameless = focus on systems and processes, not individuals.**
```markdown
# Postmortem: [Service] [Brief Description] — [Date]
## Summary
[2–3 sentences: what happened, impact, how it was resolved]
**Impact:** [number of users affected, % error rate, duration]
**Detection time:** [how long from start to detection]
**Resolution time:** [how long from detection to resolution]
## Timeline (UTC)
| Time | Event |
|-------|-------|
| 14:15 | Deploy v2.4.0 rolled out to 100% |
| 14:23 | Alert fired: error rate > 1% |
| 14:28 | On-call acknowledged, started investigation |
| 14:38 | Root cause identified: connection pool exhausted |
| 14:45 | Rollback initiated |
| 15:00 | Error rate returned to baseline |
| 15:02 | Incident declared resolved |
## Root Cause Analysis (5 Whys)
1. **Why** did payment requests fail?
→ DB connection pool was exhausted
2. **Why** was the pool exhausted?
→ v2.4.0 introduced a connection leak in the retry handler
3. **Why** did the retry handler leak connections?
→ The `defer conn.Close()` was placed inside the retry loop, closing on each attempt but not releasing the acquired connection back to the pool
4. **Why** wasn't this caught in testing?
→ Integration tests used a single-connection test DB; pool exhaustion only manifests at scale
5. **Why** wasn't this caught by the integration test DB pool?
→ Test pool size was set to 100 (no practical limit); prod pool size is 20
## Contributing Factors
- No load test run before this deploy
- No DB pool exhaustion alert existed
- Code review missed the subtle connection lifecycle issue
## What Went Well
- Alert fired within 8 minutes of degradation starting
- On-call was paged and acknowledged quickly
- Rollback decision was made in < 10 minutes
## Action Items
| Action | Owner | Due | Category |
|--------|-------|-----|----------|
| Add DB pool wait time alert (threshold: > 1s for 5 min) | @alice | 2024-01-19 | Detection |
| Add integration test that simulates pool exhaustion under concurrent load | @bob | 2024-01-26 | Prevention |
| Add `db_pool_size` check to pre-deploy checklist | @alice | 2024-01-19 | Prevention |
| Run k6 load test before all deploys touching DB connection code | @bob | 2024-01-26 | Prevention || Metric | Definition | Target |
|---|---|---|
| MTTD | Mean Time To Detect — start of incident to first alert firing | < 5 min |
| MTTA | Mean Time To Acknowledge — alert fires to on-call acks | < 5 min |
| MTTR | Mean Time To Resolve — detection to resolution | < 30 min for P0/P1 |
| Incident frequency | Number of P0/P1 incidents per month per service | Track trend; goal: decreasing |
| Repeat incidents | Incidents with the same root cause as a prior incident | Goal: 0 |
Review these monthly per service. Rising MTTR = runbooks need updating. Repeat incidents = action items not implemented.
See also:
observability,deployment-strategies
© kid-sid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/incident-response of kid-sid/claude-spellbook.
Open the folder on GitHubat commit a7c2ac9
Incident Response next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Incident Response this skillkid-sid/claude-spellbook | 190 | — | ~3k | Automated safety check: Pass | MIT | |
| Oncallpigweed-project/pigweed | 548 | — | ~992 | Automated safety check: Pass | Apache-2.0 | |
| Activation Governance Chaos RolloutAli-Marandi/DataSense | 107 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Incident Response686f6c61/alfred-dev | 117 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Superset Incident Triagesuperset-sh/superset | 15k | — | ~1k | Automated safety check: Pass | Custom licence | |
| Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan | 109 | — | ~1.9k | Automated safety check: Pass | MIT |
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
Ali-Marandi/DataSense
Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.
686f6c61/alfred-dev
Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.
superset-sh/superset
Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.
VeryGoodOpenSource/vgv-wingspan
Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.
Jeffallan/claude-skills
Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.
kid-sid/claude-spellbook
A skill your agent uses when building or reviewing UI components for keyboard and screen reader compatibility, adding ARIA to custom widgets, auditing a page for WCAG AA conformance, or preparing…
kid-sid/claude-spellbook
A skill your agent uses when building, wiring, or debugging an Agentex agent — choosing agent type, configuring acp.py and manifest.yaml, using adk.messages or adk.state, or resolving…
kid-sid/claude-spellbook
A skill your agent uses when building production LLM applications — designing RAG pipelines, choosing vector databases, implementing agent orchestration, optimizing cost, or adding AI safety…
kid-sid/claude-spellbook
A skill your agent uses when building or refactoring Angular applications — choosing between signals, RxJS, and NgRx for state, configuring routing with guards and lazy loading, optimizing change…
kid-sid/claude-spellbook
A skill your agent uses when designing new REST endpoints, reviewing an existing API contract, adding pagination or filtering, planning a versioning strategy, or building a public or partner-facing…
kid-sid/claude-spellbook
A skill your agent uses when implementing login flows, issuing or validating JWTs, setting up OAuth2/OIDC with a provider, designing role-based or attribute-based access control, securing API…
Categories
A skill your agent uses when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths. Incident Response is an agent skill from kid-sid/claude-spellbook. Use when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.
Incident Response fits situations like: triaging a production alert; writing a postmortem; updating a runbook; classifying incident severity.
Run `npx skills add kid-sid/claude-spellbook --skill incident-response -a claude-code`. Or copy the skill folder (skills/incident-response in kid-sid/claude-spellbook) into .claude/skills/incident-response in your project. Claude Code loads it when a task matches its description.
Run `npx skills add kid-sid/claude-spellbook --skill incident-response -a codex`. Or copy the skill folder (skills/incident-response in kid-sid/claude-spellbook) into .agents/skills/incident-response in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kid-sid/claude-spellbook --skill incident-response -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-response, .gemini/skills/incident-response, .github/skills/incident-response and .opencode/skills/incident-response in your project.
Going by SKILL.md and its folder, Incident Response needs the command-line tools its instructions call (kubectl).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Incident Response is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Incident Response: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
kid-sid (a GitHub user) maintains it in kid-sid/claude-spellbook, which has 190 GitHub stars. The repository holds 55 skills in this directory. The repository was last updated on August 5, 2026.
Source: kid-sid/claude-spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.