Agent skill

Oncall Runbook

by mohitagw15856 in mohitagw15856/pm-claude-skills

Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures.

MITAuto-check passedDevOps & Cloud

Install Oncall Runbook

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills oncall-runbook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/oncall-runbook .claude/skills/oncall-runbook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
oncall-runbook
GitHub stars
1.4k
Token cost
~3.4k tokens
SKILL.md length
1,387 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures.

  • Works in 4 steps: Write for the paged engineer, not the… → Turn postmortem learnings into per-alert… → Make escalation and handoff unambiguous.… → …
  • Asked to write an on-call guide
  • SKILL.md covers Where this sits — the spine's…, The loop, Required Inputs and Output Format, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Oncall Runbook is an agent skill from mohitagw15856/pm-claude-skills. Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures. Use when asked to write an on-call guide, create alert runbooks, document escalation procedures, or prepare an on-call handoff document. Produces a structured on-call runbook with per-alert response procedures, escalation matrix, diagnostic commands, and handoff template.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Runbooks and postmortems and Incident response. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to write an on-call guide
  • Create alert runbooks
  • Document escalation procedures
  • Prepare an on-call handoff document

Example prompts

  • “/oncall-runbook”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Write for the paged engineer, not the documentarian. The reader has just been
  2. Turn postmortem learnings into per-alert procedures. For each known failure (the
  3. Make escalation and handoff unambiguous. Who to page, when, and how to hand off
  4. Close the loop back to prevention. Flag where a runbook step reveals a gap that

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Oncall Runbook loads about 3.4k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 1,387 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 1,387 words, ~3,437 tokens.

Download SKILL.mdSave it as .claude/skills/oncall-runbook/SKILL.md (or your agent's skills folder).
name
oncall-runbook
description
Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures. Use when asked to write an on-call guide, create alert runbooks, document escalation procedures, or prepare an on-call handoff document. Produces a structured on-call runbook with per-alert response procedures, escalation matrix, diagnostic commands, and handoff template.

On-Call Runbook Skill

Produce a complete on-call runbook for a service — giving the on-call engineer everything they need to respond confidently to alerts at 3am, without having to ask anyone for help.

A good on-call runbook reduces mean time to resolution (MTTR) by eliminating the "what do I do first?" problem. It is written for the on-call engineer who has just been paged and needs to act, not for someone calmly reading documentation.

Where this sits — the spine's terminus

Last in the incident-response spine: /slo-error-budget (frame) → /debugging-log-analyser → /incident-postmortem → oncall-runbook. It receives the contributing factors and action items from /incident-postmortem and turns the detection/mitigation learnings into an entry that makes the next responder minutes, not hours — closing the loop so the same incident doesn't recur at full cost. Runbook entry, detection/mitigation time, and the loop are defined once in docs/craft/incident-response.md.

The loop

A runbook fails when it's written for a calm reader instead of a paged one at 3am. Phase 1 sets the audience; every later choice serves it.

  1. Write for the paged engineer, not the documentarian. The reader has just been woken and needs to act — so lead with the fastest safe mitigation, put copy-pasteable commands first, and defer background. Prose that explains before it acts fails at 3am. Done when: each alert's entry lets a non-expert take the first safe action within a minute of opening it, without reading theory.
  2. Turn postmortem learnings into per-alert procedures. For each known failure (the incident-postmortem's are the highest-value), write detect → mitigate → escalate: the exact checks, the copy-pasteable commands, the rollback, and when to page whom. Done when: every alert maps to a procedure with concrete commands and a clear mitigation, not just "investigate."
  3. Make escalation and handoff unambiguous. Who to page, when, and how to hand off mid-incident — because the second failure mode after "what do I do?" is "who do I wake, and when is it okay to?" Done when: the escalation matrix names people/rotations and the trigger for each, and the handoff template captures state so the next responder isn't starting cold.
  4. Close the loop back to prevention. Flag where a runbook step reveals a gap that should become monitoring or an /slo-error-budget action — the runbook is where the loop's learnings surface the next prevention. Done when: gaps found while writing the runbook are logged as detection/prevention improvements, not silently absorbed.

Required Inputs

Ask for these if not already provided:

  • Service name and what it does
  • Team and tech lead name
  • Alert list — names of alerts that currently page on-call
  • Monitoring setup — Datadog / Grafana / CloudWatch / PagerDuty / etc.
  • Common failure modes — what breaks most often, and what fixes it
  • Escalation contacts — who to call when on-call can't resolve it
  • Deployment setup — can on-call roll back? How?
  • Service dependencies — what does this service depend on, and what depends on it?

Output Format


On-Call Runbook: [Service Name]

Team: [Team name] | Tech lead: [Name] PagerDuty service: [Link] | Escalation policy: [Policy name] Last updated: [Date] | Next review: [Date + 90 days]

First time on-call for this service? Read the [developer onboarding doc] first — it covers the architecture and how things work. This runbook assumes you understand the service.


Quick Reference

Dashboard: [Link — the first thing to open when paged] Logs: [Link — where to find logs] Runbook index: Jump to the alert that paged you → [Alert list below] Can't resolve in 30 min? Escalate to: [Name] via [Slack / PagerDuty]

Rollback command (memorise this):

bash
[rollback command — e.g. kubectl rollout undo deployment/[service-name]]

Escalation Matrix

SituationEscalate toHowAfter how long
Can't diagnose the alert[Tech lead name]Slack DM / Phone30 minutes
Alert requires infra change[Platform team]#platform SlackImmediately
Customer-facing impact[CSM / Support lead]#incidents SlackImmediately (P1)
Database issue[DBA or data team]Slack / PagerDutyImmediately
[Specific dependency] down[[Dependency] on-call]PagerDuty / SlackImmediately
Extended outage (>1 hour)[Engineering manager]Phone1 hour

Contacts:

NameRoleSlackPhone
[Name]Tech lead@[handle][Number]
[Name]Engineering manager@[handle][Number]
[Name]Platform / infra@[handle][Number]
[Platform team]Infra on-call#platformPagerDuty

Service Architecture (Quick View)

[Upstream callers]
        │
        ▼
[This Service]
        │
        ├──→ [Primary Database]
        ├──→ [Cache — e.g. Redis]
        └──→ [Downstream Service / Queue]

If this service is down, these are affected: [List downstream consumers] If these are down, this service is affected: [List upstream dependencies]


Alert Runbooks

ALERT: [Alert Name 1 — e.g. HighErrorRate]

What it means: [Plain English — e.g. "More than 5% of API requests are returning 5xx errors in the last 5 minutes"] Severity: P1 / P2 / P3 SLO impact: Yes / No — [If yes: this alert means the error budget is burning at [X]× rate]

Step 1 — Acknowledge and assess

bash
# Check current error rate
[query or dashboard link]

# Check which endpoints are erroring
[query or command]

Step 2 — Check recent changes

bash
# Any deploys in the last hour?
[command or link to deployment log]

# Recent config changes?
[where to check]

Step 3 — Check dependencies

bash
# Is the database healthy?
[health check command or link]

# Is [downstream service] healthy?
[health check command or link]

Step 4 — Diagnose

If you seeIt meansDo this
[Error pattern 1][Cause][Action]
[Error pattern 2][Cause][Action]
[Error pattern 3][Cause][Action]
No clear patternUnknown causeEscalate to [name]

Step 5 — Fix or mitigate

bash
# If caused by bad deploy — roll back:
[rollback command]

# If caused by [specific issue]:
[fix command]

# If caused by upstream dependency:
[mitigation — e.g. enable circuit breaker, reduce traffic, etc.]

After resolving:

  • Confirm error rate has returned to baseline
  • Check no downstream services were affected
  • If P1: open a post-incident review — see [incident-postmortem skill]
  • Update #incidents with resolution summary

Show full SKILL.md (569 more words)Show less
ALERT: [Alert Name 2 — e.g. HighLatency]

What it means: [e.g. "P99 response time has exceeded 1s for more than 3 consecutive minutes"] Severity: P1 / P2 / P3 SLO impact: Yes — latency SLO breach

Step 1 — Assess scope

bash
# Check which endpoints are slow
[query or dashboard — broken down by endpoint]

# Check if latency is across all regions or localised
[query or command]

Step 2 — Common causes and fixes

CauseSignalFix
Database slow queriesDB latency spike on dashboard[Check slow query log: command]
Cache miss stormCache hit rate drops on dashboard[command or action]
Memory pressure / GCHigh memory on service dashboard[command or action — e.g. restart, scale up]
Upstream service slowTrace shows time in external callEscalate to [service] on-call
Traffic spikeRequest rate spike on dashboard[Scale up: command]

Step 3 — Escalate if unresolved in 20 minutes Page [Tech lead] via PagerDuty / Slack.


ALERT: [Alert Name 3 — e.g. DatabaseConnectionPoolExhausted]

What it means: [e.g. "The service has used all available database connections — new requests will fail"] Severity: P1 SLO impact: Yes — will cause errors immediately

Immediate mitigation:

bash
# Restart the service to flush stale connections
[restart command]

# Check current connection count
[DB connection query]

Diagnose root cause after stabilising:

bash
# Check for long-running queries holding connections
[query]

# Check if a recent deploy changed connection pool config
[where to check]

Resolution: [e.g. "Increase pool size in config / kill long-running queries / scale the service"]


ALERT: [Alert Name 4 — e.g. QueueBacklogHigh / ConsumerLag]

What it means: [e.g. "The message queue backlog exceeds 10,000 messages — consumers are not keeping up"] Severity: P2 SLO impact: Depends — if queue backs up, downstream systems will receive delayed data

Step 1 — Check consumer health

bash
# Are consumers running?
[command]

# Consumer error rate?
[dashboard or query]

Step 2 — Check message contents

bash
# Are there poison messages causing retries?
[command to inspect dead-letter queue or failed messages]

Step 3 — Options

IfThen
Consumers are downRestart consumers: [command]
Poison message in queueMove to DLQ: [command]
Consumers healthy but slowScale consumers: [command]
Upstream producing too fastEscalate to [upstream service] owner

ALERT: [Add additional alerts following the same pattern]

Diagnostic Cheat Sheet

Common commands for quick diagnosis. Paste and run without modification.

bash
# Service health
[health check command]

# Recent logs (last 100 lines)
[log command]

# Error logs only
[error log filter command]

# Current pod / instance status
[kubectl get pods / aws ecs describe-tasks / etc.]

# Restart the service
[restart command]

# Roll back to previous version
[rollback command]

# Database connection count
[DB query]

# Cache hit rate
[cache stats command]

# Current request rate
[metrics query]

DashboardURLUse it to
Service overview[Link]First stop — error rate, latency, request rate
Database[Link]Connection count, slow queries, replication lag
Infrastructure[Link]CPU, memory, disk
Queue / consumers[Link]Backlog depth, consumer throughput
Upstream dependencies[Link]Dependency health at a glance

Incident Communication

When you declare an incident:

Post to #incidents immediately:

🔴 INCIDENT — [Service Name]
Status: Investigating
Impact: [Who is affected and how]
Paged: [Your name]
Next update: [Time — max 30 min from now]

Update every 30 minutes while active:

🔴 UPDATE — [Service Name] — [Time]
Status: [Investigating / Identified / Mitigating / Resolved]
Latest: [One sentence on what you found or did]
Next update: [Time]

On resolution:

✅ RESOLVED — [Service Name] — [Time]
Duration: [X minutes]
Impact: [Summary of who was affected]
Cause: [One sentence]
Follow-up: [PIR required? Yes/No — link when created]

On-Call Handoff

Use this template at the end of every on-call shift:

--- ON-CALL HANDOFF: [Service Name] ---
Date: [Date]
Outgoing: [Your name]
Incoming: [Next on-call name]

INCIDENTS THIS SHIFT:
- [Incident summary — date, duration, cause, resolution, follow-up required]

OPEN ISSUES TO WATCH:
- [Anything not fully resolved / trending in the wrong direction]

CHANGES SINCE LAST HANDOFF:
- [Deploys, config changes, infra changes that affect on-call awareness]

RUNBOOK GAPS FOUND:
- [Anything you had to figure out that isn't documented — please add it]

ANYTHING ELSE:
- [Notes for incoming on-call]

Quality Checks

  • Every alert that pages on-call has a runbook entry — no alert is missing
  • Rollback command is accurate and tested recently
  • Escalation contacts have current phone numbers and Slack handles
  • Diagnostic commands work — they have been run by at least one person recently
  • Handoff template is used at every shift change — not just during incidents
  • "Things I had to figure out that weren't documented" are added to this runbook after every incident

Anti-Patterns

  • Do not write alert runbooks with vague diagnostic steps like "check the logs" — every step must specify the exact command, dashboard link, or query to run
  • Do not include an alert in the runbook that has no specific on-call action — an alert that pages someone with no defined response path creates panic, not resolution
  • Do not leave the rollback command undocumented or untested — a rollback procedure that has never been run will fail when needed most
  • Do not list escalation contacts without phone numbers and Slack handles — email-only escalation paths are useless during a 3am incident
  • Do not write the runbook once and treat it as permanent — runbooks go stale after incidents; every incident must trigger a review of the relevant runbook entries

Example Trigger Phrases

  • "Write an on-call guide."
  • "Create alert runbooks."
  • "Document escalation procedures."
  • "Prepare an on-call handoff document."

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/oncall-runbook of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Oncall Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Oncall Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Oncall Runbook this skillmohitagw15856/pm-claude-skills1.4k—~3.4kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0
Activation Governance Chaos RolloutAli-Marandi/DataSense107—~1.9kAutomated safety check: PassMIT
Incident Response686f6c61/alfred-dev117—~1.1kAutomated safety check: PassMIT
Superset Incident Triagesuperset-sh/superset15k—~1kAutomated safety check: PassCustom licence
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT

Similar skills

  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Incident Response

    686f6c61/alfred-dev

    Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.

    117 GitHub stars~1.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Oncall Runbook

What does Oncall Runbook do?

Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures. Oncall Runbook is an agent skill from mohitagw15856/pm-claude-skills. Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures.

When should I use Oncall Runbook?

Oncall Runbook fits situations like: asked to write an on-call guide; create alert runbooks; document escalation procedures; prepare an on-call handoff document.

How do I install Oncall Runbook in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a claude-code`. Or copy the skill folder (skills/oncall-runbook in mohitagw15856/pm-claude-skills) into .claude/skills/oncall-runbook in your project. Claude Code loads it when a task matches its description.

How do I install Oncall Runbook in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a codex`. Or copy the skill folder (skills/oncall-runbook in mohitagw15856/pm-claude-skills) into .agents/skills/oncall-runbook in your project. Codex loads it when a task matches its description.

Can I use Oncall Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/oncall-runbook, .gemini/skills/oncall-runbook, .github/skills/oncall-runbook and .opencode/skills/oncall-runbook in your project.

What does Oncall Runbook need to run?

SKILL.md names no scripts, command-line tools or credentials: Oncall Runbook is instructions for the agent only.

Does Oncall Runbook access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Oncall Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Oncall Runbook use?

Oncall Runbook is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Oncall Runbook use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Oncall Runbook?

Skills that share tags, products or a category with Oncall Runbook: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Oncall Runbook?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.