Oncall
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures.
$ npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mohitagw15856/pm-claude-skills oncall-runbook --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/oncall-runbook .claude/skills/oncall-runbook && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "oncall-runbook" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/oncall-runbook into .claude/skills/oncall-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-runbook", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/oncall-runbookType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mohitagw15856/pm-claude-skills oncall-runbook --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/oncall-runbook .agents/skills/oncall-runbook && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "oncall-runbook" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/oncall-runbook into .agents/skills/oncall-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-runbook", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mohitagw15856/pm-claude-skills oncall-runbook --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/oncall-runbook .cursor/skills/oncall-runbook && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "oncall-runbook" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/oncall-runbook into .cursor/skills/oncall-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-runbook", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mohitagw15856/pm-claude-skills.git --path skills/oncall-runbook--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mohitagw15856/pm-claude-skills oncall-runbook --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/oncall-runbook .gemini/skills/oncall-runbook && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "oncall-runbook" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/oncall-runbook into .gemini/skills/oncall-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-runbook", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mohitagw15856/pm-claude-skills oncall-runbookInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/oncall-runbook .github/skills/oncall-runbook && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "oncall-runbook" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/oncall-runbook into .github/skills/oncall-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-runbook", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mohitagw15856/pm-claude-skills oncall-runbook --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/oncall-runbook .opencode/skills/oncall-runbook && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "oncall-runbook" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/oncall-runbook into .opencode/skills/oncall-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-runbook", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
oncall-runbookWrite an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures.
Oncall Runbook is an agent skill from mohitagw15856/pm-claude-skills. Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures. Use when asked to write an on-call guide, create alert runbooks, document escalation procedures, or prepare an on-call handoff document. Produces a structured on-call runbook with per-alert response procedures, escalation matrix, diagnostic commands, and handoff template.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Runbooks and postmortems and Incident response. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Oncall Runbook loads about 3.4k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 1,387 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 1,387 words, ~3,437 tokens.
.claude/skills/oncall-runbook/SKILL.md (or your agent's skills folder).Produce a complete on-call runbook for a service — giving the on-call engineer everything they need to respond confidently to alerts at 3am, without having to ask anyone for help.
A good on-call runbook reduces mean time to resolution (MTTR) by eliminating the "what do I do first?" problem. It is written for the on-call engineer who has just been paged and needs to act, not for someone calmly reading documentation.
Last in the incident-response spine: /slo-error-budget (frame) →
/debugging-log-analyser → /incident-postmortem → oncall-runbook. It receives the
contributing factors and action items from /incident-postmortem and turns the
detection/mitigation learnings into an entry that makes the next responder minutes, not
hours — closing the loop so the same incident doesn't recur at full cost. Runbook entry,
detection/mitigation time, and the loop are defined once in
docs/craft/incident-response.md.
A runbook fails when it's written for a calm reader instead of a paged one at 3am. Phase 1 sets the audience; every later choice serves it.
/slo-error-budget action — the runbook is where the
loop's learnings surface the next prevention.
Done when: gaps found while writing the runbook are logged as detection/prevention
improvements, not silently absorbed.Ask for these if not already provided:
Team: [Team name] | Tech lead: [Name] PagerDuty service: [Link] | Escalation policy: [Policy name] Last updated: [Date] | Next review: [Date + 90 days]
First time on-call for this service? Read the [developer onboarding doc] first — it covers the architecture and how things work. This runbook assumes you understand the service.
Dashboard: [Link — the first thing to open when paged] Logs: [Link — where to find logs] Runbook index: Jump to the alert that paged you → [Alert list below] Can't resolve in 30 min? Escalate to: [Name] via [Slack / PagerDuty]
Rollback command (memorise this):
[rollback command — e.g. kubectl rollout undo deployment/[service-name]]| Situation | Escalate to | How | After how long |
|---|---|---|---|
| Can't diagnose the alert | [Tech lead name] | Slack DM / Phone | 30 minutes |
| Alert requires infra change | [Platform team] | #platform Slack | Immediately |
| Customer-facing impact | [CSM / Support lead] | #incidents Slack | Immediately (P1) |
| Database issue | [DBA or data team] | Slack / PagerDuty | Immediately |
| [Specific dependency] down | [[Dependency] on-call] | PagerDuty / Slack | Immediately |
| Extended outage (>1 hour) | [Engineering manager] | Phone | 1 hour |
Contacts:
| Name | Role | Slack | Phone |
|---|---|---|---|
| [Name] | Tech lead | @[handle] | [Number] |
| [Name] | Engineering manager | @[handle] | [Number] |
| [Name] | Platform / infra | @[handle] | [Number] |
| [Platform team] | Infra on-call | #platform | PagerDuty |
[Upstream callers]
│
▼
[This Service]
│
├──→ [Primary Database]
├──→ [Cache — e.g. Redis]
└──→ [Downstream Service / Queue]If this service is down, these are affected: [List downstream consumers] If these are down, this service is affected: [List upstream dependencies]
What it means: [Plain English — e.g. "More than 5% of API requests are returning 5xx errors in the last 5 minutes"] Severity: P1 / P2 / P3 SLO impact: Yes / No — [If yes: this alert means the error budget is burning at [X]× rate]
Step 1 — Acknowledge and assess
# Check current error rate
[query or dashboard link]
# Check which endpoints are erroring
[query or command]Step 2 — Check recent changes
# Any deploys in the last hour?
[command or link to deployment log]
# Recent config changes?
[where to check]Step 3 — Check dependencies
# Is the database healthy?
[health check command or link]
# Is [downstream service] healthy?
[health check command or link]Step 4 — Diagnose
| If you see | It means | Do this |
|---|---|---|
| [Error pattern 1] | [Cause] | [Action] |
| [Error pattern 2] | [Cause] | [Action] |
| [Error pattern 3] | [Cause] | [Action] |
| No clear pattern | Unknown cause | Escalate to [name] |
Step 5 — Fix or mitigate
# If caused by bad deploy — roll back:
[rollback command]
# If caused by [specific issue]:
[fix command]
# If caused by upstream dependency:
[mitigation — e.g. enable circuit breaker, reduce traffic, etc.]After resolving:
#incidents with resolution summaryWhat it means: [e.g. "P99 response time has exceeded 1s for more than 3 consecutive minutes"] Severity: P1 / P2 / P3 SLO impact: Yes — latency SLO breach
Step 1 — Assess scope
# Check which endpoints are slow
[query or dashboard — broken down by endpoint]
# Check if latency is across all regions or localised
[query or command]Step 2 — Common causes and fixes
| Cause | Signal | Fix |
|---|---|---|
| Database slow queries | DB latency spike on dashboard | [Check slow query log: command] |
| Cache miss storm | Cache hit rate drops on dashboard | [command or action] |
| Memory pressure / GC | High memory on service dashboard | [command or action — e.g. restart, scale up] |
| Upstream service slow | Trace shows time in external call | Escalate to [service] on-call |
| Traffic spike | Request rate spike on dashboard | [Scale up: command] |
Step 3 — Escalate if unresolved in 20 minutes Page [Tech lead] via PagerDuty / Slack.
What it means: [e.g. "The service has used all available database connections — new requests will fail"] Severity: P1 SLO impact: Yes — will cause errors immediately
Immediate mitigation:
# Restart the service to flush stale connections
[restart command]
# Check current connection count
[DB connection query]Diagnose root cause after stabilising:
# Check for long-running queries holding connections
[query]
# Check if a recent deploy changed connection pool config
[where to check]Resolution: [e.g. "Increase pool size in config / kill long-running queries / scale the service"]
What it means: [e.g. "The message queue backlog exceeds 10,000 messages — consumers are not keeping up"] Severity: P2 SLO impact: Depends — if queue backs up, downstream systems will receive delayed data
Step 1 — Check consumer health
# Are consumers running?
[command]
# Consumer error rate?
[dashboard or query]Step 2 — Check message contents
# Are there poison messages causing retries?
[command to inspect dead-letter queue or failed messages]Step 3 — Options
| If | Then |
|---|---|
| Consumers are down | Restart consumers: [command] |
| Poison message in queue | Move to DLQ: [command] |
| Consumers healthy but slow | Scale consumers: [command] |
| Upstream producing too fast | Escalate to [upstream service] owner |
Common commands for quick diagnosis. Paste and run without modification.
# Service health
[health check command]
# Recent logs (last 100 lines)
[log command]
# Error logs only
[error log filter command]
# Current pod / instance status
[kubectl get pods / aws ecs describe-tasks / etc.]
# Restart the service
[restart command]
# Roll back to previous version
[rollback command]
# Database connection count
[DB query]
# Cache hit rate
[cache stats command]
# Current request rate
[metrics query]| Dashboard | URL | Use it to |
|---|---|---|
| Service overview | [Link] | First stop — error rate, latency, request rate |
| Database | [Link] | Connection count, slow queries, replication lag |
| Infrastructure | [Link] | CPU, memory, disk |
| Queue / consumers | [Link] | Backlog depth, consumer throughput |
| Upstream dependencies | [Link] | Dependency health at a glance |
When you declare an incident:
Post to #incidents immediately:
🔴 INCIDENT — [Service Name]
Status: Investigating
Impact: [Who is affected and how]
Paged: [Your name]
Next update: [Time — max 30 min from now]Update every 30 minutes while active:
🔴 UPDATE — [Service Name] — [Time]
Status: [Investigating / Identified / Mitigating / Resolved]
Latest: [One sentence on what you found or did]
Next update: [Time]On resolution:
✅ RESOLVED — [Service Name] — [Time]
Duration: [X minutes]
Impact: [Summary of who was affected]
Cause: [One sentence]
Follow-up: [PIR required? Yes/No — link when created]Use this template at the end of every on-call shift:
--- ON-CALL HANDOFF: [Service Name] ---
Date: [Date]
Outgoing: [Your name]
Incoming: [Next on-call name]
INCIDENTS THIS SHIFT:
- [Incident summary — date, duration, cause, resolution, follow-up required]
OPEN ISSUES TO WATCH:
- [Anything not fully resolved / trending in the wrong direction]
CHANGES SINCE LAST HANDOFF:
- [Deploys, config changes, infra changes that affect on-call awareness]
RUNBOOK GAPS FOUND:
- [Anything you had to figure out that isn't documented — please add it]
ANYTHING ELSE:
- [Notes for incoming on-call]© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/oncall-runbook of mohitagw15856/pm-claude-skills.
Open the folder on GitHubat commit 1cbf1f0
Oncall Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Oncall Runbook this skillmohitagw15856/pm-claude-skills | 1.4k | — | ~3.4k | Automated safety check: Pass | MIT | |
| Oncallpigweed-project/pigweed | 548 | — | ~963 | Automated safety check: Pass | Apache-2.0 | |
| Activation Governance Chaos RolloutAli-Marandi/DataSense | 107 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Incident Response686f6c61/alfred-dev | 117 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Superset Incident Triagesuperset-sh/superset | 15k | — | ~1k | Automated safety check: Pass | Custom licence | |
| Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan | 109 | — | ~1.9k | Automated safety check: Pass | MIT |
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
Ali-Marandi/DataSense
Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.
686f6c61/alfred-dev
Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.
superset-sh/superset
Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.
VeryGoodOpenSource/vgv-wingspan
Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.
Jeffallan/claude-skills
Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.
mohitagw15856/pm-claude-skills
Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…
mohitagw15856/pm-claude-skills
Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.
mohitagw15856/pm-claude-skills
Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.
mohitagw15856/pm-claude-skills
Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.
mohitagw15856/pm-claude-skills
Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.
mohitagw15856/pm-claude-skills
Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.
Categories
Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures. Oncall Runbook is an agent skill from mohitagw15856/pm-claude-skills. Write an on-call runbook for a service — covering alert definitions, escalation paths, common incident responses, and on-call handoff procedures.
Oncall Runbook fits situations like: asked to write an on-call guide; create alert runbooks; document escalation procedures; prepare an on-call handoff document.
Run `npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a claude-code`. Or copy the skill folder (skills/oncall-runbook in mohitagw15856/pm-claude-skills) into .claude/skills/oncall-runbook in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a codex`. Or copy the skill folder (skills/oncall-runbook in mohitagw15856/pm-claude-skills) into .agents/skills/oncall-runbook in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill oncall-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/oncall-runbook, .gemini/skills/oncall-runbook, .github/skills/oncall-runbook and .opencode/skills/oncall-runbook in your project.
SKILL.md names no scripts, command-line tools or credentials: Oncall Runbook is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Oncall Runbook is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Oncall Runbook: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.
Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.