Oncall
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
Guide incident response from detection to post-mortem using SRE principles, severity classification, on-call management, blameless culture, and communication protocols.
$ npx skills add ancoleman/ai-design-components --skill managing-incidents -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ancoleman/ai-design-components managing-incidents --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/managing-incidents .claude/skills/managing-incidents && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "managing-incidents" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/managing-incidents into .claude/skills/managing-incidents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "managing-incidents", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ancoleman/ai-design-components/tree/main/skills/managing-incidentsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ancoleman/ai-design-components --skill managing-incidents -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ancoleman/ai-design-components managing-incidents --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/managing-incidents .agents/skills/managing-incidents && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "managing-incidents" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/managing-incidents into .agents/skills/managing-incidents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "managing-incidents", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill managing-incidents -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ancoleman/ai-design-components managing-incidents --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/managing-incidents .cursor/skills/managing-incidents && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "managing-incidents" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/managing-incidents into .cursor/skills/managing-incidents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "managing-incidents", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ancoleman/ai-design-components.git --path skills/managing-incidents--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ancoleman/ai-design-components --skill managing-incidents -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ancoleman/ai-design-components managing-incidents --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/managing-incidents .gemini/skills/managing-incidents && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "managing-incidents" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/managing-incidents into .gemini/skills/managing-incidents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "managing-incidents", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ancoleman/ai-design-components managing-incidentsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ancoleman/ai-design-components --skill managing-incidents -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/managing-incidents .github/skills/managing-incidents && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "managing-incidents" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/managing-incidents into .github/skills/managing-incidents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "managing-incidents", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill managing-incidents -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ancoleman/ai-design-components managing-incidents --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/managing-incidents .opencode/skills/managing-incidents && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "managing-incidents" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/managing-incidents into .opencode/skills/managing-incidents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "managing-incidents", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
managing-incidentsGuide incident response from detection to post-mortem using SRE principles, severity classification, on-call management, blameless culture, and communication protocols.
Managing Incidents is an agent skill from ancoleman/ai-design-components. Guide incident response from detection to post-mortem using SRE principles, severity classification, on-call management, blameless culture, and communication protocols. Use when setting up incident processes, designing escalation policies, or conducting post-mortems.
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 22 other files, including scripts and reference files (for example `examples/communication-templates.md`, `examples/integrations/pagerduty-slack.py` and `examples/integrations/postmortem-generator.py`).
It sits in DevOps & Cloud, covering Runbooks and postmortems and Incident response. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Managing Incidents loads about 3.6k tokens when it runs, and up to ~24k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 1,496 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 1,496 words, ~3,574 tokens.
.claude/skills/managing-incidents/SKILL.md (or your agent's skills folder). This skill also uses 18 other files; get the full folder from GitHub.Provide end-to-end incident management guidance covering detection, response, communication, and learning. Emphasizes SRE culture, blameless post-mortems, and structured processes for high-reliability operations.
Apply this skill when:
Declare Early and Often: Do not wait for certainty. Declaring an incident enables coordination, can be downgraded if needed, and prevents delayed response.
Mitigation First, Root Cause Later: Stop customer impact immediately (rollback, disable feature, failover). Debug and fix root cause after stability restored.
Blameless Culture: Assume good intentions. Focus on how systems failed, not who failed. Create psychological safety for honest learning.
Clear Command Structure: Assign Incident Commander (IC) to own coordination. IC delegates tasks but does not do hands-on debugging.
Communication is Critical: Internal coordination via dedicated channels, external transparency via status pages. Update stakeholders every 15-30 minutes during critical incidents.
Standard severity levels with response times:
SEV0 (P0) - Critical Outage:
SEV1 (P1) - Major Degradation:
SEV2 (P2) - Minor Issues:
SEV3 (P3) - Low Impact:
For detailed severity decision framework and interactive classifier, see references/severity-classification.md.
Incident Commander (IC):
Communications Lead:
Subject Matter Experts (SMEs):
Scribe:
Assign roles based on severity:
For detailed role responsibilities, see references/incident-roles.md.
Primary + Secondary:
Follow-the-Sun (24/7):
Tiered Escalation:
Standard incident lifecycle:
Detection → Triage → Declaration → Investigation
↓
Mitigation → Resolution → Monitoring → Closure
↓
Post-Mortem (within 48 hours)When to Declare: When in doubt, declare (can always downgrade severity)
When to Escalate:
When to Close:
For complete workflow details, see references/incident-workflow.md.
Incident Slack Channel:
#incident-YYYY-MM-DD-topic-descriptionWar Room: Video call for SEV0/SEV1 requiring real-time voice coordination
Status Update Cadence:
Status Page:
Customer Email:
Regulatory Notifications:
For communication templates, see examples/communication-templates.md.
Every runbook should include:
For runbook templates, see examples/runbooks/ directory.
Assume Good Intentions: Everyone made the best decision with information available.
Focus on Systems: Investigate how processes failed, not who failed.
Psychological Safety: Create environment where honesty is rewarded.
Learning Opportunity: Incidents are gifts of organizational knowledge.
1. Schedule Review (Within 48 Hours): While memory is fresh
2. Pre-Work: Reconstruct timeline, gather metrics/logs, draft document
3. Meeting Facilitation:
4. Post-Mortem Document:
5. Follow-Up: Track action items in sprint planning
For detailed facilitation guide and template, see references/blameless-postmortems.md and examples/postmortem-template.md.
Actionable Alerts Only:
Preventing Alert Fatigue:
PagerDuty:
Opsgenie:
incident.io:
For detailed tool comparison, see references/tool-comparison.md.
Statuspage.io: Most trusted, easy setup ($29-399/month) Instatus: Budget-friendly, modern design ($19-99/month)
MTTA (Mean Time To Acknowledge):
MTTR (Mean Time To Recovery):
MTBF (Mean Time Between Failures):
Incident Frequency:
Action Item Completion Rate:
Incident → Post-Mortem → Action Items → Prevention
↑ ↓
└──────────── Fewer Incidents ─────────────┘Is production completely down or critical data at risk?
├─ YES → SEV0
└─ NO → Is major functionality degraded?
├─ YES → Is there a workaround?
│ ├─ YES → SEV1
│ └─ NO → SEV0
└─ NO → Are customers impacted?
├─ YES → SEV2
└─ NO → SEV3Use interactive classifier: python scripts/classify-severity.py
For detailed escalation guidance, see references/escalation-matrix.md.
Prioritize Mitigation When:
Prioritize Root Cause When:
Default: Mitigation first (99% of cases)
Observability: Monitoring alerts trigger incidents → Use incident-management for response
Disaster Recovery: DR provides recovery procedures → Incident-management provides operational response
Security Incident Response: Similar process with added compliance/forensics
Infrastructure-as-Code: IaC enables fast recovery via automated rebuild
Performance Engineering: Performance incidents trigger response → Performance team investigates post-mitigation
Runbook Templates:
examples/runbooks/database-failover.mdexamples/runbooks/cache-invalidation.mdexamples/runbooks/ddos-mitigation.mdPost-Mortem Template:
examples/postmortem-template.md - Complete blameless post-mortem structureCommunication Templates:
examples/communication-templates.md - Status updates, customer emailsOn-Call Handoff:
examples/oncall-handoff-template.md - Weekly handoff formatIntegration Scripts:
examples/integrations/pagerduty-slack.pyexamples/integrations/statuspage-auto-update.pyexamples/integrations/postmortem-generator.pyInteractive Severity Classifier:
python scripts/classify-severity.pyAsks questions to determine appropriate severity level based on impact and urgency.
Books:
Online Resources:
Standards:
© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 18 other files (scripts, references) in skills/managing-incidents of ancoleman/ai-design-components.
Open the folder on GitHubat commit 76551b7
Managing Incidents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Managing Incidents this skillancoleman/ai-design-components | 525 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Oncallpigweed-project/pigweed | 548 | — | ~963 | Automated safety check: Pass | Apache-2.0 | |
| Activation Governance Chaos RolloutAli-Marandi/DataSense | 107 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Incident Response686f6c61/alfred-dev | 117 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Superset Incident Triagesuperset-sh/superset | 15k | — | ~1k | Automated safety check: Pass | Custom licence | |
| Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan | 109 | — | ~1.9k | Automated safety check: Pass | MIT |
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
Ali-Marandi/DataSense
Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.
686f6c61/alfred-dev
Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.
superset-sh/superset
Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.
VeryGoodOpenSource/vgv-wingspan
Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.
Jeffallan/claude-skills
Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.
ancoleman/ai-design-components
Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.
ancoleman/ai-design-components
Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.
ancoleman/ai-design-components
Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.
ancoleman/ai-design-components
Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.
ancoleman/ai-design-components
Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.
ancoleman/ai-design-components
Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.
Categories
Guide incident response from detection to post-mortem using SRE principles, severity classification, on-call management, blameless culture, and communication protocols. Managing Incidents is an agent skill from ancoleman/ai-design-components. Guide incident response from detection to post-mortem using SRE principles, severity classification, on-call management, blameless culture, and communication protocols.
Managing Incidents fits situations like: setting up incident processes; designing escalation policies; conducting post-mortems.
Run `npx skills add ancoleman/ai-design-components --skill managing-incidents -a claude-code`. Or copy the skill folder (skills/managing-incidents in ancoleman/ai-design-components) into .claude/skills/managing-incidents in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ancoleman/ai-design-components --skill managing-incidents -a codex`. Or copy the skill folder (skills/managing-incidents in ancoleman/ai-design-components) into .agents/skills/managing-incidents in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill managing-incidents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/managing-incidents, .gemini/skills/managing-incidents, .github/skills/managing-incidents and .opencode/skills/managing-incidents in your project.
Going by SKILL.md and its folder, Managing Incidents needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Managing Incidents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Managing Incidents: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 525 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.
Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.