Post-Incident Debrief
VeryGoodOpenSource/vgv-wingspan
Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.
Incident response process management following the NIST 800-61 lifecycle.
$ npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills incident-response-lifecycle --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/incident-response-lifecycle .claude/skills/incident-response-lifecycle && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "incident-response-lifecycle" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/incident-response-lifecycle into .claude/skills/incident-response-lifecycle/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response-lifecycle", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/incident-response-lifecycleType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills incident-response-lifecycle --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/incident-response-lifecycle .agents/skills/incident-response-lifecycle && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "incident-response-lifecycle" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/incident-response-lifecycle into .agents/skills/incident-response-lifecycle/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response-lifecycle", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills incident-response-lifecycle --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/incident-response-lifecycle .cursor/skills/incident-response-lifecycle && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "incident-response-lifecycle" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/incident-response-lifecycle into .cursor/skills/incident-response-lifecycle/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response-lifecycle", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/incident-response-lifecycle--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills incident-response-lifecycle --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/incident-response-lifecycle .gemini/skills/incident-response-lifecycle && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "incident-response-lifecycle" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/incident-response-lifecycle into .gemini/skills/incident-response-lifecycle/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response-lifecycle", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills incident-response-lifecycleInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/incident-response-lifecycle .github/skills/incident-response-lifecycle && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "incident-response-lifecycle" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/incident-response-lifecycle into .github/skills/incident-response-lifecycle/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response-lifecycle", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills incident-response-lifecycle --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/incident-response-lifecycle .opencode/skills/incident-response-lifecycle && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "incident-response-lifecycle" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/incident-response-lifecycle into .opencode/skills/incident-response-lifecycle/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response-lifecycle", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
incident-response-lifecycleIncident response process management following the NIST 800-61 lifecycle.
Incident Response Lifecycle is an agent skill from LeoYeAI/openclaw-master-skills. Incident response process management following the NIST 800-61 lifecycle. Covers severity classification, escalation matrices, role assignment, communication management, phased recovery coordination, blameless post-mortem facilitation, and 5-whys root cause analysis. Scoped to the process and coordination layer — for network-level evidence collection and forensic analysis, use incident-response-network instead.
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `_meta.json`, `references/communication-templates.md` and `references/rca-framework.md`).
It sits in DevOps & Cloud, covering Incident response, Runbooks and postmortems and Root cause analysis. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Incident Response Lifecycle loads about 5k tokens when it runs, and up to ~8.9k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 1,979 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its Apache-2.0 licence (© LeoYeAI). 1,979 words, ~4,974 tokens.
.claude/skills/incident-response-lifecycle/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Structured process management for network incidents from detection through post-incident review. This skill covers the organizational coordination layer: severity classification, escalation, role assignment, stakeholder communication, recovery coordination, and root cause analysis. It does not cover technical evidence collection, device forensics, or containment execution — use the incident-response-network skill for network-level evidence gathering and forensic analysis.
The procedure follows the operational lifecycle shape: detect and classify the incident, triage and escalate to the right people, coordinate the investigation across teams, manage communications to all audiences, drive resolution and recovery, then conduct a blameless post-incident review.
See references/communication-templates.md for notification templates by
audience and severity level. See references/rca-framework.md for
the 5-whys methodology, fishbone diagram guidance, and post-mortem
document structure.
Follow these six steps in sequence. Steps 3 and 4 run in parallel once
roles are assigned — investigation coordination and communication
management proceed simultaneously. Each step references templates from
references/communication-templates.md and methodology from
references/rca-framework.md where applicable.
Classify the incident by severity, type, and scope to determine the appropriate response level.
Severity assignment — apply the P1–P4 taxonomy from the Threshold Tables section. Base severity on the highest-impact criterion met. When multiple criteria apply at different levels, the highest governs.
Incident type classification — categorize as outage (service unavailable), degradation (reduced capacity), security (unauthorized access or data exposure), or data loss (corruption or deletion).
Scope determination — assess whether the incident affects a single device, a network segment, an entire site, or multiple sites. Scope drives staffing, communication breadth, and recovery complexity.
Initial impact assessment — estimate affected user count, impacted services and their business criticality, data at risk, and revenue impact per hour. Record estimates in the incident ticket.
Assign roles, notify stakeholders, and set response timeline expectations based on the severity classification from Step 1.
Role assignment — every P1 or P2 incident needs four named roles: Incident Commander (IC) — owns the incident end-to-end and makes escalation decisions; Technical Lead — coordinates diagnostics and synthesizes findings; Communications Lead — drafts stakeholder notifications and manages the status page; Scribe — maintains the real-time timeline and records bridge call decisions. For P3, IC and Technical Lead may be combined. P4 uses normal operations workflows.
Escalation matrix execution — notify by severity: P1 — all four roles plus engineering management, VP/director on-call, vendor TAC if vendor equipment is involved, executive notification within 30 minutes. P2 — all four roles plus engineering management within 1 hour. P3 — Technical Lead plus team lead within 4 hours. P4 — assigned engineer via normal ticket queue.
Response timeline expectations: P1 — bridge in 15 minutes, first update in 30 minutes, then every 30 minutes. P2 — bridge in 30 minutes, first update in 1 hour, then every 2 hours. P3 — initial assessment in 4 hours, daily updates. P4 — acknowledgment within 1 business day.
Vendor engagement criteria — engage vendor TAC when the incident involves hardware failure, software defects requiring patches, or when internal triage has not identified root cause within the severity time window.
Coordinate the technical investigation across teams and evidence sources. For network-level evidence collection (device state, routing tables, interface data, log retrieval), reference the incident-response-network skill — this step focuses on organizing the investigation, not executing forensic commands.
Evidence collection tasking — assign team members to collect evidence from relevant domains: network devices (via incident-response-network procedures), application logs, infrastructure metrics, and security tooling alerts. Each assignee reports findings to the Technical Lead.
Parallel investigation streams — for complex incidents, run multiple investigation threads simultaneously. Common parallel tracks: (1) symptom analysis — what is failing and for whom, (2) change correlation — what changed recently (deployments, config modifications, maintenance), (3) external factors — upstream provider issues, DDoS, DNS resolution failures.
Hypothesis tracking — maintain a running list of hypotheses with current status (investigating, confirmed, ruled out). Each hypothesis should have an owner and a validation method. Update the list on every bridge call.
Timeline of events (ToE) — the Scribe maintains a running chronological log of when events occurred, when they were detected, what actions were taken, and what was discovered. The ToE becomes the foundation for the post-incident review in Step 6.
Subject matter expert engagement — when investigation stalls or enters an unfamiliar domain, escalate to specialists. Define clear handoff: what has been tried, what data is available, and what specific question needs answering.
Manage stakeholder communications throughout the incident. Use the
templates in references/communication-templates.md for consistent
messaging across audiences.
Stakeholder notification by audience — executive summary (business
impact, estimated resolution, customer exposure — no technical detail),
technical detail (root cause hypothesis, diagnostics, remediation plan
— delivered on bridge call), customer-facing (service impact, workaround
if available, estimated resolution — via status page), regulatory
(formal notification per compliance framework when required). Use
templates from references/communication-templates.md.
Status update cadence — follow severity-based cadence from Step 2. Each update includes: current status, progress since last update, next planned action, and revised time-to-resolution estimate.
Bridge call management — the IC runs calls with a fixed agenda: (1) technical status from Tech Lead, (2) communication status from Comms Lead, (3) hypothesis updates, (4) decisions needed, (5) action items with owners and deadlines. Keep calls focused — park side discussions as action items.
External notification requirements — track regulatory reporting deadlines, law enforcement notification when criminal activity is suspected, customer SLA breach notification per contractual terms, and vendor escalation for ongoing support.
Drive service restoration through validated recovery steps with monitoring to confirm the fix holds.
Recovery validation criteria — before declaring resolved, confirm: (1) service health checks return normal for all affected components, (2) monitoring dashboards show green for at least 15 minutes (P1) or 30 minutes (P2), (3) no new related alerts during observation, (4) affected users confirm restoration (sample check for large populations).
Phased restoration — for multi-layer network incidents, restore in order: core infrastructure → distribution layer → access layer → end-to-end verification. Verify each phase before proceeding. Do not restore all layers simultaneously — cascading failures during recovery are worse than a phased approach.
Back-out plan execution — if the fix causes new issues, execute the pre-defined rollback. Every remediation action should have a documented rollback method before execution.
Enhanced monitoring period — maintain heightened monitoring after resolution: P1 for 24 hours, P2 for 12 hours, P3 through the next business day. This means reduced alert thresholds on affected systems, active watch by on-call, and immediate re-escalation if symptoms recur.
Incident closure — send closure notification to all stakeholders
(template in references/communication-templates.md). Update the
ticket with resolution summary, total duration, and final impact.
Schedule the post-incident review.
Conduct a blameless post-incident review to identify root cause,
contributing factors, and improvement actions. See
references/rca-framework.md for the full methodology.
Scheduling — hold the post-mortem within 72 hours of incident
resolution while details are fresh. Invite all incident participants
plus relevant stakeholders. Send the invitation using the template in
references/communication-templates.md.
5-whys root cause analysis — apply iteratively: for each "why"
answer, ask "why" again until reaching a systemic root cause (typically
3–5 iterations). See references/rca-framework.md for worked
examples and facilitation guidance.
Contributing factor categorization — classify each contributing factor as process (missing runbook, unclear escalation path), people (training gap, staffing shortage), or technology (monitoring gap, single point of failure, software defect). This categorization guides the type of remediation action needed.
Action item classification — assign each action item one of four dispositions: fix (eliminate the root cause), mitigate (reduce likelihood or impact), accept (risk is within tolerance, document rationale), or transfer (assign to another team or vendor). Every fix or mitigate action must have an owner, due date, and verification method.
Incident metrics — collect and record: Mean Time to Detect (MTTD), Mean Time to Investigate (MTTI), Mean Time to Resolve (MTTR), total incident duration, number of customers affected, and whether this is a recurrence of a previous incident. Track these metrics over time to measure improvement trends.
| Severity | User Impact | Service Impact | Data Risk | Response SLA |
|---|---|---|---|---|
| P1 Critical | >50% of users or all VIP users | Complete outage of revenue-generating service | Confirmed data breach or loss | Bridge in 15 min, updates every 30 min |
| P2 High | 10–50% of users affected | Major degradation or redundancy loss on critical path | Suspected data exposure | Bridge in 30 min, updates every 2 hr |
| P3 Medium | <10% of users, workaround exists | Partial degradation, non-critical service | No data risk identified | Assessment in 4 hr, updates daily |
| P4 Low | Minimal or no user impact | Cosmetic, non-production, or fully redundant | None | Ack in 1 business day |
| Severity | Incident Commander | Technical Lead | Comms Lead | Scribe | Management | Executive |
|---|---|---|---|---|---|---|
| P1 | Required | Required | Required | Required | Immediate | Within 30 min |
| P2 | Required | Required | Required | Optional | Within 1 hr | If SLA breached |
| P3 | Combined with Tech Lead | Required | Optional | No | Within 4 hr | No |
| P4 | No | Assigned engineer | No | No | Normal reporting | No |
| Severity | Monitoring Period | Alert Threshold | Re-escalation Trigger |
|---|---|---|---|
| P1 | 24 hours | Reduced by 20% | Any recurrence symptom |
| P2 | 12 hours | Reduced by 10% | Same failure signature |
| P3 | Next business day | Normal thresholds | Identical alert |
| P4 | None | Normal | Normal process |
Event detected or reported
├── Is the service completely unavailable?
│ ├── Yes → Is it a revenue-generating or safety-critical service?
│ │ ├── Yes → P1 Critical
│ │ └── No → P2 High
│ └── No → Service is partially available
│ ├── Are more than 10% of users affected without workaround?
│ │ ├── Yes → P2 High
│ │ └── No → Is there a workaround available?
│ │ ├── Yes → P3 Medium
│ │ └── No, but fewer than 10% of users → P3 Medium
│ └── Is this a non-production or cosmetic issue?
│ └── Yes → P4 Low
├── Is there confirmed or suspected data exposure?
│ ├── Confirmed breach → P1 Critical (regardless of service status)
│ └── Suspected exposure → P2 High minimum
└── Has redundancy been lost on a critical path?
├── Yes, no failover remaining → P2 High
└── Yes, failover still available → P3 MediumSeverity assigned
├── P1 or P2?
│ ├── Yes → Assign all four roles immediately
│ │ ├── Is vendor equipment involved in the failure?
│ │ │ ├── Yes → Open vendor TAC case immediately
│ │ │ └── No → Internal investigation first
│ │ └── Has root cause been identified within time window?
│ │ ├── P1: not identified within 30 min → Escalate to next tier
│ │ └── P2: not identified within 2 hr → Escalate to next tier
│ └── P3 or P4?
│ ├── P3 → Assign Technical Lead, monitor for escalation
│ │ └── Impact worsening? → Re-classify severity upward
│ └── P4 → Normal ticket queue, no escalation
└── At any point: if scope expands beyond initial classification
└── Re-evaluate severity from Step 1, escalate if neededINCIDENT REPORT
=====================================
Incident ID: [ticket/tracking number]
Severity: [P1/P2/P3/P4]
Incident Commander: [name]
Duration: [detection time] — [resolution time] ([total hours])
Status: [Resolved / Monitoring / Under Review]
IMPACT SUMMARY:
Users Affected: [count or percentage]
Services Affected: [list of impacted services]
Revenue Impact: [estimated or confirmed]
Data Impact: [none / suspected / confirmed — description]
TIMELINE OF EVENTS:
| # | Time (UTC) | Event | Actor | Notes |
|---|-----------|-------|-------|-------|
| 1 | [time] | [event description] | [person/system] | [context] |
ROOT CAUSE:
Category: [Process / People / Technology]
Root Cause: [description from 5-whys analysis]
Contributing Factors:
- [factor 1 — category]
- [factor 2 — category]
RESOLUTION:
Fix Applied: [description of what resolved the incident]
Validated By: [how resolution was confirmed]
Back-out Available: [yes/no — description]
METRICS:
MTTD: [time from occurrence to detection]
MTTI: [time from detection to root cause identified]
MTTR: [time from detection to resolution]
Recurrence: [yes/no — reference to prior incident if yes]
ACTION ITEMS:
| # | Action | Type | Owner | Due Date | Status |
|---|--------|------|-------|----------|--------|
| 1 | [action] | [Fix/Mitigate/Accept/Transfer] | [name] | [date] | [status] |
POST-MORTEM STATUS:
Scheduled: [date/time or "pending"]
Attendees: [roles invited]
Document Location: [link to post-mortem document]Symptom: Teams classify the same incident at different severity levels, causing confusion about response urgency.
Resolution: The IC makes the final determination using the Threshold Tables criteria. The highest applicable severity governs. Document rationale in the ticket. If the IC is not yet assigned, the first responder sets initial severity and the IC may adjust.
Symptom: Frequent P1/P2 declarations for issues that resolve quickly, eroding trust in severity classification.
Resolution: Review severity criteria quarterly. Track the false-positive rate (incidents downgraded after initial classification). If P1 downgrade rate exceeds 30%, tighten P1 criteria. Ensure P3/P4 incidents are not over-classified.
Symptom: Action items accumulate but are not completed, leading to recurring incidents from known causes.
Resolution: Assign every action item an owner and due date at the review. Track completion in the incident system, not separate documents. Review open items in weekly standups. Escalate overdue items to management and report completion rates alongside MTTD/MTTR.
Symptom: Status updates become infrequent during long incidents (>4 hours), leaving stakeholders uninformed.
Resolution: The Communications Lead maintains cadence regardless of investigation progress. If no new findings exist, state that explicitly in the update. For incidents exceeding 8 hours, rotate the Comms Lead role to prevent fatigue.
Symptom: The same incident recurs after being marked resolved.
Resolution: Check whether prior post-mortem action items were completed. If yes, the root cause analysis was incomplete — reconvene with broader scope. If not, escalate the completion failure. Tag the new incident as a recurrence and increase severity by one level to reflect accumulated impact.
© LeoYeAI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/incident-response-lifecycle of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Incident Response Lifecycle next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Incident Response Lifecycle this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~5k | Automated safety check: Pass | Apache-2.0 | |
| Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan | 109 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Conducting Post Incident Lessons Learnedmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Incident Postmortemgithub/awesome-copilot | 40k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Responding To Incidentstrilwu/secskills | 157 | — | ~3.9k | Automated safety check: Notes | MIT | |
| Incident Postmortemrevfactory/harness-100 | 1.3k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 |
VeryGoodOpenSource/vgv-wingspan
Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.
mukul975/Anthropic-Cybersecurity-Skills
Facilitate structured post-incident reviews to identify root causes, document what worked and failed, and produce actionable recommendations to improve future incident response.
github/awesome-copilot
A skill your agent uses when an outage, production incident, or significant service degradation has occurred and the team needs to write a structured blameless post-mortem.
trilwu/secskills
Run digital forensics and incident response — triage, evidence acquisition with chain of custody, host and cloud artifact analysis, timeline reconstruction, scoping, containment, eradication, and…
revfactory/harness-100
A full pipeline where an agent team collaborates to generate incident postmortem reports.
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Categories
Incident response process management following the NIST 800-61 lifecycle. Incident Response Lifecycle is an agent skill from LeoYeAI/openclaw-master-skills. Incident response process management following the NIST 800-61 lifecycle.
Incident Response Lifecycle fits situations like: tasks that involve Incident response; tasks that involve Runbooks and postmortems; tasks that involve Root cause analysis.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a claude-code`. Or copy the skill folder (skills/incident-response-lifecycle in LeoYeAI/openclaw-master-skills) into .claude/skills/incident-response-lifecycle in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a codex`. Or copy the skill folder (skills/incident-response-lifecycle in LeoYeAI/openclaw-master-skills) into .agents/skills/incident-response-lifecycle in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill incident-response-lifecycle -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-response-lifecycle, .gemini/skills/incident-response-lifecycle, .github/skills/incident-response-lifecycle and .opencode/skills/incident-response-lifecycle in your project.
SKILL.md names no scripts, command-line tools or credentials: Incident Response Lifecycle is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Incident Response Lifecycle is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Incident Response Lifecycle: Post-Incident Debrief (VeryGoodOpenSource/vgv-wingspan, 109 stars), Conducting Post Incident Lessons Learned (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Incident Postmortem (github/awesome-copilot, 40k stars) and Responding To Incidents (trilwu/secskills, 157 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,160 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.