Kubernetes Network Root Cause Analysis
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
Comprehensive incident response framework from detection through resolution and post-incident review.
$ npx skills add alirezarezvani/claude-skills --skill incident-commander -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install alirezarezvani/claude-skills incident-commander --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering-team/skills/incident-commander .claude/skills/incident-commander && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "incident-commander" agent skill from https://github.com/alirezarezvani/claude-skills/tree/main/engineering-team/skills/incident-commander into .claude/skills/incident-commander/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-commander", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/alirezarezvani/claude-skills/tree/main/engineering-team/skills/incident-commanderType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add alirezarezvani/claude-skills --skill incident-commander -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install alirezarezvani/claude-skills incident-commander --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/engineering-team/skills/incident-commander .agents/skills/incident-commander && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "incident-commander" agent skill from https://github.com/alirezarezvani/claude-skills/tree/main/engineering-team/skills/incident-commander into .agents/skills/incident-commander/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-commander", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alirezarezvani/claude-skills --skill incident-commander -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install alirezarezvani/claude-skills incident-commander --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/engineering-team/skills/incident-commander .cursor/skills/incident-commander && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "incident-commander" agent skill from https://github.com/alirezarezvani/claude-skills/tree/main/engineering-team/skills/incident-commander into .cursor/skills/incident-commander/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-commander", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/alirezarezvani/claude-skills.git --path engineering-team/skills/incident-commander--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add alirezarezvani/claude-skills --skill incident-commander -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install alirezarezvani/claude-skills incident-commander --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/engineering-team/skills/incident-commander .gemini/skills/incident-commander && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "incident-commander" agent skill from https://github.com/alirezarezvani/claude-skills/tree/main/engineering-team/skills/incident-commander into .gemini/skills/incident-commander/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-commander", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install alirezarezvani/claude-skills incident-commanderInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add alirezarezvani/claude-skills --skill incident-commander -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/engineering-team/skills/incident-commander .github/skills/incident-commander && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "incident-commander" agent skill from https://github.com/alirezarezvani/claude-skills/tree/main/engineering-team/skills/incident-commander into .github/skills/incident-commander/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-commander", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alirezarezvani/claude-skills --skill incident-commander -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install alirezarezvani/claude-skills incident-commander --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/engineering-team/skills/incident-commander .opencode/skills/incident-commander && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "incident-commander" agent skill from https://github.com/alirezarezvani/claude-skills/tree/main/engineering-team/skills/incident-commander into .opencode/skills/incident-commander/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-commander", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
incident-commanderComprehensive incident response framework from detection through resolution and post-incident review.
Incident Commander is an agent skill from alirezarezvani/claude-skills. Comprehensive incident response framework from detection through resolution and post-incident review. Battle-tested SRE/DevOps practices: severity classification, timeline reconstruction, structured post-incident analysis. Use when declaring an incident, coordinating multi-team response during an outage, leading a post-mortem, or setting up on-call practices for a new service.
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including scripts, reference files and assets (for example `README.md`, `assets/incident_report_template.md` and `assets/runbook_template.md`).
It sits in DevOps & Cloud, covering Incident response. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Incident Commander loads about 3.8k tokens when it runs, and up to ~28k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 1,085 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 1,085 words, ~3,819 tokens.
.claude/skills/incident-commander/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.Category: Engineering Team
Tier: POWERFUL
Author: Claude Skills Team
Version: 1.0.0
Last Updated: February 2026
Incident response framework for availability/reliability incidents (outages, degradations, failed deploys): severity classification, timeline reconstruction, and post-incident review.
This is NOT security incident triage. For security events (ransomware, intrusion, data exfiltration, IOC analysis, NIST SP 800-61 forensics), route to incident-response. Both skills use SEV1-SEV4 labels; this one scores operational impact (users, revenue, SLA), while incident-response classifies attack types and forensic handling.
Incident Classifier (incident_classifier.py)
Timeline Reconstructor (timeline_reconstructor.py)
PIR Generator (pir_generator.py)
Definition: Complete service failure affecting all users or critical business functions
Characteristics:
Response Requirements:
Communication Frequency: Every 15 minutes until resolution
Definition: Significant degradation affecting subset of users or non-critical functions
Characteristics:
Response Requirements:
Communication Frequency: Every 30 minutes during active response
Definition: Limited impact with workarounds available
Characteristics:
Response Requirements:
Communication Frequency: At key milestones only
Definition: Minimal impact, cosmetic issues, or planned maintenance
Characteristics:
Response Requirements:
Communication Frequency: Standard development cycle updates
Command and Control
Communication Hub
Process Management
Post-Incident Leadership
Emergency Decisions (SEV1/2):
Resource Allocation:
Technical Decisions:
Subject: [SEV{severity}] {Service Name} - {Brief Description}
Incident Details:
- Start Time: {timestamp}
- Severity: SEV{level}
- Impact: {user impact description}
- Current Status: {investigating/mitigating/resolved}
Technical Details:
- Affected Services: {service list}
- Symptoms: {what users are experiencing}
- Initial Assessment: {suspected root cause if known}
Response Team:
- Incident Commander: {name}
- Technical Lead: {name}
- SMEs Engaged: {list}
Next Update: {timestamp}
Status Page: {link}
War Room: {bridge/chat link}
---
{Incident Commander Name}
{Contact Information}Subject: URGENT - Customer-Impacting Outage - {Service Name}
Executive Summary:
{2-3 sentence description of customer impact and business implications}
Key Metrics:
- Time to Detection: {X minutes}
- Time to Engagement: {X minutes}
- Estimated Customer Impact: {number/percentage}
- Current Status: {status}
- ETA to Resolution: {time or "investigating"}
Leadership Actions Required:
- [ ] Customer communication approval
- [ ] PR/Communications coordination
- [ ] Resource allocation decisions
- [ ] External vendor engagement
Incident Commander: {name} ({contact})
Next Update: {time}
---
This is an automated alert from our incident response system.We are currently experiencing {brief description of issue} affecting {scope of impact}.
Our engineering team was alerted at {time} and is actively working to resolve the issue. We will provide updates every {frequency} until resolved.
What we know:
- {factual statement of impact}
- {factual statement of scope}
- {brief status of response}
What we're doing:
- {primary response action}
- {secondary response action}
Workaround (if available):
{workaround steps or "No workaround currently available"}
We apologize for the inconvenience and will share more information as it becomes available.
Next update: {time}
Status page: {link}Internal Stakeholders:
External Stakeholders:
| Stakeholder | SEV1 | SEV2 | SEV3 | SEV4 |
|---|---|---|---|---|
| Engineering Leadership | Real-time | 30min | 4hrs | Daily |
| Executive Team | 15min | 1hr | EOD | Weekly |
| Customer Support | Real-time | 30min | 2hrs | As needed |
| Customers | 15min | 1hr | Optional | None |
| Partners | 30min | 2hrs | Optional | None |
Detection Playbooks
Response Playbooks
Recovery Playbooks
# {Service/Component} Incident Response Runbook
## Quick Reference
- **Severity Indicators:** {list of conditions for each severity level}
- **Key Contacts:** {on-call rotations and escalation paths}
- **Critical Commands:** {list of emergency commands with descriptions}
## Detection
### Monitoring Alerts
- {Alert name}: {description and thresholds}
- {Alert name}: {description and thresholds}
### Manual Detection Signs
- {Symptom}: {what to look for and where}
- {Symptom}: {what to look for and where}
## Initial Response (0-15 minutes)
1. **Assess Severity**
- [ ] Check {primary metric}
- [ ] Verify {secondary indicator}
- [ ] Classify as SEV{level} based on {criteria}
2. **Establish Command**
- [ ] Page Incident Commander if SEV1/2
- [ ] Create incident tracking ticket
- [ ] Join war room: {link/bridge info}
3. **Initial Investigation**
- [ ] Check recent deployments: {deployment log location}
- [ ] Review error logs: {log location and queries}
- [ ] Verify dependencies: {dependency check commands}
## Mitigation Strategies
### Strategy 1: {Name}
**Use when:** {conditions}
**Steps:**
1. {detailed step with commands}
2. {detailed step with expected outcomes}
3. {validation step}
**Rollback Plan:**
1. {rollback step}
2. {verification step}
### Strategy 2: {Name}
{similar structure}
## Recovery and Validation
1. **Service Restoration**
- [ ] {restoration step}
- [ ] Wait for {metric} to return to normal
- [ ] Validate end-to-end functionality
2. **Communication**
- [ ] Update status page
- [ ] Notify stakeholders
- [ ] Schedule PIR
## Common Pitfalls
- **{Pitfall}:** {description and how to avoid}
- **{Pitfall}:** {description and how to avoid}
## Reference Information
→ See references/reference-information.md for details
## Usage Examples
### Example 1: Database Connection Pool Exhaustion
```bash
# Classify the incident
echo '{"description": "Users reporting 500 errors, database connections timing out", "affected_users": "80%", "business_impact": "high"}' | python scripts/incident_classifier.py
# Reconstruct timeline from logs
python scripts/timeline_reconstructor.py --input assets/sample_timeline_events.json --output timeline.md
# Generate PIR after resolution
python scripts/pir_generator.py --incident assets/sample_incident_data.json --timeline timeline.md --output pir.md# Quick classification from stdin
echo "API rate limits causing customer API calls to fail" | python scripts/incident_classifier.py --format text
# Build timeline from multiple sources
python scripts/timeline_reconstructor.py --input assets/simple_timeline_events.json --detect-phases --gap-analysis
# Generate comprehensive PIR
python scripts/pir_generator.py --incident assets/sample_incident_pir_data.json --rca-method fishbone --action-itemsMaintain Calm Leadership
Document Everything
Effective Communication
Technical Excellence
Blameless Culture
Action Item Discipline
Knowledge Sharing
Continuous Improvement
© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 23 other files (scripts, references, assets) in engineering-team/skills/incident-commander of alirezarezvani/claude-skills.
Open the folder on GitHubat commit 19392f7
Incident Commander next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Incident Commander this skillalirezarezvani/claude-skills | 28k | — | ~3.8k | Automated safety check: Pass | MIT | |
| Kubernetes Network Root Cause Analysiskubeshark/kubeshark | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | |
| UModel Root Cause Analysisalibaba/UnifiedModel | 415 | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Learningskortix-ai/suna | 20k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Oncallpigweed-project/pigweed | 548 | — | ~992 | Automated safety check: Pass | Apache-2.0 | |
| Loop Triage Reportcobusgreyling/loop-engineering | 11k | — | ~500 | Automated safety check: Pass | MIT |
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
kortix-ai/suna
The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
cobusgreyling/loop-engineering
Turns CI failures, open issues, recent commits and chat threads into a prioritized markdown report that an automation loop can act on without inventing architecture work.
openclaw/clawhub
Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.
alirezarezvani/claude-skills
Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.
alirezarezvani/claude-skills
OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.
alirezarezvani/claude-skills
Design AWS architectures for startups using serverless patterns and IaC templates.
alirezarezvani/claude-skills
Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.
alirezarezvani/claude-skills
Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.
Categories
Comprehensive incident response framework from detection through resolution and post-incident review. Incident Commander is an agent skill from alirezarezvani/claude-skills. Comprehensive incident response framework from detection through resolution and post-incident review.
Incident Commander fits situations like: declaring an incident; coordinating multi-team response during an outage; leading a post-mortem; setting up on-call practices for a new service.
Run `npx skills add alirezarezvani/claude-skills --skill incident-commander -a claude-code`. Or copy the skill folder (engineering-team/skills/incident-commander in alirezarezvani/claude-skills) into .claude/skills/incident-commander in your project. Claude Code loads it when a task matches its description.
Run `npx skills add alirezarezvani/claude-skills --skill incident-commander -a codex`. Or copy the skill folder (engineering-team/skills/incident-commander in alirezarezvani/claude-skills) into .agents/skills/incident-commander in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill incident-commander -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-commander, .gemini/skills/incident-commander, .github/skills/incident-commander and .opencode/skills/incident-commander in your project.
Going by SKILL.md and its folder, Incident Commander needs the command-line tools its instructions call (python).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Incident Commander is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 24k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Incident Commander: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 415 stars), Learnings (kortix-ai/suna, 20k stars) and Oncall (pigweed-project/pigweed, 548 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,891 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.
Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.