Agent skill

Runbook Gen

by EliasOulkadi in EliasOulkadi/shokunin

Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates…

MITAuto-check passedDevOps & Cloud

Install Runbook Gen

skills CLI
$ npx skills add EliasOulkadi/shokunin --skill runbook-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EliasOulkadi/shokunin runbook-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.pack/skills/runbook-gen .claude/skills/runbook-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
runbook-gen
GitHub stars
114
Token cost
~2.8k tokens
SKILL.md length
511 words
Files
1
Skills in repo
49
Repo updated
First seen
Licence
MIT

At a glance

Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates…

  • Works in 3 steps: Required Discovery → Severity Matrix → Runbook Template
  • Tasks that involve Runbooks and postmortems
  • SKILL.md covers Workflow, 2. Required Discovery, 3. Severity Matrix and 4. Runbook Template, plus 8 more sections
  • Calls kubectl

What it does

Runbook Gen is an agent skill from EliasOulkadi/shokunin. Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates, and blameless post-mortems with action items.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: opencode

It sits in DevOps & Cloud, covering Runbooks and postmortems. It works with Slack. The repository describes itself as: 職人 Shokunin 62 AI agent skills for OpenCode, Claude Code, Cursor, Windsurf. ChromaDB memory, MCP servers, declarative self-updates. Multi-model, open source, zero cost. The licence is MIT.

When your agent uses it

  • Tasks that involve Runbooks and postmortems

Example prompts

  • “/runbook-gen”

Requirements

  • Compatibility (from SKILL.md): opencode

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Required Discovery
  2. Severity Matrix
  3. Runbook Template

What it can do on your machine

Read from SKILL.md and the folder at commit 4c68e5b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    opencode

    From compatibility in the SKILL.md frontmatter.

Context cost

Runbook Gen loads about 2.8k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 511 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EliasOulkadi/shokunin at commit 4c68e5b, republished under its MIT licence (© EliasOulkadi). 511 words, ~2,772 tokens.

Download SKILL.mdSave it as .claude/skills/runbook-gen/SKILL.md (or your agent's skills folder).
name
runbook-gen
description
Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates, and blameless post-mortems with action items.
compatibility
opencode
license
MIT
metadata.workflow
operations
metadata.audience
sre
metadata.version
3.0.0
triggers
create a runbook, incident response plan, on-call guide, SRE procedure, escalation path, outage playbook, post-mortem, war room, sev1 procedure
negatives
Use **documentation** for general project docs, READMEs, API docs, Use **error-handler** for error handling patterns, retry logic, circuit breakers, Use…

Runbook Generator

Operation runbooks that on-call engineers can follow under pressure. Based on Google SRE practices, PagerDuty incident response, and post-mortem culture.

Workflow

Follow these steps in order when generating a runbook:

  1. Discover — Ask the 5 Required Discovery questions (below) to scope the runbook
  2. Classify — Map the incident to the Severity Matrix; if unclear start at Sev2
  3. Template — Instantiate the Runbook Template with service-specific commands, time-boxes, and contacts
  4. War Room — Configure Slack channel, Zoom bridge, shared doc, status page, and customer comms
  5. Verify — Run through the Production Checklist before marking complete
  6. Archive — Store the runbook in the team's on-call repo, test in a drill within the quarter

When the ask is a post-mortem (not a runbook), skip steps 3-4 and use the Post-Mortem Template instead.

2. Required Discovery

  1. Service/System: What is the runbook about? (API, database, cache, queue)
  2. Incident types: What can go wrong? (down, degraded, slow, data loss)
  3. Environment: Production, staging, or both?
  4. Team structure: Primary, secondary, escalation contacts
  5. Existing monitoring: Alerts, dashboards, runbooks

3. Severity Matrix

SeverityDefinitionResponse SLA
Sev1Service down, users impactedRespond within 15 min
Sev2Degraded performance, partial impactRespond within 1 hour
Sev3Minor issue, no user impactNext business day
Sev4Internal tooling, non-criticalPer team schedule

4. Runbook Template

## Runbook: [Incident Type]

### Detection
How this incident is typically discovered:
- Alert: [Prometheus/Grafana/Datadog alert]
- User symptom: [what users see or report]
- Automated detection: [auto-remediation]

### Initial Response (first 5 min)
1. Acknowledge alert (PagerDuty/Opsgenie). Time: < 2 min
2. Determine severity. Time: < 1 min
3. Assign owner. Create incident channel (#inc-sev1). Time: < 2 min
4. Post initial status:
   "Investigating [issue] affecting [scope]. Will update in 15 min."
5. Start diagnosis timer. Time: < 1 min

### Diagnosis (decision tree)
1. Check [primary dashboard]: expected [X], current [Y]
2. Check logs: `kubectl logs -n [ns] -l app=[service] --tail 200`
3. Check database health: `SELECT count(*) FROM pg_stat_activity`
4. Check cache/queue latency: `redis-cli --latency`
5. IF [error in logs] → Runbook A
6. IF [latency > threshold] → Runbook B
7. IF unknown → Escalate

### Resolution Procedures
Each time-boxed with exact commands and verification:

#### Runbook A: [Name] (10 min)
```bash
# Step 1 (2 min)
kubectl rollout restart deployment/[service]

# Verify (1 min)
kubectl rollout status deployment/[service]
Runbook B: [Alternative]
Verification Checklist
  • Service health endpoint returns 200
  • Alert resolved (dashboard green)
  • Error rate back to baseline (< 0.1%)
  • Latency p99 < [threshold]
Escalation
TimeboxActionContact
0-15 minPrimary on-call@name / phone
15-30 minSecondary on-call@name / phone
30-60 minEngineering manager@name
60+ minVP Engineering@name
Show full SKILL.md (242 more words)Show less
Post-Incident Recovery
  • Data integrity check
  • Deploy fix to production
  • Monitor for 30 min post-fix
  • Update runbook with lessons learned

## 5. War Room Setup

### Communications
- **Slack channel**: `#inc-sev1-[incident-name]`
- **Zoom bridge**: [link] (permanent war room)
- **Shared doc**: Google Doc or Notion page for live notes
- **Status page**: Update immediately on detection
- **Customer comms**: Template below

### Status Page Updates

Investigating: We're aware of [issue] affecting [scope]. Investigating root cause.

Monitoring: Deployed fix for [root cause]. Monitoring closely.

Resolved: [Issue] has been resolved. All systems operational.


### Customer Communication

Subject: [Service] incident — [date]

We experienced [description] from [start] to [end] ([duration]).

Root cause: [one sentence]

Impact: [specific metrics]

What we're doing:

  1. [Fix deployed]
  2. [Monitoring improvements]
  3. [Process changes]

We apologize for the disruption.


## 6. Post-Mortem Template

Post-Mortem: [Date] — [Title]

Severity: Sev[1-4] Duration: [detection] → [resolution] ([total]) Impact: [users/customers] for time

Timeline (UTC)
  • Time: [What happened]
  • Time: [Response started]
Root Cause

One paragraph. System-focused, not person-focused.

Contributing Factors
  • Monitoring gaps
  • Process gaps
  • Knowledge gaps
Action Items
ActionOwnerDueType
[action]@nameDateprevent/detect/respond
What Went Well
What Went Wrong
What We'll Do Differently

## Error Handling

| Scenario | Behaviour | Guidance |
|----------|-----------|----------|
| User doesn't know the service/platform | Ask clarifying questions one at a time (not a list) | Start with "what service is this for?" |
| User says "just give me a template" | Output a blank Runbook Template with placeholder brackets | Skip Discovery, fill later |
| User asks for both runbook + post-mortem at once | Generate runbook first, then offer post-mortem | "I'll write the runbook now. Do you have an incident in mind for the post-mortem?" |
| Service has no existing monitoring | Flag the gap, offer to add basic health-check instructions | Note alerts must be configured separately |
| Runbook for a third-party/SaaS dependency | Include vendor status page check; mark resolution as "vendor fixes" | Focus on detection + escalation, not resolution |
| User says "I don't know the escalation contacts" | Use generic roles (Primary on-call, Secondary, EM, VP Eng) | Leave bracketed placeholders |
| Multiple teams involved | Add a RACI section to the runbook | Clarify who decides vs who executes |
| Runbook already exists, user wants update | Treat as revision: read current, diff, propose changes | Flag deprecations, test outdated commands |

## 8. Production Checklist

Before marking a runbook complete, verify every item:

- [ ] **Commands are copy-paste ready** — no unsubstituted variables (`[ns]` only in section headers/notes, never in commands)
- [ ] **Every step has a time-box** — engineer knows when to escalate
- [ ] **Decision tree is shallow** — max 3-4 levels; deeper means split into sub-runbooks
- [ ] **Multiple resolution paths** — at least 2 common failure modes covered
- [ ] **One runbook per incident type** — don't combine "DB failover" and "deployment rollback"
- [ ] **Verification step after every resolution** — how to confirm the fix worked
- [ ] **Escalation contacts listed** — primary + backup, both Slack and phone
- [ ] **Status page templates included** — Investigating / Monitoring / Resolved
- [ ] **Runbook tested in a drill** — within the current quarter, not when incident strikes
- [ ] **README section added** — link to runbook from team's on-call repo
- [ ] **Secrets checked** — no passwords, API keys, or tokens in commands; use env vars
- [ ] **Reviewed by secondary on-call** — fresh eyes catch assumptions

## Anti-Patterns

| Anti-pattern | Why It Fails | Fix |
|-------------|-------------|-----|
| Steps that say "fix the issue" | Too vague under pressure | Tell HOW, not what — exact commands |
| Commands with unsubstituted variables | Copy-paste fails, engineer wastes time retyping | Every command must run as-is |
| No verification step | Fix might not work, but nobody knows | Add check after every resolution |
| Runbooks >6 months stale | Commands rot, trust erodes, people ignore them | Schedule quarterly review in team calendar |
| Runbooks nobody tested | First test happens during the actual incident | Require one drill per quarter per runbook |
| Too many steps (>15) | Cognitive overload during Sev1 | Split into sub-runbooks, keep shallow |
| No time-boxes | Engineer doesn't know when to escalate | Every diagnosis/resolution step has a max time |
| Single resolution path | Assumes first idea works | Always have Plan B in the same runbook |
| Post-mortem blames people | Blame culture kills incident reporting | "What broke the system?" not "Who broke it?" |
| Skipping customer comms | Stakeholders hear "we're down" from social media | Draft customer email as part of template |

## Incident Timeline Template

[HH:MM] Alert triggered: <alert name> from <monitoring system> [HH:MM] On-call acknowledged: <name> [HH:MM] War room opened: <Slack channel/Zoom link> [HH:MM] Initial diagnosis: <symptoms observed> [HH:MM] Impact confirmed: <users affected, services degraded> [HH:MM] Mitigation applied: <action taken> [HH:MM] Monitoring confirms recovery [HH:MM] All-clear declared [HH:MM+1] Post-mortem scheduled: <date>


## Post-Mortem Template

Incident Post-Mortem: <title>

Date: YYYY-MM-DD | Duration: Xh Ym | Severity: Sev1/2/3

Timeline

[HH:MM] ...

Root Cause

What system/process failure caused this.

Impact

Users affected, data loss, revenue impact.

Detection

How was this found? (alert, user report, manual)

Resolution

What fixed it.

Action Items

#ActionOwnerDue

Lessons Learned

What would prevent this next time.


## Sources

- Google SRE Book: "Monitoring Distributed Systems"
- Google SRE Workbook: Incident Response
- PagerDuty incident response documentation
- Atlassian post-mortem best practices
- Microsoft SRE practices
- AWS Well-Architected Framework: Operational Excellence

## Checklist

- [ ] Skill loads without errors in the AI agent
- [ ] YAML frontmatter is valid (description, compatibility, audience)
- [ ] Workflow section provides clear step-by-step instructions
- [ ] Error handling section covers common failure modes
- [ ] All referenced files (references/, scripts/, assets/) exist
- [ ] Skill triggers correctly for intended use cases
- [ ] No broken links or missing resources

© EliasOulkadi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .pack/skills/runbook-gen of EliasOulkadi/shokunin.

Open the folder on GitHubat commit 4c68e5b

Compare with similar skills

Runbook Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Runbook Gen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Runbook Gen this skillEliasOulkadi/shokunin114—~2.8kAutomated safety check: PassMIT
Ops Automation Agentmastra-ai/mastra29k—~1.9kAutomated safety check: PassCustom licence
Vercel Incident Runbookjeremylongshore/tons-of-skills-marketplace2.8k—~2kAutomated safety check: PassMIT
Incident RetrospectiveOpenHands/extensions163—~922Automated safety check: PassMIT
Trader Memory Coretradermonty/claude-trading-skills3k2 repos~4.3kAutomated safety check: PassMIT
Author Migrationnrwl/nx29k—~12kAutomated safety check: NotesMIT

Similar skills

  • Ops Automation Agent

    mastra-ai/mastra

    Authoring playbook for building agents that automate recurring internal tasks — running scheduled workflows, syncing data between systems, posting notifications, processing inbound events, or…

    29k GitHub stars~1.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Vercel Incident Runbook

    jeremylongshore/tons-of-skills-marketplace

    Vercel incident response procedures with triage, instant rollback, and postmortem.

    2.8k GitHub stars~2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Incident Retrospective

    OpenHands/extensions

    Create an automation that drafts incident retrospectives. An agent skill from OpenHands/extensions.

    163 GitHub stars~922 tokensUpdated yesterday
    Product & Project ManagementAuto-check passed
  • Trader Memory Core

    tradermonty/claude-trading-skills

    Track investment theses across their lifecycle — from screening idea to closed position with postmortem.

    3k GitHub starsUsed in 2 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Author or scope a first-party Nx migration. An agent skill from nrwl/nx.

    29k GitHub stars~12k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Write Notes Like Deepseek

    czm15053/write-notes-like-deepseek

    A skill your agent uses when a change is non-trivial by DSH standards (behavior, architecture, cross-file contracts, process/tooling, testing strategy, or on-disk/wire/config formats), when choosing…

    508 GitHub stars~2k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed

More from EliasOulkadi/shokunin

All 49 skills in this repo
  • CI CD

    EliasOulkadi/shokunin

    Design CI/CD pipelines for GitHub Actions, GitLab CI, and CircleCI with matrix builds, test sharding, caching, Docker layer caching, OIDC auth, deployment strategies (rolling, blue-green, canary)…

    114 GitHub stars~3.4k tokensUpdated 6 days ago
    Auto-check: notes
  • Component Forge

    EliasOulkadi/shokunin

    Build production-grade components for React, Vue 3, and Svelte 5 with all states (loading, empty, error, success, idle), TypeScript strict, WCAG 2.2 accessibility, server components (RSC), and…

    114 GitHub stars~3.6k tokensUpdated 6 days ago
    Auto-check: notes
  • DB Admin

    EliasOulkadi/shokunin

    PostgreSQL database administration — backup/restore (pgdump, PITR, WAL archiving), health monitoring (connections, bloat, cache hit ratio, dead tuples), connection pooling (PgBouncer), replication…

    114 GitHub stars~2k tokensUpdated 6 days ago
    Auto-check: notes
  • DB Sculptor

    EliasOulkadi/shokunin

    Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…

    114 GitHub stars~3.1k tokensUpdated 6 days ago
    Auto-check: notes
  • Docker

    EliasOulkadi/shokunin

    Optimize Docker images with multi-stage builds, distroless bases, BuildKit cache mounts, multi-arch builds, compose watch, security hardening (non-root, seccomp, capabilities drop), and…

    114 GitHub stars~3.8k tokensUpdated 6 days ago
    Auto-check: notes
  • Error Handler

    EliasOulkadi/shokunin

    Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

    114 GitHub stars~3.6k tokensUpdated 6 days ago
    Auto-check: notes

Works with

Categories

Questions about Runbook Gen

What does Runbook Gen do?

Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates…. Runbook Gen is an agent skill from EliasOulkadi/shokunin. Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates, and blameless post-mortems with action items.

When should I use Runbook Gen?

Runbook Gen fits situations like: tasks that involve Runbooks and postmortems.

How do I install Runbook Gen in Claude Code?

Run `npx skills add EliasOulkadi/shokunin --skill runbook-gen -a claude-code`. Or copy the skill folder (.pack/skills/runbook-gen in EliasOulkadi/shokunin) into .claude/skills/runbook-gen in your project. Claude Code loads it when a task matches its description.

How do I install Runbook Gen in Codex?

Run `npx skills add EliasOulkadi/shokunin --skill runbook-gen -a codex`. Or copy the skill folder (.pack/skills/runbook-gen in EliasOulkadi/shokunin) into .agents/skills/runbook-gen in your project. Codex loads it when a task matches its description.

Can I use Runbook Gen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EliasOulkadi/shokunin --skill runbook-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/runbook-gen, .gemini/skills/runbook-gen, .github/skills/runbook-gen and .opencode/skills/runbook-gen in your project.

What does Runbook Gen need to run?

Going by SKILL.md and its folder, Runbook Gen needs the command-line tools its instructions call (kubectl). Compatibility (from SKILL.md): opencode.

Does Runbook Gen access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Runbook Gen safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Runbook Gen use?

Runbook Gen is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Runbook Gen use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Runbook Gen?

Skills that share tags, products or a category with Runbook Gen: Ops Automation Agent (mastra-ai/mastra, 29k stars), Vercel Incident Runbook (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Incident Retrospective (OpenHands/extensions, 163 stars) and Trader Memory Core (tradermonty/claude-trading-skills, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Runbook Gen?

EliasOulkadi (a GitHub user) maintains it in EliasOulkadi/shokunin, which has 114 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 5, 2026.

Source: EliasOulkadi/shokunin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.