Ops Automation Agent
mastra-ai/mastra
Authoring playbook for building agents that automate recurring internal tasks — running scheduled workflows, syncing data between systems, posting notifications, processing inbound events, or…
Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates…
$ npx skills add EliasOulkadi/shokunin --skill runbook-gen -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install EliasOulkadi/shokunin runbook-gen --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.pack/skills/runbook-gen .claude/skills/runbook-gen && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "runbook-gen" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/runbook-gen into .claude/skills/runbook-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runbook-gen", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/runbook-genType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add EliasOulkadi/shokunin --skill runbook-gen -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install EliasOulkadi/shokunin runbook-gen --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.pack/skills/runbook-gen .agents/skills/runbook-gen && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "runbook-gen" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/runbook-gen into .agents/skills/runbook-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runbook-gen", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add EliasOulkadi/shokunin --skill runbook-gen -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install EliasOulkadi/shokunin runbook-gen --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.pack/skills/runbook-gen .cursor/skills/runbook-gen && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "runbook-gen" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/runbook-gen into .cursor/skills/runbook-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runbook-gen", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/EliasOulkadi/shokunin.git --path .pack/skills/runbook-gen--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add EliasOulkadi/shokunin --skill runbook-gen -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install EliasOulkadi/shokunin runbook-gen --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.pack/skills/runbook-gen .gemini/skills/runbook-gen && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "runbook-gen" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/runbook-gen into .gemini/skills/runbook-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runbook-gen", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install EliasOulkadi/shokunin runbook-genInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add EliasOulkadi/shokunin --skill runbook-gen -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .github/skills && cp -r skills-src/.pack/skills/runbook-gen .github/skills/runbook-gen && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "runbook-gen" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/runbook-gen into .github/skills/runbook-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runbook-gen", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add EliasOulkadi/shokunin --skill runbook-gen -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install EliasOulkadi/shokunin runbook-gen --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.pack/skills/runbook-gen .opencode/skills/runbook-gen && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "runbook-gen" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/runbook-gen into .opencode/skills/runbook-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runbook-gen", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
runbook-genGenerate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates…
Runbook Gen is an agent skill from EliasOulkadi/shokunin. Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates, and blameless post-mortems with action items.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: opencode
It sits in DevOps & Cloud, covering Runbooks and postmortems. It works with Slack. The repository describes itself as: 職人 Shokunin 62 AI agent skills for OpenCode, Claude Code, Cursor, Windsurf. ChromaDB memory, MCP servers, declarative self-updates. Multi-model, open source, zero cost. The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4c68e5b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
opencode
From compatibility in the SKILL.md frontmatter.
Runbook Gen loads about 2.8k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 511 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from EliasOulkadi/shokunin at commit 4c68e5b, republished under its MIT licence (© EliasOulkadi). 511 words, ~2,772 tokens.
.claude/skills/runbook-gen/SKILL.md (or your agent's skills folder).Operation runbooks that on-call engineers can follow under pressure. Based on Google SRE practices, PagerDuty incident response, and post-mortem culture.
Follow these steps in order when generating a runbook:
When the ask is a post-mortem (not a runbook), skip steps 3-4 and use the Post-Mortem Template instead.
| Severity | Definition | Response SLA |
|---|---|---|
| Sev1 | Service down, users impacted | Respond within 15 min |
| Sev2 | Degraded performance, partial impact | Respond within 1 hour |
| Sev3 | Minor issue, no user impact | Next business day |
| Sev4 | Internal tooling, non-critical | Per team schedule |
## Runbook: [Incident Type]
### Detection
How this incident is typically discovered:
- Alert: [Prometheus/Grafana/Datadog alert]
- User symptom: [what users see or report]
- Automated detection: [auto-remediation]
### Initial Response (first 5 min)
1. Acknowledge alert (PagerDuty/Opsgenie). Time: < 2 min
2. Determine severity. Time: < 1 min
3. Assign owner. Create incident channel (#inc-sev1). Time: < 2 min
4. Post initial status:
"Investigating [issue] affecting [scope]. Will update in 15 min."
5. Start diagnosis timer. Time: < 1 min
### Diagnosis (decision tree)
1. Check [primary dashboard]: expected [X], current [Y]
2. Check logs: `kubectl logs -n [ns] -l app=[service] --tail 200`
3. Check database health: `SELECT count(*) FROM pg_stat_activity`
4. Check cache/queue latency: `redis-cli --latency`
5. IF [error in logs] → Runbook A
6. IF [latency > threshold] → Runbook B
7. IF unknown → Escalate
### Resolution Procedures
Each time-boxed with exact commands and verification:
#### Runbook A: [Name] (10 min)
```bash
# Step 1 (2 min)
kubectl rollout restart deployment/[service]
# Verify (1 min)
kubectl rollout status deployment/[service]| Timebox | Action | Contact |
|---|---|---|
| 0-15 min | Primary on-call | @name / phone |
| 15-30 min | Secondary on-call | @name / phone |
| 30-60 min | Engineering manager | @name |
| 60+ min | VP Engineering | @name |
## 5. War Room Setup
### Communications
- **Slack channel**: `#inc-sev1-[incident-name]`
- **Zoom bridge**: [link] (permanent war room)
- **Shared doc**: Google Doc or Notion page for live notes
- **Status page**: Update immediately on detection
- **Customer comms**: Template below
### Status Page Updates
Investigating: We're aware of [issue] affecting [scope]. Investigating root cause.
Monitoring: Deployed fix for [root cause]. Monitoring closely.
Resolved: [Issue] has been resolved. All systems operational.
### Customer Communication
Subject: [Service] incident — [date]
We experienced [description] from [start] to [end] ([duration]).
Root cause: [one sentence]
Impact: [specific metrics]
What we're doing:
We apologize for the disruption.
## 6. Post-Mortem Template
Severity: Sev[1-4] Duration: [detection] → [resolution] ([total]) Impact: [users/customers] for time
One paragraph. System-focused, not person-focused.
| Action | Owner | Due | Type |
|---|---|---|---|
| [action] | @name | Date | prevent/detect/respond |
## Error Handling
| Scenario | Behaviour | Guidance |
|----------|-----------|----------|
| User doesn't know the service/platform | Ask clarifying questions one at a time (not a list) | Start with "what service is this for?" |
| User says "just give me a template" | Output a blank Runbook Template with placeholder brackets | Skip Discovery, fill later |
| User asks for both runbook + post-mortem at once | Generate runbook first, then offer post-mortem | "I'll write the runbook now. Do you have an incident in mind for the post-mortem?" |
| Service has no existing monitoring | Flag the gap, offer to add basic health-check instructions | Note alerts must be configured separately |
| Runbook for a third-party/SaaS dependency | Include vendor status page check; mark resolution as "vendor fixes" | Focus on detection + escalation, not resolution |
| User says "I don't know the escalation contacts" | Use generic roles (Primary on-call, Secondary, EM, VP Eng) | Leave bracketed placeholders |
| Multiple teams involved | Add a RACI section to the runbook | Clarify who decides vs who executes |
| Runbook already exists, user wants update | Treat as revision: read current, diff, propose changes | Flag deprecations, test outdated commands |
## 8. Production Checklist
Before marking a runbook complete, verify every item:
- [ ] **Commands are copy-paste ready** — no unsubstituted variables (`[ns]` only in section headers/notes, never in commands)
- [ ] **Every step has a time-box** — engineer knows when to escalate
- [ ] **Decision tree is shallow** — max 3-4 levels; deeper means split into sub-runbooks
- [ ] **Multiple resolution paths** — at least 2 common failure modes covered
- [ ] **One runbook per incident type** — don't combine "DB failover" and "deployment rollback"
- [ ] **Verification step after every resolution** — how to confirm the fix worked
- [ ] **Escalation contacts listed** — primary + backup, both Slack and phone
- [ ] **Status page templates included** — Investigating / Monitoring / Resolved
- [ ] **Runbook tested in a drill** — within the current quarter, not when incident strikes
- [ ] **README section added** — link to runbook from team's on-call repo
- [ ] **Secrets checked** — no passwords, API keys, or tokens in commands; use env vars
- [ ] **Reviewed by secondary on-call** — fresh eyes catch assumptions
## Anti-Patterns
| Anti-pattern | Why It Fails | Fix |
|-------------|-------------|-----|
| Steps that say "fix the issue" | Too vague under pressure | Tell HOW, not what — exact commands |
| Commands with unsubstituted variables | Copy-paste fails, engineer wastes time retyping | Every command must run as-is |
| No verification step | Fix might not work, but nobody knows | Add check after every resolution |
| Runbooks >6 months stale | Commands rot, trust erodes, people ignore them | Schedule quarterly review in team calendar |
| Runbooks nobody tested | First test happens during the actual incident | Require one drill per quarter per runbook |
| Too many steps (>15) | Cognitive overload during Sev1 | Split into sub-runbooks, keep shallow |
| No time-boxes | Engineer doesn't know when to escalate | Every diagnosis/resolution step has a max time |
| Single resolution path | Assumes first idea works | Always have Plan B in the same runbook |
| Post-mortem blames people | Blame culture kills incident reporting | "What broke the system?" not "Who broke it?" |
| Skipping customer comms | Stakeholders hear "we're down" from social media | Draft customer email as part of template |
## Incident Timeline Template
[HH:MM] Alert triggered: <alert name> from <monitoring system>
[HH:MM] On-call acknowledged: <name>
[HH:MM] War room opened: <Slack channel/Zoom link>
[HH:MM] Initial diagnosis: <symptoms observed>
[HH:MM] Impact confirmed: <users affected, services degraded>
[HH:MM] Mitigation applied: <action taken>
[HH:MM] Monitoring confirms recovery
[HH:MM] All-clear declared
[HH:MM+1] Post-mortem scheduled: <date>
## Post-Mortem Template
<title>Date: YYYY-MM-DD | Duration: Xh Ym | Severity: Sev1/2/3
[HH:MM] ...
What system/process failure caused this.
Users affected, data loss, revenue impact.
How was this found? (alert, user report, manual)
What fixed it.
| # | Action | Owner | Due |
|---|
What would prevent this next time.
## Sources
- Google SRE Book: "Monitoring Distributed Systems"
- Google SRE Workbook: Incident Response
- PagerDuty incident response documentation
- Atlassian post-mortem best practices
- Microsoft SRE practices
- AWS Well-Architected Framework: Operational Excellence
## Checklist
- [ ] Skill loads without errors in the AI agent
- [ ] YAML frontmatter is valid (description, compatibility, audience)
- [ ] Workflow section provides clear step-by-step instructions
- [ ] Error handling section covers common failure modes
- [ ] All referenced files (references/, scripts/, assets/) exist
- [ ] Skill triggers correctly for intended use cases
- [ ] No broken links or missing resources© EliasOulkadi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .pack/skills/runbook-gen of EliasOulkadi/shokunin.
Open the folder on GitHubat commit 4c68e5b
Runbook Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Runbook Gen this skillEliasOulkadi/shokunin | 114 | — | ~2.8k | Automated safety check: Pass | MIT | |
| Ops Automation Agentmastra-ai/mastra | 29k | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Vercel Incident Runbookjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~2k | Automated safety check: Pass | MIT | |
| Incident RetrospectiveOpenHands/extensions | 163 | — | ~922 | Automated safety check: Pass | MIT | |
| Trader Memory Coretradermonty/claude-trading-skills | 3k | 2 repos | ~4.3k | Automated safety check: Pass | MIT | |
| Author Migrationnrwl/nx | 29k | — | ~12k | Automated safety check: Notes | MIT |
mastra-ai/mastra
Authoring playbook for building agents that automate recurring internal tasks — running scheduled workflows, syncing data between systems, posting notifications, processing inbound events, or…
jeremylongshore/tons-of-skills-marketplace
Vercel incident response procedures with triage, instant rollback, and postmortem.
OpenHands/extensions
Create an automation that drafts incident retrospectives. An agent skill from OpenHands/extensions.
tradermonty/claude-trading-skills
Track investment theses across their lifecycle — from screening idea to closed position with postmortem.
nrwl/nx
Author or scope a first-party Nx migration. An agent skill from nrwl/nx.
czm15053/write-notes-like-deepseek
A skill your agent uses when a change is non-trivial by DSH standards (behavior, architecture, cross-file contracts, process/tooling, testing strategy, or on-disk/wire/config formats), when choosing…
EliasOulkadi/shokunin
Design CI/CD pipelines for GitHub Actions, GitLab CI, and CircleCI with matrix builds, test sharding, caching, Docker layer caching, OIDC auth, deployment strategies (rolling, blue-green, canary)…
EliasOulkadi/shokunin
Build production-grade components for React, Vue 3, and Svelte 5 with all states (loading, empty, error, success, idle), TypeScript strict, WCAG 2.2 accessibility, server components (RSC), and…
EliasOulkadi/shokunin
PostgreSQL database administration — backup/restore (pgdump, PITR, WAL archiving), health monitoring (connections, bloat, cache hit ratio, dead tuples), connection pooling (PgBouncer), replication…
EliasOulkadi/shokunin
Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…
EliasOulkadi/shokunin
Optimize Docker images with multi-stage builds, distroless bases, BuildKit cache mounts, multi-arch builds, compose watch, security hardening (non-root, seccomp, capabilities drop), and…
EliasOulkadi/shokunin
Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…
Works with
Categories
Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates…. Runbook Gen is an agent skill from EliasOulkadi/shokunin. Generate operations runbooks and post-mortems for incident response — severity matrix, decision trees, escalation paths, war room setup (Slack/Zoom), status page updates, customer comms templates, and blameless post-mortems with action items.
Runbook Gen fits situations like: tasks that involve Runbooks and postmortems.
Run `npx skills add EliasOulkadi/shokunin --skill runbook-gen -a claude-code`. Or copy the skill folder (.pack/skills/runbook-gen in EliasOulkadi/shokunin) into .claude/skills/runbook-gen in your project. Claude Code loads it when a task matches its description.
Run `npx skills add EliasOulkadi/shokunin --skill runbook-gen -a codex`. Or copy the skill folder (.pack/skills/runbook-gen in EliasOulkadi/shokunin) into .agents/skills/runbook-gen in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EliasOulkadi/shokunin --skill runbook-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/runbook-gen, .gemini/skills/runbook-gen, .github/skills/runbook-gen and .opencode/skills/runbook-gen in your project.
Going by SKILL.md and its folder, Runbook Gen needs the command-line tools its instructions call (kubectl). Compatibility (from SKILL.md): opencode.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Runbook Gen is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Runbook Gen: Ops Automation Agent (mastra-ai/mastra, 29k stars), Vercel Incident Runbook (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Incident Retrospective (OpenHands/extensions, 163 stars) and Trader Memory Core (tradermonty/claude-trading-skills, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
EliasOulkadi (a GitHub user) maintains it in EliasOulkadi/shokunin, which has 114 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 5, 2026.
Source: EliasOulkadi/shokunin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.