Agent skill

Omh Live Incident Response

by rlaope in rlaope/oh-my-hermes

[omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary…

MITAuto-check passedDevOps & Cloud

Install Omh Live Incident Response

skills CLI
$ npx skills add rlaope/oh-my-hermes --skill omh-live-incident-response -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rlaope/oh-my-hermes omh-live-incident-response --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rlaope/oh-my-hermes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/omh-live-incident-response .claude/skills/omh-live-incident-response && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
omh-live-incident-response
GitHub stars
3.2k
Token cost
~2.7k tokens
SKILL.md length
1,421 words
Files
2 (incl. references)
Skills in repo
143
Repo updated
First seen
Licence
MIT

At a glance

[omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary…

  • The user says: live-incident-response
  • SKILL.md covers Why This Exists, Do Not Use When, Examples and Completion Checklist, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Live incident response

What it does

Omh Live Incident Response is an agent skill from rlaope/oh-my-hermes. [omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary mitigation, verified recovery, and the customer notice. Use when the user says: live-incident-response, live incident response, incident response, active incident, ongoing incident, open incident, incident open, incident commander.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/incident-command-method.md`).

It sits in DevOps & Cloud, covering Incident response. The repository describes itself as: All in one plugin for Hermes Agent ⚚ the coding intelligence, a long-term memory system and model optimized workflow packages. The licence is MIT.

When your agent uses it

  • The user says: live-incident-response
  • Live incident response
  • Incident response
  • Active incident

Example prompts

  • “/omh-live-incident-response”

What it can do on your machine

Read from SKILL.md and the folder at commit 7cd0d02. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Omh Live Incident Response loads about 2.7k tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 112 tokens; SKILL.md has 1,421 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rlaope/oh-my-hermes at commit 7cd0d02, republished under its MIT licence (© rlaope). 1,421 words, ~2,704 tokens.

Download SKILL.mdSave it as .claude/skills/omh-live-incident-response/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
omh-live-incident-response
description
[omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary mitigation, verified recovery, and the customer notice. Use when the user says: live-incident-response, live incident response, incident response, active incident, ongoing incident, open incident, incident open, incident commander.

Live Incident Response

This is a Hermes-native live-incident-response workflow skill.

Why This Exists

live-incident-response exists because an incident that is still open had no owner. support-operations sent an active incident to reliability-review, and reliability-review reviews incident notes after the fact, so the one skill that saw the request handed it to a postmortem while the outage was still running.

Do Not Use When

  • The incident is over and the request is the postmortem, the SLO or error-budget consequence, or remediation follow-up; use reliability-review.
  • The request is one customer's support case needing a reply, a severity opinion, and an escalation path, with no incident declared; use support-operations.
  • The request is a release being rolled out and watched -- deploy checklist, health signals, rollback criteria -- and nothing has been declared broken; use deploy-and-monitor.
  • The request is to send the page, publish the status-page update, or deliver the customer notice; use connector-operator, which records a send as observed only on a returned result.
  • The request asks whether a release is ready across rollout, rollback, and observability, before anything broke; use production-audit.

Examples

Good example:

  • Prompt: we have a production outage right now, declare severity and assign an incident commander
  • Expected behavior: Prepare live_incident_record/v1: ask for the user-visible symptom and blast radius, declare the severity with the observation that set it, name the commander and the remaining roles, open the append-only timeline, and state which signal decides recovery.
  • Why: The incident is open, so severity and command are live state rather than findings to review later.

Bad example:

  • Prompt: live-incident-response write up the postmortem for last week's outage and what it cost the error budget
  • Expected behavior: Route to reliability-review: a closed incident is reviewed, never commanded.
  • Why: Severity, roles, and a running timeline have no subject once the incident is over.

Completion Checklist

  • Severity is declared, carries the observation that set it, and every change appended rather than overwrote the previous level.
  • One commander is named; operations, communications, and scribe each name a person or read unfilled.
  • Every timeline entry is timestamped, attributed, and typed, and no earlier entry was edited.
  • Each mitigation reads temporary or permanent, and a temporary one names what removes it.
  • Recovery cites the named signal, its healthy value, the observed value, and the observer, never the mitigation alone.
  • Paging, status-page, and customer-send entries read prepared unless a connector result was observed.

Recovery Notes

  • If nobody is named commander, ask for one before anything else; an incident without a commander produces opinions instead of decisions.
  • If the recovery signal is not stated, ask which signal and which value counts as healthy before calling anything recovered.
  • If the incident turns out to be closed, hand the postmortem to reliability-review and leave this record as the timeline it reads.
  • If a connector call fails or returns nothing, keep the send prepared and name the channel that is unconfirmed instead of assuming delivery.

Workflow Lane

  • Current lane: Automation and status (achievements, workspace-audit, production-audit, live-incident-response, automation-blueprint, github-event-ops, github-issue-intake, buzz, +39 more) - schedules, status, health, and ops review.
  • If intent belongs to another lane, hand back to oh-my-hermes or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: omh-routing/references/skill-common-rail.md.

Use When

Use when an incident is open right now and the user needs it commanded: severity declared as live state, a commander and the other roles assigned, an append-only timeline kept, a temporary mitigation recorded as temporary, recovery verified against a named signal, and the customer notice drafted. The incident is still running; once it is closed the work is a review.

Strong routing signals: `live-incident-response`, `live incident response`, `incident response`, `active incident`, `ongoing incident`, `open incident`, `incident open`, `incident commander`, `incident command`, `incident bridge`, `incident channel`, `incident timeline`, `incident roles`, `declare severity`, `declare an incident`, `declare the incident`, `sev1`, `sev2`, `sev3`, `production outage`, `production is down`, `the site is down`, `service is down`, `we have an outage`, `outage right now`, `war room`, `stop the bleeding`, `temporary mitigation`, `page the on-call`, `page on-call`, `who is the incident commander`, `assign an incident commander`, `verify recovery`

Catalog Metadata

Category: reliability Phase: live-incident-command Hermes role: operator Quality tier: incident-command-gated Reasoning demand: standard

Quality bar:

  • Declare severity from the observed blast radius and record the observation that set it; an undeclared severity is not a severity, and a changed one appends rather than replaces.
  • Name one commander before anything else, then name or mark unfilled each of operations, communications, and scribe.
  • Give every timeline entry a timestamp, an actor, and a type -- observation, action, or decision -- per omh-live-incident-response/references/incident-command-method.md.
  • Separate a mitigation from a fix: say what was changed, whether it is temporary, and what removes it.
  • Verify recovery against the named signal at its healthy value; when that signal is unavailable the incident stays open and says so.
  • Keep paging, status-page updates, and customer sends listed as prepared until a connector result is observed.

Handoff policy:

Keep severity, role assignment, the timeline, mitigation records, and recovery verification in Hermes. Paging, status-page updates, and customer sends are connector-operator requests recorded as observed only when the connector returns a result; rollbacks, code changes, and infrastructure operations are executor or operator work and reach the timeline as observations, never as claims.

Required inputs:

  • what is broken right now, and the user-visible behavior that shows it
  • blast radius: which customers, tenants, or regions, and since when
  • who is available for commander, operations, communications, and scribe
  • the signal that decides recovery, and the value that counts as healthy
  • mitigation state so far: nothing tried, tried and failed, or in place
Show full SKILL.md (503 more words)Show less

Expert clarification questions:

  • who is available for commander, operations, communications, and scribe
    • English: Who is the incident commander right now, and who else is available to take operations, communications, and scribe?
    • Korean: 지금 인시던트 커맨더는 누구이고, 운영·커뮤니케이션·기록 역할을 맡을 수 있는 사람은 누구인가요?
  • the signal that decides recovery, and the value that counts as healthy
    • English: Which signal decides that this is recovered, and what value does it have to reach?
    • Korean: 어떤 신호로 복구를 판정하며, 그 값이 얼마가 되어야 정상인가요?

Expected outputs:

  • live_incident_record/v1
  • declared severity: the level, the observation that set it, and when it last changed
  • role assignment naming a person per role, or recording the role unfilled
  • append-only timeline: one entry per observation, action, or decision, each timestamped and attributed
  • mitigation entries marked temporary or permanent, each with what removes it
  • recovery verification: the signal, its healthy value, the observed value, and who observed it
  • customer notice draft plus the prepared paging and status-page requests, kept apart from observed sends

Artifact expectations:

  • live_incident_record/v1 with severity, roles, timeline, mitigations, recovery verification, and a communication ledger
  • every communication entry reads prepared or observed and never both; a correction appends an entry and never edits one

Safety rules:

  • Never rewrite or delete a timeline entry. A correction is a new entry naming the entry it corrects, because the timeline is what the review reads afterwards.
  • Do not claim a page was sent, a status page was updated, or a customer was notified; those are connector-operator requests, observed only when the connector returns a result.
  • Do not call the incident recovered because a mitigation was applied; recovery needs the named signal observed at its healthy value, with the observer recorded.
  • Never leave a temporary mitigation unmarked; record what it changed and what removes it, or it becomes permanent because nobody wrote it down.
  • Never print customer records, credentials, tokens, or connection strings pulled into the timeline as evidence.

Runtime Evidence

Record observed delegation results; otherwise return not_available or not_observed. Prepared OMH routing is not execution, review, CI, merge-readiness, or merge evidence.

  • Treat wrapper memory/context summaries as advisory local context, not proof of opaque Hermes memory reads or changes. Preserve workflow intent and stop conditions; verify before claiming completion. Reply in the user's own words and the host's own voice: its SOUL.md persona owns reply language, tone, speech level, and sentence endings, progress updates included (where it sets no language, use the one the user wrote in), and OMH shapes structure and content only; OMH's record terms (surface, lane, wrapper, handoff, evidence boundary, not_observed) stay in records and tool calls, never in the sentence the user reads unless they ask about one; and when a stop condition or a decision the user owns ends the turn, offer the next action as a question rather than declaring what will not be done.

Use Hermes-native subagent/delegation features when available: native subagents -> Hermes delegation when available, otherwise sequential lanes.

Shared product, compatibility, topology, memory, harness, and execution rules: omh-routing/references/skill-common-rail.md. Load it when applicable; otherwise name an unavailable capability.

© rlaope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/omh-live-incident-response of rlaope/oh-my-hermes.

  • SKILL.md
  • references/incident-command-method.md

Open the folder on GitHubat commit 7cd0d02

Compare with similar skills

Omh Live Incident Response next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Omh Live Incident Response compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Omh Live Incident Response this skillrlaope/oh-my-hermes3.2k—~2.7kAutomated safety check: PassMIT
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0
UModel Root Cause Analysisalibaba/UnifiedModel415—~1.9kAutomated safety check: PassCustom licence
Learningskortix-ai/suna20k—~1.1kAutomated safety check: PassCustom licence
Nix Config Debugryan4yin/nix-config2.1k—~1.2kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0

Similar skills

  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    415 GitHub stars~1.9k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed
  • Learnings

    kortix-ai/suna

    The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.

    20k GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Nix Config Debug

    ryan4yin/nix-config

    A skill your agent uses when something here is broken or stops working: an eval or build error, a failed activation, a dead or restarting unit, a mihomo or DNS outage, an unreachable host or MicroVM…

    2.1k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed

More from rlaope/oh-my-hermes

All 143 skills in this repo
  • Omh Accessibility Audit

    rlaope/oh-my-hermes

    [omh] Screen-reader or keyboard accessibility gaps: prepare WCAG, keyboard, focus, screen-reader, target-size, and reflow evidence gates for UI surfaces.

    3.3k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Omh Agent Debug

    rlaope/oh-my-hermes

    [omh] Agent is stuck, looping, or drifting: capture a stuck, looping, drifting, or repeatedly failing agent run, diagnose the likely failure pattern, and prepare the smallest safe recovery action.

    3.3k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Omh Agent Evaluation

    rlaope/oh-my-hermes

    [omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics.

    3.3k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Omh Agent Instructions

    rlaope/oh-my-hermes

    [omh] Agent instruction file for a repo -- AGENTS.md, CLAUDE.md, a Cursor rule: write or update what an agent cannot derive from the code, inside a marked region, with every command verified or…

    3.3k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Omh Agent Ops Review

    rlaope/oh-my-hermes

    [omh] AI agent progress for managers: help managers inspect AI-agent progress, blockers, quality gates, and throughput levers.

    3.3k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Omh AI Slop Cleaner

    rlaope/oh-my-hermes

    [omh] Messy or AI-generated code to clean up: delete AI-generated slop, dead code, and duplication while observable behavior stays identical.

    3.3k GitHub stars~2.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Omh Live Incident Response

What does Omh Live Incident Response do?

[omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary…. Omh Live Incident Response is an agent skill from rlaope/oh-my-hermes. [omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary mitigation, verified recovery, and the customer notice.

When should I use Omh Live Incident Response?

Omh Live Incident Response fits situations like: the user says: live-incident-response; live incident response; incident response; active incident.

How do I install Omh Live Incident Response in Claude Code?

Run `npx skills add rlaope/oh-my-hermes --skill omh-live-incident-response -a claude-code`. Or copy the skill folder (skills/omh-live-incident-response in rlaope/oh-my-hermes) into .claude/skills/omh-live-incident-response in your project. Claude Code loads it when a task matches its description.

How do I install Omh Live Incident Response in Codex?

Run `npx skills add rlaope/oh-my-hermes --skill omh-live-incident-response -a codex`. Or copy the skill folder (skills/omh-live-incident-response in rlaope/oh-my-hermes) into .agents/skills/omh-live-incident-response in your project. Codex loads it when a task matches its description.

Can I use Omh Live Incident Response in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rlaope/oh-my-hermes --skill omh-live-incident-response -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/omh-live-incident-response, .gemini/skills/omh-live-incident-response, .github/skills/omh-live-incident-response and .opencode/skills/omh-live-incident-response in your project.

What does Omh Live Incident Response need to run?

SKILL.md names no scripts, command-line tools or credentials: Omh Live Incident Response is instructions for the agent only.

Does Omh Live Incident Response access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Omh Live Incident Response safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Omh Live Incident Response use?

Omh Live Incident Response is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Omh Live Incident Response use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Omh Live Incident Response?

Skills that share tags, products or a category with Omh Live Incident Response: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 415 stars), Learnings (kortix-ai/suna, 20k stars) and Nix Config Debug (ryan4yin/nix-config, 2.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Omh Live Incident Response?

rlaope (a GitHub user) maintains it in rlaope/oh-my-hermes, which has 3,243 GitHub stars. The repository holds 143 skills in this directory. The repository was last updated on October 10, 2026.

Source: rlaope/oh-my-hermes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.