Agent skill

Rescue

by Houseofmvps in Houseofmvps/ultraship

Production Incident Commander — diagnose and recover from production incidents.

MITAuto-check passedDevOps & Cloud

Install Rescue

skills CLI
$ npx skills add Houseofmvps/ultraship --skill rescue -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Houseofmvps/ultraship rescue --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Houseofmvps/ultraship.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rescue .claude/skills/rescue && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rescue
GitHub stars
123
Token cost
~2k tokens
SKILL.md length
852 words
Files
1
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Production Incident Commander — diagnose and recover from production incidents.

  • Works in 8 steps: Gather Context → Run Diagnostics → Triage → …
  • Something is broken in production
  • SKILL.md covers Severity Classification, Process and Key Principles
  • Calls node, git and railway

What it does

Rescue is an agent skill from Houseofmvps/ultraship. Production Incident Commander — diagnose and recover from production incidents. Use when something is broken in production, site is down, errors spiking, or user reports a critical bug.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Incident response. It works with Sentry. The repository describes itself as: "ULTRASHIP" Claude Code plugin — 39 skills, 33 tools, 11 agents for ship-ready workflows: planning, review, pentesting, safety guardrails, canary monitoring, SEO/AI-readiness… The licence is MIT.

When your agent uses it

  • Something is broken in production
  • User reports a critical bug

Example prompts

  • “/rescue”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Gather Context
  2. Run Diagnostics
  3. Triage
  4. Recovery Options
  5. Verify Recovery
  6. Communication
  7. Post-Mortem
  8. Prevention

What it can do on your machine

Read from SKILL.md and the folder at commit ed232cb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node
    • git
    • railway

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rescue loads about 2k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 852 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Houseofmvps/ultraship at commit ed232cb, republished under its MIT licence (© Houseofmvps). 852 words, ~1,961 tokens.

Download SKILL.mdSave it as .claude/skills/rescue/SKILL.md (or your agent's skills folder).
name
rescue
description
Production Incident Commander — diagnose and recover from production incidents. Use when something is broken in production, site is down, errors spiking, or user reports a critical bug.
argument-hint
<url-or-error-description>

Production Incident Commander

When production is down, every minute costs trust. This skill runs an incident like a principal SRE — fast triage, clear decision-making, structured recovery, and prevention so it never happens again.

Severity Classification

Before doing anything, classify the incident:

SeverityDefinitionResponse TimeExample
SEV-1Service completely down, all users affectedImmediatelySite returns 500, database unreachable
SEV-2Major feature broken, many users affectedWithin 15 minAuth broken, payments failing, data loss
SEV-3Minor feature broken, some users affectedWithin 1 hourOne API endpoint slow, email not sending
SEV-4Cosmetic or edge caseNext business dayUI glitch on one browser, non-critical error log

Severity determines urgency. SEV-1/2: restore first, investigate later. SEV-3/4: investigate first, then fix.

Process

Phase 1: Gather Context

Ask for:

  1. Production URL (if not already known)
  2. What's happening? (down, slow, errors, specific feature broken)
  3. When did it start? (narrows the commit search window)
  4. What changed recently? (deploy, config change, dependency update, traffic spike)

If the user is panicking, skip questions and use whatever info is available. Speed > completeness for SEV-1.

Phase 2: Run Diagnostics
bash
node ${CLAUDE_PLUGIN_ROOT}/tools/incident-commander.mjs <project-directory> --url=<production-url>

Parse the JSON output.

If a Sentry MCP server is connected (check your available tools — search for sentry tools), pull live production errors before guessing at code: list the most recent / most frequent issues since the incident window, read the top stack traces, and map each frame back to a file and line in this repo. A real stack trace from production beats inferring the culprit from recent commits. Use the actual error signature to narrow the suspect commit. If no Sentry server is connected, continue with the static diagnostics above (and mention that connecting Sentry would sharpen this step).

Phase 3: Triage

Present findings in order of urgency:

Site Status:

  • UP / DOWN / DEGRADED
  • Response time and status code
  • Health endpoint status
  • SSL certificate validity

Likely Culprit:

  • Most recent commit with significant changes
  • Files changed in that commit
  • When it was deployed
  • Correlation: did the issue start after this deploy?

Error Patterns Found:

  • Unhandled promises, missing error handlers
  • Environment variable issues (missing, placeholder values)
  • Database connection problems
  • Third-party service failures

Resource Issues:

  • Memory pressure signals (process.memoryUsage patterns in code)
  • Unbounded data growth (arrays that grow without cleanup)
  • Connection pool exhaustion (too many DB connections)
Phase 4: Recovery Options

Present in order of speed — for SEV-1/2, always recommend Option 1 first:

Option 1: Rollback (fastest — 2-5 min)

bash
git revert <culprit-hash> --no-edit && git push

This is almost always the right first move. Restore service, then investigate.

When NOT to rollback:

  • The rollback would cause data loss (destructive migration already ran)
  • The issue isn't in the latest deploy (pre-existing problem that suddenly surfaced)
  • The rollback is bigger than the fix (e.g., reverting 50 files when the fix is 1 line)

Option 2: Hot Fix (5-15 min) If the error pattern is clear and the fix is small:

  • Apply the fix using Edit tool
  • Run tests locally
  • Push the fix with a clear commit message: fix: [what was broken] — incident [date]
  • Verify with health check

Option 3: Traffic Management (immediate) If the issue is load-related:

  • Enable maintenance mode if available
  • Scale up infrastructure if possible (Railway: increase instance count)
  • Add rate limiting to affected endpoints
  • Redirect traffic away from broken feature

Option 4: Investigate Further If the cause isn't clear:

  • Check application logs (Railway: railway logs, Vercel: function logs)
  • Check database connectivity and query performance
  • Check third-party service status pages (Stripe, Resend, Supabase, etc.)
  • Check recent environment variable changes
  • Check if DNS/SSL certificate expired
Show full SKILL.md (272 more words)Show less
Phase 5: Verify Recovery

After applying a fix:

bash
node ${CLAUDE_PLUGIN_ROOT}/tools/health-check.mjs <production-url>

Confirm the site is back to healthy status. Check:

  • Status code 200
  • Response time within normal range
  • SSL still valid
  • Key functionality working (not just the homepage)
Phase 6: Communication

For SEV-1/2, the user needs to communicate with their users:

Status page update template:

[Investigating] We're aware of [issue description] and are actively working on a fix.
[Identified] We've identified the cause and are deploying a fix.
[Resolved] The issue has been resolved. [Brief explanation]. We apologize for the disruption.

If the user has a status page: help them post the update. If they don't: suggest setting up a simple one (Instatus, Betteruptime, or a static page).

Phase 7: Post-Mortem

Generate a post-mortem document from the incident-commander output:

markdown
# Incident Post-Mortem — [Date]

## Summary
- **What happened:** [One sentence]
- **Severity:** SEV-[N]
- **Duration:** [start time] to [end time] ([N] minutes)
- **Impact:** [Who was affected, what they experienced]
- **Root cause:** [One sentence]

## Timeline
| Time | Event |
|---|---|
| HH:MM | Issue detected (how: monitoring/user report/deploy) |
| HH:MM | Investigation started |
| HH:MM | Root cause identified |
| HH:MM | Fix deployed |
| HH:MM | Service restored |

## Root Cause Analysis
[Detailed explanation of what went wrong and why]

## What Went Well
- [Fast detection, quick recovery, etc.]

## What Went Wrong
- [Missed in review, no test coverage, no monitoring, etc.]

## Action Items
| Action | Priority | Owner | Deadline |
|---|---|---|---|
| Add test for this failure case | High | [user] | This week |
| Add monitoring for [pattern] | High | [user] | This week |
| Add pre-deploy check that would have caught this | Medium | [user] | This sprint |
| [Update runbook/docs] | Low | [user] | This month |

Save to docs/incidents/YYYY-MM-DD-incident.md.

Phase 8: Prevention

Based on the incident, suggest concrete preventive measures:

Immediate (today):

  • Add a test that reproduces the exact failure
  • Add the specific check to the /ship pre-deploy audit

This week:

  • Set up uptime monitoring (Betteruptime, UptimeRobot — free tiers available)
  • Add health check endpoint if one doesn't exist (/health or /api/health)
  • Set up error alerting (Sentry free tier, or a simple error webhook)

This month:

  • Add the failure pattern to code review checklist
  • Document the runbook for this type of incident
  • If this was a database issue: add connection pool monitoring
  • If this was a deployment issue: add canary deployments or staged rollouts

Key Principles

  • Speed over perfection. Restore service FIRST, investigate AFTER. Rollback is almost always the right first move.
  • No blame. Post-mortems are about systems, not people. "The deploy process didn't catch this" not "Developer X broke production."
  • Every incident is a gift. It reveals a gap in your system. The post-mortem action items are how you prevent the next incident.
  • Communicate early and often. Silence during an outage erodes trust faster than the outage itself.

© Houseofmvps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/rescue of Houseofmvps/ultraship.

Open the folder on GitHubat commit ed232cb

Compare with similar skills

Rescue next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rescue compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rescue this skillHouseofmvps/ultraship123—~2kAutomated safety check: PassMIT
Axiom SRE Investigatoropenclaw/clawhub9.5k—~7.1kAutomated safety check: PassMIT
Superset Incident Triagesuperset-sh/superset15k—~1kAutomated safety check: PassCustom licence
Sentry Incident Runbookjeremylongshore/tons-of-skills-marketplace2.8k—~3.7kAutomated safety check: PassMIT
Sentry Reliability Patternsjeremylongshore/tons-of-skills-marketplace2.8k—~1.7kAutomated safety check: PassMIT
Sentry Alert TunerLeoYeAI/openclaw-master-skills2.2k—~7.3kAutomated safety check: PassMIT

Similar skills

  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Sentry Incident Runbook

    jeremylongshore/tons-of-skills-marketplace

    Execute incident response procedures using Sentry error monitoring.

    2.8k GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Sentry Reliability Patterns

    jeremylongshore/tons-of-skills-marketplace

    Build reliable Sentry integrations with graceful degradation, circuit breakers, and offline queuing.

    2.8k GitHub stars~1.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Sentry Alert Tuner

    LeoYeAI/openclaw-master-skills

    Reduce Sentry alert fatigue by surgically tuning issue grouping, fingerprint rules, severity mapping, sample rates, before-send filters, sourcemap pipelines, and release-health gates.

    2.2k GitHub stars~7.3k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from Houseofmvps/ultraship

All 28 skills in this repo
  • Using Ultraship

    Houseofmvps/ultraship

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    123 GitHub stars~2.2k tokensUpdated 3 mo ago
    Auto-check passed
  • A11y

    Houseofmvps/ultraship

    Accessibility audit + auto-fix (WCAG 2.2 A/AA). An agent skill from Houseofmvps/ultraship.

    123 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check: notes
  • Architecture

    Houseofmvps/ultraship

    Living Architecture Map — auto-generate Mermaid diagrams of your codebase.

    123 GitHub stars~708 tokensUpdated 3 mo ago
    Auto-check: notes
  • Clone Patterns

    Houseofmvps/ultraship

    Learn From the Best — analyze patterns from any codebase and apply them to yours.

    123 GitHub stars~682 tokensUpdated 3 mo ago
    Auto-check: notes
  • Code Review

    Houseofmvps/ultraship

    Code review with principal-engineer-level depth. An agent skill from Houseofmvps/ultraship.

    123 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Compete

    Houseofmvps/ultraship

    Competitive X-Ray — analyze any competitor URL vs your site.

    123 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check: notes

Works with

Categories

Questions about Rescue

What does Rescue do?

Production Incident Commander — diagnose and recover from production incidents. Rescue is an agent skill from Houseofmvps/ultraship. Production Incident Commander — diagnose and recover from production incidents.

When should I use Rescue?

Rescue fits situations like: something is broken in production; user reports a critical bug.

How do I install Rescue in Claude Code?

Run `npx skills add Houseofmvps/ultraship --skill rescue -a claude-code`. Or copy the skill folder (skills/rescue in Houseofmvps/ultraship) into .claude/skills/rescue in your project. Claude Code loads it when a task matches its description.

How do I install Rescue in Codex?

Run `npx skills add Houseofmvps/ultraship --skill rescue -a codex`. Or copy the skill folder (skills/rescue in Houseofmvps/ultraship) into .agents/skills/rescue in your project. Codex loads it when a task matches its description.

Can I use Rescue in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Houseofmvps/ultraship --skill rescue -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rescue, .gemini/skills/rescue, .github/skills/rescue and .opencode/skills/rescue in your project.

What does Rescue need to run?

Going by SKILL.md and its folder, Rescue needs the command-line tools its instructions call (node, git and railway).

Does Rescue access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Rescue safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Rescue use?

Rescue is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rescue use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Rescue?

Skills that share tags, products or a category with Rescue: Axiom SRE Investigator (openclaw/clawhub, 9.5k stars), Superset Incident Triage (superset-sh/superset, 15k stars), Sentry Incident Runbook (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Sentry Reliability Patterns (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rescue?

Houseofmvps (a GitHub user) maintains it in Houseofmvps/ultraship, which has 123 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on July 8, 2026.

Source: Houseofmvps/ultraship on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.