Incident response — diagnose production issues, find root cause, propose fix with rollback.

MITAuto-check: notesDevOps & Cloud

Install Vigil Incident

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill vigil-incident -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace vigil-incident --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ai-agency/tonone/skills/vigil-incident .claude/skills/vigil-incident && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vigil-incident
GitHub stars
2.8k
Token cost
~1.5k tokens
SKILL.md length
636 words
Files
2
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Incident response — diagnose production issues, find root cause, propose fix with rollback.

  • Works in 8 steps: Detect Environment → Gather Symptoms → Read Logs → …
  • Asked about something is broken
  • SKILL.md covers Steps and Delivery
  • Calls git, gcloud and fly

What it does

Vigil Incident is an agent skill from jeremylongshore/tons-of-skills-marketplace. Incident response — diagnose production issues, find root cause, propose fix with rollback. Use when asked about "something is broken", "production issue", "why is this down", "incident", or "debug production".

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `.claude-plugin/plugin.json`).

It sits in DevOps & Cloud, covering Incident response and Root cause analysis. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Asked about something is broken
  • Production issue
  • Why is this down
  • Debug production

Example prompts

  • “something is broken”
  • “production issue”
  • “why is this down”
  • “/vigil-incident”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Detect Environment
  2. Gather Symptoms
  3. Read Logs
  4. Check Metrics
  5. Trace the Request Path
  6. Identify Root Cause
  7. Propose Fix and Rollback Plan
  8. Generate Postmortem Template

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • Glob
    • Grep
    • WebFetch
    • WebSearch
    • Task
    • TodoWrite

    …and 1 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • gcloud
    • fly
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, gcloud and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vigil Incident loads about 1.5k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 636 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 636 words, ~1,468 tokens.

Download SKILL.mdSave it as .claude/skills/vigil-incident/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vigil-incident
description
Incident response — diagnose production issues, find root cause, propose fix with rollback. Use when asked about "something is broken", "production issue", "why is this down", "incident", or "debug production".
allowed-tools
Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion
version
0.6.4
author
tonone-ai <hello@tonone.ai>
license
MIT

Incident Response

You are Vigil — the observability and reliability engineer from the Engineering Team.

Steps

Step 0: Detect Environment

Discover the project's infrastructure and observability stack:

  • Check deployment platform: fly.toml, app.yaml, Dockerfile, Kubernetes manifests, render.yaml, serverless configs
  • Check for logging: look for log configuration files, logging libraries in dependencies
  • Check for monitoring: Prometheus configs, Datadog agent, Cloud Monitoring setup, APM configs
  • Check for recent deployments: git log --oneline -20, CI/CD configs, deployment history
  • Check for existing runbooks: search docs for runbook, incident, playbook

Establish what tools are available for diagnosis before proceeding.

Step 1: Gather Symptoms

Collect the facts before diagnosing:

  • What's broken? — which service, endpoint, or functionality is affected
  • When did it start? — check deployment history, git log --since, recent config changes
  • What changed? — recent commits, deployments, config changes, dependency updates, infrastructure changes
  • What's the blast radius? — is it all users, some users, one region, one endpoint
  • Is it intermittent or constant? — this narrows the cause significantly

Ask the user for any symptoms they haven't shared. Don't guess — gather data.

Step 2: Read Logs

Search for errors in the available logging system:

  • Look for ERROR and WARN level logs in the timeframe the issue started
  • Search for stack traces, exception messages, timeout errors
  • Check for patterns: are errors correlated with specific endpoints, users, or regions
  • Look for upstream dependency errors: database connection failures, API timeouts, DNS resolution failures
  • Check for resource-related messages: OOM kills, CPU throttling, disk full, connection pool exhaustion

Use Grep and Read to search log files, or use platform-specific CLI commands (gcloud logging read, fly logs, kubectl logs) to fetch recent logs.

Step 3: Check Metrics

Look for anomalies in the timeframe:

  • Request rate: did traffic spike or drop suddenly
  • Error rate: when did 5xx errors start, what's the rate vs. baseline
  • Latency: did P50/P99 latency spike — this often precedes errors
  • Resources: CPU, memory, disk, connection count — is anything at capacity
  • Dependencies: are downstream services healthy, are database queries slow

If metrics are available via CLI or config files, check them. If dashboards exist, reference them.

Step 4: Trace the Request Path

Follow the failing request through the system:

  • Identify the entry point: which endpoint or service receives the failing request
  • Trace through each hop: load balancer → service → database/cache/API
  • At each hop, check: is the request arriving? Is it processed correctly? Is the response correct?
  • Find the exact point of failure: where does the request succeed upstream but fail downstream
  • If distributed tracing is available, use trace IDs to follow the exact path
Show full SKILL.md (219 more words)Show less
Step 5: Identify Root Cause

Based on evidence gathered, determine root cause:

  • Correlate the timeline: what changed just before the issue started
  • Distinguish between trigger and root cause — a deployment may be the trigger, but the root cause is what the deployment changed
  • Consider common causes: bad deploy, config change, dependency failure, resource exhaustion, traffic spike, data corruption
  • State your confidence level: confirmed (evidence proves it), likely (evidence strongly suggests it), possible (one of several hypotheses)
Step 6: Propose Fix and Rollback Plan

Provide a concrete fix:

  • Immediate mitigation: what to do right now to stop the bleeding (e.g., rollback, scale up, disable feature flag, redirect traffic)
  • Root cause fix: what code/config change addresses the underlying issue
  • Rollback plan: if the fix makes things worse, how to revert — include exact commands
  • Verification: how to confirm the fix worked — what metrics/logs to check
Step 7: Generate Postmortem Template

Follow the output format defined in docs/output-kit.md — 40-line CLI max, box-drawing skeleton, unified severity indicators, compressed prose.

Create a postmortem document:

markdown
# Incident Postmortem: [Title]

**Date:** [date]
**Duration:** [start time] — [resolution time]
**Severity:** [S1/S2/S3/S4]
**Author:** [name]

## Summary

[1-2 sentence summary of what happened and impact]

## Timeline

- [HH:MM] — [event]
- [HH:MM] — [event]

## Root Cause

[What actually broke and why]

## Impact

- **Users affected:** [number/percentage]
- **Duration:** [minutes]
- **Revenue impact:** [if applicable]

## Resolution

[What was done to fix it]

## What Went Well

- [thing that helped]

## What Went Poorly

- [thing that made it worse or slower to resolve]

## Action Items

- [ ] [preventive action] — owner: [name] — due: [date]
- [ ] [detective action] — owner: [name] — due: [date]
- [ ] [mitigative action] — owner: [name] — due: [date]

## Lessons Learned

[What the team should internalize from this incident]

Postmortems are blameless. Blame a person and you lose the truth.

Delivery

If output exceeds the 40-line CLI budget, invoke /atlas-report with the full findings. The HTML report is the output. CLI is the receipt — box header, one-line verdict, top 3 findings, and the report path. Never dump analysis to CLI.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/ai-agency/tonone/skills/vigil-incident of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • .claude-plugin/plugin.json

Open the folder on GitHubat commit cfae287

Compare with similar skills

Vigil Incident next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vigil Incident compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vigil Incident this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.5kAutomated safety check: NotesMIT
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0
UModel Root Cause Analysisalibaba/UnifiedModel415—~1.9kAutomated safety check: PassCustom licence
Axiom SRE Investigatoropenclaw/clawhub9.5k—~7.1kAutomated safety check: PassMIT
Incident Triage Harnessmadebyaris/advance-minimax-m3-cursor-rules126—~984Automated safety check: PassMIT
Broken API InterviewerPrepLabsAI/InterviewMentor112—~2.6kAutomated safety check: PassMIT

Similar skills

  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    415 GitHub stars~1.9k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed
  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Incident Triage Harness

    madebyaris/advance-minimax-m3-cursor-rules

    Walks an agent through an evidence-first incident investigation across logs, metrics, code and screenshots, from first symptom to the smallest safe mitigation.

    126 GitHub stars~984 tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • Broken API Interviewer

    PrepLabsAI/InterviewMentor

    An on-call SRE interviewer who just got paged about a broken checkout API.

    112 GitHub stars~2.6k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Vigil Incident

What does Vigil Incident do?

Incident response — diagnose production issues, find root cause, propose fix with rollback. Vigil Incident is an agent skill from jeremylongshore/tons-of-skills-marketplace. Incident response — diagnose production issues, find root cause, propose fix with rollback.

When should I use Vigil Incident?

Vigil Incident fits situations like: asked about something is broken; production issue; why is this down; debug production.

How do I install Vigil Incident in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vigil-incident -a claude-code`. Or copy the skill folder (plugins/ai-agency/tonone/skills/vigil-incident in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/vigil-incident in your project. Claude Code loads it when a task matches its description.

How do I install Vigil Incident in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vigil-incident -a codex`. Or copy the skill folder (plugins/ai-agency/tonone/skills/vigil-incident in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/vigil-incident in your project. Codex loads it when a task matches its description.

Can I use Vigil Incident in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vigil-incident -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vigil-incident, .gemini/skills/vigil-incident, .github/skills/vigil-incident and .opencode/skills/vigil-incident in your project.

What does Vigil Incident need to run?

Going by SKILL.md and its folder, Vigil Incident needs the command-line tools its instructions call (git, gcloud, fly and kubectl). Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion.

Does Vigil Incident access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vigil Incident safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Vigil Incident use?

Vigil Incident is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vigil Incident use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vigil Incident?

Skills that share tags, products or a category with Vigil Incident: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 415 stars), Axiom SRE Investigator (openclaw/clawhub, 9.5k stars) and Incident Triage Harness (madebyaris/advance-minimax-m3-cursor-rules, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vigil Incident?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.