Agent skill

Incident Response

by agulli in agulli/atlas-agents

Diagnose and respond to production incidents. An agent skill from agulli/atlas-agents.

MITAuto-check passedDevOps & Cloud

Install Incident Response

skills CLI
$ npx skills add agulli/atlas-agents --skill incident-response -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agulli/atlas-agents incident-response --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agulli/atlas-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ch09_agent_skills/skills/incident-response .claude/skills/incident-response && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
incident-response
GitHub stars
579
Token cost
~587 tokens
SKILL.md length
297 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Diagnose and respond to production incidents. An agent skill from agulli/atlas-agents.

  • Works in 7 steps: Assess severity. Ask or determine → Gather signals. Before forming any… → Form ONE hypothesis. Based on the… → …
  • A service is down
  • SKILL.md covers Overview, Process, Rationalizations and Verification
  • Calls git

What it does

Incident Response is an agent skill from agulli/atlas-agents. Diagnose and respond to production incidents. Use when a service is down, errors are spiking, latency is degraded, or the user reports a production issue.

Its SKILL.md is about 590 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Incident response. The licence is MIT.

When your agent uses it

  • A service is down
  • Errors are spiking
  • Latency is degraded
  • The user reports a production issue

Example prompts

  • “/incident-response”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Assess severity. Ask or determine
  2. Gather signals. Before forming any hypothesis, collect
  3. Form ONE hypothesis. Based on the signals, state your best guess in one sentence. Do not enumerate multiple possibilities — pick the most…
  4. Test the hypothesis. Run exactly one diagnostic command or query that would confirm or refute your hypothesis. Read the output.
  5. If confirmed: Propose a fix. If the fix involves restarting a service or rolling back a deploy, state the exact command. Do not improvise…
  6. If refuted: Return to step 2 with the new information. Form a new hypothesis.
  7. Post-mortem. After the incident is resolved, write a brief post-mortem with

What it can do on your machine

Read from SKILL.md and the folder at commit 2b21998. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Incident Response loads about 587 tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 297 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~587

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agulli/atlas-agents at commit 2b21998, republished under its MIT licence (© agulli). 297 words, ~587 tokens.

Download SKILL.mdSave it as .claude/skills/incident-response/SKILL.md (or your agent's skills folder).
name
incident-response
description
Diagnose and respond to production incidents. Use when a service is down, errors are spiking, latency is degraded, or the user reports a production issue.
license
MIT

Overview

You are an on-call engineer triaging a live incident. Speed matters, but reckless changes make things worse. Follow the process.

Process

  1. Assess severity. Ask or determine:

    • Is the service fully down, partially degraded, or experiencing elevated errors?
    • How many users are affected?
    • Is data being lost or corrupted?
  2. Gather signals. Before forming any hypothesis, collect:

    • Recent deployments (git log --oneline -10)
    • Error logs (last 100 lines of the relevant log file)
    • Resource utilization (CPU, memory, disk, connections)
    • Recent configuration changes
  3. Form ONE hypothesis. Based on the signals, state your best guess in one sentence. Do not enumerate multiple possibilities — pick the most likely one.

  4. Test the hypothesis. Run exactly one diagnostic command or query that would confirm or refute your hypothesis. Read the output.

  5. If confirmed: Propose a fix. If the fix involves restarting a service or rolling back a deploy, state the exact command. Do not improvise commands.

  6. If refuted: Return to step 2 with the new information. Form a new hypothesis.

  7. Post-mortem. After the incident is resolved, write a brief post-mortem with:

    • Timeline (when it started, when it was detected, when it was resolved)
    • Root cause (one sentence)
    • Fix applied
    • Follow-up actions to prevent recurrence

Rationalizations

ExcuseRebuttal
"Let me just restart the service first"Restarting without diagnosis destroys evidence. Gather signals first.
"I have three theories"Pick one. Test it. If wrong, pick another. Parallel investigation wastes time.
"It's probably fine now"Confirm with metrics. "Probably" is not a resolution status.
"We can skip the post-mortem, it was minor"Minor incidents reveal systemic issues. Write the post-mortem.

Verification

  • Signals were gathered before any remediation was attempted
  • The root cause was identified (not assumed)
  • Service health was confirmed after the fix (not assumed)
  • A post-mortem was written

© agulli, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in ch09_agent_skills/skills/incident-response of agulli/atlas-agents.

Open the folder on GitHubat commit 2b21998

Compare with similar skills

Incident Response next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Incident Response compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Incident Response this skillagulli/atlas-agents579—~587Automated safety check: PassMIT
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0
Nix Config Debugryan4yin/nix-config2.1k—~1.2kAutomated safety check: PassMIT
UModel Root Cause Analysisalibaba/UnifiedModel415—~1.9kAutomated safety check: PassCustom licence
Learningskortix-ai/suna20k—~1.1kAutomated safety check: PassCustom licence
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0

Similar skills

  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Nix Config Debug

    ryan4yin/nix-config

    A skill your agent uses when something here is broken or stops working: an eval or build error, a failed activation, a dead or restarting unit, a mihomo or DNS outage, an unreachable host or MicroVM…

    2.1k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    415 GitHub stars~1.9k tokensUpdated 16 days ago
    DevOps & CloudAuto-check passed
  • Learnings

    kortix-ai/suna

    The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.

    20k GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Loop Triage Report

    cobusgreyling/loop-engineering

    Turns CI failures, open issues, recent commits and chat threads into a prioritized markdown report that an automation loop can act on without inventing architecture work.

    11k GitHub stars~500 tokensUpdated today
    DevOps & CloudAuto-check passed

More from agulli/atlas-agents

All 12 skills in this repo
  • API Design

    agulli/atlas-agents

    Design or review REST and GraphQL API interfaces. An agent skill from agulli/atlas-agents.

    579 GitHub stars~539 tokensUpdated 2 mo ago
    Auto-check passed
  • Data Pipeline

    agulli/atlas-agents

    Design, build, or debug data processing pipelines. An agent skill from agulli/atlas-agents.

    579 GitHub stars~714 tokensUpdated 2 mo ago
    Auto-check passed
  • Database Migration

    agulli/atlas-agents

    Safely run database schema migrations. An agent skill from agulli/atlas-agents.

    579 GitHub stars~702 tokensUpdated 2 mo ago
    Auto-check passed
  • Deploy Checklist

    agulli/atlas-agents

    Execute a structured deployment to staging or production. An agent skill from agulli/atlas-agents.

    579 GitHub stars~756 tokensUpdated 2 mo ago
    Auto-check passed
  • Documentation Writer

    agulli/atlas-agents

    Write or update technical documentation for code, APIs, or systems.

    579 GitHub stars~639 tokensUpdated 2 mo ago
    Auto-check passed
  • Git Commit

    agulli/atlas-agents

    Create well-structured git commits with conventional commit messages.

    579 GitHub stars~463 tokensUpdated 2 mo ago
    Auto-check: notes

Categories

Questions about Incident Response

What does Incident Response do?

Diagnose and respond to production incidents. An agent skill from agulli/atlas-agents. Incident Response is an agent skill from agulli/atlas-agents. Diagnose and respond to production incidents.

When should I use Incident Response?

Incident Response fits situations like: A service is down; errors are spiking; latency is degraded; the user reports a production issue.

How do I install Incident Response in Claude Code?

Run `npx skills add agulli/atlas-agents --skill incident-response -a claude-code`. Or copy the skill folder (ch09_agent_skills/skills/incident-response in agulli/atlas-agents) into .claude/skills/incident-response in your project. Claude Code loads it when a task matches its description.

How do I install Incident Response in Codex?

Run `npx skills add agulli/atlas-agents --skill incident-response -a codex`. Or copy the skill folder (ch09_agent_skills/skills/incident-response in agulli/atlas-agents) into .agents/skills/incident-response in your project. Codex loads it when a task matches its description.

Can I use Incident Response in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agulli/atlas-agents --skill incident-response -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-response, .gemini/skills/incident-response, .github/skills/incident-response and .opencode/skills/incident-response in your project.

What does Incident Response need to run?

Going by SKILL.md and its folder, Incident Response needs the command-line tools its instructions call (git).

Does Incident Response access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Incident Response safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Incident Response use?

Incident Response is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Incident Response use?

About 587 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Incident Response?

Skills that share tags, products or a category with Incident Response: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), Nix Config Debug (ryan4yin/nix-config, 2.1k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 415 stars) and Learnings (kortix-ai/suna, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Incident Response?

agulli (a GitHub user) maintains it in agulli/atlas-agents, which has 579 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on July 17, 2026.

Source: agulli/atlas-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.