Agent skill

Agent Incident Postmortem

by mohitagw15856 in mohitagw15856/pm-claude-skills

Run a blameless postmortem for an incident caused by an AI agent or LLM feature — hallucinated facts shipped to users, runaway tool use, prompt injection, cost blowouts, or wrong actions taken…

MITAuto-check passedDevOps & Cloud

Install Agent Incident Postmortem

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill agent-incident-postmortem -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills agent-incident-postmortem --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-incident-postmortem .claude/skills/agent-incident-postmortem && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-incident-postmortem
GitHub stars
1.4k
Token cost
~1.4k tokens
SKILL.md length
687 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Run a blameless postmortem for an incident caused by an AI agent or LLM feature — hallucinated facts shipped to users, runaway tool use, prompt injection, cost blowouts, or wrong actions taken…

  • Asked to write up an AI incident
  • SKILL.md covers What This Skill Produces, Required Inputs, Root-Cause Layers and Nondeterminism Discipline, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Analyse why an agent did something wrong

What it does

Agent Incident Postmortem is an agent skill from mohitagw15856/pm-claude-skills. Run a blameless postmortem for an incident caused by an AI agent or LLM feature — hallucinated facts shipped to users, runaway tool use, prompt injection, cost blowouts, or wrong actions taken autonomously. Use when asked to write up an AI incident, analyse why an agent did something wrong, or produce corrective actions after an LLM failure. Produces a structured postmortem with trace reconstruction, a root-cause layer analysis, and corrective actions including a permanent regression case. For non-AI production…

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Runbooks and postmortems. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to write up an AI incident
  • Analyse why an agent did something wrong
  • Produce corrective actions after an LLM failure

Example prompts

  • “/agent-incident-postmortem”

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Incident Postmortem loads about 1.4k tokens when it runs. Until then it costs about 144 tokens; SKILL.md has 687 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 687 words, ~1,382 tokens.

Download SKILL.mdSave it as .claude/skills/agent-incident-postmortem/SKILL.md (or your agent's skills folder).
name
agent-incident-postmortem
description
Run a blameless postmortem for an incident caused by an AI agent or LLM feature — hallucinated facts shipped to users, runaway tool use, prompt injection, cost blowouts, or wrong actions taken autonomously. Use when asked to write up an AI incident, analyse why an agent did something wrong, or produce corrective actions after an LLM failure. Produces a structured postmortem with trace reconstruction, a root-cause layer analysis, and corrective actions including a permanent regression case. For non-AI production incidents use incident-postmortem.

Agent Incident Postmortem Skill

AI incidents differ from outages: the system didn't go down — it did something wrong, confidently, and maybe only once. This skill adapts blameless postmortem practice to nondeterministic systems, where "can we reproduce it?" needs traces, not just steps.

What This Skill Produces

  • A blameless postmortem document with timeline and user/business impact
  • A trace reconstruction of what the agent saw, decided, and did
  • A root-cause analysis across the AI failure layers (not "the model hallucinated" as a conclusion)
  • Corrective actions — always including a new permanent case in the regression suite

Required Inputs

Ask for (if not already provided):

  • What the agent did and what it should have done
  • The trace — the full request: system prompt, context, tool calls and results, output. If no trace exists, that absence is itself a finding
  • Blast radius — how many users/requests, over what window, and whether it's ongoing
  • Detection — how it was noticed (user report? monitor? luck?) and how long after it started

Root-Cause Layers

Walk the layers in order; the root cause is usually the earliest layer that could have prevented the outcome. "The model was wrong" is a starting point, never the conclusion — models are known to be fallible, so the question is what let a fallible output become an incident.

LayerAsk
Input / contextWas the context wrong, stale, contradictory, or poisoned (injection)? Did retrieval feed it bad ground truth?
Model behaviourGiven that context, was the output a foreseeable failure mode (fabrication under missing data, over-compliance with injected text)?
GuardrailsWhat check should have caught this output and didn't exist / didn't fire? (schema validation, groundedness check, action allow-list)
Action layerWhy could the wrong output become a real action or reach a user without the appropriate gate for its risk level?
DetectionWhy did we learn about it this way, this late? What signal would have caught it in minutes?

Nondeterminism Discipline

  • Reproduce with the trace, not the anecdote: replay the exact context; then re-run N times to measure frequency — a 1-in-20 failure at 10k requests/day is 500 incidents/day.
  • Pin everything when replaying: model version, prompt version, temperature, tool results.
  • If it can't be reproduced: say so, keep the trace as the evidence, and treat frequency as unknown — not as "rare".

Output Format

Show full SKILL.md (312 more words)Show less
AI Incident Postmortem: [title] — [date]

Severity: [level] · Status: [resolved/monitoring] · Owner: [name]

Summary: [3 sentences: what the agent did, impact, root cause layer]

Impact: [users/requests affected, window, cost, trust/regulatory dimension]

Timeline: [first bad output → detection → mitigation → resolution, with the detection gap called out]

Trace reconstruction: [what was in the window; which tool calls ran; where the path diverged from intended behaviour]

Root cause by layer:

LayerFinding
Input/context
Model behaviour
Guardrails
Action layer
Detection

Reproduction: [replayed? failure frequency over N runs / not reproducible — evidence is the trace]

Corrective actions:

ActionLayerOwnerDue
Add this trace as a permanent regression caseeval
[guardrail/monitor/context fix]

What went well / what got lucky: [both, honestly]

Quality Checks

  • The postmortem is blameless toward humans and useful about the system — "prompt engineer error" and "model hallucinated" are both banned conclusions
  • Root cause identifies the earliest layer that could have prevented impact, not just the layer that misbehaved
  • The trace (or its absence) is in the document; findings cite it
  • Failure frequency was measured or explicitly marked unknown
  • Corrective actions include the permanent regression case and at least one detection improvement

Anti-Patterns

  • Do not close with "improved the prompt" as the only action — the same class of output must also be caught by a guardrail or gate next time
  • Do not assess frequency from one replay — nondeterministic failures hide at low temperatures and reappear at scale
  • Do not skip the injection question when any untrusted text (web, user docs, tickets) was in the window
  • Do not let "the model will be better next version" close an action item — upgrades are migrations (see model-migration-plan), not fixes
  • Do not write it as an outage report — the system was up; the failure was behavioural, and the doc must analyse behaviour

Example Trigger Phrases

  • "Write up an AI incident."
  • "Analyse why an agent did something wrong."
  • "Produce corrective actions after an LLM failure."

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agent-incident-postmortem of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Agent Incident Postmortem next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Incident Postmortem compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Incident Postmortem this skillmohitagw15856/pm-claude-skills1.4k—~1.4kAutomated safety check: PassMIT
Trader Memory Coretradermonty/claude-trading-skills3k2 repos~4.3kAutomated safety check: PassMIT
Author Migrationnrwl/nx29k—~12kAutomated safety check: NotesMIT
Write Notes Like Deepseekczm15053/write-notes-like-deepseek497—~2kAutomated safety check: PassNone
OpenRig Upgrade Proceduremvschwarz/openrig6.2k—~2.9kAutomated safety check: PassApache-2.0
GreptimeDB Release RunbookGreptimeTeam/greptimedb6.7k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Trader Memory Core

    tradermonty/claude-trading-skills

    Track investment theses across their lifecycle — from screening idea to closed position with postmortem.

    3k GitHub starsUsed in 2 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Author or scope a first-party Nx migration. An agent skill from nrwl/nx.

    29k GitHub stars~12k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Write Notes Like Deepseek

    czm15053/write-notes-like-deepseek

    A skill your agent uses when a change is non-trivial by DSH standards (behavior, architecture, cross-file contracts, process/tooling, testing strategy, or on-disk/wire/config formats), when choosing…

    497 GitHub stars~2k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • OpenRig Upgrade Procedure

    mvschwarz/openrig

    Walks an agent through upgrading the OpenRig CLI and daemon one observed step at a time, keeping live seats alive and reconciling managed plugin files.

    6.2k GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • GreptimeDB Release Runbook

    GreptimeTeam/greptimedb

    Runbook for publishing a GreptimeDB version: pick the release branch, verify the Cargo version, then tag, create the GitHub release and open the docs note PR.

    6.7k GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Statem

    henryqin1997/statem

    A skill your agent uses when a long coding or research task should be managed with statem state-machine runbooks, including creating specs, starting or resuming runs, checking current state…

    1.3k GitHub stars~1.2k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Agent Incident Postmortem

What does Agent Incident Postmortem do?

Run a blameless postmortem for an incident caused by an AI agent or LLM feature — hallucinated facts shipped to users, runaway tool use, prompt injection, cost blowouts, or wrong actions taken…. Agent Incident Postmortem is an agent skill from mohitagw15856/pm-claude-skills. Run a blameless postmortem for an incident caused by an AI agent or LLM feature — hallucinated facts shipped to users, runaway tool use, prompt injection, cost blowouts, or wrong actions taken autonomously.

When should I use Agent Incident Postmortem?

Agent Incident Postmortem fits situations like: asked to write up an AI incident; analyse why an agent did something wrong; produce corrective actions after an LLM failure.

How do I install Agent Incident Postmortem in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill agent-incident-postmortem -a claude-code`. Or copy the skill folder (skills/agent-incident-postmortem in mohitagw15856/pm-claude-skills) into .claude/skills/agent-incident-postmortem in your project. Claude Code loads it when a task matches its description.

How do I install Agent Incident Postmortem in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill agent-incident-postmortem -a codex`. Or copy the skill folder (skills/agent-incident-postmortem in mohitagw15856/pm-claude-skills) into .agents/skills/agent-incident-postmortem in your project. Codex loads it when a task matches its description.

Can I use Agent Incident Postmortem in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill agent-incident-postmortem -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-incident-postmortem, .gemini/skills/agent-incident-postmortem, .github/skills/agent-incident-postmortem and .opencode/skills/agent-incident-postmortem in your project.

What does Agent Incident Postmortem need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Incident Postmortem is instructions for the agent only.

Does Agent Incident Postmortem access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Incident Postmortem safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Incident Postmortem use?

Agent Incident Postmortem is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Incident Postmortem use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Incident Postmortem?

Skills that share tags, products or a category with Agent Incident Postmortem: Trader Memory Core (tradermonty/claude-trading-skills, 3k stars), Author Migration (nrwl/nx, 29k stars), Write Notes Like Deepseek (czm15053/write-notes-like-deepseek, 497 stars) and OpenRig Upgrade Procedure (mvschwarz/openrig, 6.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Incident Postmortem?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,433 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 8, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.