Agent skill

Incident Triage Harness

by madebyaris in madebyaris/advance-minimax-m3-cursor-rules

Walks an agent through an evidence-first incident investigation across logs, metrics, code and screenshots, from first symptom to the smallest safe mitigation.

MITAuto-check passedDevOps & Cloud

Install Incident Triage Harness

skills CLI
$ npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill incident-triage-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install madebyaris/advance-minimax-m3-cursor-rules incident-triage-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/madebyaris/advance-minimax-m3-cursor-rules.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/incident-triage-harness .claude/skills/incident-triage-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
incident-triage-harness
GitHub stars
126
Token cost
~984 tokens
SKILL.md length
432 words
Files
2
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Walks an agent through an evidence-first incident investigation across logs, metrics, code and screenshots, from first symptom to the smallest safe mitigation.

  • Works in 5 steps: Triage The Situation → Evidence Loop → What To Inspect → …
  • Debugging an outage, alert or regression in a production-style system
  • SKILL.md covers When to Use, Step 0: Triage The Situation, Step 1: Evidence Loop and Step 2: What To Inspect, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill opens with a triage step that pins down the user-visible symptom, the blast radius, the recent change surface such as deploys, migrations, flags and dependency bumps, and the evidence on hand. No code edits happen until there is at least one concrete failure hypothesis. When you attach a screenshot or screen recording, it counts as primary evidence: the agent reads the actual file, quotes the visible text and names the file path in its report.

The investigation then loops through confirming the symptom, correlating changes with the failure window, forming the smallest plausible hypothesis, checking, narrowing, mitigating and verifying. Inspection runs in order from user symptom to runtime evidence (logs, traces, metrics, console and network errors), change history, code surface and the persistence layer. For UI regressions a re-rendered frame after the fix serves as proof, and animation bugs favor a short clip with a cited timestamp. A reference file holds deeper prompts, evidence templates and mitigation checklists.

When your agent uses it

  • Debugging an outage, alert or regression in a production-style system
  • Investigating a failure after a deploy, migration or config change
  • Triaging a broken UI from a screenshot or screen recording

Example prompts

  • “Our checkout API started timing out after the last deploy; triage it from the logs.”
  • “Here is a screenshot of the broken dashboard; find what regressed and propose the smallest safe fix.”
  • “Investigate the intermittent 500s on the upload route and list your hypotheses first.”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Triage The Situation
  2. Evidence Loop
  3. What To Inspect
  4. Mitigation Bias
  5. Closeout Expectations

What it can do on your machine

Read from SKILL.md and the folder at commit 4d6c552. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Incident Triage Harness loads about 984 tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 432 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~984

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from madebyaris/advance-minimax-m3-cursor-rules at commit 4d6c552, republished under its MIT licence (© madebyaris). 432 words, ~984 tokens.

Download SKILL.mdSave it as .claude/skills/incident-triage-harness/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
incident-triage-harness
description
Production-style incident triage workflow for logs, metrics, code, safe mitigations, and M3 multimodal visual evidence (screenshots, screen recordings). Use when debugging alerts, regressions, outages, or suspicious runtime behavior.
license
MIT
metadata.version
1.1.0
metadata.category
debugging
metadata.sources
Incident response and SRE triage practice, Evidence-first debugging workflows, Production mitigation and verification patterns

Incident Triage Harness

Investigate production-style failures with an evidence-first loop: confirm symptoms, narrow the blast radius, correlate signals, inspect code, and prove the smallest safe mitigation.

When to Use

  • User asks to debug an outage, incident, alert, regression, or mysterious failure
  • Logs, traces, deployments, migrations, config drift, or runtime behavior are involved
  • You need a structured loop for "hypothesis -> check -> narrow -> mitigate -> verify"

For deeper prompts, evidence templates, and mitigation checklists, also read reference.md in this skill directory.


Step 0: Triage The Situation

Before doing anything else, identify:

  1. User-visible symptom: What is broken right now?
  2. Blast radius: Which route, job, service, or user segment is affected?
  3. Recent change surface: Deploys, migrations, feature flags, config changes, dependency bumps
  4. Available evidence: logs, traces, dashboards, DB data, repo history, local repro, tests

Do not jump into code edits until you have at least one concrete failure hypothesis.

Step 0a: Visual Evidence (M3)

When the user attaches a screenshot of a broken UI, a screen recording of the failure, or a short clip of the symptom, the visual is primary evidence — not a supplement.

  • Read the actual file in the current session. Quote the visible text (error message, broken layout, empty state) directly in the report.
  • Name the file path in the closeout; do not describe the image from memory.
  • For a UI regression, treat the screenshot as the "what is broken" input. Treat a re-rendered post-fix frame as the "what is fixed" proof (multimodal-grounded).
  • For animation / interaction bugs, prefer a short clip over a single frame; cite the timestamp region of the relevant behavior.

Show full SKILL.md (167 more words)Show less

Step 1: Evidence Loop

Use this sequence:

text
1. Confirm the symptom
2. Correlate recent changes with the failure window
3. Form the smallest plausible hypothesis
4. Check the hypothesis with the strongest available evidence
5. Narrow the cause before changing code
6. Prefer safe mitigation before broad refactors
7. Verify the mitigation at the affected surface

Step 2: What To Inspect

Inspect in this order when relevant:

  1. User symptom
    • broken page, failing endpoint, timeout, data corruption, auth issue
  2. Runtime evidence
    • logs, traces, metrics, console errors, network failures
  3. Recent change history
    • deploy timeline, feature flags, migrations, dependency changes
  4. Code surface
    • likely entry points, data flow, integration boundaries
  5. Persistence layer
    • DB schema, indexes, missing migrations, stale data, queue state

Step 3: Mitigation Bias

Prefer the smallest safe mitigation that reduces impact:

  • disable or isolate the failing path
  • revert a narrow change if the evidence is strong
  • add or restore a missing migration or config
  • patch one integration boundary rather than refactoring the whole subsystem

Do not claim a root cause until the evidence supports it.


Step 4: Closeout Expectations

When reporting back, make clear:

  • symptom confirmed
  • strongest hypothesis
  • evidence checked
  • mitigation applied or recommended
  • verification run
  • remaining unknowns

If the system is safer but the root cause is still incomplete, say so explicitly.

© madebyaris, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .cursor/skills/incident-triage-harness of madebyaris/advance-minimax-m3-cursor-rules.

  • SKILL.md
  • reference.md

Open the folder on GitHubat commit 4d6c552

Compare with similar skills

Incident Triage Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Incident Triage Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Incident Triage Harness this skillmadebyaris/advance-minimax-m3-cursor-rules126—~984Automated safety check: PassMIT
Axiom SRE Investigatoropenclaw/clawhub9.5k—~7.1kAutomated safety check: PassMIT
Broken API InterviewerPrepLabsAI/InterviewMentor112—~2.6kAutomated safety check: PassMIT
Root-Cause TroubleshootingdavidYichengWei/agentic-engineering-framework158—~646Automated safety check: PassMIT
GitHub Actions Failure Analysisykdojo/claude-code-tips10k—~639Automated safety check: PassCustom licence
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT

Similar skills

  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Broken API Interviewer

    PrepLabsAI/InterviewMentor

    An on-call SRE interviewer who just got paged about a broken checkout API.

    112 GitHub stars~2.6k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Root-Cause Troubleshooting

    davidYichengWei/agentic-engineering-framework

    Diagnoses compile errors, runtime exceptions, failing tests, pipeline failures and production alerts from code and logs, giving a root cause before any fix.

    158 GitHub stars~646 tokensUpdated 6 mo ago
    DevelopmentAuto-check passed
  • GitHub Actions Failure Analysis

    ykdojo/claude-code-tips

    Investigates a failed GitHub Actions run from its URL: pinpoints the real failure, checks the job's history for flakiness, and finds the breaking commit and any existing fix PR.

    10k GitHub stars~639 tokensUpdated 16 days ago
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    415 GitHub stars~1.9k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed

More from madebyaris/advance-minimax-m3-cursor-rules

  • 3D Web Experiences

    madebyaris/advance-minimax-m3-cursor-rules

    Builds 3D web scenes with Three.js and React Three Fiber, with attention to art direction, a performance budget, mobile behavior and fallbacks.

    126 GitHub stars~3k tokensUpdated 3 mo ago
    Auto-check passed
  • Deep Research Loop

    madebyaris/advance-minimax-m3-cursor-rules

    Runs multi-step research with a loop of search, compress, reflect and synthesize, scaling effort from a quick sourced answer to an exhaustive cited report.

    126 GitHub stars~2.9k tokensUpdated 3 mo ago
    Auto-check passed
  • MiniMax M3 Long-Context Discipline

    madebyaris/advance-minimax-m3-cursor-rules

    Teaches how to work within MiniMax M3's 1M-token context: decide per source what to keep, summarize or drop, plan the loading, and cap raw blocks across iterations.

    126 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed
  • MiniMax M3 Multimodal Input

    madebyaris/advance-minimax-m3-cursor-rules

    Teaches an agent on MiniMax M3 to ground visual claims in attached images, screenshots, mockups and clips, and to re-read results after a visual change.

    126 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed
  • MiniMax Multimodal Toolkit

    madebyaris/advance-minimax-m3-cursor-rules

    Routes image, video, voice, music and media-processing requests to the smallest suitable MiniMax or local path, with FFmpeg-style processing around generated media.

    126 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Questions about Incident Triage Harness

What does Incident Triage Harness do?

Walks an agent through an evidence-first incident investigation across logs, metrics, code and screenshots, from first symptom to the smallest safe mitigation. The skill opens with a triage step that pins down the user-visible symptom, the blast radius, the recent change surface such as deploys, migrations, flags and dependency bumps, and the evidence on hand. No code edits happen until there is at least one concrete failure hypothesis.

When should I use Incident Triage Harness?

Incident Triage Harness fits situations like: debugging an outage, alert or regression in a production-style system; investigating a failure after a deploy, migration or config change; triaging a broken UI from a screenshot or screen recording.

How do I install Incident Triage Harness in Claude Code?

Run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill incident-triage-harness -a claude-code`. Or copy the skill folder (.cursor/skills/incident-triage-harness in madebyaris/advance-minimax-m3-cursor-rules) into .claude/skills/incident-triage-harness in your project. Claude Code loads it when a task matches its description.

How do I install Incident Triage Harness in Codex?

Run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill incident-triage-harness -a codex`. Or copy the skill folder (.cursor/skills/incident-triage-harness in madebyaris/advance-minimax-m3-cursor-rules) into .agents/skills/incident-triage-harness in your project. Codex loads it when a task matches its description.

Can I use Incident Triage Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill incident-triage-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-triage-harness, .gemini/skills/incident-triage-harness, .github/skills/incident-triage-harness and .opencode/skills/incident-triage-harness in your project.

What does Incident Triage Harness need to run?

SKILL.md names no scripts, command-line tools or credentials: Incident Triage Harness is instructions for the agent only.

Does Incident Triage Harness access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Incident Triage Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Incident Triage Harness use?

Incident Triage Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Incident Triage Harness use?

About 984 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Incident Triage Harness?

Skills that share tags, products or a category with Incident Triage Harness: Axiom SRE Investigator (openclaw/clawhub, 9.5k stars), Broken API Interviewer (PrepLabsAI/InterviewMentor, 112 stars), Root-Cause Troubleshooting (davidYichengWei/agentic-engineering-framework, 158 stars) and GitHub Actions Failure Analysis (ykdojo/claude-code-tips, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Incident Triage Harness?

madebyaris (a GitHub user) maintains it in madebyaris/advance-minimax-m3-cursor-rules, which has 126 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 16, 2026.

Source: madebyaris/advance-minimax-m3-cursor-rules on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.