Agent skill

Root-Cause Troubleshooting

by davidYichengWei in davidYichengWei/agentic-engineering-framework

Diagnoses compile errors, runtime exceptions, failing tests, pipeline failures and production alerts from code and logs, giving a root cause before any fix.

MITAuto-check passedDevelopment

SKILL.md written in Chinese; this summary is our English description.

Install Root-Cause Troubleshooting

skills CLI
$ npx skills add davidYichengWei/agentic-engineering-framework --skill troubleshooting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davidYichengWei/agentic-engineering-framework troubleshooting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davidYichengWei/agentic-engineering-framework.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/troubleshooting .claude/skills/troubleshooting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
troubleshooting
GitHub stars
158
Token cost
~646 tokens
SKILL.md length
156 words
Files
3
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Diagnoses compile errors, runtime exceptions, failing tests, pipeline failures and production alerts from code and logs, giving a root cause before any fix.

  • Works in 4 steps: 收集信息 → 假设-验证循环 → 历史案例(深度排查时) → …
  • Locating the cause of a compile error or runtime exception
  • SKILL.md covers Iron Law, 核心约束, Log 不足时的处理 and 排查记录, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill enforces one rule: no assumptions, and no fix suggestion until a root cause is found. Every conclusion must rest on the code plus the logs you provide. The agent cannot reach your environment, so it asks you to run commands, shown in copy-ready code blocks, and it delivers a root cause with a fix recommendation instead of applying the fix itself.

When logs are too thin, it finds the log statements in the code that separate the failure scenarios and gives you grep commands to pull them. It collects error logs, stack traces, reproduction steps and recent changes, calls a codebase-researcher subagent for the relevant call chains, then loops through hypothesis, your verification and revision, widening its search after 3 failed rounds in a row. For pipeline failures and production alerts it keeps a live troubleshooting log from a template and can match symptoms against past cases under `reference/cases`. The skill text is written in Chinese.

When your agent uses it

  • Locating the cause of a compile error or runtime exception
  • Diagnosing a failing unit test or a broken pipeline run
  • Investigating a production alert from logs and code
  • Working out what to grep when the available logs are not enough

Example prompts

  • “The CI pipeline fails at the integration test stage; here is the log, so find the root cause.”
  • “Our service throws a null pointer after the last deploy, and the stack trace is attached.”
  • “A production alert fired overnight. Start a troubleshooting record and tell me which logs to pull.”

Requirements

  • A codebase-researcher subagent

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 收集信息
  2. 假设-验证循环
  3. 历史案例(深度排查时)
  4. 模块专项排查

What it can do on your machine

Read from SKILL.md and the folder at commit 1f7ac0f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Root-Cause Troubleshooting loads about 646 tokens when it runs. Until then it costs about 16 tokens; SKILL.md has 156 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~16
When it runs · the whole SKILL.md, loaded when a task matches
~646

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davidYichengWei/agentic-engineering-framework at commit 1f7ac0f, republished under its MIT licence (© davidYichengWei). 156 words, ~646 tokens.

Download SKILL.mdSave it as .claude/skills/troubleshooting/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
troubleshooting
description
问题排查。当用户遇到编译错误、运行时异常、单测失败、流水线报错、现网告警等需要定位问题时触发。

问题排查 (Troubleshooting)

Iron Law

NO ASSUMPTIONS. NO FIX WITHOUT ROOT CAUSE.
  1. 禁止假设:所有结论必须基于代码 + 用户提供的 log,不得凭经验猜测
  2. 未找到根因,禁止提修复建议

核心约束

  1. AI 无法访问用户环境:只能看用户发来的信息 + codebase 代码
  2. 让用户执行命令:命令必须用代码块输出,方便复制
  3. 输出目标:根因 + 修复建议(不自己修复)

Log 不足时的处理

禁止在缺少 log 的情况下猜测根因。

当用户提供的 log 不足以定位问题时:

  1. 在代码中找关键 log:定位相关代码路径,找出能区分不同故障场景的 log 语句
  2. 提供 grep 命令:告诉用户如何从日志文件中提取关键信息

示例:

bash
# 查找某个错误码相关的日志
grep -E "error_code|ErrorCode" /path/to/log | head -50

# 查找某个函数调用前后的上下文
grep -B5 -A10 "FunctionName" /path/to/log

# 按时间范围过滤
grep "2025-01-29 10:3[0-9]" /path/to/log | grep -i error

要点:

  • 明确告诉用户 grep 什么关键字
  • 说明这个 log 能帮助确认/排除什么

排查记录

流水线/现网问题排查时,创建排查记录文档实时跟踪进度。

判断是否创建:询问用户问题类型:

  • 流水线报错 / 现网告警 → 创建排查记录
  • 开发调试中的问题 → 不创建,直接排查

模板位置:reference/troubleshooting-log-template.md

创建方式:

bash
cp skills/troubleshooting/reference/troubleshooting-log-template.md \
   troubleshooting-[问题简述]-$(date +%Y%m%d).md

记录要点:

  • 每个重要发现立即记录(日志、代码位置、中间结论)
  • 每次有新进展必须更新文档:新发现的 log、代码分析结果、排除的假设
  • 同步更新待确认点:哪些假设已验证、哪些还需确认、下一步要做什么
  • 定位后补充根因和证据链

Red Flags:瞎猜信号

危险想法正确做法
"看起来像是 X"有什么证据?让用户验证
"试试改 Y 看看"这是猜测,不是诊断
"应该是 Z 导致的""应该"不是证据

排查流程

1. 收集信息
必须收集深度排查额外收集
错误日志、堆栈、错误码时间线、环境差异
复现条件、触发步骤是否间歇性发生
代码版本、最近变更完整服务拓扑

代码上下文调研(必须):调用 codebase-researcher subagent 调研问题相关的代码上下文,包括:

  • 报错涉及的函数/模块的实现逻辑和调用链
  • 相关数据结构和状态流转
  • 上下游模块的交互方式

信息不足时主动追问,不要猜测。

2. 假设-验证循环
形成假设 → 让用户验证 → 确认或否定 → 迭代

3+ 轮失败规则:连续 3 轮假设被否定 → 停止猜测,扩大信息收集范围。

3. 历史案例(深度排查时)

流水线/现网问题时,在 reference/cases/ 搜索匹配案例:

  • 提取错误关键字(错误码、异常类型、模块名)
  • 匹配 symptoms.keywords
  • 按案例诊断步骤验证
4. 模块专项排查

根据项目需要,可在 reference/ 下为特定模块添加专项排查资料。


输出格式

  • 开发调试问题:直接在对话中输出根因和修复建议
  • 流水线/现网问题:更新排查记录文档,格式参见 troubleshooting-log-template.md

案例沉淀

复杂/代表性问题排查后,按 case_template.md 沉淀到 reference/cases/<module>/。


强制规则

规则说明
禁止假设结论必须基于代码 + log,不得凭经验猜测
Iron Law未找到根因,禁止提修复建议
禁止缺 log 猜根因log 不足时,从代码中找关键 log 并提供 grep 命令,不得猜测
代码上下文调研必须调用 codebase-researcher subagent 调研问题相关代码
实时更新文档每次新进展/新结论必须更新排查记录,同步维护待确认点
假设验证让用户验证,不脑补结果
3+ 轮规则连续失败则扩大范围

© davidYichengWei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/troubleshooting of davidYichengWei/agentic-engineering-framework.

  • SKILL.md
  • reference/cases/case_template.md
  • reference/troubleshooting-log-template.md

Open the folder on GitHubat commit 1f7ac0f

Compare with similar skills

Root-Cause Troubleshooting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Root-Cause Troubleshooting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Root-Cause Troubleshooting this skilldavidYichengWei/agentic-engineering-framework158—~646Automated safety check: PassMIT
Incident Triage Harnessmadebyaris/advance-minimax-m3-cursor-rules126—~984Automated safety check: PassMIT
Axiom SRE Investigatoropenclaw/clawhub9.5k—~7.1kAutomated safety check: PassMIT
Broken API InterviewerPrepLabsAI/InterviewMentor112—~2.6kAutomated safety check: PassMIT
Production Error Huntdifferent-ai/openwork24k—~803Automated safety check: PassCustom licence
gh-aw Workflow Diagnosisgithub/gh-aw5.4k—~4kAutomated safety check: PassMIT

Similar skills

  • Incident Triage Harness

    madebyaris/advance-minimax-m3-cursor-rules

    Walks an agent through an evidence-first incident investigation across logs, metrics, code and screenshots, from first symptom to the smallest safe mitigation.

    126 GitHub stars~984 tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Broken API Interviewer

    PrepLabsAI/InterviewMentor

    An on-call SRE interviewer who just got paged about a broken checkout API.

    112 GitHub stars~2.6k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Production Error Hunt

    different-ai/openwork

    Traces an opaque production error in an OpenWork build to its cause using local server logs and Sentry, names the regressing PR and files a report.

    24k GitHub stars~803 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Official

    Guides your agent through diagnosing failed GitHub Agentic Workflows by downloading run logs, auditing individual runs and reading the artifacts they leave behind.

    5.4k GitHub stars~4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • GitHub Actions Failure Analysis

    ykdojo/claude-code-tips

    Investigates a failed GitHub Actions run from its URL: pinpoints the real failure, checks the job's history for flakiness, and finds the breaking commit and any existing fix PR.

    10k GitHub stars~639 tokensUpdated 14 days ago
    DevOps & CloudAuto-check passed

More from davidYichengWei/agentic-engineering-framework

All 14 skills in this repo
  • Architecture Design Principles

    davidYichengWei/agentic-engineering-framework

    Gives architecture design principles for system design discussions and code review: module boundaries, dependency direction, data ownership and interface rules, plus a checklist.

    158 GitHub stars~545 tokensUpdated 6 mo ago
    Auto-check passed
  • General Coding Best Practices

    davidYichengWei/agentic-engineering-framework

    A checklist of language-neutral rules for writing and reviewing code: naming, function design, control flow, resource safety, comments and logging.

    158 GitHub stars~631 tokensUpdated 6 mo ago
    Auto-check passed
  • Component Design Principles

    davidYichengWei/agentic-engineering-framework

    Chinese-language checklists for component-level design: class and module structure, public interfaces, data models, concurrency and error handling.

    158 GitHub stars~1k tokensUpdated 6 mo ago
    Auto-check passed
  • Skill Authoring Guide (Chinese)

    davidYichengWei/agentic-engineering-framework

    Chinese-language guide to writing and improving SKILL.md files: frontmatter rules, concise writing, progressive disclosure, common patterns and a pre-release checklist.

    158 GitHub stars~788 tokensUpdated 6 mo ago
    Auto-check passed
  • Self-Refinement from Corrections

    davidYichengWei/agentic-engineering-framework

    Turns mistakes you correct into proposed updates to persistent Rules and Skills so the same error does not recur in later sessions, triggered automatically or with /reflect.

    158 GitHub stars~575 tokensUpdated 6 mo ago
    Auto-check passed
  • Workflow Code Generation

    davidYichengWei/agentic-engineering-framework

    代码文件修改的统一入口。当用户请求任何代码变更(新功能、优化、Bug 修复、重构)时必须首先调用此 skill。仅适用于代码文件(如 .cc/.cpp/.h/.go/.py 等),修改 .md 等非代码文件时不需要调用。它会评估复杂度、检查 spec.md、生成 tasks.md、并逐个任务执行。

    158 GitHub stars~733 tokensUpdated 6 mo ago
    Auto-check passed

Questions about Root-Cause Troubleshooting

What does Root-Cause Troubleshooting do?

Diagnoses compile errors, runtime exceptions, failing tests, pipeline failures and production alerts from code and logs, giving a root cause before any fix. The skill enforces one rule: no assumptions, and no fix suggestion until a root cause is found. Every conclusion must rest on the code plus the logs you provide.

When should I use Root-Cause Troubleshooting?

Root-Cause Troubleshooting fits situations like: locating the cause of a compile error or runtime exception; diagnosing a failing unit test or a broken pipeline run; investigating a production alert from logs and code; working out what to grep when the available logs are not enough.

How do I install Root-Cause Troubleshooting in Claude Code?

Run `npx skills add davidYichengWei/agentic-engineering-framework --skill troubleshooting -a claude-code`. Or copy the skill folder (skills/troubleshooting in davidYichengWei/agentic-engineering-framework) into .claude/skills/troubleshooting in your project. Claude Code loads it when a task matches its description.

How do I install Root-Cause Troubleshooting in Codex?

Run `npx skills add davidYichengWei/agentic-engineering-framework --skill troubleshooting -a codex`. Or copy the skill folder (skills/troubleshooting in davidYichengWei/agentic-engineering-framework) into .agents/skills/troubleshooting in your project. Codex loads it when a task matches its description.

Can I use Root-Cause Troubleshooting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davidYichengWei/agentic-engineering-framework --skill troubleshooting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/troubleshooting, .gemini/skills/troubleshooting, .github/skills/troubleshooting and .opencode/skills/troubleshooting in your project.

What does Root-Cause Troubleshooting need to run?

SKILL.md names no scripts, command-line tools or credentials: Root-Cause Troubleshooting is instructions for the agent only. Our summary lists: A codebase-researcher subagent.

Does Root-Cause Troubleshooting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Root-Cause Troubleshooting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Root-Cause Troubleshooting use?

Root-Cause Troubleshooting is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Root-Cause Troubleshooting use?

About 646 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Root-Cause Troubleshooting?

Skills that share tags, products or a category with Root-Cause Troubleshooting: Incident Triage Harness (madebyaris/advance-minimax-m3-cursor-rules, 126 stars), Axiom SRE Investigator (openclaw/clawhub, 9.5k stars), Broken API Interviewer (PrepLabsAI/InterviewMentor, 112 stars) and Production Error Hunt (different-ai/openwork, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Root-Cause Troubleshooting?

davidYichengWei (a GitHub user) maintains it in davidYichengWei/agentic-engineering-framework, which has 158 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on March 25, 2026.

Source: davidYichengWei/agentic-engineering-framework on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.