Agent skill

Debug

by fcakyon in fcakyon/phd-skills

Evidence-before-action diagnosis of failing ML experiments. An agent skill from fcakyon/phd-skills.

MITAuto-check passedDevelopment

Install Debug

skills CLI
$ npx skills add fcakyon/phd-skills --skill debug -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fcakyon/phd-skills debug --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fcakyon/phd-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/debug .claude/skills/debug && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug
GitHub stars
414
Token cost
~1.3k tokens
SKILL.md length
672 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Evidence-before-action diagnosis of failing ML experiments. An agent skill from fcakyon/phd-skills.

  • Works in 5 steps: cheap probes → hypothesis (labeled as hypothesis) → smoke run → …
  • The user asks why a run is failing
  • SKILL.md covers When to run, Five-step protocol, What to avoid and Output
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Debug is an agent skill from fcakyon/phd-skills. Evidence-before-action diagnosis of failing ML experiments. Probes the system before guessing causes, process list, dmesg, GPU stats, log scrollback, checkpoint state, then states a hypothesis as a hypothesis and runs a smoke before claiming a root cause. Use when the user asks why a run is failing, diverging, OOMing, hanging, slow, producing weird metrics, has crashed, or asks to debug, diagnose, troubleshoot, or investigate a training issue.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Root cause analysis and Debugging. The repository describes itself as: PhD Research Skills for Claude Code: paper reproduction, experiment design, paper review, result comparison and more. The licence is MIT.

When your agent uses it

  • The user asks why a run is failing
  • Producing weird metrics
  • Investigate a training issue

Example prompts

  • “/debug”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. cheap probes
  2. hypothesis (labeled as hypothesis)
  3. smoke run
  4. controls
  5. claim cause

What it can do on your machine

Read from SKILL.md and the folder at commit 67acd61. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug loads about 1.3k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 672 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fcakyon/phd-skills at commit 67acd61, republished under its MIT licence (© fcakyon). 672 words, ~1,314 tokens.

Download SKILL.mdSave it as .claude/skills/debug/SKILL.md (or your agent's skills folder).
name
debug
description
Evidence-before-action diagnosis of failing ML experiments. Probes the system before guessing causes, process list, dmesg, GPU stats, log scrollback, checkpoint state, then states a hypothesis as a hypothesis and runs a smoke before claiming a root cause. Use when the user asks why a run is failing, diverging, OOMing, hanging, slow, producing weird metrics, has crashed, or asks to debug, diagnose, troubleshoot, or investigate a training issue.

Debug: evidence-before-action investigation

The most expensive class of mistake in ML debugging is asserting a cause based on plausibility, then attempting a "fix" that masks the real problem. This skill enforces the discipline of probe → hypothesis → smoke → controls → claim, in that order.

The agentic Stop hook routes here from reason when an assistant claims a cause without backing tool output.

When to run

The user just said any of:

  • "why is X failing / diverging / NaN / OOM / hung / slow / crashed"
  • "the loss is going up", "metrics look weird", "GPU util is 0"
  • "debug this", "diagnose", "troubleshoot", "investigate this run"
  • pasted a log excerpt asking what's wrong

Five-step protocol

Step 1: cheap probes

Before forming any hypothesis, gather the cheap evidence. None of these cost more than a few seconds:

Process state:

bash
ps aux | grep -E '(python|train|torchrun|accelerate)' | grep -v grep

Is the process still running? Zombie? Defunct? Multiple instances?

Kernel / system events:

bash
dmesg | tail -100 # OOM kills, hardware errors, NFS errors
journalctl -xe --since "1 hour ago" | tail -50

GPU state:

bash
nvidia-smi
nvidia-smi --query-gpu=utilization.gpu,memory.used,temperature.gpu --format=csv

Is the GPU even being used? Idle GPU during "training" means the process is blocked on data loading or has died.

Disk / filesystem:

bash
df -h /path/to/run-dir
du -sh /path/to/run-dir/*

Out of disk? Checkpoints not being written?

Log scrollback: Read the last few hundred lines of the training log. Don't trust the user's summary, they may have skimmed. Look for:

  • exception tracebacks
  • repeated "loss=NaN" or "grad_norm=Inf"
  • early-stop announcements (the run may have completed normally)
  • the last successful epoch / step (where did progress stop)

Checkpoint state:

bash
ls -la /path/to/run-dir/checkpoints/

When was the last checkpoint written? What does its size suggest? An empty .pt is different from a 2GB one cut short.

Step 2: hypothesis (labeled as hypothesis)

After the probe, state what might be happening, explicitly framed as a hypothesis:

"Hypothesis: the run is OOMing because dmesg shows oom-kill 3 minutes ago and the process is gone. Alternative hypotheses I haven't ruled out: (a) NFS write timeout, (b) explicit kill from a sibling process."

Never skip to "the cause is X." The hypothesis labels what you don't yet know.

Step 3: smoke run

The cheapest way to confirm or refute a hypothesis is to reproduce the failure shape under a controlled condition:

  • OOM hypothesis: rerun with batch_size=1 for 1 step. If it survives, OOM is confirmed; if it fails the same way, OOM is wrong.
  • Data hypothesis: rerun with a synthetic in-memory dataset. If it works, the data path is implicated.
  • Model hypothesis: forward pass only on a single batch with eval() mode. Loss finite? Outputs sane?
  • Optimizer hypothesis: rerun with lr=0. If the loss still explodes, the loss itself is broken (not the optimizer).
  • Distributed hypothesis: rerun on 1 GPU. If it works, DDP / NCCL is implicated.

A 30-second smoke beats a 30-minute restart-and-pray.

Show full SKILL.md (244 more words)Show less
Step 4: controls

If the smoke is ambiguous, run a control: change exactly one variable from the failing config and rerun the smoke. The differences narrow what mechanism is responsible.

Common control axes (change one at a time):

  • single-source vs multi-source data
  • default workers vs adjusted workers
  • mixed-precision on vs off
  • gradient checkpointing on vs off
  • torch.compile on vs off
Step 5: claim cause

Only after evidence stacks up, probe, smoke, control, do you assert a cause. The claim should cite the specific tool output that proves it:

"Root cause: NFS write timeout. Evidence: dmesg shows nfs server X not responding at 14:23 (the same minute the last checkpoint was written), and the smoke with batch=1 reproduces the timeout. Recommended fix: bind-mount a local scratch dir for checkpoints and rsync to NFS at end of epoch."

If the evidence isn't stacking up, do not promote a hypothesis to a cause. Say "I don't yet know" and propose the next probe.

What to avoid

  • "It's probably X, let me try Y" → no. Probe first.
  • Restarting the run with a small change as the diagnostic. Smoke first, then restart deliberately.
  • Citing only the user's narrative as evidence: re-read the actual log.
  • Stopping at the first plausible cause when artifacts contradict it.

Output

A concise diagnostic report: (1) what the probes showed, (2) the hypothesis, (3) the smoke outcome, (4) the cause-or-uncertain verdict, (5) the recommended next action. Each claim cites the tool output that backs it.

© fcakyon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/debug of fcakyon/phd-skills.

Open the folder on GitHubat commit 67acd61

Compare with similar skills

Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug this skillfcakyon/phd-skills414—~1.3kAutomated safety check: PassMIT
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Bug Finder for daisyUIsaadeghi/daisyui43k—~2.3kAutomated safety check: PassMIT
Root Cause Debugginggarrytan/gstack136k—~1.4kAutomated safety check: PassMIT
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT
Systematic DebuggingChrisWiles/claude-code-showcase6.1k3 repos~1.2kAutomated safety check: PassNone

Similar skills

  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Bug Finder for daisyUI

    saadeghi/daisyui

    Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.

    43k GitHub stars~2.3k tokensUpdated 8 days ago
    DevelopmentAuto-check passed
  • Root Cause Debugging

    garrytan/gstack

    Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.

    136k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Systematic Debugging

    ChrisWiles/claude-code-showcase

    Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.

    6.1k GitHub starsUsed in 3 repos~1.2k tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    103k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed

More from fcakyon/phd-skills

All 12 skills in this repo
  • Reproduce

    fcakyon/phd-skills

    End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Compare

    fcakyon/phd-skills

    Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

    414 GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Experiment Design

    fcakyon/phd-skills

    A skill your agent uses when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies.

    414 GitHub stars~987 tokensUpdated 21 days ago
    Auto-check passed
  • Latex Setup

    fcakyon/phd-skills

    A skill your agent uses when the user wants to set up or troubleshoot a LaTeX environment, choose between biber and bibtex, install packages for a specific venue template, or configure compilation.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check: notes
  • Launch

    fcakyon/phd-skills

    Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup.

    414 GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Literature Research

    fcakyon/phd-skills

    A skill your agent uses when the user wants to find related work, survey a research area, identify literature gaps, or discover open-source implementations.

    414 GitHub stars~999 tokensUpdated 21 days ago
    Auto-check passed

Categories

Questions about Debug

What does Debug do?

Evidence-before-action diagnosis of failing ML experiments. An agent skill from fcakyon/phd-skills. Debug is an agent skill from fcakyon/phd-skills. Evidence-before-action diagnosis of failing ML experiments.

When should I use Debug?

Debug fits situations like: the user asks why a run is failing; producing weird metrics; investigate a training issue.

How do I install Debug in Claude Code?

Run `npx skills add fcakyon/phd-skills --skill debug -a claude-code`. Or copy the skill folder (plugin/skills/debug in fcakyon/phd-skills) into .claude/skills/debug in your project. Claude Code loads it when a task matches its description.

How do I install Debug in Codex?

Run `npx skills add fcakyon/phd-skills --skill debug -a codex`. Or copy the skill folder (plugin/skills/debug in fcakyon/phd-skills) into .agents/skills/debug in your project. Codex loads it when a task matches its description.

Can I use Debug in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fcakyon/phd-skills --skill debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug, .gemini/skills/debug, .github/skills/debug and .opencode/skills/debug in your project.

What does Debug need to run?

SKILL.md names no scripts, command-line tools or credentials: Debug is instructions for the agent only.

Does Debug access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debug safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debug use?

Debug is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug?

Skills that share tags, products or a category with Debug: OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars), Root Cause Debugging (garrytan/gstack, 136k stars) and Graph-Based Bug Tracing (tirth8205/code-review-graph, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug?

fcakyon (a GitHub user) maintains it in fcakyon/phd-skills, which has 414 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 16, 2026.

Source: fcakyon/phd-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.