The Art of Debugging
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ByteDance-Seed/VeOmni veomni-debug --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/veomni-debug .claude/skills/veomni-debug && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "veomni-debug" agent skill from https://github.com/ByteDance-Seed/VeOmni/tree/main/.agents/skills/veomni-debug into .claude/skills/veomni-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "veomni-debug", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ByteDance-Seed/VeOmni/tree/main/.agents/skills/veomni-debugType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ByteDance-Seed/VeOmni veomni-debug --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/veomni-debug .agents/skills/veomni-debug && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "veomni-debug" agent skill from https://github.com/ByteDance-Seed/VeOmni/tree/main/.agents/skills/veomni-debug into .agents/skills/veomni-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "veomni-debug", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ByteDance-Seed/VeOmni veomni-debug --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/veomni-debug .cursor/skills/veomni-debug && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "veomni-debug" agent skill from https://github.com/ByteDance-Seed/VeOmni/tree/main/.agents/skills/veomni-debug into .cursor/skills/veomni-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "veomni-debug", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ByteDance-Seed/VeOmni.git --path .agents/skills/veomni-debug--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ByteDance-Seed/VeOmni veomni-debug --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/veomni-debug .gemini/skills/veomni-debug && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "veomni-debug" agent skill from https://github.com/ByteDance-Seed/VeOmni/tree/main/.agents/skills/veomni-debug into .gemini/skills/veomni-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "veomni-debug", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ByteDance-Seed/VeOmni veomni-debugInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/veomni-debug .github/skills/veomni-debug && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "veomni-debug" agent skill from https://github.com/ByteDance-Seed/VeOmni/tree/main/.agents/skills/veomni-debug into .github/skills/veomni-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "veomni-debug", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ByteDance-Seed/VeOmni veomni-debug --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/veomni-debug .opencode/skills/veomni-debug && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "veomni-debug" agent skill from https://github.com/ByteDance-Seed/VeOmni/tree/main/.agents/skills/veomni-debug into .opencode/skills/veomni-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "veomni-debug", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
veomni-debugA skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…
Veomni Debug is an agent skill from ByteDance-Seed/VeOmni. Use this skill for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected behavior. Covers both quick fixes (clear root cause) and complex debugging (unclear cause). Trigger: 'fix bug', 'fix error', 'broken', 'crash', 'doesn't work', 'fails with', 'loss NaN', 'training hangs', 'FSDP error', 'OOM'.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development, covering Deep learning, Debugging and Root cause analysis. It works with CUDA. The repository describes itself as: VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 16c94aa. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvmakegitpytestFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Veomni Debug loads about 2.8k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 1,149 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ByteDance-Seed/VeOmni at commit 16c94aa, republished under its Apache-2.0 licence (© ByteDance-Seed). 1,149 words, ~2,800 tokens.
.claude/skills/veomni-debug/SKILL.md (or your agent's skills folder).| Situation | Path |
|---|---|
| Clear error, obvious root cause, fix in <15 min | Quick Path (below) |
| Root cause unclear, multiple hypotheses | Full Protocol (Phase 1–5) |
| Distributed training issue (hang, wrong loss, sharding) | Full Protocol |
| Numerical accuracy / loss divergence | Full Protocol |
| 2+ failed fix attempts | Full Protocol |
.agents/knowledge/constraints.md for known pitfalls.pytest tests/<module>/ passes, no regressions across modalities.make quality, commit. Run /veomni-review before opening the PR or pushing a substantive update, not per commit.If not resolved in 15 min → switch to Full Protocol.
Track the phases with whatever todo/plan tool the running agent provides:
Phase 1: Investigate <symptom> -> in_progress
Phase 2: Pattern analysis -> pending
Phase 3: Hypothesis & test -> pending
Phase 4: Implement fix -> pending
Phase 5: Knowledge capture -> pending.agents/knowledge/constraints.md — many issues are known constraint violations.git log --oneline -10 — what changed recently?veomni/distributed/parallel_plan.py).veomni/distributed/sequence_parallel/).veomni/distributed/moe/).Find a working example (previous commit, different config, reference implementation).
Compare completely — diff line by line, not skim. Include config YAML, environment vars, and launcher scripts.
Identify ALL differences between working and broken code.
Check dependencies — different transformers version? Different PyTorch version?
If a package version upgrade is suspected, create isolated uv environments to bisect:
# Env A: the current default pin (the `transformers-stable` group).
uv venv .venv-a
VIRTUAL_ENV=.venv-a uv sync --active --extra gpu --dev
# Env B: the same tree with exactly one package moved.
uv venv .venv-b
VIRTUAL_ENV=.venv-b uv sync --active --extra gpu --dev
VIRTUAL_ENV=.venv-b uv pip install "<package>==<other-version>"--active is load-bearing. Without it uv sync runs in project mode and
targets .venv/, ignoring VIRTUAL_ENV — so both commands would rebuild
the main environment instead of the two you just created, which is the
opposite of what this is for. (UV_PROJECT_ENVIRONMENT works too.)
Confirm that installing the alternate version did not change other packages:
uv pip freeze --python .venv-a/bin/python > /tmp/veomni-bisect-a.freeze
uv pip freeze --python .venv-b/bin/python > /tmp/veomni-bisect-b.freeze
diff -u /tmp/veomni-bisect-a.freeze /tmp/veomni-bisect-b.freezeOnly the target package may differ. Pin or restore every non-target difference in Env B to Env A's version, then compare again before running the reproducer. If the target cannot run with that dependency set, report the compatibility conflict; a multi-package change is not a one-package bisect.
Then run the same reproducer in both envs, each with its own env
activated — the VIRTUAL_ENV= prefixes above apply only to the uv sync
lines they are attached to, not to whatever you run next:
(source .venv-a/bin/activate && <reproducer>)
(source .venv-b/bin/activate && <reproducer>)If the suspect package is transformers, you need two worktrees and two
venvs — one venv per worktree. They isolate different things and neither
substitutes for the other: a venv isolates the installed packages, a
worktree isolates the checkout. generated/ modeling lives in the
checkout, so two venvs in one worktree share a single generated/ and
regenerating it for Env B silently changes what Env A runs. Two worktrees
without separate venvs share one transformers install, which defeats the
bisect outright.
git worktree add ../bisect-a HEAD && (cd ../bisect-a && uv venv .venv && VIRTUAL_ENV=.venv uv sync --active --extra gpu --dev)
git worktree add ../bisect-b HEAD && (cd ../bisect-b && uv venv .venv && VIRTUAL_ENV=.venv uv sync --active --extra gpu --dev && VIRTUAL_ENV=.venv uv pip install "transformers==<other-version>")Regenerate generated/ inside each worktree against its own pin
(make patchgen) before running the reproducer — it is produced against
the pinned version, and a stale generated/ is itself a source of
failures.
Compare the package sets here too, using uv pip freeze --python with
../bisect-a/.venv/bin/python and ../bisect-b/.venv/bin/python, and
reconcile non-transformers dependency version differences as above. The
editable VeOmni and patchgen paths must point to their respective worktrees;
normalize those corresponding paths only when comparing the freeze output,
without changing either environment's editable installs. Run codegen and
the reproducer from each worktree with its own environment activated:
(cd ../bisect-a && source .venv/bin/activate && make patchgen && <reproducer>)
(cd ../bisect-b && source .venv/bin/activate && make patchgen && <reproducer>)Verification gate — before acting on a conclusion, check:
/veomni-review over the branch diff.Do this immediately after the fix is verified. Knowledge decays fast.
.agents/knowledge/constraints.md.agents/knowledge/architecture.md.agents/knowledge/testing.md. Only paths not
already covered by a directory-level CI entry need a new workflow line.
Use the workflow that owns the path, including the e2e workflows for
end-to-end tests; account for the GPU/NPU differences in that table.docs/ if the fix changes API behavior, config semantics, or usage patternsIf none apply, explicitly note "no new knowledge to capture."
Restart from Phase 1 if you catch yourself thinking "let me just try changing X and see", "quick fix for now, clean up later", or "it probably works, moving on".
After 3 consecutive failed fix attempts, stop fixing symptoms. Question whether the underlying approach is wrong, re-examine whether you are solving the right problem, and report the analysis to the user before continuing.
data_collator type matches the dataset.veomni/models/transformers/*/ are auto-generated — editing generated files directly will be overwritten.Include the relevant checklist when investigating.
When confidence is low or evidence is ambiguous, launch a subagent to challenge your conclusion:
You are a critical reviewer. Your job is to find flaws in the following conclusion.
## Conclusion Under Review
<the specific claim or decision>
## Evidence Presented
<the data, logs, experiments supporting the conclusion>
## Your Task
1. Does the evidence actually support the conclusion, or just correlate?
2. Generate 2+ alternative explanations consistent with the same evidence.
3. What specific observation would DISPROVE this conclusion? Has it been checked?
4. Was the experiment controlled (one variable changed at a time)?
## Output
Verdict: CONFIRMED / CHALLENGED / INSUFFICIENT_EVIDENCE
Findings: [issues found, counter-hypotheses, missing evidence]© ByteDance-Seed, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/veomni-debug of ByteDance-Seed/VeOmni.
Open the folder on GitHubat commit 16c94aa
Veomni Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Veomni Debug this skillByteDance-Seed/VeOmni | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| The Art of Debuggingstas00/the-art-of-debugging | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | |
| Aoti Debugpytorch/pytorch | 104k | 1 repos | ~1.7k | Automated safety check: Pass | Custom licence | |
| Ascendcascend-ai-coding/awesome-ascend-skills | 174 | — | ~3.5k | Automated safety check: Pass | None | |
| Systematic DebuggingChrisWiles/claude-code-showcase | 6.1k | 3 repos | ~1.2k | Automated safety check: Pass | None | |
| Debugging and Error Recoveryaddyosmani/agent-skills | 103k | 1 repos | ~2.6k | Automated safety check: Pass | MIT |
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
ascend-ai-coding/awesome-ascend-skills
End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.
ChrisWiles/claude-code-showcase
Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.
addyosmani/agent-skills
Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.
ed3dai/ed3d-plugins
A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes - four-phase framework (root cause investigation, pattern analysis, hypothesis…
ByteDance-Seed/VeOmni
Create a pull request for the current branch. An agent skill from ByteDance-Seed/VeOmni.
ByteDance-Seed/VeOmni
A skill your agent uses when adding support for a new model to VeOmni.
ByteDance-Seed/VeOmni
A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/.
ByteDance-Seed/VeOmni
Author or refresh a VeOmni model's patchgen-generated modeling under generated/ — GPU and/or NPU config, dense or MoE, text / VLM / Omni.
ByteDance-Seed/VeOmni
A skill your agent uses for performance profiling and optimization.
ByteDance-Seed/VeOmni
Pre-PR code review gate. An agent skill from ByteDance-Seed/VeOmni.
Works with
Categories
A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…. Veomni Debug is an agent skill from ByteDance-Seed/VeOmni. Use this skill for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected behavior.
Veomni Debug fits situations like: loss divergence; gradient explosion; distributed training hang; checkpoint load failure.
Run `npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a claude-code`. Or copy the skill folder (.agents/skills/veomni-debug in ByteDance-Seed/VeOmni) into .claude/skills/veomni-debug in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a codex`. Or copy the skill folder (.agents/skills/veomni-debug in ByteDance-Seed/VeOmni) into .agents/skills/veomni-debug in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/veomni-debug, .gemini/skills/veomni-debug, .github/skills/veomni-debug and .opencode/skills/veomni-debug in your project.
Going by SKILL.md and its folder, Veomni Debug needs the command-line tools its instructions call (uv, make, git and pytest). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Veomni Debug is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Veomni Debug: The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars), Aoti Debug (pytorch/pytorch, 104k stars), Ascendc (ascend-ai-coding/awesome-ascend-skills, 174 stars) and Systematic Debugging (ChrisWiles/claude-code-showcase, 6.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ByteDance-Seed (a GitHub organization) maintains it in ByteDance-Seed/VeOmni, which has 2,235 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 9, 2026.
Source: ByteDance-Seed/VeOmni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.