Agent skill

Veomni Debug

by ByteDance-Seed in ByteDance-Seed/VeOmni

A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

Apache-2.0Auto-check passedDevelopment

Install Veomni Debug

skills CLI
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ByteDance-Seed/VeOmni veomni-debug --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/veomni-debug .claude/skills/veomni-debug && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
veomni-debug
GitHub stars
2.2k
Token cost
~2.8k tokens
SKILL.md length
1,149 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

  • Works in 5 steps: Root Cause Investigation → Pattern Analysis → Hypothesis and Testing → …
  • Loss divergence
  • SKILL.md covers Quick Path vs Full Protocol, Full Protocol, Stop Conditions and Common Pitfalls, plus 2 more sections
  • Calls uv, make and git

What it does

Veomni Debug is an agent skill from ByteDance-Seed/VeOmni. Use this skill for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected behavior. Covers both quick fixes (clear root cause) and complex debugging (unclear cause). Trigger: 'fix bug', 'fix error', 'broken', 'crash', 'doesn't work', 'fails with', 'loss NaN', 'training hangs', 'FSDP error', 'OOM'.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Deep learning, Debugging and Root cause analysis. It works with CUDA. The repository describes itself as: VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo. The licence is Apache-2.0.

When your agent uses it

  • Loss divergence
  • Gradient explosion
  • Distributed training hang
  • Checkpoint load failure

Example prompts

  • “fix bug”
  • “fix error”
  • “broken”
  • “/veomni-debug”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Root Cause Investigation
  2. Pattern Analysis
  3. Hypothesis and Testing
  4. Implementation
  5. Knowledge Capture (mandatory)

What it can do on your machine

Read from SKILL.md and the folder at commit 16c94aa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • make
    • git
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Veomni Debug loads about 2.8k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 1,149 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ByteDance-Seed/VeOmni at commit 16c94aa, republished under its Apache-2.0 licence (© ByteDance-Seed). 1,149 words, ~2,800 tokens.

Download SKILL.mdSave it as .claude/skills/veomni-debug/SKILL.md (or your agent's skills folder).
name
veomni-debug
description
Use this skill for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected behavior. Covers both quick fixes (clear root cause) and complex debugging (unclear cause). Trigger: 'fix bug', 'fix error', 'broken', 'crash', 'doesn't work', 'fails with', 'loss NaN', 'training hangs', 'FSDP error', 'OOM'.

Quick Path vs Full Protocol

SituationPath
Clear error, obvious root cause, fix in <15 minQuick Path (below)
Root cause unclear, multiple hypothesesFull Protocol (Phase 1–5)
Distributed training issue (hang, wrong loss, sharding)Full Protocol
Numerical accuracy / loss divergenceFull Protocol
2+ failed fix attemptsFull Protocol
Quick Path
  1. Reproduce the error. Read the full traceback.
  2. Check .agents/knowledge/constraints.md for known pitfalls.
  3. Write a reproducer test if feasible.
  4. Minimal fix — root cause only, don't touch surrounding code.
  5. Verify: reproducer passes, pytest tests/<module>/ passes, no regressions across modalities.
  6. Run make quality, commit. Run /veomni-review before opening the PR or pushing a substantive update, not per commit.

If not resolved in 15 min → switch to Full Protocol.


Full Protocol

Before You Start

Track the phases with whatever todo/plan tool the running agent provides:

Phase 1: Investigate <symptom>       -> in_progress
Phase 2: Pattern analysis            -> pending
Phase 3: Hypothesis & test           -> pending
Phase 4: Implement fix               -> pending
Phase 5: Knowledge capture           -> pending
Phase 1: Root Cause Investigation
  1. Read the FULL error message / symptom. Don't skim. Extract 2-3 keywords.
  2. Check constraints first: Read .agents/knowledge/constraints.md — many issues are known constraint violations.
  3. Reproduce consistently. If you can't reproduce, you don't understand it.
  4. git log --oneline -10 — what changed recently?
  5. Trace data flow backward through the call stack.
  6. Distributed training specifics:
    • Check if error appears on all ranks or just rank 0.
    • FSDP2: verify sharding plan matches model structure (veomni/distributed/parallel_plan.py).
    • Sequence parallel: check that attention inputs are properly split/gathered (veomni/distributed/sequence_parallel/).
    • MoE: verify expert routing and load balancing (veomni/distributed/moe/).
Phase 2: Pattern Analysis
  1. Find a working example (previous commit, different config, reference implementation).

  2. Compare completely — diff line by line, not skim. Include config YAML, environment vars, and launcher scripts.

  3. Identify ALL differences between working and broken code.

  4. Check dependencies — different transformers version? Different PyTorch version?

  5. If a package version upgrade is suspected, create isolated uv environments to bisect:

    bash
    # Env A: the current default pin (the `transformers-stable` group).
    uv venv .venv-a
    VIRTUAL_ENV=.venv-a uv sync --active --extra gpu --dev
    
    # Env B: the same tree with exactly one package moved.
    uv venv .venv-b
    VIRTUAL_ENV=.venv-b uv sync --active --extra gpu --dev
    VIRTUAL_ENV=.venv-b uv pip install "<package>==<other-version>"

    --active is load-bearing. Without it uv sync runs in project mode and targets .venv/, ignoring VIRTUAL_ENV — so both commands would rebuild the main environment instead of the two you just created, which is the opposite of what this is for. (UV_PROJECT_ENVIRONMENT works too.)

    Confirm that installing the alternate version did not change other packages:

    bash
    uv pip freeze --python .venv-a/bin/python > /tmp/veomni-bisect-a.freeze
    uv pip freeze --python .venv-b/bin/python > /tmp/veomni-bisect-b.freeze
    diff -u /tmp/veomni-bisect-a.freeze /tmp/veomni-bisect-b.freeze

    Only the target package may differ. Pin or restore every non-target difference in Env B to Env A's version, then compare again before running the reproducer. If the target cannot run with that dependency set, report the compatibility conflict; a multi-package change is not a one-package bisect.

    Then run the same reproducer in both envs, each with its own env activated — the VIRTUAL_ENV= prefixes above apply only to the uv sync lines they are attached to, not to whatever you run next:

    bash
    (source .venv-a/bin/activate && <reproducer>)
    (source .venv-b/bin/activate && <reproducer>)

    If the suspect package is transformers, you need two worktrees and two venvs — one venv per worktree. They isolate different things and neither substitutes for the other: a venv isolates the installed packages, a worktree isolates the checkout. generated/ modeling lives in the checkout, so two venvs in one worktree share a single generated/ and regenerating it for Env B silently changes what Env A runs. Two worktrees without separate venvs share one transformers install, which defeats the bisect outright.

    bash
    git worktree add ../bisect-a HEAD && (cd ../bisect-a && uv venv .venv && VIRTUAL_ENV=.venv uv sync --active --extra gpu --dev)
    git worktree add ../bisect-b HEAD && (cd ../bisect-b && uv venv .venv && VIRTUAL_ENV=.venv uv sync --active --extra gpu --dev && VIRTUAL_ENV=.venv uv pip install "transformers==<other-version>")

    Regenerate generated/ inside each worktree against its own pin (make patchgen) before running the reproducer — it is produced against the pinned version, and a stale generated/ is itself a source of failures.

    Compare the package sets here too, using uv pip freeze --python with ../bisect-a/.venv/bin/python and ../bisect-b/.venv/bin/python, and reconcile non-transformers dependency version differences as above. The editable VeOmni and patchgen paths must point to their respective worktrees; normalize those corresponding paths only when comparing the freeze output, without changing either environment's editable installs. Run codegen and the reproducer from each worktree with its own environment activated:

    bash
    (cd ../bisect-a && source .venv/bin/activate && make patchgen && <reproducer>)
    (cd ../bisect-b && source .venv/bin/activate && make patchgen && <reproducer>)
Show full SKILL.md (537 more words)Show less
Phase 3: Hypothesis and Testing
  1. Form ONE specific, falsifiable hypothesis.
  2. Design a MINIMAL experiment (change one thing only).
  3. Run the experiment. Record the result.
  4. If wrong, update understanding and form new hypothesis. No random guess-and-check.

Verification gate — before acting on a conclusion, check:

  • Does the evidence actually support this cause, or just correlate?
  • Could a different root cause produce the same symptoms?
  • What observation would disprove this hypothesis? Have you looked for it?
  • If confidence < 80% or the evidence is ambiguous, launch a verification subagent (see Appendix).
Phase 4: Implementation
  1. Write a failing test that demonstrates the bug (if feasible).
  2. Implement a SINGLE targeted fix addressing the root cause.
  3. Verify: test passes, training runs correctly, no regressions.
  4. Check for collateral — did the fix break other modalities or trainers?
  5. Before opening the PR or pushing a substantive update: run /veomni-review over the branch diff.
Phase 5: Knowledge Capture (mandatory)

Do this immediately after the fix is verified. Knowledge decays fast.

  • New hard constraint? → add to .agents/knowledge/constraints.md
  • Architecture insight? → add to .agents/knowledge/architecture.md
  • Regression guard needed? → prefer adding a case to an existing CI-enumerated test; see .agents/knowledge/testing.md. Only paths not already covered by a directory-level CI entry need a new workflow line. Use the workflow that owns the path, including the e2e workflows for end-to-end tests; account for the GPU/NPU differences in that table.
  • Docs outdated? → update docs/ if the fix changes API behavior, config semantics, or usage patterns

If none apply, explicitly note "no new knowledge to capture."


Stop Conditions

Restart from Phase 1 if you catch yourself thinking "let me just try changing X and see", "quick fix for now, clean up later", or "it probably works, moving on".

After 3 consecutive failed fix attempts, stop fixing symptoms. Question whether the underlying approach is wrong, re-examine whether you are solving the right problem, and report the analysis to the user before continuing.

Common Pitfalls

  • FSDP2 + gradient accumulation: gradients must be accumulated in the unsharded space — accumulating sharded gradients produces wrong results.
  • DCP checkpoint format: model state dict keys must match exactly between save and load — renamed parameters break checkpoint loading silently.
  • Multi-modality data collators: text-only collators crash on multimodal data and vice versa — always check data_collator type matches the dataset.
  • Sequence parallel: attention outputs must be gathered before loss computation — partial outputs produce incorrect loss values.
  • Patchgen: model patches in veomni/models/transformers/*/ are auto-generated — editing generated files directly will be overwritten.

Domain-Specific Checklists

Include the relevant checklist when investigating.

Distributed Training Correctness
  • Is the loss identical (within tolerance) between 1-GPU and multi-GPU runs?
  • Are ALL model parameters sharded correctly? (check parallel_plan)
  • Is gradient clipping applied in the correct coordinate space?
  • For sequence parallel: are attention masks split consistently across ranks?
  • For MoE: are expert assignments deterministic across runs with the same seed?
Numerical Correctness
  • Is there a reference implementation showing the SAME numbers?
  • Are ALL weights loaded? (check logs for missing/unexpected keys)
  • Is the comparison fair? (same inputs, same dtype, same parallelism)
  • Could there be a dtype mismatch? (float32 vs bfloat16 in computation)
  • Are there NaN/Inf values being silently masked or replaced?

Appendix: Verification Subagent

When confidence is low or evidence is ambiguous, launch a subagent to challenge your conclusion:

You are a critical reviewer. Your job is to find flaws in the following conclusion.

## Conclusion Under Review
<the specific claim or decision>

## Evidence Presented
<the data, logs, experiments supporting the conclusion>

## Your Task
1. Does the evidence actually support the conclusion, or just correlate?
2. Generate 2+ alternative explanations consistent with the same evidence.
3. What specific observation would DISPROVE this conclusion? Has it been checked?
4. Was the experiment controlled (one variable changed at a time)?

## Output
Verdict: CONFIRMED / CHALLENGED / INSUFFICIENT_EVIDENCE
Findings: [issues found, counter-hypotheses, missing evidence]

© ByteDance-Seed, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/veomni-debug of ByteDance-Seed/VeOmni.

Open the folder on GitHubat commit 16c94aa

Compare with similar skills

Veomni Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Veomni Debug compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Veomni Debug this skillByteDance-Seed/VeOmni2.2k—~2.8kAutomated safety check: PassApache-2.0
The Art of Debuggingstas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.0
Aoti Debugpytorch/pytorch104k1 repos~1.7kAutomated safety check: PassCustom licence
Ascendcascend-ai-coding/awesome-ascend-skills174—~3.5kAutomated safety check: PassNone
Systematic DebuggingChrisWiles/claude-code-showcase6.1k3 repos~1.2kAutomated safety check: PassNone
Debugging and Error Recoveryaddyosmani/agent-skills103k1 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes
  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • Ascendc

    ascend-ai-coding/awesome-ascend-skills

    End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.

    174 GitHub stars~3.5k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Systematic Debugging

    ChrisWiles/claude-code-showcase

    Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.

    6.1k GitHub starsUsed in 3 repos~1.2k tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    103k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed
  • Systematic Debugging

    ed3dai/ed3d-plugins

    A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes - four-phase framework (root cause investigation, pattern analysis, hypothesis…

    250 GitHub starsUsed in 3 repos~2.4k tokens
    DevelopmentAuto-check passed

More from ByteDance-Seed/VeOmni

All 10 skills in this repo
  • Create PR

    ByteDance-Seed/VeOmni

    Create a pull request for the current branch. An agent skill from ByteDance-Seed/VeOmni.

    2.2k GitHub stars~1.6k tokensUpdated today
    Auto-check: notes
  • Veomni New Model

    ByteDance-Seed/VeOmni

    A skill your agent uses when adding support for a new model to VeOmni.

    2.2k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Veomni New Op

    ByteDance-Seed/VeOmni

    A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/.

    2.2k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Veomni Patchgen Model

    ByteDance-Seed/VeOmni

    Author or refresh a VeOmni model's patchgen-generated modeling under generated/ — GPU and/or NPU config, dense or MoE, text / VLM / Omni.

    2.2k GitHub stars~9.6k tokensUpdated today
    Auto-check passed
  • Veomni Profile

    ByteDance-Seed/VeOmni

    A skill your agent uses for performance profiling and optimization.

    2.2k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Veomni Review

    ByteDance-Seed/VeOmni

    Pre-PR code review gate. An agent skill from ByteDance-Seed/VeOmni.

    2.2k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Veomni Debug

What does Veomni Debug do?

A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…. Veomni Debug is an agent skill from ByteDance-Seed/VeOmni. Use this skill for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected behavior.

When should I use Veomni Debug?

Veomni Debug fits situations like: loss divergence; gradient explosion; distributed training hang; checkpoint load failure.

How do I install Veomni Debug in Claude Code?

Run `npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a claude-code`. Or copy the skill folder (.agents/skills/veomni-debug in ByteDance-Seed/VeOmni) into .claude/skills/veomni-debug in your project. Claude Code loads it when a task matches its description.

How do I install Veomni Debug in Codex?

Run `npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a codex`. Or copy the skill folder (.agents/skills/veomni-debug in ByteDance-Seed/VeOmni) into .agents/skills/veomni-debug in your project. Codex loads it when a task matches its description.

Can I use Veomni Debug in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ByteDance-Seed/VeOmni --skill veomni-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/veomni-debug, .gemini/skills/veomni-debug, .github/skills/veomni-debug and .opencode/skills/veomni-debug in your project.

What does Veomni Debug need to run?

Going by SKILL.md and its folder, Veomni Debug needs the command-line tools its instructions call (uv, make, git and pytest). Our summary lists: Python 3.

Does Veomni Debug access the network?

SKILL.md contains no URLs. Its commands use uv and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Veomni Debug safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Veomni Debug use?

Veomni Debug is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Veomni Debug use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Veomni Debug?

Skills that share tags, products or a category with Veomni Debug: The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars), Aoti Debug (pytorch/pytorch, 104k stars), Ascendc (ascend-ai-coding/awesome-ascend-skills, 174 stars) and Systematic Debugging (ChrisWiles/claude-code-showcase, 6.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Veomni Debug?

ByteDance-Seed (a GitHub organization) maintains it in ByteDance-Seed/VeOmni, which has 2,235 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 9, 2026.

Source: ByteDance-Seed/VeOmni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.