ML Failure Debugger
Leeroo-AI/superml
Diagnoses failing ML and AI work, such as OOM, NaN, divergence, crashes, slow throughput, wrong outputs and dependency conflicts, with every claim backed by documentation citations.
Route verl-omni training/inference consistency checks through MindStudio's MSProbe collection and root-cause analysis skills.
$ npx skills add verl-project/verl-omni --skill train-infer-consistency -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install verl-project/verl-omni train-infer-consistency --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/train-infer-consistency .claude/skills/train-infer-consistency && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "train-infer-consistency" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/train-infer-consistency into .claude/skills/train-infer-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-infer-consistency", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/verl-project/verl-omni/tree/main/.agents/skills/train-infer-consistencyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add verl-project/verl-omni --skill train-infer-consistency -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install verl-project/verl-omni train-infer-consistency --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/train-infer-consistency .agents/skills/train-infer-consistency && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "train-infer-consistency" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/train-infer-consistency into .agents/skills/train-infer-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-infer-consistency", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add verl-project/verl-omni --skill train-infer-consistency -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install verl-project/verl-omni train-infer-consistency --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/train-infer-consistency .cursor/skills/train-infer-consistency && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "train-infer-consistency" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/train-infer-consistency into .cursor/skills/train-infer-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-infer-consistency", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/verl-project/verl-omni.git --path .agents/skills/train-infer-consistency--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add verl-project/verl-omni --skill train-infer-consistency -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install verl-project/verl-omni train-infer-consistency --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/train-infer-consistency .gemini/skills/train-infer-consistency && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "train-infer-consistency" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/train-infer-consistency into .gemini/skills/train-infer-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-infer-consistency", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install verl-project/verl-omni train-infer-consistencyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add verl-project/verl-omni --skill train-infer-consistency -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/train-infer-consistency .github/skills/train-infer-consistency && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "train-infer-consistency" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/train-infer-consistency into .github/skills/train-infer-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-infer-consistency", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add verl-project/verl-omni --skill train-infer-consistency -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install verl-project/verl-omni train-infer-consistency --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/train-infer-consistency .opencode/skills/train-infer-consistency && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "train-infer-consistency" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/train-infer-consistency into .opencode/skills/train-infer-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-infer-consistency", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
train-infer-consistencyRoute verl-omni training/inference consistency checks through MindStudio's MSProbe collection and root-cause analysis skills.
Train Infer Consistency is an agent skill from verl-project/verl-omni. Route verl-omni training/inference consistency checks through MindStudio's MSProbe collection and root-cause analysis skills. Use when collecting paired rollout/actor dumps or investigating numerical differences in diffusion or omni models.
Its SKILL.md is about 860 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Root cause analysis. The repository describes itself as: Multimodal RL training framework for diffusion & omni models. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 022b110. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Train Infer Consistency loads about 860 tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 361 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from verl-project/verl-omni at commit 022b110, republished under its Apache-2.0 licence (© verl-project). 361 words, ~860 tokens.
.claude/skills/train-infer-consistency/SKILL.md (or your agent's skills folder).The MindStudio skills below own collection and analysis. Read them in order; this skill connects their inputs and outputs without repeating their procedures.
| Stage | Skill |
|---|---|
| Collect dumps | verl-omni-msprobe-dump |
| Analyze differences | rl-consistency-analysis |
MSProbe (mindstudio-probe) is required for data collection. If missing, identify
the Python environment used for the user's verl-omni task and install it using
that environment's package manager and workflow (e.g., uv or pip).
If the user provides an experiment directory or logs, check the saved config for
calculate_log_probs=true and inspect that run's rollout_corr/* metrics for
initial evidence of differences. Record the source and affected steps/timesteps.
Missing metrics or bypassed actor log-prob recomputation is inconclusive.
Without previous artifacts, proceed directly to collection.
Prefer installed skills or an existing local msagent checkout. Read each
SKILL.md; if missing, fetch the complete skill directory and referenced resources
from the linked repository, preserving scripts/, references/, and relative
paths. Record the source/version used. If unavailable, report what is missing
rather than reconstructing the procedure from memory.
Follow the collection skill using the user's launch script and current verl-omni source. Collect full tensor data for both sides within the diagnostic window, using Step 0's findings, if available, to guide reproduction and fine-grained analysis. Establish sample pairing through correlation logs, not aggregate metrics. Produce the diagnostic wrapper, both dumps, and correlation logs. Existing artifacts may be reused after passing that skill's checks; if either side is missing, fix collection and rerun.
Confirm the dumps represent the same sample and computation, with comparable inputs and weights. Pass the paired dumps, correlation evidence, and layout differences to the analysis skill. Follow its module-mapping and script workflow, then interpret the results against current source. Statistics support screening; elementwise conclusions require tensor evidence.
Provide actual paths to the diagnostic script, dumps, module mapping, and
output_5_root_cause_report.md. Include the screening metrics and diagnostic
config changes. Explain the pairing, evidence, limitations, and
next steps. If execution is unavailable, complete the feasible integration work
and identify unverified steps without claiming collection or analysis succeeded.
<!--
MAINTAINER GUIDE — Keep this skill a router. Collection and analysis procedures
belong to the linked MindStudio skills. Recheck links and artifact handoff when
their directory layout or output contract changes.
-->
© verl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/train-infer-consistency of verl-project/verl-omni.
Open the folder on GitHubat commit 022b110
Train Infer Consistency next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Train Infer Consistency this skillverl-project/verl-omni | 1.2k | — | ~860 | Automated safety check: Pass | Apache-2.0 | |
| ML Failure DebuggerLeeroo-AI/superml | 195 | — | ~9.5k | Automated safety check: Pass | Apache-2.0 | |
| Do Not Retry Without Diagnosisaiming-lab/MetaClaw | 3.5k | — | ~189 | Automated safety check: Pass | MIT | |
| The Art of Debuggingstas00/the-art-of-debugging | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | |
| Xpu CI Health Checkintel/torch-xpu-ops | 115 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Datadog Query Recipeslangfuse/langfuse | 36k | — | ~824 | Automated safety check: Pass | Custom licence |
Leeroo-AI/superml
Diagnoses failing ML and AI work, such as OOM, NaN, divergence, crashes, slow throughput, wrong outputs and dependency conflicts, with every claim backed by documentation citations.
aiming-lab/MetaClaw
Common mistake — retrying the same failing command or API call without understanding why it failed.
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
intel/torch-xpu-ops
Check PyTorch ciflow/xpu (xpu.yml) on the main branch, collect the failing XPU test cases from the most recent completed run(s), analyze the ROOT CAUSE of each failure with AI, and produce a list…
langfuse/langfuse
Research Langfuse production telemetry with reusable Datadog queries.
facebookexperimental/triton
Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…
verl-project/verl-omni
Router for adding a diffusion or omni pipeline to verl-omni.
verl-project/verl-omni
Guide for adding a new reward scorer to verl-omni and wiring it into a run.
verl-project/verl-omni
Route a verl-omni performance investigation to the right tool and capture a usable trace.
verl-project/verl-omni
How to write and run verl-omni CPU tests (testoncpu.py) that exercise adapters, rewards, and configs without a GPU or model weights.
verl-project/verl-omni
verl-omni commit message + PR conventions and the mandatory contribution policy.
verl-project/verl-omni
Review your own verl-omni branch against the project rubric before opening or updating a PR.
Categories
Route verl-omni training/inference consistency checks through MindStudio's MSProbe collection and root-cause analysis skills. Train Infer Consistency is an agent skill from verl-project/verl-omni. Route verl-omni training/inference consistency checks through MindStudio's MSProbe collection and root-cause analysis skills.
Train Infer Consistency fits situations like: collecting paired rollout/actor dumps; investigating numerical differences in diffusion.
Run `npx skills add verl-project/verl-omni --skill train-infer-consistency -a claude-code`. Or copy the skill folder (.agents/skills/train-infer-consistency in verl-project/verl-omni) into .claude/skills/train-infer-consistency in your project. Claude Code loads it when a task matches its description.
Run `npx skills add verl-project/verl-omni --skill train-infer-consistency -a codex`. Or copy the skill folder (.agents/skills/train-infer-consistency in verl-project/verl-omni) into .agents/skills/train-infer-consistency in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add verl-project/verl-omni --skill train-infer-consistency -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/train-infer-consistency, .gemini/skills/train-infer-consistency, .github/skills/train-infer-consistency and .opencode/skills/train-infer-consistency in your project.
SKILL.md names no scripts, command-line tools or credentials: Train Infer Consistency is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Train Infer Consistency is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 860 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Train Infer Consistency: ML Failure Debugger (Leeroo-AI/superml, 195 stars), Do Not Retry Without Diagnosis (aiming-lab/MetaClaw, 3.5k stars), The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars) and Xpu CI Health Check (intel/torch-xpu-ops, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
verl-project (a GitHub organization) maintains it in verl-project/verl-omni, which has 1,210 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.
Source: verl-project/verl-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.