Agent skill

Validate Correctness

by mlc-ai in mlc-ai/pith-train

Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

Apache-2.0Auto-check passedTesting & QA

Install Validate Correctness

skills CLI
$ npx skills add mlc-ai/pith-train --skill validate-correctness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mlc-ai/pith-train validate-correctness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/validate-correctness .claude/skills/validate-correctness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
validate-correctness
GitHub stars
355
Token cost
~2.3k tokens
SKILL.md length
1,097 words
Files
4 (incl. scripts)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

  • Works in 4 steps: Create the three run files → Run the base arm twice → Run the feature arm → …
  • User asks to validate correctness
  • SKILL.md covers Method, Prerequisites, Step 1: Create the three run… and Step 2: Run the base arm twice, plus 5 more sections
  • Runs Python and Shell scripts from its folder; calls git, bash and python3

What it does

Validate Correctness is an agent skill from mlc-ai/pith-train. Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope. Use when user asks to "validate correctness", "check if changes break training", "compare loss curves", "run a regression test", or "verify my changes are correct". For throughput use validate-performance instead. The user specifies which model to validate, at which parallelism mesh (PP/EP/CP), and at which sequence length — do not infer any of it from git diff.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/compare.py`, `scripts/launch.sh` and `templates/validate.py`).

It sits in Testing & QA. It works with Git. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.

When your agent uses it

  • User asks to validate correctness
  • Check if changes break training
  • Compare loss curves
  • Run a regression test

Example prompts

  • “validate correctness”
  • “check if changes break training”
  • “compare loss curves”
  • “/validate-correctness”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Create the three run files
  2. Run the base arm twice
  3. Run the feature arm
  4. Compare

What it can do on your machine

Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • bash
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Validate Correctness loads about 2.3k tokens when it runs. Until then it costs about 129 tokens; SKILL.md has 1,097 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~129
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 1,097 words, ~2,324 tokens.

Download SKILL.mdSave it as .claude/skills/validate-correctness/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
validate-correctness
description
Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope. Use when user asks to "validate correctness", "check if changes break training", "compare loss curves", "run a regression test", or "verify my changes are correct". For throughput use validate-performance instead. The user specifies which model to validate, at which parallelism mesh (PP/EP/CP), and at which sequence length — do not infer any of it from git diff.

Validate Correctness

Runs the same short training three times — the base branch twice, the feature branch once — and judges the base-vs-feature difference against the base-vs-base difference.

The user picks the model, the mesh and the sequence length, e.g. "validate deepseek-v2-lite at pp=2 ep=2 seq=4096". Sequence length is filled per run, defaulting to 2048 when the user expresses no preference. Dtype is bf16 unless asked otherwise, and is the one dimension carried in the template as a real value. If the model or the mesh is vague, ask.

Runs live under <model> as a directory, named by a <tag> that spells out every dimension: pp2-dp1-cp1-ep8-seq2048-bf16. Include the defaults, so a log identifies itself. DP is not a config field — it is inferred from the world size and appears in the tag for identification only. The three runs are base0, base1 and feat0, and the wandb group is correctness/<model>.

Method

Two runs of identical code do not agree past step 1. FlashAttention's backward is not reproducible even under torch.use_deterministic_algorithms(True), which reaches neither it nor the Triton kernels; measured on this codebase the forward is bit-reproducible while gradient norms differ ~8% between two identical runs.

So a base-vs-feature delta means nothing on its own. base0 and base1 are identical code, so their difference is the run-to-run floor, and the feature difference is read as a ratio against it. A fixed tolerance cannot substitute: measured drift spans 3.6x across meshes, and --tolerance 5e-3 both rejected a change that four meshes later showed to be indistinguishable and rejected identical code on two of four meshes.

Prerequisites

  • Activate .venv in the repo root: source .venv/bin/activate.
  • A tokenized corpus for the model — run setup-benchmark-inputs if workspace/datasets/dclm-baseline/toktxt/<tokenizer> is missing. No checkpoint is needed; runs start from fresh weights.
  • world_size % PP == 0 with CP and EP each dividing world_size / PP.
  • A clean tree. The arms are switched with git checkout, so commit the feature work first — pushed or not. Uncommitted changes either follow you onto the base branch or block the checkout.

Runs must be strictly sequential, since one tree serves both arms. Never queue all three and let the scheduler interleave them: a checkout landing while a run is pending or in flight silently executes the wrong branch. Each run must finish before the next starts.

<tokenizer> is not <model>, and <moe-load-balance-type> differs per model:

<model><tokenizer><moe-load-balance-type>
qwen3-30b-a3bqwen3global-batch
qwen3.5-35b-a3bqwen3.5global-batch
deepseek-v2-litedeepseek-v2sequence
gpt-oss-20b, gpt-oss-120bgpt-ossglobal-batch

Step 1: Create the three run files

Copy the template to base0, fill in every <placeholder>, then copy that filled file to base1 and feat0. Filling once and copying afterwards makes the three identical by construction rather than by discipline. An unfilled copy fails unevenly, so scan the file before launching: the numeric fills — the mesh sizes, <sequence-length> and <global-batch-size> — are unquoted, so a leftover there is a SyntaxError on every rank, while the quoted ones — <model>, <tokenizer>, <moe-load-balance-type>, <wandb-project> — parse and surface far later, a forgotten <wandb-project> only at wandb.init.

bash
R=$(git rev-parse --show-toplevel)
W=$R/workspace/validate-correctness/<model>
mkdir -p $W $R/logging/validate-correctness/<model>
cp $R/.agents/skills/validate-correctness/templates/validate.py $W/<tag>-base0.py
cp $R/.agents/skills/validate-correctness/scripts/launch.sh $W/launch.sh
git rev-parse --abbrev-ref HEAD > $W/FEATURE

The launcher is copied into workspace/ because the checkouts below switch branches, and .agents/ is tracked.

Fill the placeholders in that file, then:

bash
cp $W/<tag>-base0.py $W/<tag>-base1.py
cp $W/<tag>-base0.py $W/<tag>-feat0.py

Each run takes its wandb name from its own filename, so nothing inside the three files differs. Confirm that, because it is the whole basis of the comparison:

bash
diff $W/<tag>-base0.py $W/<tag>-base1.py
diff $W/<tag>-base0.py $W/<tag>-feat0.py

Both must print nothing. workspace/ and *.log are both gitignored, so the run files and the logs survive every checkout below without ever making the tree dirty.

Step 2: Run the base arm twice

Each fence re-derives its own paths, because shell variables do not survive between commands.

bash
R=$(git rev-parse --show-toplevel); W=$R/workspace/validate-correctness/<model>
G=$R/logging/validate-correctness/<model>/<tag>
git checkout main
bash $W/launch.sh $W/<tag>-base0.py 2>&1 | tee $G-base0.log
bash $W/launch.sh $W/<tag>-base1.py 2>&1 | tee $G-base1.log

Under SLURM, wrap each in srun — see launch-with-slurm for the flags that matter:

bash
srun -N <nodes> -W 0 -o $G-base0.log bash $W/launch.sh $W/<tag>-base0.py
srun -N <nodes> -W 0 -o $G-base1.log bash $W/launch.sh $W/<tag>-base1.py

Step 3: Run the feature arm

bash
R=$(git rev-parse --show-toplevel); W=$R/workspace/validate-correctness/<model>
G=$R/logging/validate-correctness/<model>/<tag>
FEATURE=$(cat $W/FEATURE)
git checkout $FEATURE
bash $W/launch.sh $W/<tag>-feat0.py 2>&1 | tee $G-feat0.log
Show full SKILL.md (485 more words)Show less

Step 4: Compare

bash
R=$(git rev-parse --show-toplevel); G=$R/logging/validate-correctness/<model>/<tag>
python3 $R/.agents/skills/validate-correctness/scripts/compare.py $G-base0.log $G-base1.log $G-feat0.log

The verdict is the ratio of mean |delta| for base-vs-feature over base-vs-base:

  • below 3x — PASS, indistinguishable from run-to-run drift.
  • 3x to 5x — INVESTIGATE. Add a base2 run: one envelope is a point estimate, and a 2.92x reading on one mesh was contradicted by 0.25x, 0.59x and 1.14x on three others.
  • above 5x — FAIL.

The step-1 row is reported, not gated. It precedes any optimizer update, so its floor is normally zero and its signal is exactly the numerical difference the change makes to the forward — nonzero means the forward moved, which is expected for a reordered reduction or a swapped kernel and unexpected otherwise. Judge it yourself; a real forward regression also shows up in the 32-step ratio.

Constraints

max_steps * global_batch_size must not exceed the corpus sample count, which scales inversely with sequence length — one tokenized DCLM shard yields ~18,000 samples at 4096 and roughly twice that at 2048. A run that does not fit raises at startup, but only after the JIT warmup.

sequence_length % (2 * cp_size) == 0 for the zigzag split, and global_batch / dp >= 2 * pp (where dp = world_size / (pp * cp)).

Keep sequence length and global batch identical across meshes, so meshes stay comparable to each other and not only within themselves.

The fixture

templates/validate.py sets 32 steps and a linear warmup from 1e-6 to 1e-5 with no decay. Change these before any run, or not at all: runs are only comparable if they agree on them.

Warmup is load-bearing rather than cosmetic. From fresh weights a constant LR spikes hard in the first steps — on qwen3-30b-a3b at pp=2 ep=8, 12.33 to 19.90 with a pre-clip gradient norm of 1388 at 1e-4 — and lowering the LR tenfold barely helps, because AdamW's second moment is near zero early so the effective step lr/(sqrt(v)+eps) saturates whatever lr is. With warmup the same run descends monotonically and peaks at a gradient norm of 49, which is what makes a ~0.01 envelope measurable at all.

Correctness uses real routing, so the template leaves training.benchmark = False. Never turn it on here: force-balanced routing overwrites the router's top-k with a round-robin over token index, so the routing path goes untested and load-balance-loss pins at exactly 1.000000 — a gate that can never fail. Throughput measurement wants it on, which is why it lives in validate-performance.

Multi-node notes

If all three runs are on one machine this section does not apply.

Prefer all three on the same nodes, but do not require it. If they differ, the base-vs-base envelope absorbs node variation too, which makes the test more conservative: a within-envelope verdict stays trustworthy, and only an outside-envelope verdict needs a same-node re-run before attributing it to the code.

The three runs are independent, so one that fails can be re-run on its own and the others stay valid — check out the matching branch first. Re-running the whole set instead makes it only as durable as its unluckiest run.

© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in .agents/skills/validate-correctness of mlc-ai/pith-train.

  • SKILL.md
  • scripts/compare.py
  • scripts/launch.sh
  • templates/validate.py

Open the folder on GitHubat commit c7c8b1d

Compare with similar skills

Validate Correctness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Validate Correctness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Validate Correctness this skillmlc-ai/pith-train355—~2.3kAutomated safety check: PassApache-2.0
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
Evaluate PR Testsdotnet/maui23k—~2.9kAutomated safety check: PassMIT
Update Megatron Golden ValuesNVIDIA/Megatron-LM18k—~2.8kAutomated safety check: PassCustom licence
TiDB Test Diff Triagepingcap/tidb41k—~498Automated safety check: PassApache-2.0
SimpleITK Binary Data UploadSimpleITK/SimpleITK1.1k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • Official

    Reviews the tests added in a pull request for fix coverage, quality, edge cases and test type, and recommends lighter test types where they would do.

    23k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description.

    18k GitHub stars~2.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Investigates TiDB plan or test-result diffs that the change does not explain, ruling out failpoint setup and merge effects before expected outputs are updated.

    41k GitHub stars~498 tokensUpdated today
    Testing & QAAuto-check passed
  • SimpleITK Binary Data Upload

    SimpleITK/SimpleITK

    Uploads a binary test file to the SimpleITK ExternalData repository by hashing it with SHA-512, staging it in the object store, writing a content-link file and opening a draft PR.

    1.1k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Continuously maintain automatic NemoClaw main E2E results through coordinated repairs.

    23k GitHub stars~2.6k tokensUpdated today
    Testing & QAAuto-check passed

More from mlc-ai/pith-train

All 10 skills in this repo
  • Analyze Nsys Profile

    mlc-ai/pith-train

    Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

    355 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Capture Nsys Profile

    mlc-ai/pith-train

    Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.

    355 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed
  • Validate Performance

    mlc-ai/pith-train

    Measures the throughput difference between two branches with force-balanced routing.

    355 GitHub stars~1.4k tokensUpdated 3 days ago
    Auto-check passed
  • Setup Benchmark Inputs

    mlc-ai/pith-train

    Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.

    355 GitHub stars~399 tokensUpdated 3 days ago
    Auto-check passed
  • Add New Model

    mlc-ai/pith-train

    Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.

    355 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed
  • Wandb Tracking

    mlc-ai/pith-train

    Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.

    355 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check: warnings

Works with

Categories

Questions about Validate Correctness

What does Validate Correctness do?

Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope. Validate Correctness is an agent skill from mlc-ai/pith-train. Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

When should I use Validate Correctness?

Validate Correctness fits situations like: user asks to validate correctness; check if changes break training; compare loss curves; run a regression test.

How do I install Validate Correctness in Claude Code?

Run `npx skills add mlc-ai/pith-train --skill validate-correctness -a claude-code`. Or copy the skill folder (.agents/skills/validate-correctness in mlc-ai/pith-train) into .claude/skills/validate-correctness in your project. Claude Code loads it when a task matches its description.

How do I install Validate Correctness in Codex?

Run `npx skills add mlc-ai/pith-train --skill validate-correctness -a codex`. Or copy the skill folder (.agents/skills/validate-correctness in mlc-ai/pith-train) into .agents/skills/validate-correctness in your project. Codex loads it when a task matches its description.

Can I use Validate Correctness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill validate-correctness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-correctness, .gemini/skills/validate-correctness, .github/skills/validate-correctness and .opencode/skills/validate-correctness in your project.

What does Validate Correctness need to run?

Going by SKILL.md and its folder, Validate Correctness needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (git, bash and python3). Our summary lists: Python 3; A Bash shell.

Does Validate Correctness access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Validate Correctness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Validate Correctness use?

Validate Correctness is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Validate Correctness use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Validate Correctness?

Skills that share tags, products or a category with Validate Correctness: OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars), Evaluate PR Tests (dotnet/maui, 23k stars), Update Megatron Golden Values (NVIDIA/Megatron-LM, 18k stars) and TiDB Test Diff Triage (pingcap/tidb, 41k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Validate Correctness?

mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.

Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.