OpenHarness End-to-End Evals
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.
$ npx skills add mlc-ai/pith-train --skill validate-correctness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mlc-ai/pith-train validate-correctness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/validate-correctness .claude/skills/validate-correctness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "validate-correctness" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/validate-correctness into .claude/skills/validate-correctness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-correctness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/validate-correctnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mlc-ai/pith-train --skill validate-correctness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mlc-ai/pith-train validate-correctness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/validate-correctness .agents/skills/validate-correctness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "validate-correctness" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/validate-correctness into .agents/skills/validate-correctness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-correctness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mlc-ai/pith-train --skill validate-correctness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mlc-ai/pith-train validate-correctness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/validate-correctness .cursor/skills/validate-correctness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "validate-correctness" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/validate-correctness into .cursor/skills/validate-correctness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-correctness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mlc-ai/pith-train.git --path .agents/skills/validate-correctness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mlc-ai/pith-train --skill validate-correctness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mlc-ai/pith-train validate-correctness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/validate-correctness .gemini/skills/validate-correctness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "validate-correctness" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/validate-correctness into .gemini/skills/validate-correctness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-correctness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mlc-ai/pith-train validate-correctnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mlc-ai/pith-train --skill validate-correctness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/validate-correctness .github/skills/validate-correctness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "validate-correctness" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/validate-correctness into .github/skills/validate-correctness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-correctness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mlc-ai/pith-train --skill validate-correctness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mlc-ai/pith-train validate-correctness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/validate-correctness .opencode/skills/validate-correctness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "validate-correctness" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/validate-correctness into .opencode/skills/validate-correctness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-correctness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
validate-correctnessValidates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.
Validate Correctness is an agent skill from mlc-ai/pith-train. Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope. Use when user asks to "validate correctness", "check if changes break training", "compare loss curves", "run a regression test", or "verify my changes are correct". For throughput use validate-performance instead. The user specifies which model to validate, at which parallelism mesh (PP/EP/CP), and at which sequence length — do not infer any of it from git diff.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/compare.py`, `scripts/launch.sh` and `templates/validate.py`).
It sits in Testing & QA. It works with Git. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
gitbashpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Validate Correctness loads about 2.3k tokens when it runs. Until then it costs about 129 tokens; SKILL.md has 1,097 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 1,097 words, ~2,324 tokens.
.claude/skills/validate-correctness/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Runs the same short training three times — the base branch twice, the feature branch once — and judges the base-vs-feature difference against the base-vs-base difference.
The user picks the model, the mesh and the sequence length, e.g. "validate deepseek-v2-lite at pp=2 ep=2 seq=4096". Sequence length is filled per run, defaulting to 2048 when the user expresses no preference. Dtype is bf16 unless asked otherwise, and is the one dimension carried in the template as a real value. If the model or the mesh is vague, ask.
Runs live under <model> as a directory, named by a <tag> that spells out every dimension: pp2-dp1-cp1-ep8-seq2048-bf16. Include the defaults, so a log identifies itself. DP is not a config field — it is inferred from the world size and appears in the tag for identification only. The three runs are base0, base1 and feat0, and the wandb group is correctness/<model>.
Two runs of identical code do not agree past step 1. FlashAttention's backward is not reproducible even under torch.use_deterministic_algorithms(True), which reaches neither it nor the Triton kernels; measured on this codebase the forward is bit-reproducible while gradient norms differ ~8% between two identical runs.
So a base-vs-feature delta means nothing on its own. base0 and base1 are identical code, so their difference is the run-to-run floor, and the feature difference is read as a ratio against it. A fixed tolerance cannot substitute: measured drift spans 3.6x across meshes, and --tolerance 5e-3 both rejected a change that four meshes later showed to be indistinguishable and rejected identical code on two of four meshes.
.venv in the repo root: source .venv/bin/activate.workspace/datasets/dclm-baseline/toktxt/<tokenizer> is missing. No checkpoint is needed; runs start from fresh weights.world_size % PP == 0 with CP and EP each dividing world_size / PP.git checkout, so commit the feature work first — pushed or not. Uncommitted changes either follow you onto the base branch or block the checkout.Runs must be strictly sequential, since one tree serves both arms. Never queue all three and let the scheduler interleave them: a checkout landing while a run is pending or in flight silently executes the wrong branch. Each run must finish before the next starts.
<tokenizer> is not <model>, and <moe-load-balance-type> differs per model:
<model> | <tokenizer> | <moe-load-balance-type> |
|---|---|---|
qwen3-30b-a3b | qwen3 | global-batch |
qwen3.5-35b-a3b | qwen3.5 | global-batch |
deepseek-v2-lite | deepseek-v2 | sequence |
gpt-oss-20b, gpt-oss-120b | gpt-oss | global-batch |
Copy the template to base0, fill in every <placeholder>, then copy that filled file to base1 and feat0. Filling once and copying afterwards makes the three identical by construction rather than by discipline. An unfilled copy fails unevenly, so scan the file before launching: the numeric fills — the mesh sizes, <sequence-length> and <global-batch-size> — are unquoted, so a leftover there is a SyntaxError on every rank, while the quoted ones — <model>, <tokenizer>, <moe-load-balance-type>, <wandb-project> — parse and surface far later, a forgotten <wandb-project> only at wandb.init.
R=$(git rev-parse --show-toplevel)
W=$R/workspace/validate-correctness/<model>
mkdir -p $W $R/logging/validate-correctness/<model>
cp $R/.agents/skills/validate-correctness/templates/validate.py $W/<tag>-base0.py
cp $R/.agents/skills/validate-correctness/scripts/launch.sh $W/launch.sh
git rev-parse --abbrev-ref HEAD > $W/FEATUREThe launcher is copied into workspace/ because the checkouts below switch branches, and .agents/ is tracked.
Fill the placeholders in that file, then:
cp $W/<tag>-base0.py $W/<tag>-base1.py
cp $W/<tag>-base0.py $W/<tag>-feat0.pyEach run takes its wandb name from its own filename, so nothing inside the three files differs. Confirm that, because it is the whole basis of the comparison:
diff $W/<tag>-base0.py $W/<tag>-base1.py
diff $W/<tag>-base0.py $W/<tag>-feat0.pyBoth must print nothing. workspace/ and *.log are both gitignored, so the run files and the logs survive every checkout below without ever making the tree dirty.
Each fence re-derives its own paths, because shell variables do not survive between commands.
R=$(git rev-parse --show-toplevel); W=$R/workspace/validate-correctness/<model>
G=$R/logging/validate-correctness/<model>/<tag>
git checkout main
bash $W/launch.sh $W/<tag>-base0.py 2>&1 | tee $G-base0.log
bash $W/launch.sh $W/<tag>-base1.py 2>&1 | tee $G-base1.logUnder SLURM, wrap each in srun — see launch-with-slurm for the flags that matter:
srun -N <nodes> -W 0 -o $G-base0.log bash $W/launch.sh $W/<tag>-base0.py
srun -N <nodes> -W 0 -o $G-base1.log bash $W/launch.sh $W/<tag>-base1.pyR=$(git rev-parse --show-toplevel); W=$R/workspace/validate-correctness/<model>
G=$R/logging/validate-correctness/<model>/<tag>
FEATURE=$(cat $W/FEATURE)
git checkout $FEATURE
bash $W/launch.sh $W/<tag>-feat0.py 2>&1 | tee $G-feat0.logR=$(git rev-parse --show-toplevel); G=$R/logging/validate-correctness/<model>/<tag>
python3 $R/.agents/skills/validate-correctness/scripts/compare.py $G-base0.log $G-base1.log $G-feat0.logThe verdict is the ratio of mean |delta| for base-vs-feature over base-vs-base:
base2 run: one envelope is a point estimate, and a 2.92x reading on one mesh was contradicted by 0.25x, 0.59x and 1.14x on three others.The step-1 row is reported, not gated. It precedes any optimizer update, so its floor is normally zero and its signal is exactly the numerical difference the change makes to the forward — nonzero means the forward moved, which is expected for a reordered reduction or a swapped kernel and unexpected otherwise. Judge it yourself; a real forward regression also shows up in the 32-step ratio.
max_steps * global_batch_size must not exceed the corpus sample count, which scales inversely with sequence length — one tokenized DCLM shard yields ~18,000 samples at 4096 and roughly twice that at 2048. A run that does not fit raises at startup, but only after the JIT warmup.
sequence_length % (2 * cp_size) == 0 for the zigzag split, and global_batch / dp >= 2 * pp (where dp = world_size / (pp * cp)).
Keep sequence length and global batch identical across meshes, so meshes stay comparable to each other and not only within themselves.
templates/validate.py sets 32 steps and a linear warmup from 1e-6 to 1e-5 with no decay. Change these before any run, or not at all: runs are only comparable if they agree on them.
Warmup is load-bearing rather than cosmetic. From fresh weights a constant LR spikes hard in the first steps — on qwen3-30b-a3b at pp=2 ep=8, 12.33 to 19.90 with a pre-clip gradient norm of 1388 at 1e-4 — and lowering the LR tenfold barely helps, because AdamW's second moment is near zero early so the effective step lr/(sqrt(v)+eps) saturates whatever lr is. With warmup the same run descends monotonically and peaks at a gradient norm of 49, which is what makes a ~0.01 envelope measurable at all.
Correctness uses real routing, so the template leaves training.benchmark = False. Never turn it on here: force-balanced routing overwrites the router's top-k with a round-robin over token index, so the routing path goes untested and load-balance-loss pins at exactly 1.000000 — a gate that can never fail. Throughput measurement wants it on, which is why it lives in validate-performance.
If all three runs are on one machine this section does not apply.
Prefer all three on the same nodes, but do not require it. If they differ, the base-vs-base envelope absorbs node variation too, which makes the test more conservative: a within-envelope verdict stays trustworthy, and only an outside-envelope verdict needs a same-node re-run before attributing it to the code.
The three runs are independent, so one that fails can be re-run on its own and the others stay valid — check out the matching branch first. Re-running the whole set instead makes it only as durable as its unluckiest run.
© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in .agents/skills/validate-correctness of mlc-ai/pith-train.
Open the folder on GitHubat commit c7c8b1d
Validate Correctness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Validate Correctness this skillmlc-ai/pith-train | 355 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| Evaluate PR Testsdotnet/maui | 23k | — | ~2.9k | Automated safety check: Pass | MIT | |
| Update Megatron Golden ValuesNVIDIA/Megatron-LM | 18k | — | ~2.8k | Automated safety check: Pass | Custom licence | |
| TiDB Test Diff Triagepingcap/tidb | 41k | — | ~498 | Automated safety check: Pass | Apache-2.0 | |
| SimpleITK Binary Data UploadSimpleITK/SimpleITK | 1.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 |
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
dotnet/maui
Reviews the tests added in a pull request for fix coverage, quality, edge cases and test type, and recommends lighter test types where they would do.
NVIDIA/Megatron-LM
Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description.
pingcap/tidb
Investigates TiDB plan or test-result diffs that the change does not explain, ruling out failpoint setup and merge effects before expected outputs are updated.
SimpleITK/SimpleITK
Uploads a binary test file to the SimpleITK ExternalData repository by hashing it with SHA-512, staging it in the object store, writing a content-link file and opening a draft PR.
NVIDIA/NemoClaw
Continuously maintain automatic NemoClaw main E2E results through coordinated repairs.
mlc-ai/pith-train
Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.
mlc-ai/pith-train
Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.
mlc-ai/pith-train
Measures the throughput difference between two branches with force-balanced routing.
mlc-ai/pith-train
Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.
mlc-ai/pith-train
Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.
mlc-ai/pith-train
Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.
Works with
Categories
Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope. Validate Correctness is an agent skill from mlc-ai/pith-train. Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.
Validate Correctness fits situations like: user asks to validate correctness; check if changes break training; compare loss curves; run a regression test.
Run `npx skills add mlc-ai/pith-train --skill validate-correctness -a claude-code`. Or copy the skill folder (.agents/skills/validate-correctness in mlc-ai/pith-train) into .claude/skills/validate-correctness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mlc-ai/pith-train --skill validate-correctness -a codex`. Or copy the skill folder (.agents/skills/validate-correctness in mlc-ai/pith-train) into .agents/skills/validate-correctness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill validate-correctness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-correctness, .gemini/skills/validate-correctness, .github/skills/validate-correctness and .opencode/skills/validate-correctness in your project.
Going by SKILL.md and its folder, Validate Correctness needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (git, bash and python3). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Validate Correctness is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Validate Correctness: OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars), Evaluate PR Tests (dotnet/maui, 23k stars), Update Megatron Golden Values (NVIDIA/Megatron-LM, 18k stars) and TiDB Test Diff Triage (pingcap/tidb, 41k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.
Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.