Agent skill

Validate Performance

by mlc-ai in mlc-ai/pith-train

Measures the throughput difference between two branches with force-balanced routing.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Validate Performance

skills CLI
$ npx skills add mlc-ai/pith-train --skill validate-performance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mlc-ai/pith-train validate-performance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/validate-performance .claude/skills/validate-performance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
validate-performance
GitHub stars
355
Token cost
~1.4k tokens
SKILL.md length
619 words
Files
4 (incl. scripts)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Measures the throughput difference between two branches with force-balanced routing.

  • Works in 3 steps: Create the two run files → Run both arms → Report
  • The user asks to benchmark the performance
  • SKILL.md covers Prerequisites, Step 1: Create the two run files, Step 2: Run both arms and Step 3: Report, plus 1 more section
  • Runs Python and Shell scripts from its folder; calls git, bash and python3

What it does

Validate Performance is an agent skill from mlc-ai/pith-train. Measures the throughput difference between two branches with force-balanced routing. Use when the user asks to "benchmark the performance", "measure throughput", "check for a slowdown", "check throughput did not regress", or "compare step time". The user specifies the model and the mesh; sequence length defaults to 2048 and dtype to bf16 — do not infer any of it from git diff.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/compare.py`, `scripts/launch.sh` and `templates/validate.py`).

It sits in AI & LLM Engineering. It works with Git and Qwen. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.

When your agent uses it

  • The user asks to benchmark the performance
  • Measure throughput
  • Check for a slowdown
  • Check throughput did not regress

Example prompts

  • “benchmark the performance”
  • “measure throughput”
  • “check for a slowdown”
  • “/validate-performance”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Create the two run files
  2. Run both arms
  3. Report

What it can do on your machine

Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • bash
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Validate Performance loads about 1.4k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 619 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 619 words, ~1,446 tokens.

Download SKILL.mdSave it as .claude/skills/validate-performance/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
validate-performance
description
Measures the throughput difference between two branches with force-balanced routing. Use when the user asks to "benchmark the performance", "measure throughput", "check for a slowdown", "check throughput did not regress", or "compare step time". The user specifies the model and the mesh; sequence length defaults to 2048 and dtype to bf16 — do not infer any of it from git diff.

Validate Performance

Runs the same short training on the base branch and on the feature branch with force-balanced routing, then reports the step-time difference. It measures; it does not gate. For loss correctness see validate-correctness.

The user picks the model and the mesh, e.g. "benchmark qwen3-30b-a3b at pp=2 ep=8". Sequence length defaults to 2048 and dtype to bf16; take either from the user when given. If the model or the mesh is vague, ask.

Runs live under <model> as a directory, named by a <tag> that spells out every dimension: pp2-dp1-cp1-ep8-seq2048-bf16. Include the defaults, so a log identifies itself. DP is not a config field — it is inferred from the world size and appears in the tag for identification only.

The wandb run is grouped as performance/<model>, so several meshes for one model land in one group and stay directly comparable, while the correctness pass groups separately and can never be mistaken for it.

Prerequisites

  • Activate .venv in the repo root: source .venv/bin/activate.
  • A tokenized corpus for the model — run setup-benchmark-inputs if workspace/datasets/dclm-baseline/toktxt/<tokenizer> is missing. No checkpoint is needed.
  • world_size % PP == 0 with CP and EP each dividing world_size / PP, sequence_length % (2 * cp_size) == 0, and global_batch / dp >= 2 * pp (where dp = world_size / (pp * cp)).
  • Keep sequence length, dtype and global batch identical across the two arms, and across meshes you intend to compare to each other.
  • A clean tree. The arms are switched with git checkout, so commit the feature work first.

The two runs must be strictly sequential. One tree serves both arms, so a checkout landing while a run is pending or in flight silently benchmarks the wrong branch.

<tokenizer> is not <model>, and <moe-load-balance-type> differs per model:

<model><tokenizer><moe-load-balance-type>
qwen3-30b-a3bqwen3global-batch
qwen3.5-35b-a3bqwen3.5global-batch
deepseek-v2-litedeepseek-v2sequence
gpt-oss-20b, gpt-oss-120bgpt-ossglobal-batch

Step 1: Create the two run files

bash
R=$(git rev-parse --show-toplevel)
W=$R/workspace/validate-performance/<model>
mkdir -p $W $R/logging/validate-performance/<model>
cp $R/.agents/skills/validate-performance/templates/validate.py $W/<tag>-base.py
cp $R/.agents/skills/validate-performance/scripts/launch.sh $W/launch.sh
git rev-parse --abbrev-ref HEAD > $W/FEATURE

The launcher is copied into workspace/ because the checkouts below switch branches, and .agents/ is tracked.

The template carries the defaults as real values, so only a non-default request needs editing: training.sequence_length = 2048 and training.fp8 = False. The <tag> suffixes name them — -fp8 is training.fp8 = True, and fp8 is the only dtype knob there is.

Fill every <placeholder> in that file, then copy it so the two arms are identical by construction:

bash
cp $W/<tag>-base.py $W/<tag>-feat.py
diff $W/<tag>-base.py $W/<tag>-feat.py

The diff must print nothing. Each run takes its wandb name from its own filename, so nothing inside the two files differs. workspace/ and *.log are both gitignored, so the run files and the logs survive every checkout below without ever making the tree dirty.

Show full SKILL.md (200 more words)Show less

Step 2: Run both arms

Each fence re-derives its own paths, because shell variables do not survive between commands.

bash
R=$(git rev-parse --show-toplevel); W=$R/workspace/validate-performance/<model>
G=$R/logging/validate-performance/<model>/<tag>
git checkout main
bash $W/launch.sh $W/<tag>-base.py 2>&1 | tee $G-base.log
bash
R=$(git rev-parse --show-toplevel); W=$R/workspace/validate-performance/<model>
G=$R/logging/validate-performance/<model>/<tag>
FEATURE=$(cat $W/FEATURE)
git checkout $FEATURE
bash $W/launch.sh $W/<tag>-feat.py 2>&1 | tee $G-feat.log

Under SLURM, wrap each bash $W/launch.sh … in srun -N <nodes> -W 0 -o <log> — see launch-with-slurm.

Step 3: Report

bash
R=$(git rev-parse --show-toplevel); G=$R/logging/validate-performance/<model>/<tag>
python3 $R/.agents/skills/validate-performance/scripts/compare.py $G-base.log $G-feat.log

It prints median step time, tokens per second and peak memory for each arm, plus the percentage difference. Quote those numbers; do not convert them into a pass or fail.

Why the fixture looks like this

templates/validate.py sets training.benchmark = True, which force-balances routing by overwriting the router's top-k with a round-robin over token index. Expert imbalance is the largest source of step-time variance, and removing it is what makes a short run readable. It also makes the loss meaningless, which is why correctness lives in a separate skill.

training.max_steps = 8, because step time settles within a few steps once routing is balanced.

compare.py drops step 1: it carries the JIT warmup and the torch.compile trace, which dwarf the steady-state step time.

It compares medians of whole runs and never pools individual steps. Within a run step times are near-duplicates, while two runs of identical code sit at different offsets, so pooling counts correlated measurements as independent and reports significance for that offset.

© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in .agents/skills/validate-performance of mlc-ai/pith-train.

  • SKILL.md
  • scripts/compare.py
  • scripts/launch.sh
  • templates/validate.py

Open the folder on GitHubat commit c7c8b1d

Compare with similar skills

Validate Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Validate Performance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Validate Performance this skillmlc-ai/pith-train355—~1.4kAutomated safety check: PassApache-2.0
Fix Art IssuesOpenPipe/ART11k—~840Automated safety check: NotesApache-2.0
Secondary Architecture ReviewerHsienW/chat-gun143—~775Automated safety check: PassCustom licence
Release Prepfinch-xu/cc-router272—~1.5kAutomated safety check: PassMIT
AutofixQwenLM/qwen-code28k—~10kAutomated safety check: WarnApache-2.0
Git Guardrails Claude Codefossasia/eventyay-interpretation1.6k13 repos~578Automated safety check: PassApache-2.0

Similar skills

  • Fix Art Issues

    OpenPipe/ART

    Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

    11k GitHub stars~840 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • 對 OpenSpec、Git Diff、frontend、bff、backend、LangGraph、Tool、MCP 與安全邊界進行獨立唯讀審查,輸出 Blocker、Major、Minor 與標準 Verdict。

    143 GitHub stars~775 tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Release Prep

    finch-xu/cc-router

    cc-router 发版准备一条龙:升版本号 → 根据上一个 tag 以来的提交写 release-notes/<版本/ 的中英日三份更新内容 → 校验 → 给用户审 → 本地提交「Bump version to X.Y.Z」,停在打 tag 之前。当用户说「准备发版」「发个版」「发 6.1.0」「写发版说明 / 更新内容 / release notes」「bump 版本」时必须走本…

    272 GitHub stars~1.5k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Autofix

    QwenLM/qwen-code

    Review and repair current local changes until they converge, or run Qwen Code Autofix issue and review workflows from GitHub Actions.

    28k GitHub stars~10k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • Git Guardrails Claude Code

    fossasia/eventyay-interpretation

    Set up Claude Code hooks to block dangerous git commands (push, reset --hard, clean, branch -D, etc.) before they execute.

    1.6k GitHub starsUsed in 13 repos~578 tokens
    AI & LLM EngineeringAuto-check passed
  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from mlc-ai/pith-train

All 10 skills in this repo
  • Analyze Nsys Profile

    mlc-ai/pith-train

    Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

    355 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Capture Nsys Profile

    mlc-ai/pith-train

    Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.

    355 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed
  • Validate Correctness

    mlc-ai/pith-train

    Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

    355 GitHub stars~2.3k tokensUpdated 3 days ago
    Auto-check passed
  • Setup Benchmark Inputs

    mlc-ai/pith-train

    Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.

    355 GitHub stars~399 tokensUpdated 3 days ago
    Auto-check passed
  • Add New Model

    mlc-ai/pith-train

    Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.

    355 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed
  • Wandb Tracking

    mlc-ai/pith-train

    Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.

    355 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check: warnings

Works with

Questions about Validate Performance

What does Validate Performance do?

Measures the throughput difference between two branches with force-balanced routing. Validate Performance is an agent skill from mlc-ai/pith-train. Measures the throughput difference between two branches with force-balanced routing.

When should I use Validate Performance?

Validate Performance fits situations like: the user asks to benchmark the performance; measure throughput; check for a slowdown; check throughput did not regress.

How do I install Validate Performance in Claude Code?

Run `npx skills add mlc-ai/pith-train --skill validate-performance -a claude-code`. Or copy the skill folder (.agents/skills/validate-performance in mlc-ai/pith-train) into .claude/skills/validate-performance in your project. Claude Code loads it when a task matches its description.

How do I install Validate Performance in Codex?

Run `npx skills add mlc-ai/pith-train --skill validate-performance -a codex`. Or copy the skill folder (.agents/skills/validate-performance in mlc-ai/pith-train) into .agents/skills/validate-performance in your project. Codex loads it when a task matches its description.

Can I use Validate Performance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill validate-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-performance, .gemini/skills/validate-performance, .github/skills/validate-performance and .opencode/skills/validate-performance in your project.

What does Validate Performance need to run?

Going by SKILL.md and its folder, Validate Performance needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (git, bash and python3). Our summary lists: Python 3; A Bash shell.

Does Validate Performance access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Validate Performance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Validate Performance use?

Validate Performance is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Validate Performance use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Validate Performance?

Skills that share tags, products or a category with Validate Performance: Fix Art Issues (OpenPipe/ART, 11k stars), Secondary Architecture Reviewer (HsienW/chat-gun, 143 stars), Release Prep (finch-xu/cc-router, 272 stars) and Autofix (QwenLM/qwen-code, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Validate Performance?

mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.

Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.