Add Uint Support
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.
$ npx skills add mlc-ai/pith-train --skill add-new-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mlc-ai/pith-train add-new-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-new-model .claude/skills/add-new-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-new-model" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/add-new-model into .claude/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/add-new-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mlc-ai/pith-train --skill add-new-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mlc-ai/pith-train add-new-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/add-new-model .agents/skills/add-new-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-new-model" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/add-new-model into .agents/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mlc-ai/pith-train --skill add-new-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mlc-ai/pith-train add-new-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/add-new-model .cursor/skills/add-new-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-new-model" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/add-new-model into .cursor/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mlc-ai/pith-train.git --path .agents/skills/add-new-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mlc-ai/pith-train --skill add-new-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mlc-ai/pith-train add-new-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/add-new-model .gemini/skills/add-new-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-new-model" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/add-new-model into .gemini/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mlc-ai/pith-train add-new-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mlc-ai/pith-train --skill add-new-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/add-new-model .github/skills/add-new-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-new-model" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/add-new-model into .github/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mlc-ai/pith-train --skill add-new-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mlc-ai/pith-train add-new-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/add-new-model .opencode/skills/add-new-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-new-model" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/add-new-model into .opencode/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-new-modelAdds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.
Add New Model is an agent skill from mlc-ai/pith-train. Adds support for a new MoE language model to PithTrain. Use when the user asks to "add support for model X", "implement model Y in pithtrain", "port model Z", or otherwise integrate a new MoE architecture. Scope covers the model file, all framework wiring (setupmodel, applyfsdp, testdualpipev), optional checkpoint conversion, and running training + inference tests from pp=1/ep=1 up to pp=2/ep=2.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files (for example `reference/checkpoint.md`, `reference/compile.md` and `reference/conventions.md`).
It sits in AI & LLM Engineering, covering Deep learning. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonhfbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add New Model loads about 4.6k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 2,179 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 2,179 words, ~4,633 tokens.
.claude/skills/add-new-model/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.End-to-end workflow for integrating a new MoE language model. This file is the entry point: it tells you the phase order, the gates between phases, and which reference/*.md to load before each phase. Do not try to do everything from memory - load the reference file for the phase you're in.
One of the following:
"mistralai/Mixtral-8x7B-v0.1"). Used as the --hf-id for snapshot_download, and as the from_pretrained source for config and tokenizer.config.json, model*.safetensors, tokenizer* etc.). Treat it exactly like the HF ID case - AutoConfig.from_pretrained(path) works on both. The snapshot itself came from HF, so online reference material (HF's modeling_<model>.py, TorchTitan, OpenAI release repo, upstream papers) is still fair game and should be consulted.Optionally a model short name (filename stem, e.g. mixtral_8x7b). If the user doesn't give one, derive it from the HF ID by lowercasing and replacing / and - with _. Confirm with the user before using.
These are non-negotiable. Violating any of them will cost time later - in most cases that's exactly how past bugs landed.
modeling_<model>.py. Do not base the new model on Qwen3 / DeepSeek-V2 / GPT-OSS and rename - that path produced the GPT-OSS gate/router mismatch. See reference/conventions.md.fullgraph=True for the compiled hot regions. forward_stage1_compute and forward_stage5 must each carry @torch.compile(fullgraph=True) (plus the attention forward itself when the model uses flex_attention, as GPT-OSS does). Never reach for fullgraph=False. See reference/compile.md.forward_stage1, not forward_stage5. If the model has shared experts (e.g. DeepSeek-V2), fold them into the residual at the end of forward_stage1_compute. See reference/protocol.md.nvidia-smi before every GPU command. This is a shared cluster. Free GPUs can change between commands; don't reuse indices. nvidia-smi --query-gpu=index,memory.used,memory.free --format=csv and pick indices with memory.used < 1000 MiB. Do this once per invocation, not once at the start of the whole skill.__init__ accepts a group it does not actually implement (e.g. cp_group on a model with no ring-attention path), silently ignoring it will produce wrong results when a real group is eventually passed. Raise NotImplementedError when group is not None and group.size() > 1. See reference/protocol.md (init-requirements).__init__; don't hardcode them. Any value that appears in HF's config.json is a per-checkpoint knob - read it from the config even if every released checkpoint currently ships the same value. Only true architectural constants (paper coefficients, spec-defined magic numbers) stay as module-level literals, and they get a one-line comment naming the source. See reference/conventions.md (thread-config)..agents/, .claude/, AGENTS.md, CLAUDE.md, or docs/ in commits.Phases 1-4 are modeling + training correctness. Phases 5-6 are real-weight inference (only needed if the user wants to generate from trained / released weights). Phase 5 is skippable when the user only cares about training from scratch.
| Phase | Goal | Gate to next phase |
|---|---|---|
| 0 | Analyze HF's reference implementation | Have class/attribute/shape/config inventory |
| 1 | Write pithtrain/models/<model>.py | Imports cleanly; reference_forward runs |
| 2 | Wire into pithtrain/modules/training.py + tests/test_dualpipev.py + example config | Example config mirrors upstream; imports clean |
| 3 | Single-GPU sanity test | pp=1/ep=1 test_dualpipev.py passes (loss allclose 1e-3, grad calc_diff < 1e-2) |
| 4 | FSDP scaling (pp=1/ep=1 -> 2/2) | All 4 configs pass |
| 5 | (If needed) Checkpoint converter + round-trip | hf -> dcp -> hf -> transformers.load succeeds |
| 6 | (If needed) Ad-hoc inference test | Coherent text from real weights |
Do not skip ahead. Each phase is a gate: if phase N fails, do not move to phase N+1. If you're tempted to, stop and read reference/pitfalls.md.
Before writing any code, inventory what you need to match.
Load reference/conventions.md before starting this phase. It has the diagnostic commands (grep class names, grep attribute names, safetensors shape dump) under the Quick diagnostic commands section.
Work through these sources in order:
modeling_<model>.py - class names, attribute names, fused vs split projections, special-case features (shared experts, sinks, sliding window, YaRN RoPE, clamped SwiGLU, attention biases).[E, out, in] vs [E, in, out]).configuration_<model>.py - every default in <Model>Config.__init__. When model-specific defaults disagree with a generic fallback path, match the model-specific default (see reference/conventions.md (example-config)).config.hidden_act to tell you the whole activation. See reference/conventions.md (activation-math).Record in a scratch doc (not a committed file): class names, attribute names, expert tensor layout, fused/split projections, per-checkpoint knobs (thread through __init__) vs architectural constants (module literals with a source comment), process groups the model accepts but doesn't implement (reject via NotImplementedError; see reference/protocol.md (init-requirements)), and any special-case features that map to entries in reference/pitfalls.md.
Gate: you can articulate exactly which class names, attribute names, tensor shapes, and config knobs you will wire. If any item is a guess, go back and print() it from the actual data.
pithtrain/models/<model>.pyLoad reference/protocol.md and reference/compile.md before starting.
templates/model_skeleton.py. It is a structural outline (NOT a copy of Qwen3). Fill in the TODO placeholders with the HF- derived names and shapes from phase 0.model_forward, posemb / prolog / epilog, reference_forward)self.idx = layer_id and self.mlp assigned (satisfies LayerProtocol in pithtrain/models/interface.py)distributed context (distributed.ep_size / distributed.ep_group), with the local expert count on self.mlp.experts_per_rank@torch.compile(fullgraph=True) on forward_stage1_compute and forward_stage5forward_stage1_compute, before the returnreference_forward runs eager (no compile) and is numerically equivalent to forward_stage1 -> forward_stage3 -> forward_stage5forward_stage3 truncates expert input by sum(ks) if the expert block has biases or elementwise post-ops (prevents 0*NaN=NaN in backward)forward_stage3 uses padded_index_gather (not raw indexing) for both expand and reverse shufflelayer_partition(config.num_hidden_layers, stage_count, stage_index) from pithtrain/pipeline/dualpipev.pyforward delegates to model_forward(self, hidden_states, self.chunk_record, cu_seqlens); the engine records each stage into self.chunk_record (a ChunkRecord from pithtrain/pipeline/execution.py) for the pipeline backward, which the engine drives via model_backward.forward_posemb, forward_prolog, forward_epilog, and reference_forward implemented per ModelProtocol.Gate: file imports cleanly (python -c "from pithtrain.models.<model> import <Model>").
No new reference file needed - the changes are small and mechanical.
pithtrain/modules/training.py:<Model>Model class.setup_model:elif module_config.model_type == "<model_type>":
ModelClass = <Model>Model
model_kwargs = {"cp_group": cp_group} # or {} if no CP supportapply_fsdp isinstance assertion tuple.TrainingCfg.model Literal[...] union (if the user wants the HF ID to be an accepted value; config-path usage doesn't require this).tests/test_dualpipev.py:nn.Parameter expert weights - see reference/pitfalls.md).apply_fsdp isinstance assertion tuple.config.model_type switch in main. Slice num_hidden_layers down to 8 (and any parallel arrays like layer_types) to keep the test fast.fill_weights branch if:nn.Parameter (not GroupedLinear). Without this, expert weights default to zero and the MoE subtree silently produces all-zero outputs - see reference/pitfalls.md.weight (e.g. a per-expert bias).shard_experts can detect the experts module. If using raw nn.Parameter, the fallback gate on gate_up_proj already handles it. If the Parameter name is different, extend the fallbacknum_experts alone (the router has num_experts too and must not be sharded).models list at the bottom of tests/test_dualpipev.py.examples/pretrain_lm/<model>/config.json. Mirror upstream HF's config.json field-by-field - including every nested block (rope_scaling, quantization_config, etc.). See reference/conventions.md (example-config) for the diff command and the three layered defaults you need to reconcile.Gate: python -c "import tests.test_dualpipev" imports cleanly AND the example-config diff is either empty or has a documented reason for each remaining difference.
Load reference/testing.md. Tier 1 there is the whole phase: wiring the new model into tests/test_dualpipev.py, then running that committed harness on one GPU at the lightest rung:
CUDA_VISIBLE_DEVICES=<g0> timeout 180 torchrun --nproc-per-node=1 $RDZV \
tests/test_dualpipev.py --pp-size 1 --ep-size 1 --model $CFG(or bash tests/test_dualpipev.sh <config>). It builds the model at phase=-1 (reference) and phase=0/phase=1 (the two DualPipeV chunks) and compares the pipelined 5-stage forward against reference_forward. Single GPU, timeout 180.
Gate: loss matches (allclose, rtol=atol=1e-3) and every parameter gradient passes (calc_diff < 1e-2); logits and gradients are finite.
Keep reference/testing.md loaded. It owns the ladder: full torchrun commands for pp=1/ep=1 -> pp=2/ep=1 -> pp=1/ep=2 -> pp=2/ep=2, what each config adds, thresholds, and the failure decision tree.
Run the ladder in that order. After each step, stop and diagnose before continuing if anything fails. nvidia-smi before each run. Timeouts 120-180s; past that it's hanging (compile retrace or deadlocked all-to-all) - kill, don't raise.
Gate: all 4 configs pass (loss rtol=atol=1e-3, per-param calc_diff < 1e-2).
Skip this phase entirely if the user only wants training from scratch (no real released weights involved). The generic path in pithtrain/tasks/convert_checkpoint/_core.py already handles un-quantized, un-transposed HF checkpoints - Qwen3 and DeepSeek-V2 work with no model-specific converter.
Add a converter only if one of the following applies:
[E, out, in] vs HF's [E, in, out]).Load reference/checkpoint.md before starting this phase.
pithtrain/tasks/convert_checkpoint/<model>.py with a <Model>Converter class (see gpt_oss.py for the pattern). Implement detect_hf / detect_dcp probes, hf2dcp, and postprocess_canonical.pithtrain/tasks/convert_checkpoint/_registry.py (append to CONVERTERS).examples/convert_checkpoint/<model>/script.py that downloads + converts. Mirror examples/convert_checkpoint/gpt-oss-20b/script.py.transformers.AutoModelForCausalLM.from_pretrained. Compare state_dict() element-wise against HF's own BF16 dequant. Expected max_abs_diff == 0.Gate: round-trip succeeds, one expert weight compares element-wise equal (not just norms!) against HF's live tensor.
Only needed if the user wants to verify that real weights produce coherent text. This test is not committed - it's model-specific and lives as a scratch file.
templates/inference_test.py - the DualPipeV autoregressive harness, parameterized for any <Model>Model. Fill in the model-class import and HF ID default.tests/test_dualpipev.py with tests/test_<model>_inference.py and drop --model <cfg>). Each config should print coherent continuations.Gate: coherent text from real released weights, identical (bf16-noise equivalent) across pp/ep configurations.
Three sweeps on the new files before opening the PR - low-noise, high-signal self-reviews that save a review round-trip:
import/from in the new model file; move them to module level.docs/, AGENTS.md, CLAUDE.md, .agents/, .claude/ pointers in comments or docstrings. Those paths aren't committed, so any pointer is a broken link. Grep the new files and inline the derivation or delete.cp_group for protocol parity), then either prefix with _ or raise NotImplementedError when size() > 1 (Hard Rule 7). Bare unused params trip pyright/pylance.| Symptom | First thing to read |
|---|---|
| Single-GPU loss/grad mismatch at pp=1/ep=1 | reference/testing.md (why the gradient threshold is loose) |
| All-zero gradient warnings on MoE params | reference/pitfalls.md (fill-weights) |
| FSDP loss matches but grads don't | reference/testing.md (label-scaling) + reference/pitfalls.md (nan-padding) |
RuntimeError: tensor data is not allocated yet | Wrong reshard settings - check apply_fsdp |
| Inference gibberish but FSDP passed | reference/checkpoint.md (weight-norm-comparison) + reference/conventions.md (example-config) + (thread-config) + (activation-math) |
Wrong results only when a real cp_group is passed | reference/protocol.md (init-requirements) (silent-ignore of unused groups) |
| "invalid gradient shape" in stage 4 backward | reference/protocol.md (stage-record-copy) |
compile-inside-compile on attention | reference/compile.md (flex-unwrap) |
| Left-padded prompts give gibberish on short inputs | reference/pitfalls.md (trim-to-shortest) |
reference/protocol.md - 5-stage protocol, Model.forward/backward, stage-record copyreference/conventions.md - naming, tensor layout, canonical keysreference/compile.md - three @torch.compile(fullgraph=True) hot regions, unwrap patternsreference/checkpoint.md - hf2dcp/dcp2hf recipes, when to add, round-trip validationreference/testing.md - pp/ep scaling ladder, test_dualpipev wiring, label scalingreference/pitfalls.md - NaN padding, .view() vs .transpose(), silent-zero experts, etc.templates/model_skeleton.py - structural outline with HF-derived placeholderstemplates/inference_test.py - DualPipeV autoregressive harness© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files in .agents/skills/add-new-model of mlc-ai/pith-train.
Open the folder on GitHubat commit c7c8b1d
Add New Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add New Model this skillmlc-ai/pith-train | 355 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3.3k | Automated safety check: Pass | MIT | |
| Add Oponnx/onnx | 22k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Add Function Bodyonnx/onnx | 22k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 |
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
onnx/onnx
Add a new ONNX operator or update an existing operator to a new opset version.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
onnx/onnx
Add a function body definition to an ONNX operator, defining how it decomposes into simpler ops.
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…
mlc-ai/pith-train
Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.
mlc-ai/pith-train
Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.
mlc-ai/pith-train
Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.
mlc-ai/pith-train
Measures the throughput difference between two branches with force-balanced routing.
mlc-ai/pith-train
Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.
mlc-ai/pith-train
Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.
Categories
Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train. Add New Model is an agent skill from mlc-ai/pith-train. Adds support for a new MoE language model to PithTrain.
Add New Model fits situations like: the user asks to add support for model X; implement model Y in pithtrain; otherwise integrate a new MoE architecture.
Run `npx skills add mlc-ai/pith-train --skill add-new-model -a claude-code`. Or copy the skill folder (.agents/skills/add-new-model in mlc-ai/pith-train) into .claude/skills/add-new-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mlc-ai/pith-train --skill add-new-model -a codex`. Or copy the skill folder (.agents/skills/add-new-model in mlc-ai/pith-train) into .agents/skills/add-new-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill add-new-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-new-model, .gemini/skills/add-new-model, .github/skills/add-new-model and .opencode/skills/add-new-model in your project.
Going by SKILL.md and its folder, Add New Model needs Python for the scripts in its folder and the command-line tools its instructions call (python, hf and bash). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Add New Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Add New Model: Add Uint Support (pytorch/pytorch, 104k stars), Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Add Op (onnx/onnx, 22k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.
Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.