Agent skill

Add New Model

by mlc-ai in mlc-ai/pith-train

Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Add New Model

skills CLI
$ npx skills add mlc-ai/pith-train --skill add-new-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mlc-ai/pith-train add-new-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-new-model .claude/skills/add-new-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-new-model
GitHub stars
355
Token cost
~4.6k tokens
SKILL.md length
2,179 words
Files
9
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.

  • Works in 7 steps: Analyze the HF reference → Write pithtrain/models/.py → Wire into the training framework → …
  • The user asks to add support for model X
  • SKILL.md covers Input, Hard rules (apply in every…, Phase overview and Phase 0 - Analyze the HF…, plus 10 more sections
  • Runs Python scripts from its folder; calls python, hf and bash

What it does

Add New Model is an agent skill from mlc-ai/pith-train. Adds support for a new MoE language model to PithTrain. Use when the user asks to "add support for model X", "implement model Y in pithtrain", "port model Z", or otherwise integrate a new MoE architecture. Scope covers the model file, all framework wiring (setupmodel, applyfsdp, testdualpipev), optional checkpoint conversion, and running training + inference tests from pp=1/ep=1 up to pp=2/ep=2.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files (for example `reference/checkpoint.md`, `reference/compile.md` and `reference/conventions.md`).

It sits in AI & LLM Engineering, covering Deep learning. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.

When your agent uses it

  • The user asks to add support for model X
  • Implement model Y in pithtrain
  • Otherwise integrate a new MoE architecture

Example prompts

  • “add support for model X”
  • “implement model Y in pithtrain”
  • “port model Z”
  • “/add-new-model”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Analyze the HF reference
  2. Write pithtrain/models/.py
  3. Wire into the training framework
  4. Single-GPU sanity test
  5. FSDP scaling (training correctness)
  6. Checkpoint converter (only if needed)
  7. Ad-hoc inference test (only if needed)

What it can do on your machine

Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • hf
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add New Model loads about 4.6k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 2,179 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 2,179 words, ~4,633 tokens.

Download SKILL.mdSave it as .claude/skills/add-new-model/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
add-new-model
description
Adds support for a new MoE language model to PithTrain. Use when the user asks to "add support for model X", "implement model Y in pithtrain", "port model Z", or otherwise integrate a new MoE architecture. Scope covers the model file, all framework wiring (setup_model, apply_fsdp, test_dualpipev), optional checkpoint conversion, and running training + inference tests from pp=1/ep=1 up to pp=2/ep=2.
argument-hint
<hf-id-or-snapshot-path> [model-short-name]

Add a New Model to PithTrain

End-to-end workflow for integrating a new MoE language model. This file is the entry point: it tells you the phase order, the gates between phases, and which reference/*.md to load before each phase. Do not try to do everything from memory - load the reference file for the phase you're in.

Input

One of the following:

  • HuggingFace model ID (e.g. "mistralai/Mixtral-8x7B-v0.1"). Used as the --hf-id for snapshot_download, and as the from_pretrained source for config and tokenizer.
  • Local snapshot path (a directory containing config.json, model*.safetensors, tokenizer* etc.). Treat it exactly like the HF ID case - AutoConfig.from_pretrained(path) works on both. The snapshot itself came from HF, so online reference material (HF's modeling_<model>.py, TorchTitan, OpenAI release repo, upstream papers) is still fair game and should be consulted.

Optionally a model short name (filename stem, e.g. mixtral_8x7b). If the user doesn't give one, derive it from the HF ID by lowercasing and replacing / and - with _. Confirm with the user before using.

Hard rules (apply in every phase)

These are non-negotiable. Violating any of them will cost time later - in most cases that's exactly how past bugs landed.

  1. Mirror HuggingFace, not our existing models. Class names, attribute names, and tensor structure (fused vs split) must match HF's modeling_<model>.py. Do not base the new model on Qwen3 / DeepSeek-V2 / GPT-OSS and rename - that path produced the GPT-OSS gate/router mismatch. See reference/conventions.md.
  2. fullgraph=True for the compiled hot regions. forward_stage1_compute and forward_stage5 must each carry @torch.compile(fullgraph=True) (plus the attention forward itself when the model uses flex_attention, as GPT-OSS does). Never reach for fullgraph=False. See reference/compile.md.
  3. Shared experts live in forward_stage1, not forward_stage5. If the model has shared experts (e.g. DeepSeek-V2), fold them into the residual at the end of forward_stage1_compute. See reference/protocol.md.
  4. Check nvidia-smi before every GPU command. This is a shared cluster. Free GPUs can change between commands; don't reuse indices. nvidia-smi --query-gpu=index,memory.used,memory.free --format=csv and pick indices with memory.used < 1000 MiB. Do this once per invocation, not once at the start of the whole skill.
  5. Test timeouts stay short (120-180s). Do not set a 10-minute timeout and walk away - if a test hasn't progressed in 3 minutes, it's hanging (likely torch.compile retrace). Kill and diagnose.
  6. When tests fail, diagnose before relaxing anything. Print actual magnitudes. A 6-order-of-magnitude gradient discrepancy is not bf16 noise. Never loosen thresholds or add name-based skips as a first move.
  7. Reject unused process-group arguments explicitly. If a model's __init__ accepts a group it does not actually implement (e.g. cp_group on a model with no ring-attention path), silently ignoring it will produce wrong results when a real group is eventually passed. Raise NotImplementedError when group is not None and group.size() > 1. See reference/protocol.md (init-requirements).
  8. Thread config values through __init__; don't hardcode them. Any value that appears in HF's config.json is a per-checkpoint knob - read it from the config even if every released checkpoint currently ships the same value. Only true architectural constants (paper coefficients, spec-defined magic numbers) stay as module-level literals, and they get a one-line comment naming the source. See reference/conventions.md (thread-config).
  9. Never stage .agents/, .claude/, AGENTS.md, CLAUDE.md, or docs/ in commits.

Phase overview

Phases 1-4 are modeling + training correctness. Phases 5-6 are real-weight inference (only needed if the user wants to generate from trained / released weights). Phase 5 is skippable when the user only cares about training from scratch.

PhaseGoalGate to next phase
0Analyze HF's reference implementationHave class/attribute/shape/config inventory
1Write pithtrain/models/<model>.pyImports cleanly; reference_forward runs
2Wire into pithtrain/modules/training.py + tests/test_dualpipev.py + example configExample config mirrors upstream; imports clean
3Single-GPU sanity testpp=1/ep=1 test_dualpipev.py passes (loss allclose 1e-3, grad calc_diff < 1e-2)
4FSDP scaling (pp=1/ep=1 -> 2/2)All 4 configs pass
5(If needed) Checkpoint converter + round-triphf -> dcp -> hf -> transformers.load succeeds
6(If needed) Ad-hoc inference testCoherent text from real weights

Do not skip ahead. Each phase is a gate: if phase N fails, do not move to phase N+1. If you're tempted to, stop and read reference/pitfalls.md.


Phase 0 - Analyze the HF reference

Before writing any code, inventory what you need to match.

Load reference/conventions.md before starting this phase. It has the diagnostic commands (grep class names, grep attribute names, safetensors shape dump) under the Quick diagnostic commands section.

Work through these sources in order:

  1. modeling_<model>.py - class names, attribute names, fused vs split projections, special-case features (shared experts, sinks, sliding window, YaRN RoPE, clamped SwiGLU, attention biases).
  2. A safetensors shard - actual expert-weight shapes and dtypes, not comments.
  3. TorchTitan / Megatron-LM reference for this model - the training-framework-consensus expert layout ([E, out, in] vs [E, in, out]).
  4. configuration_<model>.py - every default in <Model>Config.__init__. When model-specific defaults disagree with a generic fallback path, match the model-specific default (see reference/conventions.md (example-config)).
  5. HF's MLP / activation forward - read the math directly. Don't trust config.hidden_act to tell you the whole activation. See reference/conventions.md (activation-math).

Record in a scratch doc (not a committed file): class names, attribute names, expert tensor layout, fused/split projections, per-checkpoint knobs (thread through __init__) vs architectural constants (module literals with a source comment), process groups the model accepts but doesn't implement (reject via NotImplementedError; see reference/protocol.md (init-requirements)), and any special-case features that map to entries in reference/pitfalls.md.

Gate: you can articulate exactly which class names, attribute names, tensor shapes, and config knobs you will wire. If any item is a guess, go back and print() it from the actual data.


Phase 1 - Write pithtrain/models/<model>.py

Load reference/protocol.md and reference/compile.md before starting.

  1. Start from templates/model_skeleton.py. It is a structural outline (NOT a copy of Qwen3). Fill in the TODO placeholders with the HF- derived names and shapes from phase 0.
  2. Implement in this order:
    • RotaryEmbedding (mirror HF's)
    • Attention (mirror HF's kernel choice - flash_attn for standard MHA/GQA, flex_attention for sinks/sliding)
    • Experts module
    • Router / Gate
    • MLP (the MoE block that wires router + experts)
    • DecoderLayer (the 5-stage split)
    • Model (forward via model_forward, posemb / prolog / epilog, reference_forward)
  3. Checklist for the decoder layer:
    • self.idx = layer_id and self.mlp assigned (satisfies LayerProtocol in pithtrain/models/interface.py)
    • EP sizing read from the distributed context (distributed.ep_size / distributed.ep_group), with the local expert count on self.mlp.experts_per_rank
    • @torch.compile(fullgraph=True) on forward_stage1_compute and forward_stage5
    • Shared experts (if any) fold into residual at the end of forward_stage1_compute, before the return
    • reference_forward runs eager (no compile) and is numerically equivalent to forward_stage1 -> forward_stage3 -> forward_stage5
    • forward_stage3 truncates expert input by sum(ks) if the expert block has biases or elementwise post-ops (prevents 0*NaN=NaN in backward)
    • forward_stage3 uses padded_index_gather (not raw indexing) for both expand and reverse shuffle
  4. Checklist for the model class:
    • Layers built via layer_partition(config.num_hidden_layers, stage_count, stage_index) from pithtrain/pipeline/dualpipev.py
    • forward delegates to model_forward(self, hidden_states, self.chunk_record, cu_seqlens); the engine records each stage into self.chunk_record (a ChunkRecord from pithtrain/pipeline/execution.py) for the pipeline backward, which the engine drives via model_backward.
    • forward_posemb, forward_prolog, forward_epilog, and reference_forward implemented per ModelProtocol.

Gate: file imports cleanly (python -c "from pithtrain.models.<model> import <Model>").


Phase 2 - Wire into the training framework

No new reference file needed - the changes are small and mechanical.

  1. pithtrain/modules/training.py:
    • Import the new <Model>Model class.
    • Add a branch in setup_model:
      python
      elif module_config.model_type == "<model_type>":
          ModelClass = <Model>Model
          model_kwargs = {"cp_group": cp_group}  # or {} if no CP support
    • Add the new class to the apply_fsdp isinstance assertion tuple.
    • Add the HF ID to the TrainingCfg.model Literal[...] union (if the user wants the HF ID to be an accepted value; config-path usage doesn't require this).
  2. tests/test_dualpipev.py:
    • Import the new model + router/gate class (+ Experts class if it stores raw nn.Parameter expert weights - see reference/pitfalls.md).
    • Add the new class to the apply_fsdp isinstance assertion tuple.
    • Add a branch in the config.model_type switch in main. Slice num_hidden_layers down to 8 (and any parallel arrays like layer_types) to keep the test fast.
    • Add a fill_weights branch if:
      • The expert module stores raw nn.Parameter (not GroupedLinear). Without this, expert weights default to zero and the MoE subtree silently produces all-zero outputs - see reference/pitfalls.md.
      • The router/gate has new Parameters beyond weight (e.g. a per-expert bias).
    • Verify shard_experts can detect the experts module. If using raw nn.Parameter, the fallback gate on gate_up_proj already handles it. If the Parameter name is different, extend the fallback
      • gate on the distinctive weight name, not on num_experts alone (the router has num_experts too and must not be sharded).
  3. Add the model config and HF ID to the models list at the bottom of tests/test_dualpipev.py.
  4. Write examples/pretrain_lm/<model>/config.json. Mirror upstream HF's config.json field-by-field - including every nested block (rope_scaling, quantization_config, etc.). See reference/conventions.md (example-config) for the diff command and the three layered defaults you need to reconcile.

Gate: python -c "import tests.test_dualpipev" imports cleanly AND the example-config diff is either empty or has a documented reason for each remaining difference.


Show full SKILL.md (750 more words)Show less

Phase 3 - Single-GPU sanity test

Load reference/testing.md. Tier 1 there is the whole phase: wiring the new model into tests/test_dualpipev.py, then running that committed harness on one GPU at the lightest rung:

bash
CUDA_VISIBLE_DEVICES=<g0> timeout 180 torchrun --nproc-per-node=1 $RDZV \
  tests/test_dualpipev.py --pp-size 1 --ep-size 1 --model $CFG

(or bash tests/test_dualpipev.sh <config>). It builds the model at phase=-1 (reference) and phase=0/phase=1 (the two DualPipeV chunks) and compares the pipelined 5-stage forward against reference_forward. Single GPU, timeout 180.

Gate: loss matches (allclose, rtol=atol=1e-3) and every parameter gradient passes (calc_diff < 1e-2); logits and gradients are finite.


Phase 4 - FSDP scaling (training correctness)

Keep reference/testing.md loaded. It owns the ladder: full torchrun commands for pp=1/ep=1 -> pp=2/ep=1 -> pp=1/ep=2 -> pp=2/ep=2, what each config adds, thresholds, and the failure decision tree.

Run the ladder in that order. After each step, stop and diagnose before continuing if anything fails. nvidia-smi before each run. Timeouts 120-180s; past that it's hanging (compile retrace or deadlocked all-to-all) - kill, don't raise.

Gate: all 4 configs pass (loss rtol=atol=1e-3, per-param calc_diff < 1e-2).


Phase 5 - Checkpoint converter (only if needed)

Skip this phase entirely if the user only wants training from scratch (no real released weights involved). The generic path in pithtrain/tasks/convert_checkpoint/_core.py already handles un-quantized, un-transposed HF checkpoints - Qwen3 and DeepSeek-V2 work with no model-specific converter.

Add a converter only if one of the following applies:

  • The released weights are quantized (MXFP4, GPTQ, AWQ, FP8, etc.).
  • The HF live tensor layout differs from your model's in-memory layout (e.g. our [E, out, in] vs HF's [E, in, out]).
  • HF's key structure differs from ours (e.g. per-expert indexed vs stacked, fused vs split projections). This should be rare if you followed phase 0 faithfully - ideally our model mirrors HF's structure so the converter is trivial.

Load reference/checkpoint.md before starting this phase.

  1. Create pithtrain/tasks/convert_checkpoint/<model>.py with a <Model>Converter class (see gpt_oss.py for the pattern). Implement detect_hf / detect_dcp probes, hf2dcp, and postprocess_canonical.
  2. Register the converter instance in pithtrain/tasks/convert_checkpoint/_registry.py (append to CONVERTERS).
  3. Write an examples/convert_checkpoint/<model>/script.py that downloads + converts. Mirror examples/convert_checkpoint/gpt-oss-20b/script.py.
  4. Run the round-trip: hf -> dcp -> hf -> transformers.AutoModelForCausalLM.from_pretrained. Compare state_dict() element-wise against HF's own BF16 dequant. Expected max_abs_diff == 0.

Gate: round-trip succeeds, one expert weight compares element-wise equal (not just norms!) against HF's live tensor.


Phase 6 - Ad-hoc inference test (only if needed)

Only needed if the user wants to verify that real weights produce coherent text. This test is not committed - it's model-specific and lives as a scratch file.

  1. Start from templates/inference_test.py - the DualPipeV autoregressive harness, parameterized for any <Model>Model. Fill in the model-class import and HF ID default.
  2. Run the same pp/ep scaling ladder as phase 4 (same torchrun form, replace tests/test_dualpipev.py with tests/test_<model>_inference.py and drop --model <cfg>). Each config should print coherent continuations.
  3. Compare outputs across configurations - they should produce identical tokens (within bf16 noise). If a specific config produces gibberish, diagnose by layer/stage - do not loosen expectations.

Gate: coherent text from real released weights, identical (bf16-noise equivalent) across pp/ep configurations.


Pre-PR self-review

Three sweeps on the new files before opening the PR - low-noise, high-signal self-reviews that save a review round-trip:

  1. Function-scope imports. Only justified for circular imports or heavy optional deps. Ruff doesn't flag it; reviewers will. Grep for indented import/from in the new model file; move them to module level.
  2. Dangling docs/, AGENTS.md, CLAUDE.md, .agents/, .claude/ pointers in comments or docstrings. Those paths aren't committed, so any pointer is a broken link. Grep the new files and inline the derivation or delete.
  3. Unused parameters for interface compatibility. Accept them (e.g. cp_group for protocol parity), then either prefix with _ or raise NotImplementedError when size() > 1 (Hard Rule 7). Bare unused params trip pyright/pylance.

Common failure modes -> where to look

SymptomFirst thing to read
Single-GPU loss/grad mismatch at pp=1/ep=1reference/testing.md (why the gradient threshold is loose)
All-zero gradient warnings on MoE paramsreference/pitfalls.md (fill-weights)
FSDP loss matches but grads don'treference/testing.md (label-scaling) + reference/pitfalls.md (nan-padding)
RuntimeError: tensor data is not allocated yetWrong reshard settings - check apply_fsdp
Inference gibberish but FSDP passedreference/checkpoint.md (weight-norm-comparison) + reference/conventions.md (example-config) + (thread-config) + (activation-math)
Wrong results only when a real cp_group is passedreference/protocol.md (init-requirements) (silent-ignore of unused groups)
"invalid gradient shape" in stage 4 backwardreference/protocol.md (stage-record-copy)
compile-inside-compile on attentionreference/compile.md (flex-unwrap)
Left-padded prompts give gibberish on short inputsreference/pitfalls.md (trim-to-shortest)

Reference files

  • reference/protocol.md - 5-stage protocol, Model.forward/backward, stage-record copy
  • reference/conventions.md - naming, tensor layout, canonical keys
  • reference/compile.md - three @torch.compile(fullgraph=True) hot regions, unwrap patterns
  • reference/checkpoint.md - hf2dcp/dcp2hf recipes, when to add, round-trip validation
  • reference/testing.md - pp/ep scaling ladder, test_dualpipev wiring, label scaling
  • reference/pitfalls.md - NaN padding, .view() vs .transpose(), silent-zero experts, etc.

Templates

  • templates/model_skeleton.py - structural outline with HF-derived placeholders
  • templates/inference_test.py - DualPipeV autoregressive harness

© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files in .agents/skills/add-new-model of mlc-ai/pith-train.

  • SKILL.md
  • reference/checkpoint.md
  • reference/compile.md
  • reference/conventions.md
  • reference/pitfalls.md
  • reference/protocol.md
  • reference/testing.md
  • templates/inference_test.py
  • templates/model_skeleton.py

Open the folder on GitHubat commit c7c8b1d

Compare with similar skills

Add New Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add New Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add New Model this skillmlc-ai/pith-train355—~4.6kAutomated safety check: PassApache-2.0
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
Add Oponnx/onnx22k—~1.2kAutomated safety check: PassApache-2.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Add Function Bodyonnx/onnx22k—~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Op

    onnx/onnx

    Add a new ONNX operator or update an existing operator to a new opset version.

    22k GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Add a function body definition to an ONNX operator, defining how it decomposes into simpler ops.

    22k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Distributed

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

    24k GitHub stars~660 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from mlc-ai/pith-train

All 10 skills in this repo
  • Analyze Nsys Profile

    mlc-ai/pith-train

    Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

    355 GitHub stars~1.9k tokensUpdated 5 days ago
    Auto-check passed
  • Capture Nsys Profile

    mlc-ai/pith-train

    Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.

    355 GitHub stars~1k tokensUpdated 5 days ago
    Auto-check passed
  • Validate Correctness

    mlc-ai/pith-train

    Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

    355 GitHub stars~2.3k tokensUpdated 5 days ago
    Auto-check passed
  • Validate Performance

    mlc-ai/pith-train

    Measures the throughput difference between two branches with force-balanced routing.

    355 GitHub stars~1.4k tokensUpdated 5 days ago
    Auto-check passed
  • Setup Benchmark Inputs

    mlc-ai/pith-train

    Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.

    355 GitHub stars~399 tokensUpdated 5 days ago
    Auto-check passed
  • Wandb Tracking

    mlc-ai/pith-train

    Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.

    355 GitHub stars~1.1k tokensUpdated 5 days ago
    Auto-check: warnings

Questions about Add New Model

What does Add New Model do?

Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train. Add New Model is an agent skill from mlc-ai/pith-train. Adds support for a new MoE language model to PithTrain.

When should I use Add New Model?

Add New Model fits situations like: the user asks to add support for model X; implement model Y in pithtrain; otherwise integrate a new MoE architecture.

How do I install Add New Model in Claude Code?

Run `npx skills add mlc-ai/pith-train --skill add-new-model -a claude-code`. Or copy the skill folder (.agents/skills/add-new-model in mlc-ai/pith-train) into .claude/skills/add-new-model in your project. Claude Code loads it when a task matches its description.

How do I install Add New Model in Codex?

Run `npx skills add mlc-ai/pith-train --skill add-new-model -a codex`. Or copy the skill folder (.agents/skills/add-new-model in mlc-ai/pith-train) into .agents/skills/add-new-model in your project. Codex loads it when a task matches its description.

Can I use Add New Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill add-new-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-new-model, .gemini/skills/add-new-model, .github/skills/add-new-model and .opencode/skills/add-new-model in your project.

What does Add New Model need to run?

Going by SKILL.md and its folder, Add New Model needs Python for the scripts in its folder and the command-line tools its instructions call (python, hf and bash). Our summary lists: Python 3.

Does Add New Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add New Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add New Model use?

Add New Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add New Model use?

About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add New Model?

Skills that share tags, products or a category with Add New Model: Add Uint Support (pytorch/pytorch, 104k stars), Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Add Op (onnx/onnx, 22k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add New Model?

mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.

Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.