Agent skill

Estimate Memory

by mlc-ai in mlc-ai/pith-train

Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Estimate Memory

skills CLI
$ npx skills add mlc-ai/pith-train --skill estimate-memory -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mlc-ai/pith-train estimate-memory --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/estimate-memory .claude/skills/estimate-memory && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
estimate-memory
GitHub stars
355
Token cost
~1.9k tokens
SKILL.md length
890 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train.

  • Works in 3 steps: Gather Parameters → Run the Estimator → Interpret Results
  • The user asks to estimate memory
  • SKILL.md covers How It Works, Step 1: Gather Parameters, Step 2: Run the Estimator and Step 3: Interpret Results, plus 1 more section
  • Calls python

What it does

Estimate Memory is an agent skill from mlc-ai/pith-train. Estimate peak GPU memory for a DualPipeV training run. Use when the user asks to "estimate memory", "will this fit in memory", "how much GPU memory", "check if this OOMs", "memory for training X on Y GPUs", or mentions memory planning for a training configuration. Translates natural-language descriptions of hardware, model, and training setup into the exact CLI arguments for python -m tools.memoryestimator.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with Python. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.

When your agent uses it

  • The user asks to estimate memory
  • Will this fit in memory
  • How much GPU memory
  • Check if this OOMs

Example prompts

  • “estimate memory”
  • “will this fit in memory”
  • “how much GPU memory”
  • “/estimate-memory”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Gather Parameters
  2. Run the Estimator
  3. Interpret Results

What it can do on your machine

Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Estimate Memory loads about 1.9k tokens when it runs. Until then it costs about 107 tokens; SKILL.md has 890 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 890 words, ~1,873 tokens.

Download SKILL.mdSave it as .claude/skills/estimate-memory/SKILL.md (or your agent's skills folder).
name
estimate-memory
description
Estimate peak GPU memory for a DualPipeV training run. Use when the user asks to "estimate memory", "will this fit in memory", "how much GPU memory", "check if this OOMs", "memory for training X on Y GPUs", or mentions memory planning for a training configuration. Translates natural-language descriptions of hardware, model, and training setup into the exact CLI arguments for `python -m tools.memory_estimator`.

Memory Estimation Skill

Estimates peak GPU memory usage for DualPipeV MoE training runs using an analytical simulator. Takes natural-language descriptions and translates them into the correct CLI invocation.

How It Works

The memory estimator at tools/memory_estimator/ simulates the full DualPipeV 8-step pipeline schedule, tracking activations, autograd saved tensors, gradients, optimizer states, communication buffers, and non-PyTorch overhead (CUDA context, NCCL, torch.compile) at every event boundary. It reports peak memory, a per-component breakdown, and whether the config fits in GPU memory.

Step 1: Gather Parameters

Extract these parameters from the user's message. If any required parameters are missing, do not guess or fill in defaults — list ALL required parameters with a brief description of each and ask the user to provide the missing ones before proceeding. If the user provided no parameters at all (e.g., just /estimate-memory), show the full list of required parameters.

Required parameters
ParameterCLI flagHow to determine
Model config--modelPath to a HuggingFace-style config.json. Available models: examples/pretrain_lm/qwen3-30b-a3b/config.json, examples/pretrain_lm/deepseek-v2-lite/config.json. If the user names a model, find its config under examples/.
PP size--pp-sizePipeline parallel degree. If the user says "2-way pipeline", use 2.
EP size--ep-sizeExpert parallel degree.
CP size--cp-sizeContext parallel degree. Use 1 if not using context parallelism.
Total GPUs--total-gpusTotal GPU count. E.g., "4x8 H100" = 32. "16 GPUs" = 16. FSDP dimensions are derived: dp = (total_gpus / pp) / cp for attention, expt_dp = (total_gpus / pp) / ep for the experts.
Micro batch size--micro-batch-sizePer-GPU micro batch size.
Global batch size--global-batch-sizeGlobal batch size across all GPUs.
Sequence length--sequence-lengthToken count per sequence.
GPU type/memory--gpu-memory-gbGPU memory capacity in GB. See common hardware table below.
Optional parameters (with defaults)
ParameterCLI flagDefaultNotes
FP8 training--fp8-trainingdisableddeep-gemm if user mentions FP8. Caveat: FP8 training increases memory (additional quantized weight and input tensor caches), but this overhead is not yet modeled. When FP8 is enabled, warn the user that the estimate is a lower bound — real usage will be higher.
PP rank--pp-rank-1 (scan all)-1 finds worst case automatically.
EP imbalance--ep-imbalance1.0Increase for skewed expert routing.
Fragmentation--fragmentation0.10CUDA allocator overhead.
CP accuracy caveat

Important: When cp_size > 1, the estimator reduces the per-rank sequence length (S = seq_len / cp_size) and adjusts FSDP sharding, but it does not model ring attention communication buffers or any changes to the autograd saved tensors from ring attention. CP memory estimates have not been validated against real measurements. When reporting results with CP > 1, warn the user: "Note: CP > 1 estimates are approximate — ring attention memory overhead is not fully modeled and has not been validated against real measurements."

Deriving dp_size

The tool computes stage_size = total_gpus / pp_size, then dp_size = stage_size / cp_size and expt_dp_size = stage_size / ep_size. Verify each division is exact. If not, the config is invalid — tell the user.

Common hardware specs
Hardware--gpu-memory-gb
H100 SXM80
H200 SXM141
B200192
Constraint: num_chunks >= pp_size * 2

The tool validates that num_chunks = global_batch_size / (micro_batch_size * dp_size) >= pp_size * 2. EP shards experts, not data, so it does not divide the batch. If this fails, suggest increasing global_batch_size or decreasing micro_batch_size.

Show full SKILL.md (374 more words)Show less

Step 2: Run the Estimator

Build the command and always show the exact command to the user before running it, so they can copy-paste it for manual runs.

bash
python -m tools.memory_estimator \
    --model <config_path> \
    --pp-size <N> --ep-size <N> --cp-size <N> --total-gpus <N> \
    --micro-batch-size <N> --global-batch-size <N> --sequence-length <N> \
    --gpu-memory-gb <N>

Add --detail if the user asks for a detailed breakdown. Add --timeline if the user wants to see the schedule progression.

Step 3: Interpret Results

The output has four sections:

  1. Static Memory — parameters, FSDP shards, optimizer states. Constant during training.
  2. Peak Dynamic Memory — activations, autograd, gradients, comm buffers at the peak event. This is the memory high-water mark.
  3. Grand Total — model tensors + non-PyTorch overhead + fragmentation.
  4. Suggestions — actionable advice if memory is tight.

Status interpretation:

  • OK — fits with >5% headroom.
  • TIGHT — fits but <5% headroom. May OOM under non-uniform routing or memory spikes.
  • OOM — estimated peak exceeds GPU capacity.

Examples

Example 1: Missing parameters

User: "How much memory does Qwen3-30B-A3B need on 32 H100s with pp=4, ep=8?"

Response: "I need a few more details to run the estimate:

  • CP size — context parallel degree (1 if not using CP)
  • Micro batch size — per-GPU micro batch size
  • Global batch size — total batch size across all GPUs
  • Sequence length — tokens per sequence

Could you provide these?"

Example 2: Complete query

User: "Qwen3-30B-A3B, 32 H100s, pp=4, ep=8, cp=1, micro_bs=1, gbs=1024, seq_len=4096"

bash
python -m tools.memory_estimator \
    --model examples/pretrain_lm/qwen3-30b-a3b/config.json \
    --pp-size 4 --ep-size 8 --cp-size 1 --total-gpus 32 \
    --micro-batch-size 1 --global-batch-size 1024 --sequence-length 4096 \
    --gpu-memory-gb 80
Example 3: How many nodes do I need?

User: "Qwen3-235B-A22B, ep=8, cp=1, fsdp=1, micro_bs=1, gbs=1024, seq_len=2048, pp size is the number of B200 nodes, how many B200 nodes do I need?"

This is a search problem. Each B200 node has 8 GPUs (192 GB each). With ep=8, cp=1, fsdp=1, total_gpus = pp * 8. Try pp=2, pp=4, pp=8, etc. until the estimate shows OK status. Run the estimator for each pp size and report the minimum that fits.

Note: this requires a config.json for the model. If the model doesn't have one under examples/, tell the user you need the model's HuggingFace config.json.

Example 4: Comparing configs

User: "Will training fit if I increase sequence length to 8192?"

Run the estimator twice — once with --sequence-length 4096 and once with --sequence-length 8192 — and compare the peak memory and headroom.

Example 5: Exploring parallelism

User: "What's the best parallelism config for 16 GPUs?"

Try multiple configs (e.g., pp=2/ep=8, pp=4/ep=4, pp=2/ep=4/dp=2) and compare peak memory. Report which has the most headroom.

© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/estimate-memory of mlc-ai/pith-train.

Open the folder on GitHubat commit c7c8b1d

Compare with similar skills

Estimate Memory next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Estimate Memory compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Estimate Memory this skillmlc-ai/pith-train355—~1.9kAutomated safety check: PassApache-2.0
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0
PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Axolotl Fine-Tuning ReferenceOrchestra-Research/AI-Research-SKILLs13k9 repos~1.2kAutomated safety check: PassMIT
Hugging Face AccelerateOrchestra-Research/AI-Research-SKILLs13k6 repos~2.1kAutomated safety check: PassMIT
bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs13k3 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Axolotl Fine-Tuning Reference

    Orchestra-Research/AI-Research-SKILLs

    Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats.

    13k GitHub starsUsed in 9 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Accelerate

    Orchestra-Research/AI-Research-SKILLs

    Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups.

    13k GitHub starsUsed in 6 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • bitsandbytes Model Quantization

    Orchestra-Research/AI-Research-SKILLs

    Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

    13k GitHub starsUsed in 3 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Ray Train Distributed Training

    Orchestra-Research/AI-Research-SKILLs

    Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

    13k GitHub starsUsed in 3 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed

More from mlc-ai/pith-train

All 10 skills in this repo
  • Analyze Nsys Profile

    mlc-ai/pith-train

    Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

    355 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Capture Nsys Profile

    mlc-ai/pith-train

    Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.

    355 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed
  • Validate Correctness

    mlc-ai/pith-train

    Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

    355 GitHub stars~2.3k tokensUpdated 3 days ago
    Auto-check passed
  • Validate Performance

    mlc-ai/pith-train

    Measures the throughput difference between two branches with force-balanced routing.

    355 GitHub stars~1.4k tokensUpdated 3 days ago
    Auto-check passed
  • Setup Benchmark Inputs

    mlc-ai/pith-train

    Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.

    355 GitHub stars~399 tokensUpdated 3 days ago
    Auto-check passed
  • Add New Model

    mlc-ai/pith-train

    Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.

    355 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Estimate Memory

What does Estimate Memory do?

Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train. Estimate Memory is an agent skill from mlc-ai/pith-train. Estimate peak GPU memory for a DualPipeV training run.

When should I use Estimate Memory?

Estimate Memory fits situations like: the user asks to estimate memory; will this fit in memory; how much GPU memory; check if this OOMs.

How do I install Estimate Memory in Claude Code?

Run `npx skills add mlc-ai/pith-train --skill estimate-memory -a claude-code`. Or copy the skill folder (.agents/skills/estimate-memory in mlc-ai/pith-train) into .claude/skills/estimate-memory in your project. Claude Code loads it when a task matches its description.

How do I install Estimate Memory in Codex?

Run `npx skills add mlc-ai/pith-train --skill estimate-memory -a codex`. Or copy the skill folder (.agents/skills/estimate-memory in mlc-ai/pith-train) into .agents/skills/estimate-memory in your project. Codex loads it when a task matches its description.

Can I use Estimate Memory in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill estimate-memory -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/estimate-memory, .gemini/skills/estimate-memory, .github/skills/estimate-memory and .opencode/skills/estimate-memory in your project.

What does Estimate Memory need to run?

Going by SKILL.md and its folder, Estimate Memory needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Estimate Memory access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Estimate Memory safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Estimate Memory use?

Estimate Memory is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Estimate Memory use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Estimate Memory?

Skills that share tags, products or a category with Estimate Memory: Hugging Face Vision Trainer (huggingface/skills, 11k stars), PyTorch Lightning Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), Axolotl Fine-Tuning Reference (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Hugging Face Accelerate (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Estimate Memory?

mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.

Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.