ML Experiment Iteration
Leeroo-AI/superml
Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues.
Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add mlc-ai/pith-train --skill wandb-tracking -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mlc-ai/pith-train wandb-tracking --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/wandb-tracking .claude/skills/wandb-tracking && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "wandb-tracking" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/wandb-tracking into .claude/skills/wandb-tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wandb-tracking", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/wandb-trackingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mlc-ai/pith-train --skill wandb-tracking -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mlc-ai/pith-train wandb-tracking --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/wandb-tracking .agents/skills/wandb-tracking && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "wandb-tracking" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/wandb-tracking into .agents/skills/wandb-tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wandb-tracking", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mlc-ai/pith-train --skill wandb-tracking -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mlc-ai/pith-train wandb-tracking --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/wandb-tracking .cursor/skills/wandb-tracking && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "wandb-tracking" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/wandb-tracking into .cursor/skills/wandb-tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wandb-tracking", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mlc-ai/pith-train.git --path .agents/skills/wandb-tracking--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mlc-ai/pith-train --skill wandb-tracking -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mlc-ai/pith-train wandb-tracking --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/wandb-tracking .gemini/skills/wandb-tracking && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "wandb-tracking" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/wandb-tracking into .gemini/skills/wandb-tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wandb-tracking", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mlc-ai/pith-train wandb-trackingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mlc-ai/pith-train --skill wandb-tracking -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/wandb-tracking .github/skills/wandb-tracking && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "wandb-tracking" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/wandb-tracking into .github/skills/wandb-tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wandb-tracking", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mlc-ai/pith-train --skill wandb-tracking -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mlc-ai/pith-train wandb-tracking --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/wandb-tracking .opencode/skills/wandb-tracking && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "wandb-tracking" agent skill from https://github.com/mlc-ai/pith-train/tree/main/.agents/skills/wandb-tracking into .opencode/skills/wandb-tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wandb-tracking", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
wandb-trackingRead, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.
Wandb Tracking is an agent skill from mlc-ai/pith-train. Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs. Use when the user pastes a wandb.ai URL, or asks to list runs, compare runs or loss curves (per-step and final delta), check whether a variant matches a baseline, read training throughput (tokens-per-second), pull a run's console log after a crash (output.log, stdout+stderr with tracebacks), or build and manage a saved view of which runs to show.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/views.md`, `scripts/runs.py` and `scripts/views.py`).
It sits in AI & LLM Engineering. It works with Weights & Biases. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Wandb Tracking loads about 1.1k tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 535 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
`wandb.Api()` reads credentials from `~/.netrc`.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 535 words, ~1,120 tokens.
.claude/skills/wandb-tracking/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Work with wandb experiment data for PithTrain via the scripts below: runs, metrics, console output, and saved views.
The scripts below are the primitives; compose them for the question asked. Do not produce an unsolicited full report. Answer the specific question.
wandb.Api() reads credentials from ~/.netrc.scripts/ beside this file.| Question | Command |
|---|---|
| What runs / metric keys exist? | runs.py list <entity/project> |
| Per-step + final delta of a metric across runs | runs.py compare <entity/project> --metric KEY --ref RUN --runs RUN[,RUN...] |
Download console output (output.log, stdout+stderr) for runs | runs.py logs <entity/project> --runs RUN[,RUN...] |
| List views / what each view shows | views.py list <entity/project> |
What does one ?nw= view show? | views.py show <entity/project> <nw_slug> |
| Create / update a view | views.py set <entity/project> --show RUN[,RUN...] --name NAME [--update NW] |
| Delete a view | views.py delete <entity/project> <nw_slug> |
RUN is a run id or display name everywhere. entity/project example: pithtrain/pr74.
Run runs.py list first. It prints id | state | steps | name and the union of logged metric keys to feed runs.py compare --metric. PithTrain logs train/cross-entropy-loss, train/gradient-norm, train/load-balance-loss, train/learning-rate, train/step, infra/step-time, infra/tokens-per-second, infra/peak-gpu-memory.
runs.py compare is the canonical report: it matters for loss curves because a raw final number hides the trajectory. Each --runs entry is compared against --ref (delta = run - ref); for a pairwise A-vs-B use --ref B --runs A. It reports, per run:
These are facts, not a verdict. Report them to the user as prose plus a small table, and always state the step count / horizon: a 64-step run only rules out large effects, so say so, and judge the deltas against that horizon yourself.
Throughput is the logged metric infra/tokens-per-second. Pass --drop-first: step 0 is a warmup outlier (first-step compile/alloc). Steady-state fp8-vs-bf16 and pow2-vs-full-mantissa (the two fp8 scale formats) deltas then fall out of runs.py compare. (At small scale, expect fp8 slower than bf16, since quant overhead outweighs the GEMM matrix-multiply win.)
runs.py logs downloads output.log per run: the run's console output (both stdout and stderr), captured and synced upstream. A crashed run's traceback lands here too, so this is the first place to look when a run died. --file selects a different run file: config.yaml, wandb-metadata.json (host/git/command/GPU), wandb-summary.json, requirements.txt. Offset gotcha: output.log is 1-indexed and starts at step 2; the wandb scalar _step is 0-indexed (0..N-1). Cross-reference by content, not by line number.
?nw=<slug>)A view (the ?nw=<slug> saved workspace) narrows which of a project's runs are shown. Driving views.py:
?nw=<slug> URL, pass the <slug> to show (or list to see every view's shown / hidden runs).set --show RUN[,...] --name NAME creates a view of those runs and prints its ?nw= URL. --from-view NW starts from an existing view's layout/panels; --update NW edits a view in place (same URL), otherwise a new slug is minted.delete.See references/views.md if you need to modify views.py itself.
© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts, references) in .agents/skills/wandb-tracking of mlc-ai/pith-train.
Open the folder on GitHubat commit c7c8b1d
Wandb Tracking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Wandb Tracking this skillmlc-ai/pith-train | 355 | — | ~1.1k | Automated safety check: Warn | Apache-2.0 | |
| ML Experiment IterationLeeroo-AI/superml | 195 | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | |
| nanoGPT Training GuideOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Perforatedai WandbPerforatedAI/PerforatedAI | 237 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Weights & Biases Experiment TrackingOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.1k | Automated safety check: Pass | MIT | |
| DashboardLegoX/Lego-RL | 111 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 |
Leeroo-AI/superml
Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues.
Orchestra-Research/AI-Research-SKILLs
Walks through nanoGPT, Karpathy's compact GPT implementation: training on Shakespeare, reproducing GPT-2, fine-tuning GPT-2 checkpoints and training on your own text.
PerforatedAI/PerforatedAI
WandB-specific PerforatedAI integration guardrail skill. An agent skill from PerforatedAI/PerforatedAI.
Orchestra-Research/AI-Research-SKILLs
Guides an agent through tracking ML experiments with W&B: run logging, config capture, hyperparameter sweeps, artifacts and a model registry.
LegoX/Lego-RL
Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths.
open-thoughts/OpenThoughts-Agent
Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…
mlc-ai/pith-train
Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.
mlc-ai/pith-train
Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.
mlc-ai/pith-train
Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.
mlc-ai/pith-train
Measures the throughput difference between two branches with force-balanced routing.
mlc-ai/pith-train
Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.
mlc-ai/pith-train
Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.
Works with
Categories
Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs. Wandb Tracking is an agent skill from mlc-ai/pith-train. Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.
Wandb Tracking fits situations like: the user pastes a wandb.ai URL; asks to list runs; loss curves (per-step and final delta); check whether a variant matches a baseline.
Run `npx skills add mlc-ai/pith-train --skill wandb-tracking -a claude-code`. Or copy the skill folder (.agents/skills/wandb-tracking in mlc-ai/pith-train) into .claude/skills/wandb-tracking in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mlc-ai/pith-train --skill wandb-tracking -a codex`. Or copy the skill folder (.agents/skills/wandb-tracking in mlc-ai/pith-train) into .agents/skills/wandb-tracking in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill wandb-tracking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wandb-tracking, .gemini/skills/wandb-tracking, .github/skills/wandb-tracking and .opencode/skills/wandb-tracking in your project.
Going by SKILL.md and its folder, Wandb Tracking needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Wandb Tracking is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 559 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Wandb Tracking: ML Experiment Iteration (Leeroo-AI/superml, 195 stars), nanoGPT Training Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Perforatedai Wandb (PerforatedAI/PerforatedAI, 237 stars) and Weights & Biases Experiment Tracking (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.
Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.