Agent skill

Wandb Tracking

by mlc-ai in mlc-ai/pith-train

Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.

Apache-2.0Auto-check: warningsAI & LLM Engineering

Install Wandb Tracking

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add mlc-ai/pith-train --skill wandb-tracking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mlc-ai/pith-train wandb-tracking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/wandb-tracking .claude/skills/wandb-tracking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wandb-tracking
GitHub stars
355
Token cost
~1.1k tokens
SKILL.md length
535 words
Files
4 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.

  • The user pastes a wandb.ai URL
  • SKILL.md covers Prerequisites, Primitives, Runs and metrics and Throughput, plus 2 more sections
  • Runs Python scripts from its folder
  • Asks to list runs

What it does

Wandb Tracking is an agent skill from mlc-ai/pith-train. Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs. Use when the user pastes a wandb.ai URL, or asks to list runs, compare runs or loss curves (per-step and final delta), check whether a variant matches a baseline, read training throughput (tokens-per-second), pull a run's console log after a crash (output.log, stdout+stderr with tracebacks), or build and manage a saved view of which runs to show.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/views.md`, `scripts/runs.py` and `scripts/views.py`).

It sits in AI & LLM Engineering. It works with Weights & Biases. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.

When your agent uses it

  • The user pastes a wandb.ai URL
  • Asks to list runs
  • Loss curves (per-step and final delta)
  • Check whether a variant matches a baseline

Example prompts

  • “/wandb-tracking”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wandb Tracking loads about 1.1k tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 535 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:14
    `wandb.Api()` reads credentials from `~/.netrc`.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 535 words, ~1,120 tokens.

Download SKILL.mdSave it as .claude/skills/wandb-tracking/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
wandb-tracking
description
Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs. Use when the user pastes a wandb.ai URL, or asks to list runs, compare runs or loss curves (per-step and final delta), check whether a variant matches a baseline, read training throughput (tokens-per-second), pull a run's console log after a crash (output.log, stdout+stderr with tracebacks), or build and manage a saved view of which runs to show.

wandb-tracking

Work with wandb experiment data for PithTrain via the scripts below: runs, metrics, console output, and saved views.

The scripts below are the primitives; compose them for the question asked. Do not produce an unsolicited full report. Answer the specific question.

Prerequisites

  • wandb must be authenticated: wandb.Api() reads credentials from ~/.netrc.
  • Scripts live in scripts/ beside this file.

Primitives

QuestionCommand
What runs / metric keys exist?runs.py list <entity/project>
Per-step + final delta of a metric across runsruns.py compare <entity/project> --metric KEY --ref RUN --runs RUN[,RUN...]
Download console output (output.log, stdout+stderr) for runsruns.py logs <entity/project> --runs RUN[,RUN...]
List views / what each view showsviews.py list <entity/project>
What does one ?nw= view show?views.py show <entity/project> <nw_slug>
Create / update a viewviews.py set <entity/project> --show RUN[,RUN...] --name NAME [--update NW]
Delete a viewviews.py delete <entity/project> <nw_slug>

RUN is a run id or display name everywhere. entity/project example: pithtrain/pr74.

Runs and metrics

Run runs.py list first. It prints id | state | steps | name and the union of logged metric keys to feed runs.py compare --metric. PithTrain logs train/cross-entropy-loss, train/gradient-norm, train/load-balance-loss, train/learning-rate, train/step, infra/step-time, infra/tokens-per-second, infra/peak-gpu-memory.

Delta reporting (comparing loss curves)

runs.py compare is the canonical report: it matters for loss curves because a raw final number hides the trajectory. Each --runs entry is compared against --ref (delta = run - ref); for a pairwise A-vs-B use --ref B --runs A. It reports, per run:

  • per-step delta range across all steps (the extreme is usually an early transient).
  • final delta at the last step, plus settling (mean of the last 5): where it's headed.

These are facts, not a verdict. Report them to the user as prose plus a small table, and always state the step count / horizon: a 64-step run only rules out large effects, so say so, and judge the deltas against that horizon yourself.

Show full SKILL.md (231 more words)Show less

Throughput

Throughput is the logged metric infra/tokens-per-second. Pass --drop-first: step 0 is a warmup outlier (first-step compile/alloc). Steady-state fp8-vs-bf16 and pow2-vs-full-mantissa (the two fp8 scale formats) deltas then fall out of runs.py compare. (At small scale, expect fp8 slower than bf16, since quant overhead outweighs the GEMM matrix-multiply win.)

Console output (stdout + stderr)

runs.py logs downloads output.log per run: the run's console output (both stdout and stderr), captured and synced upstream. A crashed run's traceback lands here too, so this is the first place to look when a run died. --file selects a different run file: config.yaml, wandb-metadata.json (host/git/command/GPU), wandb-summary.json, requirements.txt. Offset gotcha: output.log is 1-indexed and starts at step 2; the wandb scalar _step is 0-indexed (0..N-1). Cross-reference by content, not by line number.

Views (?nw=<slug>)

A view (the ?nw=<slug> saved workspace) narrows which of a project's runs are shown. Driving views.py:

  • From a pasted ?nw=<slug> URL, pass the <slug> to show (or list to see every view's shown / hidden runs).
  • set --show RUN[,...] --name NAME creates a view of those runs and prints its ?nw= URL. --from-view NW starts from an existing view's layout/panels; --update NW edits a view in place (same URL), otherwise a new slug is minted.
  • Creating/updating a view is an outward-facing write to the user's project, so confirm the target runs + name first. Reversible with delete.

See references/views.md if you need to modify views.py itself.

© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in .agents/skills/wandb-tracking of mlc-ai/pith-train.

  • SKILL.md
  • references/views.md
  • scripts/runs.py
  • scripts/views.py

Open the folder on GitHubat commit c7c8b1d

Compare with similar skills

Wandb Tracking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wandb Tracking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wandb Tracking this skillmlc-ai/pith-train355—~1.1kAutomated safety check: WarnApache-2.0
ML Experiment IterationLeeroo-AI/superml195—~4.8kAutomated safety check: PassApache-2.0
nanoGPT Training GuideOrchestra-Research/AI-Research-SKILLs13k2 repos~1.7kAutomated safety check: PassMIT
Perforatedai WandbPerforatedAI/PerforatedAI237—~2.8kAutomated safety check: PassApache-2.0
Weights & Biases Experiment TrackingOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
DashboardLegoX/Lego-RL111—~4.2kAutomated safety check: PassApache-2.0

Similar skills

  • ML Experiment Iteration

    Leeroo-AI/superml

    Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues.

    195 GitHub stars~4.8k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • nanoGPT Training Guide

    Orchestra-Research/AI-Research-SKILLs

    Walks through nanoGPT, Karpathy's compact GPT implementation: training on Shakespeare, reproducing GPT-2, fine-tuning GPT-2 checkpoints and training on your own text.

    13k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Perforatedai Wandb

    PerforatedAI/PerforatedAI

    WandB-specific PerforatedAI integration guardrail skill. An agent skill from PerforatedAI/PerforatedAI.

    237 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Weights & Biases Experiment Tracking

    Orchestra-Research/AI-Research-SKILLs

    Guides an agent through tracking ML experiments with W&B: run logging, config capture, hyperparameter sweeps, artifacts and a model registry.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    DevOps & CloudAuto-check passed
  • Dashboard

    LegoX/Lego-RL

    Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths.

    111 GitHub stars~4.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Crud Archive Run

    open-thoughts/OpenThoughts-Agent

    Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

    301 GitHub stars~1.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed

More from mlc-ai/pith-train

All 10 skills in this repo
  • Analyze Nsys Profile

    mlc-ai/pith-train

    Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

    355 GitHub stars~1.9k tokensUpdated 5 days ago
    Auto-check passed
  • Capture Nsys Profile

    mlc-ai/pith-train

    Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.

    355 GitHub stars~1k tokensUpdated 5 days ago
    Auto-check passed
  • Validate Correctness

    mlc-ai/pith-train

    Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

    355 GitHub stars~2.3k tokensUpdated 5 days ago
    Auto-check passed
  • Validate Performance

    mlc-ai/pith-train

    Measures the throughput difference between two branches with force-balanced routing.

    355 GitHub stars~1.4k tokensUpdated 5 days ago
    Auto-check passed
  • Setup Benchmark Inputs

    mlc-ai/pith-train

    Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.

    355 GitHub stars~399 tokensUpdated 5 days ago
    Auto-check passed
  • Add New Model

    mlc-ai/pith-train

    Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.

    355 GitHub stars~4.6k tokensUpdated 5 days ago
    Auto-check passed

Questions about Wandb Tracking

What does Wandb Tracking do?

Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs. Wandb Tracking is an agent skill from mlc-ai/pith-train. Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.

When should I use Wandb Tracking?

Wandb Tracking fits situations like: the user pastes a wandb.ai URL; asks to list runs; loss curves (per-step and final delta); check whether a variant matches a baseline.

How do I install Wandb Tracking in Claude Code?

Run `npx skills add mlc-ai/pith-train --skill wandb-tracking -a claude-code`. Or copy the skill folder (.agents/skills/wandb-tracking in mlc-ai/pith-train) into .claude/skills/wandb-tracking in your project. Claude Code loads it when a task matches its description.

How do I install Wandb Tracking in Codex?

Run `npx skills add mlc-ai/pith-train --skill wandb-tracking -a codex`. Or copy the skill folder (.agents/skills/wandb-tracking in mlc-ai/pith-train) into .agents/skills/wandb-tracking in your project. Codex loads it when a task matches its description.

Can I use Wandb Tracking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill wandb-tracking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wandb-tracking, .gemini/skills/wandb-tracking, .github/skills/wandb-tracking and .opencode/skills/wandb-tracking in your project.

What does Wandb Tracking need to run?

Going by SKILL.md and its folder, Wandb Tracking needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Wandb Tracking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Wandb Tracking safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Wandb Tracking use?

Wandb Tracking is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wandb Tracking use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 559 tokens, read only when the agent opens those files.

What are the alternatives to Wandb Tracking?

Skills that share tags, products or a category with Wandb Tracking: ML Experiment Iteration (Leeroo-AI/superml, 195 stars), nanoGPT Training Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Perforatedai Wandb (PerforatedAI/PerforatedAI, 237 stars) and Weights & Biases Experiment Tracking (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wandb Tracking?

mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.

Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.