Agent skill

Compare

by fcakyon in fcakyon/phd-skills

Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

MITAuto-check passedResearch & Science

Install Compare

skills CLI
$ npx skills add fcakyon/phd-skills --skill compare -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fcakyon/phd-skills compare --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fcakyon/phd-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/compare .claude/skills/compare && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compare
GitHub stars
414
Token cost
~1.2k tokens
SKILL.md length
578 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

  • Works in 6 steps: Identify the runs → Fetch metric history (not just final… → Find the student's current step → …
  • The user asks to compare runs
  • SKILL.md covers When to run, Auto-detect the tracker, The protocol and Anti-patterns, plus 1 more section
  • Needs WANDB_API_KEY and NEPTUNE_API_TOKEN

What it does

Compare is an agent skill from fcakyon/phd-skills. Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow. Aligns runs at the student's current step (never current-vs-final-of-baseline) and separates proxy metrics from downstream targets. Use when the user asks to compare runs, check if a run is improving, track lag against a baseline, rank experiments, or evaluate run-vs-run performance.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science. It works with Weights & Biases and MLflow. The repository describes itself as: PhD Research Skills for Claude Code: paper reproduction, experiment design, paper review, result comparison and more. The licence is MIT.

When your agent uses it

  • The user asks to compare runs
  • Check if a run is improving
  • Track lag against a baseline
  • Rank experiments

Example prompts

  • “/compare”

Requirements

  • Python 3
  • A credential in WANDB_API_KEY
  • A credential in NEPTUNE_API_TOKEN

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Identify the runs
  2. Fetch metric history (not just final value)
  3. Find the student's current step
  4. Slice the baseline at the same step
  5. Separate proxy metrics from downstream
  6. Run names in output

What it can do on your machine

Read from SKILL.md and the folder at commit 67acd61. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • WANDB_API_KEY
    • NEPTUNE_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compare loads about 1.2k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 578 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fcakyon/phd-skills at commit 67acd61, republished under its MIT licence (© fcakyon). 578 words, ~1,197 tokens.

Download SKILL.mdSave it as .claude/skills/compare/SKILL.md (or your agent's skills folder).
name
compare
description
Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow. Aligns runs at the student's current step (never current-vs-final-of-baseline) and separates proxy metrics from downstream targets. Use when the user asks to compare runs, check if a run is improving, track lag against a baseline, rank experiments, or evaluate run-vs-run performance.

Compare: same-epoch run comparison across trackers

The most common comparison error is reporting "run A is 4 percentage points behind baseline" when run A is at epoch 11 of 100 and the baseline number is from epoch 100. The student is still training; the comparison is meaningless. This skill enforces same-epoch alignment.

The agentic Stop hook routes here from reason when an assistant reports a delta without aligning the runs.

When to run

The user just said any of:

  • "compare run A to baseline / to run B"
  • "is my run improving / catching up / falling behind"
  • "rank these experiments"
  • "X vs Y wandb / neptune"
  • "track lag against baseline"

Auto-detect the tracker

Check in this order:

  1. WANDB_API_KEY env var set, or wandb imports in the project → wandb
  2. NEPTUNE_API_TOKEN env var set → neptune
  3. MLFLOW_TRACKING_URI env var set, or mlruns/ dir present → mlflow
  4. runs/ or lightning_logs/ dir present → tensorboard
  5. *results*.json / *meta*.json files in run dirs → local file format

If none, ask the user where metrics live before guessing.

The protocol

1. Identify the runs

Get full names (no shortcodes). If the user says "fvs-fm vs the baseline", clarify:

  • which fvs-fm run (project + entity + run-id)
  • which baseline (full run name; baselines often have several variants)
2. Fetch metric history (not just final value)

You need the full curve, not the last reported value. Final-value-only comparisons hide convergence dynamics.

For wandb:

python
import wandb
api = wandb.Api()
run = api.run("entity/project/run-id")
history = run.history(samples=10000)  # full history, not just summary

For tensorboard, parse the event files (tensorboard.backend.event_processing.event_accumulator.EventAccumulator).

For neptune / mlflow, use their respective APIs.

3. Find the student's current step

The student is the run still in progress (or the one being evaluated). Get its current epoch / step from the latest history row.

4. Slice the baseline at the same step

This is the critical step. The baseline went all the way to (say) epoch 100. The student is at epoch 11. Pull the baseline's metrics at epoch 11, not at epoch 100.

python
student_step = student_history['epoch'].max()
baseline_at_same_step = baseline_history[baseline_history['epoch'] == student_step]

If the baseline doesn't have an exactly-matching step, interpolate or pick the nearest. State which.

Show full SKILL.md (250 more words)Show less
5. Separate proxy metrics from downstream

Most ML pipelines have a proxy metric (cheap, computed during training, kNN accuracy on features, loss, perplexity) and a target downstream metric (expensive, computed periodically or only at the end, finetuned linear probe accuracy, downstream task F1).

The proxy is for tracking convergence; the target is what the project is actually optimizing. Reporting only the proxy can mislead, a run that lags on kNN may close the gap on downstream finetune. Report both, separately:

                        | student (ep 11) | baseline (ep 11) | delta |
| proxy (kNN top-1)     | 36.4%           | 38.9%            | -2.5  |
| downstream (linear)   | not yet         | 42.1%            | n/a   |

If the user only has proxy data, say so explicitly. Never declare a winner from proxy alone.

6. Run names in output

In every line of the report, use full run names. Never cs-ad vs fvs-fm; always phase1-7src-conv-s-adaptor-mlp vs phase1-7src-fastvit-s-featmap-mlp. Future-you reading this will not remember the shortcode.

Anti-patterns

  • "X is behind baseline by 4pp": without saying at what step. Almost always wrong.
  • "X has converged": without showing the last 5 epochs of the curve.
  • "Best run is Y": based on a metric that was logged differently across runs (different reduction, different eval set).
  • Single-seed comparison treated as definitive. Note variance if known; otherwise label as single-seed.

Output

Compact comparison table per metric pair (proxy + downstream). Each row aligned at the student's current step. Each cell traceable to a specific tracker run-id and step. End with one or two sentences interpreting the comparison, student is on track to catch up at step N, projected from current slope is a useful framing; student is winning / losing is rarely warranted before convergence.

© fcakyon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/compare of fcakyon/phd-skills.

Open the folder on GitHubat commit 67acd61

Compare with similar skills

Compare next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compare compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compare this skillfcakyon/phd-skills414—~1.2kAutomated safety check: PassMIT
LaminDB Biological Data Managementdavila7/claude-code-templates32k12 repos~3.6kAutomated safety check: PassMIT
Experiment Tracking Setuprevfactory/harness-1001.3k—~1.4kAutomated safety check: PassApache-2.0
ML Pipeline ExpertJeffallan/claude-skills12k1 repos~1.9kAutomated safety check: PassMIT
Implementing Mlopsancoleman/ai-design-components5261 repos~9.2kAutomated safety check: PassMIT
Setting Up Experiment Trackingjeremylongshore/tons-of-skills-marketplace2.8k—~954Automated safety check: PassMIT

Similar skills

  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    32k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Experiment Tracking Setup

    revfactory/harness-100

    Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology.

    1.3k GitHub stars~1.4k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • ML Pipeline Expert

    Jeffallan/claude-skills

    Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

    12k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Implementing Mlops

    ancoleman/ai-design-components

    Strategic guidance for operationalizing machine learning models from experimentation to production.

    526 GitHub starsUsed in 1 repo~9.2k tokens
    DevOps & CloudAuto-check passed
  • Setting Up Experiment Tracking

    jeremylongshore/tons-of-skills-marketplace

    Implement machine learning experiment tracking using MLflow or Weights & Biases.

    2.8k GitHub stars~954 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Official

    Exports sanitized metadata, parameters, reproducibility details, quality metrics, and optional review artifacts from Medical AI inference runs or evidence packs to MLflow.

    3.5k GitHub stars~1.3k tokensUpdated today
    Research & ScienceAuto-check: notes

More from fcakyon/phd-skills

All 12 skills in this repo
  • Reproduce

    fcakyon/phd-skills

    End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Debug

    fcakyon/phd-skills

    Evidence-before-action diagnosis of failing ML experiments. An agent skill from fcakyon/phd-skills.

    414 GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed
  • Experiment Design

    fcakyon/phd-skills

    A skill your agent uses when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies.

    414 GitHub stars~987 tokensUpdated 21 days ago
    Auto-check passed
  • Latex Setup

    fcakyon/phd-skills

    A skill your agent uses when the user wants to set up or troubleshoot a LaTeX environment, choose between biber and bibtex, install packages for a specific venue template, or configure compilation.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check: notes
  • Launch

    fcakyon/phd-skills

    Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup.

    414 GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Literature Research

    fcakyon/phd-skills

    A skill your agent uses when the user wants to find related work, survey a research area, identify literature gaps, or discover open-source implementations.

    414 GitHub stars~999 tokensUpdated 21 days ago
    Auto-check passed

Questions about Compare

What does Compare do?

Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow. Compare is an agent skill from fcakyon/phd-skills. Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

When should I use Compare?

Compare fits situations like: the user asks to compare runs; check if a run is improving; track lag against a baseline; rank experiments.

How do I install Compare in Claude Code?

Run `npx skills add fcakyon/phd-skills --skill compare -a claude-code`. Or copy the skill folder (plugin/skills/compare in fcakyon/phd-skills) into .claude/skills/compare in your project. Claude Code loads it when a task matches its description.

How do I install Compare in Codex?

Run `npx skills add fcakyon/phd-skills --skill compare -a codex`. Or copy the skill folder (plugin/skills/compare in fcakyon/phd-skills) into .agents/skills/compare in your project. Codex loads it when a task matches its description.

Can I use Compare in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fcakyon/phd-skills --skill compare -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compare, .gemini/skills/compare, .github/skills/compare and .opencode/skills/compare in your project.

What does Compare need to run?

Going by SKILL.md and its folder, Compare needs credentials named WANDB_API_KEY and NEPTUNE_API_TOKEN. Our summary lists: Python 3; A credential in WANDB_API_KEY; A credential in NEPTUNE_API_TOKEN.

Does Compare access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Compare safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Compare use?

Compare is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compare use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Compare?

Skills that share tags, products or a category with Compare: LaminDB Biological Data Management (davila7/claude-code-templates, 32k stars), Experiment Tracking Setup (revfactory/harness-100, 1.3k stars), ML Pipeline Expert (Jeffallan/claude-skills, 12k stars) and Implementing Mlops (ancoleman/ai-design-components, 526 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compare?

fcakyon (a GitHub user) maintains it in fcakyon/phd-skills, which has 414 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 16, 2026.

Source: fcakyon/phd-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.