Agent skill

Experiment Tracking Setup

by revfactory in revfactory/harness-100

Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology.

Apache-2.0Auto-check passedDevOps & Cloud

Install Experiment Tracking Setup

skills CLI
$ npx skills add revfactory/harness-100 --skill experiment-tracking-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install revfactory/harness-100 experiment-tracking-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .claude/skills && cp -r skills-src/en/31-ml-experiment/.claude/skills/experiment-tracking-setup .claude/skills/experiment-tracking-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-tracking-setup
GitHub stars
1.3k
Token cost
~1.4k tokens
SKILL.md length
53 words
Files
1
Skills in repo
464
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology.

  • ML experiment management involving experiment tracking
  • SKILL.md covers MLflow Setup, Reproducibility Assurance…, Model Registry and Experiment Comparison Framework, plus 1 more section
  • Calls pip and conda
  • Weights and Biases

What it does

Experiment Tracking Setup is an agent skill from revfactory/harness-100. Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology. Use this skill for ML experiment management involving 'experiment tracking', 'MLflow', 'W&B', 'Weights and Biases', 'reproducibility', 'model registry', 'experiment comparison', 'hyperparameter logging', etc. Enhances the training-manager's experiment management capabilities. Note: model architecture design and feature engineering are outside this skill's…

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering MLOps, Reproducible research and Machine learning. It works with Weights & Biases and MLflow. The licence is Apache-2.0.

When your agent uses it

  • ML experiment management involving experiment tracking
  • Weights and Biases
  • Reproducibility
  • Experiment comparison

Example prompts

  • “experiment tracking”
  • “MLflow”
  • “Weights and Biases”
  • “/experiment-tracking-setup”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 8e8d35c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Tracking Setup loads about 1.4k tokens when it runs. Until then it costs about 137 tokens; SKILL.md has 53 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from revfactory/harness-100 at commit 8e8d35c, republished under its Apache-2.0 licence (© revfactory). 53 words, ~1,413 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-tracking-setup/SKILL.md (or your agent's skills folder).
name
experiment-tracking-setup
description
Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology. Use this skill for ML experiment management involving 'experiment tracking', 'MLflow', 'W&B', 'Weights and Biases', 'reproducibility', 'model registry', 'experiment comparison', 'hyperparameter logging', etc. Enhances the training-manager's experiment management capabilities. Note: model architecture design and feature engineering are outside this skill's scope.

Experiment Tracking Setup — Experiment Tracking and Reproducibility Guide

A practical guide for ML experiment tracking, reproducibility assurance, and model version management.

MLflow Setup

Basic Structure
python
import mlflow

mlflow.set_tracking_uri("http://localhost:5000")
mlflow.set_experiment("order-prediction")

with mlflow.start_run(run_name="xgboost-v2"):
    # Parameter logging
    mlflow.log_params({
        "model": "XGBClassifier",
        "n_estimators": 500,
        "max_depth": 6,
        "learning_rate": 0.1,
    })

    # Training
    model.fit(X_train, y_train)
    predictions = model.predict(X_test)

    # Metric logging
    mlflow.log_metrics({
        "accuracy": accuracy_score(y_test, predictions),
        "f1": f1_score(y_test, predictions),
        "precision": precision_score(y_test, predictions),
        "recall": recall_score(y_test, predictions),
    })

    # Save model
    mlflow.sklearn.log_model(model, "model")

    # Save artifacts
    mlflow.log_artifact("confusion_matrix.png")
    mlflow.log_artifact("feature_importance.csv")
Auto-logging
python
# Framework-specific auto-logging
mlflow.sklearn.autolog()     # scikit-learn
mlflow.xgboost.autolog()     # XGBoost
mlflow.lightgbm.autolog()    # LightGBM
mlflow.pytorch.autolog()     # PyTorch
mlflow.tensorflow.autolog()  # TensorFlow

Reproducibility Assurance Checklist

Required Recording Items
python
import platform, sys

reproducibility_info = {
    # Environment
    "python_version": sys.version,
    "os": platform.platform(),
    "gpu": torch.cuda.get_device_name(0) if torch.cuda.is_available() else "N/A",

    # Seeds
    "random_seed": 42,
    "numpy_seed": 42,
    "torch_seed": 42,

    # Data
    "data_version": "v2.1",
    "data_hash": hashlib.md5(open('data.csv','rb').read()).hexdigest(),
    "train_size": len(X_train),
    "test_size": len(X_test),
    "split_method": "StratifiedKFold(5)",

    # Code
    "git_commit": subprocess.check_output(['git', 'rev-parse', 'HEAD']).decode().strip(),
    "git_branch": subprocess.check_output(['git', 'branch', '--show-current']).decode().strip(),
}
mlflow.log_params(reproducibility_info)
Seed Fixing
python
import random, numpy as np, torch

def set_seed(seed=42):
    random.seed(seed)
    np.random.seed(seed)
    torch.manual_seed(seed)
    torch.cuda.manual_seed_all(seed)
    torch.backends.cudnn.deterministic = True
    torch.backends.cudnn.benchmark = False
    os.environ['PYTHONHASHSEED'] = str(seed)
Dependency Pinning
bash
# requirements.txt with exact versions
pip freeze > requirements.txt

# pip-compile (recommended)
pip-compile requirements.in --generate-hashes

# conda
conda env export --no-builds > environment.yml

Model Registry

MLflow Model Registry Workflow
Experiment
└── Run
    └── Model Artifact
        └── Model Registration (Model Registry)
            ├── Stage: Staging → Validation
            ├── Stage: Production → Deployment
            └── Stage: Archived → Archive
python
# Register model
mlflow.register_model(
    model_uri=f"runs:/{run_id}/model",
    name="order-prediction-model"
)

# Stage transition
client = mlflow.tracking.MlflowClient()
client.transition_model_version_stage(
    name="order-prediction-model",
    version=3,
    stage="Production"
)

# Load Production model
model = mlflow.pyfunc.load_model("models:/order-prediction-model/Production")

Experiment Comparison Framework

Statistical Verification
python
from scipy import stats

# Compare 5-fold CV results
model_a_scores = [0.85, 0.87, 0.84, 0.86, 0.88]
model_b_scores = [0.82, 0.84, 0.83, 0.81, 0.85]

# Paired t-test
t_stat, p_value = stats.ttest_rel(model_a_scores, model_b_scores)
print(f"p-value: {p_value:.4f}")
if p_value < 0.05:
    print("Statistically significant difference")
Experiment Comparison Table
markdown
| Experiment | Model | F1 | Precision | Recall | Training Time | Inference Time |
|-----------|-------|-----|-----------|--------|--------------|---------------|
| exp-001 | LogReg (baseline) | 0.78 | 0.80 | 0.76 | 2s | 0.1ms |
| exp-002 | XGBoost | 0.85 | 0.87 | 0.83 | 45s | 0.5ms |
| exp-003 | LightGBM | 0.86 | 0.88 | 0.84 | 20s | 0.3ms |
| exp-004 | LightGBM + Optuna | 0.88 | 0.89 | 0.87 | 2h | 0.3ms |
| exp-005 | Stacking (top3) | 0.89 | 0.90 | 0.88 | 3h | 1.2ms |

Project Structure Template

ml-project/
├── data/
│   ├── raw/              # Original data (do not modify)
│   ├── processed/        # Preprocessed
│   └── external/         # External data
├── notebooks/            # Exploratory analysis
├── src/
│   ├── data/             # Data loading/preprocessing
│   ├── features/         # Feature engineering
│   ├── models/           # Model definitions
│   └── evaluation/       # Evaluation logic
├── configs/              # Hyperparameter YAML
├── models/               # Trained models
├── reports/              # Analysis reports
├── requirements.txt
└── Makefile              # Reproducible execution

© revfactory, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in en/31-ml-experiment/.claude/skills/experiment-tracking-setup of revfactory/harness-100.

Open the folder on GitHubat commit 8e8d35c

Compare with similar skills

Experiment Tracking Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Tracking Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Tracking Setup this skillrevfactory/harness-1001.3k—~1.4kAutomated safety check: PassApache-2.0
ML Pipeline ExpertJeffallan/claude-skills12k1 repos~1.9kAutomated safety check: PassMIT
Implementing Mlopsancoleman/ai-design-components5261 repos~9.2kAutomated safety check: PassMIT
MLflow Experiment TrackingOrchestra-Research/AI-Research-SKILLs13k2 repos~3.9kAutomated safety check: PassMIT
Weights & Biases Experiment TrackingOrchestra-Research/AI-Research-SKILLs13k10 repos~3.1kAutomated safety check: PassMIT
Build ML Pipelineprobabl-ai/skills135—~4kAutomated safety check: PassBSD-3-Clause

Similar skills

  • ML Pipeline Expert

    Jeffallan/claude-skills

    Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

    12k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Implementing Mlops

    ancoleman/ai-design-components

    Strategic guidance for operationalizing machine learning models from experimentation to production.

    526 GitHub starsUsed in 1 repo~9.2k tokens
    DevOps & CloudAuto-check passed
  • MLflow Experiment Tracking

    Orchestra-Research/AI-Research-SKILLs

    Tracks ML experiments, versions models in the MLflow registry and covers deployment and reproducibility, with autologging for common frameworks.

    13k GitHub starsUsed in 2 repos~3.9k tokens
    DevOps & CloudAuto-check passed
  • Weights & Biases Experiment Tracking

    Orchestra-Research/AI-Research-SKILLs

    Guides an agent through tracking ML experiments with W&B: run logging, config capture, hyperparameter sweeps, artifacts and a model registry.

    13k GitHub starsUsed in 10 repos~3.1k tokens
    DevOps & CloudAuto-check passed
  • Build ML Pipeline

    probabl-ai/skills

    Declare the pipeline from data source to predictor as a skrub DataOps graph.

    135 GitHub stars~4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    32k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed

More from revfactory/harness-100

All 464 skills in this repo
  • Anti Bot Analyzer

    revfactory/harness-100

    A skill for analyzing website anti-bot defense mechanisms and developing legitimate evasion strategies.

    1.3k GitHub stars~1.1k tokensUpdated 6 mo ago
    Auto-check passed
  • API Error Design Patterns

    revfactory/harness-100

    Reference for designing how an API reports failures: structured error codes, response shapes, client-friendly messages, an error catalog and retry or fallback advice.

    1.3k GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • API Security Checklist

    revfactory/harness-100

    Walks a backend-dev agent through OWASP API Top 10 checks, authentication and authorization patterns, and defense code during API design.

    1.3k GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Arg Parser Generator

    revfactory/harness-100

    Methodology for systematically designing and generating CLI tool argument parser structures.

    1.3k GitHub stars~1.2k tokensUpdated 6 mo ago
    Auto-check passed
  • Audience Segmentation

    revfactory/harness-100

    Audience segmentation skill used by the analyst and curator agents.

    1.3k GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • Audio Storytelling

    revfactory/harness-100

    Audio storytelling skill used by the podcast scriptwriter and show note editor.

    1.3k GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Experiment Tracking Setup

What does Experiment Tracking Setup do?

Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology. Experiment Tracking Setup is an agent skill from revfactory/harness-100.), reproducibility assurance, model registry, and experiment comparison methodology.

When should I use Experiment Tracking Setup?

Experiment Tracking Setup fits situations like: ML experiment management involving experiment tracking; weights and Biases; reproducibility; experiment comparison.

How do I install Experiment Tracking Setup in Claude Code?

Run `npx skills add revfactory/harness-100 --skill experiment-tracking-setup -a claude-code`. Or copy the skill folder (en/31-ml-experiment/.claude/skills/experiment-tracking-setup in revfactory/harness-100) into .claude/skills/experiment-tracking-setup in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Tracking Setup in Codex?

Run `npx skills add revfactory/harness-100 --skill experiment-tracking-setup -a codex`. Or copy the skill folder (en/31-ml-experiment/.claude/skills/experiment-tracking-setup in revfactory/harness-100) into .agents/skills/experiment-tracking-setup in your project. Codex loads it when a task matches its description.

Can I use Experiment Tracking Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add revfactory/harness-100 --skill experiment-tracking-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-tracking-setup, .gemini/skills/experiment-tracking-setup, .github/skills/experiment-tracking-setup and .opencode/skills/experiment-tracking-setup in your project.

What does Experiment Tracking Setup need to run?

Going by SKILL.md and its folder, Experiment Tracking Setup needs the command-line tools its instructions call (pip and conda). Our summary lists: Python 3.

Does Experiment Tracking Setup access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Experiment Tracking Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Tracking Setup use?

Experiment Tracking Setup is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Tracking Setup use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiment Tracking Setup?

Skills that share tags, products or a category with Experiment Tracking Setup: ML Pipeline Expert (Jeffallan/claude-skills, 12k stars), Implementing Mlops (ancoleman/ai-design-components, 526 stars), MLflow Experiment Tracking (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Weights & Biases Experiment Tracking (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Tracking Setup?

revfactory (a GitHub user) maintains it in revfactory/harness-100, which has 1,290 GitHub stars. The repository holds 464 skills in this directory. The repository was last updated on March 22, 2026.

Source: revfactory/harness-100 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.