Agent skill

Experiment Tracking Setup

by revfactory in revfactory/harness-100

MLflow, Weights & Biases 등 실험 추적 도구 설정, 재현성 보장, 모델 레지스트리, 실험 비교 방법론 가이드.

Apache-2.0Auto-check passed

Install Experiment Tracking Setup

skills CLI
$ npx skills add revfactory/harness-100 --skill experiment-tracking-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install revfactory/harness-100 experiment-tracking-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ko/31-ml-experiment/.claude/skills/experiment-tracking-setup .claude/skills/experiment-tracking-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-tracking-setup
GitHub stars
1.3k
Token cost
~1.2k tokens
SKILL.md length
52 words
Files
1
Skills in repo
464
Repo updated
First seen
Licence
Apache-2.0

At a glance

MLflow, Weights & Biases 등 실험 추적 도구 설정, 재현성 보장, 모델 레지스트리, 실험 비교 방법론 가이드.

  • SKILL.md covers MLflow 설정, 재현성 보장 체크리스트, 모델 레지스트리 and 실험 비교 프레임워크, plus 1 more section
  • Calls pip and conda

What it does

Experiment Tracking Setup is an agent skill from revfactory/harness-100. MLflow, Weights & Biases 등 실험 추적 도구 설정, 재현성 보장, 모델 레지스트리, 실험 비교 방법론 가이드. '실험 추적', 'MLflow', 'W&B', 'Weights and Biases', '재현성', '모델 레지스트리', '실험 비교', '하이퍼파라미터 로깅' 등 ML 실험 관리 시 이 스킬을 사용한다. training-manager의 실험 관리 역량을 강화한다. 단, 모델 아키텍처 설계나 피처 엔지니어링은 이 스킬의 범위가 아니다.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with Weights & Biases and MLflow. The licence is Apache-2.0.

Example prompts

  • “MLflow”
  • “Weights and Biases”
  • “/experiment-tracking-setup”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 8e8d35c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Tracking Setup loads about 1.2k tokens when it runs. Until then it costs about 72 tokens; SKILL.md has 52 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from revfactory/harness-100 at commit 8e8d35c, republished under its Apache-2.0 licence (© revfactory). 52 words, ~1,200 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-tracking-setup/SKILL.md (or your agent's skills folder).
name
experiment-tracking-setup
description
MLflow, Weights & Biases 등 실험 추적 도구 설정, 재현성 보장, 모델 레지스트리, 실험 비교 방법론 가이드. '실험 추적', 'MLflow', 'W&B', 'Weights and Biases', '재현성', '모델 레지스트리', '실험 비교', '하이퍼파라미터 로깅' 등 ML 실험 관리 시 이 스킬을 사용한다. training-manager의 실험 관리 역량을 강화한다. 단, 모델 아키텍처 설계나 피처 엔지니어링은 이 스킬의 범위가 아니다.

Experiment Tracking Setup — 실험 추적 및 재현성 가이드

ML 실험의 추적, 재현성 보장, 모델 버전 관리를 위한 실전 가이드.

MLflow 설정

기본 구조
python
import mlflow

mlflow.set_tracking_uri("http://localhost:5000")
mlflow.set_experiment("order-prediction")

with mlflow.start_run(run_name="xgboost-v2"):
    # 파라미터 로깅
    mlflow.log_params({
        "model": "XGBClassifier",
        "n_estimators": 500,
        "max_depth": 6,
        "learning_rate": 0.1,
    })

    # 학습
    model.fit(X_train, y_train)
    predictions = model.predict(X_test)

    # 메트릭 로깅
    mlflow.log_metrics({
        "accuracy": accuracy_score(y_test, predictions),
        "f1": f1_score(y_test, predictions),
        "precision": precision_score(y_test, predictions),
        "recall": recall_score(y_test, predictions),
    })

    # 모델 저장
    mlflow.sklearn.log_model(model, "model")

    # 아티팩트 저장
    mlflow.log_artifact("confusion_matrix.png")
    mlflow.log_artifact("feature_importance.csv")
자동 로깅
python
# 프레임워크별 자동 로깅
mlflow.sklearn.autolog()     # scikit-learn
mlflow.xgboost.autolog()     # XGBoost
mlflow.lightgbm.autolog()    # LightGBM
mlflow.pytorch.autolog()     # PyTorch
mlflow.tensorflow.autolog()  # TensorFlow

재현성 보장 체크리스트

필수 기록 항목
python
import platform, sys

reproducibility_info = {
    # 환경
    "python_version": sys.version,
    "os": platform.platform(),
    "gpu": torch.cuda.get_device_name(0) if torch.cuda.is_available() else "N/A",

    # 시드
    "random_seed": 42,
    "numpy_seed": 42,
    "torch_seed": 42,

    # 데이터
    "data_version": "v2.1",
    "data_hash": hashlib.md5(open('data.csv','rb').read()).hexdigest(),
    "train_size": len(X_train),
    "test_size": len(X_test),
    "split_method": "StratifiedKFold(5)",

    # 코드
    "git_commit": subprocess.check_output(['git', 'rev-parse', 'HEAD']).decode().strip(),
    "git_branch": subprocess.check_output(['git', 'branch', '--show-current']).decode().strip(),
}
mlflow.log_params(reproducibility_info)
시드 고정
python
import random, numpy as np, torch

def set_seed(seed=42):
    random.seed(seed)
    np.random.seed(seed)
    torch.manual_seed(seed)
    torch.cuda.manual_seed_all(seed)
    torch.backends.cudnn.deterministic = True
    torch.backends.cudnn.benchmark = False
    os.environ['PYTHONHASHSEED'] = str(seed)
의존성 고정
bash
# requirements.txt 정확한 버전
pip freeze > requirements.txt

# pip-compile (권장)
pip-compile requirements.in --generate-hashes

# conda
conda env export --no-builds > environment.yml

모델 레지스트리

MLflow Model Registry 워크플로우
실험(Experiment)
└── 실행(Run)
    └── 모델 아티팩트
        └── 모델 등록 (Model Registry)
            ├── Stage: Staging → 검증
            ├── Stage: Production → 배포
            └── Stage: Archived → 보관
python
# 모델 등록
mlflow.register_model(
    model_uri=f"runs:/{run_id}/model",
    name="order-prediction-model"
)

# 스테이지 전환
client = mlflow.tracking.MlflowClient()
client.transition_model_version_stage(
    name="order-prediction-model",
    version=3,
    stage="Production"
)

# Production 모델 로드
model = mlflow.pyfunc.load_model("models:/order-prediction-model/Production")

실험 비교 프레임워크

통계적 검증
python
from scipy import stats

# 5-fold CV 결과 비교
model_a_scores = [0.85, 0.87, 0.84, 0.86, 0.88]
model_b_scores = [0.82, 0.84, 0.83, 0.81, 0.85]

# 대응 t-검정 (Paired t-test)
t_stat, p_value = stats.ttest_rel(model_a_scores, model_b_scores)
print(f"p-value: {p_value:.4f}")
if p_value < 0.05:
    print("통계적으로 유의미한 차이 있음")
실험 비교 테이블
markdown
| 실험 | 모델 | F1 | Precision | Recall | 학습 시간 | 추론 시간 |
|------|------|-----|-----------|--------|----------|----------|
| exp-001 | LogReg (baseline) | 0.78 | 0.80 | 0.76 | 2s | 0.1ms |
| exp-002 | XGBoost | 0.85 | 0.87 | 0.83 | 45s | 0.5ms |
| exp-003 | LightGBM | 0.86 | 0.88 | 0.84 | 20s | 0.3ms |
| exp-004 | LightGBM + Optuna | 0.88 | 0.89 | 0.87 | 2h | 0.3ms |
| exp-005 | Stacking (top3) | 0.89 | 0.90 | 0.88 | 3h | 1.2ms |

프로젝트 구조 템플릿

ml-project/
├── data/
│   ├── raw/              # 원본 데이터 (수정 금지)
│   ├── processed/        # 전처리 완료
│   └── external/         # 외부 데이터
├── notebooks/            # 탐색적 분석
├── src/
│   ├── data/             # 데이터 로딩/전처리
│   ├── features/         # 피처 엔지니어링
│   ├── models/           # 모델 정의
│   └── evaluation/       # 평가 로직
├── configs/              # 하이퍼파라미터 YAML
├── models/               # 학습된 모델
├── reports/              # 분석 보고서
├── requirements.txt
└── Makefile              # 재현 가능한 실행

© revfactory, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in ko/31-ml-experiment/.claude/skills/experiment-tracking-setup of revfactory/harness-100.

Open the folder on GitHubat commit 8e8d35c

Compare with similar skills

Experiment Tracking Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Tracking Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Tracking Setup this skillrevfactory/harness-1001.3k—~1.2kAutomated safety check: PassApache-2.0
ML Pipeline ExpertJeffallan/claude-skills12k1 repos~1.9kAutomated safety check: PassMIT
Comparefcakyon/phd-skills414—~1.2kAutomated safety check: PassMIT
LaminDB Biological Data Managementdavila7/claude-code-templates32k12 repos~3.6kAutomated safety check: PassMIT
Implementing Mlopsancoleman/ai-design-components5261 repos~9.2kAutomated safety check: PassMIT
Skill Testdatabricks-solutions/ai-dev-kit1.9k—~1.9kAutomated safety check: PassCustom licence

Similar skills

  • ML Pipeline Expert

    Jeffallan/claude-skills

    Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

    12k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Compare

    fcakyon/phd-skills

    Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

    414 GitHub stars~1.2k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed
  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    32k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Implementing Mlops

    ancoleman/ai-design-components

    Strategic guidance for operationalizing machine learning models from experimentation to production.

    526 GitHub starsUsed in 1 repo~9.2k tokens
    DevOps & CloudAuto-check passed
  • Skill Test

    databricks-solutions/ai-dev-kit

    Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.

    1.9k GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Marimo Batch

    koaning/gitcharts

    An opintionated skill to prepare a marimo notebook to make it ready for a scheduled run.

    145 GitHub starsUsed in 1 repo~819 tokens
    Data & AnalyticsAuto-check: notes

More from revfactory/harness-100

All 464 skills in this repo
  • Anti Bot Analyzer

    revfactory/harness-100

    A skill for analyzing website anti-bot defense mechanisms and developing legitimate evasion strategies.

    1.3k GitHub stars~1.1k tokensUpdated 6 mo ago
    Auto-check passed
  • API Error Design Patterns

    revfactory/harness-100

    Reference for designing how an API reports failures: structured error codes, response shapes, client-friendly messages, an error catalog and retry or fallback advice.

    1.3k GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • API Security Checklist

    revfactory/harness-100

    Walks a backend-dev agent through OWASP API Top 10 checks, authentication and authorization patterns, and defense code during API design.

    1.3k GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Arg Parser Generator

    revfactory/harness-100

    Methodology for systematically designing and generating CLI tool argument parser structures.

    1.3k GitHub stars~1.2k tokensUpdated 6 mo ago
    Auto-check passed
  • Audience Segmentation

    revfactory/harness-100

    Audience segmentation skill used by the analyst and curator agents.

    1.3k GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • Audio Storytelling

    revfactory/harness-100

    Audio storytelling skill used by the podcast scriptwriter and show note editor.

    1.3k GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Experiment Tracking Setup

What does Experiment Tracking Setup do?

MLflow, Weights & Biases 등 실험 추적 도구 설정, 재현성 보장, 모델 레지스트리, 실험 비교 방법론 가이드. Experiment Tracking Setup is an agent skill from revfactory/harness-100. MLflow, Weights & Biases 등 실험 추적 도구 설정, 재현성 보장, 모델 레지스트리, 실험 비교 방법론 가이드.

How do I install Experiment Tracking Setup in Claude Code?

Run `npx skills add revfactory/harness-100 --skill experiment-tracking-setup -a claude-code`. Or copy the skill folder (ko/31-ml-experiment/.claude/skills/experiment-tracking-setup in revfactory/harness-100) into .claude/skills/experiment-tracking-setup in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Tracking Setup in Codex?

Run `npx skills add revfactory/harness-100 --skill experiment-tracking-setup -a codex`. Or copy the skill folder (ko/31-ml-experiment/.claude/skills/experiment-tracking-setup in revfactory/harness-100) into .agents/skills/experiment-tracking-setup in your project. Codex loads it when a task matches its description.

Can I use Experiment Tracking Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add revfactory/harness-100 --skill experiment-tracking-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-tracking-setup, .gemini/skills/experiment-tracking-setup, .github/skills/experiment-tracking-setup and .opencode/skills/experiment-tracking-setup in your project.

What does Experiment Tracking Setup need to run?

Going by SKILL.md and its folder, Experiment Tracking Setup needs the command-line tools its instructions call (pip and conda). Our summary lists: Python 3.

Does Experiment Tracking Setup access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Experiment Tracking Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Tracking Setup use?

Experiment Tracking Setup is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Tracking Setup use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiment Tracking Setup?

Skills that share tags, products or a category with Experiment Tracking Setup: ML Pipeline Expert (Jeffallan/claude-skills, 12k stars), Compare (fcakyon/phd-skills, 414 stars), LaminDB Biological Data Management (davila7/claude-code-templates, 32k stars) and Implementing Mlops (ancoleman/ai-design-components, 526 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Tracking Setup?

revfactory (a GitHub user) maintains it in revfactory/harness-100, which has 1,290 GitHub stars. The repository holds 464 skills in this directory. The repository was last updated on March 22, 2026.

Source: revfactory/harness-100 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.