Official agent skill

Trackio Experiment Tracking

by huggingface in huggingface/skills

Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Trackio Experiment Tracking

skills CLI
$ npx skills add huggingface/skills --skill huggingface-trackio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huggingface/skills huggingface-trackio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/huggingface-trackio .claude/skills/huggingface-trackio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
huggingface-trackio
GitHub stars
11k
Used in
2 other repos
Token cost
~1.3k tokens
SKILL.md length
424 words
Files
4 (incl. references)
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space.

  • Works in 5 steps: Set up training with alerts — insert… → Launch training — run the script in the… → Poll for alerts — use trackio list… → …
  • Logging training metrics from a Python training script
  • SKILL.md covers Three Interfaces, When to Use Each, Minimal Logging Setup and Autonomous ML Experiment…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Three interfaces are covered through separate reference files. The Python API logs metrics with `trackio.init()`, `trackio.log()` and `trackio.finish()`, or through TRL's `report_to="trackio"`, and for remote or cloud training a `space_id` syncs metrics to a Space dashboard so they outlive the training instance; auto-created Spaces are public by default unless `private=True` is passed.

The same Python API fires alerts with a call like `trackio.alert(title="loss spike", level=trackio.AlertLevel.WARN)` at INFO, WARN or ERROR severity, which print to the terminal, store in the database, show on the dashboard and can go to Slack or Discord webhooks. The skill frames alerts as the main way an agent handles autonomous iteration on training, inserting them for conditions like loss spikes, NaN gradients or stalls, and polling for them over CLI on background runs. A third interface, the `trackio` command, retrieves logged metrics and alerts, for example listing available projects, runs and metrics.

When your agent uses it

  • Logging training metrics from a Python training script
  • Setting up alerts for loss spikes or NaN gradients during training
  • Checking training metrics or alerts from the command line

Example prompts

  • “Add Trackio logging to this training loop and sync it to a private Space.”
  • “Set up a Trackio alert that fires when validation loss spikes.”
  • “List the metrics logged for the latest run with the trackio CLI.”

Requirements

  • Python with the trackio package
  • A Hugging Face account for Space syncing

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Set up training with alerts — insert trackio.alert() calls for diagnostic conditions
  2. Launch training — run the script in the background
  3. Poll for alerts — use trackio list alerts --project --json --since to check for new alerts
  4. Read metrics — use trackio get metric ... to inspect specific values
  5. Iterate — based on alerts and metrics, stop the run, adjust hyperparameters, and launch a new run

What it can do on your machine

Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Trackio Experiment Tracking loads about 1.3k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 424 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 424 words, ~1,272 tokens.

Download SKILL.mdSave it as .claude/skills/huggingface-trackio/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
huggingface-trackio
description
Track and visualize ML training experiments with Trackio. Use when logging metrics during training (Python API), firing alerts for training diagnostics, or retrieving/analyzing logged metrics (CLI). Supports real-time dashboard visualization, alerts with webhooks, HF Space syncing, and JSON output for automation.

Trackio - Experiment Tracking for ML Training

Trackio is an experiment tracking library for logging and visualizing ML training metrics. It syncs to Hugging Face Spaces for real-time monitoring dashboards.

Three Interfaces

TaskInterfaceReference
Logging metrics during trainingPython APIreferences/logging_metrics.md
Firing alerts for training diagnosticsPython APIreferences/alerts.md
Retrieving metrics & alerts after/during trainingCLIreferences/retrieving_metrics.md

When to Use Each

Python API → Logging

Use import trackio in your training scripts to log metrics:

  • Initialize tracking with trackio.init()
  • Log metrics with trackio.log() or use TRL's report_to="trackio"
  • Finalize with trackio.finish()

Key concept: For remote/cloud training, pass space_id — metrics sync to a Space dashboard so they persist after the instance terminates. Auto-created Spaces are public by default — pass private=True if the metrics should not be public.

→ See references/logging_metrics.md for setup, TRL integration, and configuration options.

Python API → Alerts

Insert trackio.alert() calls in training code to flag important events — like inserting print statements for debugging, but structured and queryable:

  • trackio.alert(title="...", level=trackio.AlertLevel.WARN) — fire an alert
  • Three severity levels: INFO, WARN, ERROR
  • Alerts are printed to terminal, stored in the database, shown in the dashboard, and optionally sent to webhooks (Slack/Discord)

Key concept for LLM agents: Alerts are the primary mechanism for autonomous experiment iteration. An agent should insert alerts into training code for diagnostic conditions (loss spikes, NaN gradients, low accuracy, training stalls). Since alerts are printed to the terminal, an agent that is watching the training script's output will see them automatically. For background or detached runs, the agent can poll via CLI instead.

→ See references/alerts.md for the full alerts API, webhook setup, and autonomous agent workflows.

Show full SKILL.md (161 more words)Show less
CLI → Retrieving

Use the trackio command to query logged metrics and alerts:

  • trackio list projects/runs/metrics — discover what's available
  • trackio get project/run/metric — retrieve summaries and values
  • trackio list alerts --project <name> --json — retrieve alerts
  • trackio show — launch the dashboard
  • trackio sync — sync to HF Space

Key concept: Add --json for programmatic output suitable for automation and LLM agents.

→ See references/retrieving_metrics.md for all commands, workflows, and JSON output formats.

Minimal Logging Setup

python
import trackio

# Spaces are PUBLIC by default (good for shareable dashboards);
# pass private=True if the metrics should not be public
trackio.init(project="my-project", space_id="username/trackio", private=True)
trackio.log({"loss": 0.1, "accuracy": 0.9})
trackio.log({"loss": 0.09, "accuracy": 0.91})
trackio.finish()
Minimal Retrieval
bash
trackio list projects --json
trackio get metric --project my-project --run my-run --metric loss --json

Autonomous ML Experiment Workflow

When running experiments autonomously as an LLM agent, the recommended workflow is:

  1. Set up training with alerts — insert trackio.alert() calls for diagnostic conditions
  2. Launch training — run the script in the background
  3. Poll for alerts — use trackio list alerts --project <name> --json --since <timestamp> to check for new alerts
  4. Read metrics — use trackio get metric ... to inspect specific values
  5. Iterate — based on alerts and metrics, stop the run, adjust hyperparameters, and launch a new run
python
import trackio

trackio.init(project="my-project", config={"lr": 1e-4})

for step in range(num_steps):
    loss = train_step()
    trackio.log({"loss": loss, "step": step})

    if step > 100 and loss > 5.0:
        trackio.alert(
            title="Loss divergence",
            text=f"Loss {loss:.4f} still high after {step} steps",
            level=trackio.AlertLevel.ERROR,
        )
    if step > 0 and abs(loss) < 1e-8:
        trackio.alert(
            title="Vanishing loss",
            text="Loss near zero — possible gradient collapse",
            level=trackio.AlertLevel.WARN,
        )

trackio.finish()

Then poll from a separate terminal/process:

bash
trackio list alerts --project my-project --json --since "2025-01-01T00:00:00"

© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/huggingface-trackio of huggingface/skills.

  • SKILL.md
  • references/alerts.md
  • references/logging_metrics.md
  • references/retrieving_metrics.md

Open the folder on GitHubat commit ca0325b

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Trackio Experiment Tracking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Trackio Experiment Tracking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Trackio Experiment Tracking this skillhuggingface/skills11k2 repos~1.3kAutomated safety check: PassApache-2.0
Phoenix LLM ObservabilityOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence

Similar skills

  • Phoenix LLM Observability

    Orchestra-Research/AI-Research-SKILLs

    Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from huggingface/skills

All 25 skills in this repo
  • Official

    Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

    11k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    Auto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed

Questions about Trackio Experiment Tracking

What does Trackio Experiment Tracking do?

Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space. Three interfaces are covered through separate reference files.finish()`, or through TRL's `report_to="trackio"`, and for remote or cloud training a `space_id` syncs metrics to a Space dashboard so they outlive the training instance; auto-created Spaces are public by default unless `private=True` is passed.

When should I use Trackio Experiment Tracking?

Trackio Experiment Tracking fits situations like: logging training metrics from a Python training script; setting up alerts for loss spikes or NaN gradients during training; checking training metrics or alerts from the command line.

How do I install Trackio Experiment Tracking in Claude Code?

Run `npx skills add huggingface/skills --skill huggingface-trackio -a claude-code`. Or copy the skill folder (skills/huggingface-trackio in huggingface/skills) into .claude/skills/huggingface-trackio in your project. Claude Code loads it when a task matches its description.

How do I install Trackio Experiment Tracking in Codex?

Run `npx skills add huggingface/skills --skill huggingface-trackio -a codex`. Or copy the skill folder (skills/huggingface-trackio in huggingface/skills) into .agents/skills/huggingface-trackio in your project. Codex loads it when a task matches its description.

Can I use Trackio Experiment Tracking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill huggingface-trackio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huggingface-trackio, .gemini/skills/huggingface-trackio, .github/skills/huggingface-trackio and .opencode/skills/huggingface-trackio in your project.

What does Trackio Experiment Tracking need to run?

SKILL.md names no scripts, command-line tools or credentials: Trackio Experiment Tracking is instructions for the agent only. Our summary lists: Python with the trackio package; A Hugging Face account for Space syncing.

Does Trackio Experiment Tracking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Trackio Experiment Tracking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Trackio Experiment Tracking use?

Trackio Experiment Tracking is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Trackio Experiment Tracking use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.1k tokens, read only when the agent opens those files.

What are the alternatives to Trackio Experiment Tracking?

Skills that share tags, products or a category with Trackio Experiment Tracking: Phoenix LLM Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars), Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars) and LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Trackio Experiment Tracking?

huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,148 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.

Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.