Agent skill

Huggingface Vision Trainer

by waybarrios in waybarrios/opencode-power-pack

Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Huggingface Vision Trainer

skills CLI
$ npx skills add waybarrios/opencode-power-pack --skill huggingface-vision-trainer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install waybarrios/opencode-power-pack huggingface-vision-trainer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/huggingface-vision-trainer .claude/skills/huggingface-vision-trainer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
huggingface-vision-trainer
GitHub stars
534
Token cost
~2.7k tokens
SKILL.md length
1,043 words
Files
12 (incl. scripts, references)
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

  • Works in 5 steps: Verify prerequisites (account, token,… → Validate dataset format with the… → Ask the user about dataset size (quick… → …
  • Vision fine-tuning and evaluation
  • SKILL.md covers When to Use, Local Script Execution, Prerequisites Checklist and Dataset Validation, plus 8 more sections
  • Runs Python scripts from its folder; calls uv and hf; reaches huggingface.co; needs HF_TOKEN

What it does

Huggingface Vision Trainer is an agent skill from waybarrios/opencode-power-pack. Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs. Use for vision fine-tuning and evaluation; use huggingface-llm-trainer for language models.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `references/finetune_sam2_trainer.md`, `references/hub_saving.md` and `references/image_classification_training_notebook.md`).

It sits in AI & LLM Engineering, covering Model hubs and datasets, Computer vision and Fine-tuning. It works with Hugging Face. The repository describes itself as: 54 rigorous skills for Codex, OpenCode, and Pi: code review, security audit, feature development, frontend design, MCP tools, Hugging Face ML/training, and more. The licence is Apache-2.0.

When your agent uses it

  • Vision fine-tuning and evaluation
  • Use huggingface-llm-trainer for language models

Example prompts

  • “/huggingface-vision-trainer”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Verify prerequisites (account, token, dataset).
  2. Validate dataset format with the inspector, before spending GPU time.
  3. Ask the user about dataset size (quick 10% test vs. full) and whether to create a validation split, and which GPU hardware to use…
  4. Prepare the training script: scripts/object_detection_training.py (OD), scripts/image_classification_training.py (IC), or…
  5. Save the script to submitted_jobs/_.py, submit the job, and report the job ID, monitoring URL, Trackio dashboard…

What it can do on your machine

Read from SKILL.md and the folder at commit 9dccb6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Huggingface Vision Trainer loads about 2.7k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 53 tokens; SKILL.md has 1,043 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from waybarrios/opencode-power-pack at commit 9dccb6d, republished under its Apache-2.0 licence (© waybarrios). 1,043 words, ~2,704 tokens.

Download SKILL.mdSave it as .claude/skills/huggingface-vision-trainer/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
huggingface-vision-trainer
description
Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs. Use for vision fine-tuning and evaluation; use huggingface-llm-trainer for language models.
license
Apache-2.0 (modified; see UPSTREAMS.json)

Vision Model Training on Hugging Face Jobs

Train object detection, image classification, and SAM/SAM2 segmentation models on managed cloud GPUs. No local GPU setup required — results are automatically saved to the Hugging Face Hub. For text/language model fine-tuning (SFT/DPO/GRPO via TRL), use this pack's huggingface-llm-trainer skill instead.

When to Use

Fine-tuning object detection models (D-FINE, RT-DETR v2, DETR, YOLOS), image classification models (any timm/ model or Transformers classifier), or SAM/SAM2 segmentation models (bbox or point prompts) on custom datasets — locally or on Hugging Face Jobs.

Local Script Execution

Helper scripts use PEP 723 inline dependencies:

bash
uv run scripts/dataset_inspector.py --dataset username/dataset-name --split train
uv run scripts/estimate_cost.py --help

Prerequisites Checklist

  • Hugging Face account with Pro/Team/Enterprise plan (Jobs require a paid plan). Authenticated login (hf auth whoami), token with write permissions passed in job secrets.
  • Object detection: dataset on the Hub with an objects column (bbox, category, optional area). Bboxes in xywh (COCO) or xyxy (Pascal VOC) — auto-detected/converted. Categories can be integers or strings (auto-remapped). image_id optional, auto-generated.
  • Image classification: an image column (PIL images) and a label column (integer or string class IDs, ClassLabel or plain — auto-remapped). Common alt names (labels, class, fine_label) auto-detected.
  • SAM/SAM2 segmentation: an image column, a mask column (binary ground-truth mask), and a prompt — either a prompt column with JSON ({"bbox": [...]} or {"point": [...]}), or dedicated bbox/point columns (xyxy, absolute pixels). Example dataset: merve/MicroMat-mini.
  • Always validate unknown datasets first (see Dataset Validation below).
  • Timeout must exceed expected training time — default 30min is too short, use 2-4h minimum for vision training.
  • Hub push enabled: push_to_hub=True, hub_model_id="username/model-name", token in secrets.

Dataset Validation

Validate BEFORE launching GPU training — the #1 cause of training failures is format mismatches. Skip only for well-known defaults (e.g. cppe-5). Run via Jobs (avoids local SSL/dependency issues), locally with uv run scripts/dataset_inspector.py --dataset ... --split train, or via HfApi().run_uv_job(script="scripts/dataset_inspector.py", script_args=[...], flavor="cpu-basic", timeout=300). Output markers: ✓ READY or ✗ NEEDS FORMATTING (with mapping code).

The object detection training script auto-handles bbox format detection/conversion, sanitization, image_id generation, and category remapping — no manual preprocessing needed beyond having objects.bbox/objects.category.

Training Workflow

  1. Verify prerequisites (account, token, dataset).
  2. Validate dataset format with the inspector, before spending GPU time.
  3. Ask the user about dataset size (quick 10% test vs. full) and whether to create a validation split, and which GPU hardware to use — present as explicit options rather than assuming.
  4. Prepare the training script: scripts/object_detection_training.py (OD), scripts/image_classification_training.py (IC), or scripts/sam_segmentation_training.py (SAM). All use HfArgumentParser — configure via CLI-style script_args, not by editing Python variables. See references/timm_trainer.md for timm details and references/finetune_sam2_trainer.md for SAM2 details.
  5. Save the script to submitted_jobs/<dataset>_<timestamp>.py, submit the job, and report the job ID, monitoring URL, Trackio dashboard (https://huggingface.co/spaces/{username}/trackio), expected time, and estimated cost. Wait for the user to request status checks — don't poll; jobs are asynchronous and can take hours.

Job Submission

Submit via the hf jobs uv run CLI, an hf_jobs() MCP tool if the Hugging Face MCP server is configured, or the Python API directly:

python
from huggingface_hub import HfApi, get_token
api = HfApi()
job_info = api.run_uv_job(
    script="scripts/object_detection_training.py",  # file PATH, not inline content, for the Python API
    script_args=["--dataset_name", "cppe-5", "--push_to_hub", "--hub_model_id", "username/model-name", ...],
    flavor="a10g-large",
    timeout=14400,  # seconds
    env={"PYTHONUNBUFFERED": "1"},
    secrets={"HF_TOKEN": get_token()},  # use get_token(), not the literal string "$HF_TOKEN"
)
print(f"Job ID: {job_info.id}")  # .id, not .job_id or .name

If using an MCP hf_jobs() tool instead, the script parameter accepts inline code or a URL (not local paths), timeout is a string ("4h"), and secrets use the literal "$HF_TOKEN" placeholder (auto-replaced) rather than get_token(). Either way, the training script must include PEP 723 inline dependency metadata and must NOT use image/command parameters (those belong to a different job type).

Token injection is required in custom scripts: the Transformers Trainer calls create_repo(token=self.args.hub_token) when push_to_hub=True, so the script must set training_args.hub_token from os.environ.get("HF_TOKEN") after parsing args but before constructing Trainer — scripts/object_detection_training.py already does this; replicate it in custom scripts. Don't call login() unless replicating that same pattern, and don't rely on implicit token resolution.

Show full SKILL.md (454 more words)Show less
Required flags per modality

Object detection: --no_remove_unused_columns (preserves the image column), --no_eval_do_concat_batches (variable box counts per image), --push_to_hub, --hub_model_id, --metric_for_best_model eval_map, --greater_is_better True (must be explicit — it's Optional[bool]), --do_train, --do_eval.

Image classification: --no_remove_unused_columns, --push_to_hub, --hub_model_id, --metric_for_best_model eval_accuracy, --greater_is_better True, --do_train, --do_eval.

SAM/SAM2: --remove_unused_columns False (preserves input_boxes/input_points), --push_to_hub, --hub_model_id, --do_train, --prompt_type bbox (or point), --dataloader_pin_memory False (avoids pin_memory issues with the custom collator).

Bare bool flags (push_to_hub, do_train) can be negated with --no_ prefix; Optional[bool] fields (greater_is_better) require an explicit True/False value.

Timeout Management

Default 30min is too short for vision training. Minimum 2-4h, with a 30% buffer for loading/preprocessing/Hub push: quick test (100-200 images) 1h, development (500-1K images) 2-3h, production (1K-5K images) 4-6h, large (5K+) 6-12h.

Trackio Monitoring

Always enabled in the object detection script (calls trackio.init()/trackio.finish() automatically, project name from --output_dir, run name from --run_name). For image classification, pass --report_to trackio explicitly. Dashboard: https://huggingface.co/spaces/{username}/trackio.

Model & Hardware Selection

Object detection (all under 100M params — t4-small, 16GB/$0.40/hr, is sufficient): start with ustc-community/dfine-small-coco (10.4M, fast/cheap SOTA), move up to ustc-community/dfine-large-coco (31.4M) or PekingU/rtdetr_v2_r50vd (43M) for accuracy; ustc-community/dfine-xlarge-obj365 (63.5M) and PekingU/rtdetr_v2_r101vd (76M) for the largest variants.

Image classification (timm/ models work out of the box via AutoModelForImageClassification, see references/timm_trainer.md): start with timm/mobilenetv3_small_100.lamb_in1k (2.5M, mobile/edge), move to timm/resnet50.a1_in1k (25.6M) or timm/vit_base_patch16_dinov3.lvd1689m (86.6M, best accuracy).

SAM/SAM2 (only the mask decoder trains by default — vision/prompt encoders frozen): start with facebook/sam2.1-hiera-small (46.0M); facebook/sam2.1-hiera-tiny (38.9M) for speed, facebook/sam2.1-hiera-large (224.4M) or the original facebook/sam-vit-* family for best accuracy at higher VRAM cost.

t4-small handles all recommended OD/IC models and SAM2 up to hiera-base-plus; use l4x1 ($0.80/hr) or a10g-large ($1.50/hr) for sam2.1-hiera-large or SAM v1 models, or if you hit OOM (reduce batch size first). Run scripts/estimate_cost.py for a cost estimate.

Checking Job Status

Via MCP tool if available: hf_jobs("ps"), hf_jobs("logs", {"job_id": "..."}), hf_jobs("inspect", {"job_id": "..."}). Via Python API: HfApi().list_jobs(), .get_job_logs(job_id=...), .get_job(job_id=...).

Common Failure Modes

  • CUDA OOM: reduce per_device_train_batch_size (try 4, then 2), reduce image size, or upgrade hardware.
  • Dataset format errors: run scripts/dataset_inspector.py first; ensure objects.bbox/objects.category are well-formed.
  • Hub push failures (401): confirm job secrets include the token, the script sets training_args.hub_token before constructing Trainer, push_to_hub=True, correct hub_model_id, and write permissions.
  • Job timeout: increase timeout, reduce epochs/dataset, or checkpoint with hub_strategy="every_save".
  • KeyError: 'test': the OD script falls back to the validation split automatically — use the latest template.
  • Single-class "iteration over a 0-d tensor": torchmetrics.MeanAveragePrecision returns scalar tensors for one-class datasets — the OD template already .unsqueeze(0)s these.
  • Poor mAP (<0.15): more epochs (30-50), 500+ images, check per-class mAP for imbalance, try learning rates 1e-5 to 1e-4, larger image size.

See references/reliability_principles.md for the full guide.

Resources

Scripts: scripts/object_detection_training.py, image_classification_training.py, sam_segmentation_training.py, dataset_inspector.py, estimate_cost.py.

References: references/object_detection_training_notebook.md, image_classification_training_notebook.md, finetune_sam2_trainer.md, timm_trainer.md, hub_saving.md, reliability_principles.md.

External: Object Detection Guide, Image Classification Guide, HF Jobs Guide, HF Jobs Configuration, SAM2 docs, SAM docs.

© waybarrios, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references) in skills/huggingface-vision-trainer of waybarrios/opencode-power-pack.

  • SKILL.md
  • references/finetune_sam2_trainer.md
  • references/hub_saving.md
  • references/image_classification_training_notebook.md
  • references/object_detection_training_notebook.md
  • references/reliability_principles.md
  • references/timm_trainer.md
  • scripts/dataset_inspector.py
  • scripts/estimate_cost.py
  • scripts/image_classification_training.py
  • scripts/object_detection_training.py
  • scripts/sam_segmentation_training.py

Open the folder on GitHubat commit 9dccb6d

Compare with similar skills

Huggingface Vision Trainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Huggingface Vision Trainer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Huggingface Vision Trainer this skillwaybarrios/opencode-power-pack534—~2.7kAutomated safety check: PassApache-2.0
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0
Hugging Face Transformers Usagedavila7/claude-code-templates33k11 repos~1.2kAutomated safety check: PassMIT
Tao Finetune Huggingface ModelNVIDIA/skills3.6k—~4.9kAutomated safety check: NotesApache-2.0
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0
Hugging Face Vision Trainerhenryalouf/ruflow157—~7.4kAutomated safety check: PassMIT

Similar skills

  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Transformers Usage

    davila7/claude-code-templates

    Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

    33k GitHub starsUsed in 11 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.

    3.6k GitHub stars~4.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Train or fine-tune vision models on Hugging Face Jobs for detection, classification, and SAM or SAM2 segmentation.

    157 GitHub stars~7.4k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Vision Trainer

    sickn33/agentic-awesome-skills

    Train object detection, image classification, and SAM or SAM2 segmentation models locally or on Hugging Face Jobs, with dataset validation and results saved to the Hub.

    47k GitHub starsUsed in 1 repo~1.1k tokens
    AI & LLM EngineeringAuto-check passed

More from waybarrios/opencode-power-pack

All 32 skills in this repo
  • Hf Cloud Sagemaker Iam Preflight

    waybarrios/opencode-power-pack

    Verify or select a SageMaker execution role before creating models, endpoints, or training jobs.

    534 GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • Huggingface LLM Trainer

    waybarrios/opencode-power-pack

    Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

    534 GitHub stars~3k tokensUpdated 4 days ago
    Auto-check passed
  • Codeql

    waybarrios/opencode-power-pack

    Run CodeQL database creation and security queries, add data-extension models, or process CodeQL SARIF.

    534 GitHub starsUsed in 2 repos~3.7k tokens
    Auto-check passed
  • Semgrep

    waybarrios/opencode-power-pack

    Run Semgrep static analysis across a codebase, optionally using Semgrep Pro for cross-file taint analysis.

    534 GitHub stars~2.4k tokensUpdated 4 days ago
    Auto-check passed
  • Insecure Defaults

    waybarrios/opencode-power-pack

    Detects fail-open insecure defaults (hardcoded secrets, weak auth, permissive security) that allow apps to run insecurely in production.

    534 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Train Sentence Transformers

    waybarrios/opencode-power-pack

    Train or fine-tune SentenceTransformer bi-encoders, CrossEncoder rerankers, or SparseEncoder models, including losses, negatives, evaluation, distillation, LoRA, and Matryoshka.

    534 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Huggingface Vision Trainer

What does Huggingface Vision Trainer do?

Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs. Huggingface Vision Trainer is an agent skill from waybarrios/opencode-power-pack. Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

When should I use Huggingface Vision Trainer?

Huggingface Vision Trainer fits situations like: vision fine-tuning and evaluation; use huggingface-llm-trainer for language models.

How do I install Huggingface Vision Trainer in Claude Code?

Run `npx skills add waybarrios/opencode-power-pack --skill huggingface-vision-trainer -a claude-code`. Or copy the skill folder (skills/huggingface-vision-trainer in waybarrios/opencode-power-pack) into .claude/skills/huggingface-vision-trainer in your project. Claude Code loads it when a task matches its description.

How do I install Huggingface Vision Trainer in Codex?

Run `npx skills add waybarrios/opencode-power-pack --skill huggingface-vision-trainer -a codex`. Or copy the skill folder (skills/huggingface-vision-trainer in waybarrios/opencode-power-pack) into .agents/skills/huggingface-vision-trainer in your project. Codex loads it when a task matches its description.

Can I use Huggingface Vision Trainer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add waybarrios/opencode-power-pack --skill huggingface-vision-trainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huggingface-vision-trainer, .gemini/skills/huggingface-vision-trainer, .github/skills/huggingface-vision-trainer and .opencode/skills/huggingface-vision-trainer in your project.

What does Huggingface Vision Trainer need to run?

Going by SKILL.md and its folder, Huggingface Vision Trainer needs Python for the scripts in its folder, the command-line tools its instructions call (uv and hf) and credentials named HF_TOKEN. Our summary lists: Python 3.

Does Huggingface Vision Trainer access the network?

SKILL.md names 1 domain. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Huggingface Vision Trainer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Huggingface Vision Trainer use?

Huggingface Vision Trainer is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Huggingface Vision Trainer use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to Huggingface Vision Trainer?

Skills that share tags, products or a category with Huggingface Vision Trainer: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Hugging Face Transformers Usage (davila7/claude-code-templates, 33k stars), Tao Finetune Huggingface Model (NVIDIA/skills, 3.6k stars) and Hugging Face LLM Trainer (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Huggingface Vision Trainer?

waybarrios (a GitHub user) maintains it in waybarrios/opencode-power-pack, which has 534 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 6, 2026.

Source: waybarrios/opencode-power-pack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.