Agent skill

Hugging Face Vision Trainer

by sickn33 in sickn33/agentic-awesome-skills

Train object detection, image classification, and SAM or SAM2 segmentation models locally or on Hugging Face Jobs, with dataset validation and results saved to the Hub.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Hugging Face Vision Trainer

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill hugging-face-vision-trainer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills hugging-face-vision-trainer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hugging-face-vision-trainer .claude/skills/hugging-face-vision-trainer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hugging-face-vision-trainer
GitHub stars
47k
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
522 words
Files
13 (incl. scripts, references)
Skills in repo
1,497
Repo updated
First seen
Licence
Apache-2.0

At a glance

Train object detection, image classification, and SAM or SAM2 segmentation models locally or on Hugging Face Jobs, with dataset validation and results saved to the Hub.

  • Tasks that involve Computer vision
  • SKILL.md covers Detailed Guide, When to Use This Skill, Prerequisites Checklist and Limitations
  • Runs Python scripts from its folder; calls hf
  • Tasks that involve Model hubs and datasets

What it does

Hugging Face Vision Trainer is an agent skill from sickn33/agentic-awesome-skills. Train object detection, image classification, and SAM or SAM2 segmentation models locally or on Hugging Face Jobs, with dataset validation and results saved to the Hub.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts and reference files (for example `references/detailed-guide.md`, `references/finetune_sam2_trainer.md` and `references/hub_saving.md`).

It sits in AI & LLM Engineering, covering Computer vision and Model hubs and datasets. It works with Hugging Face. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Computer vision
  • Tasks that involve Model hubs and datasets

Example prompts

  • “/hugging-face-vision-trainer”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • hf.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hugging Face Vision Trainer loads about 1.1k tokens when it runs, and up to ~27k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 522 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~27k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its Apache-2.0 licence (© sickn33). 522 words, ~1,141 tokens.

Download SKILL.mdSave it as .claude/skills/hugging-face-vision-trainer/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
hugging-face-vision-trainer
description
Train object detection, image classification, and SAM or SAM2 segmentation models locally or on Hugging Face Jobs, with dataset validation and results saved to the Hub.
risk
critical
source
https://github.com/huggingface/skills/tree/main/skills/huggingface-vision-trainer
source_repo
huggingface/skills
source_type
official
date_added
2026-07-01
license
Apache-2.0
license_source
https://github.com/huggingface/skills/blob/main/LICENSE

Vision Model Training on Hugging Face Jobs

Train object detection, image classification, and SAM/SAM2 segmentation models on managed cloud GPUs. No local GPU setup required—results are automatically saved to the Hugging Face Hub.

Detailed Guide

Read the detailed guide before executing this skill. It retains the complete procedure and reference material. Treat its safety, prerequisites, and validation requirements as mandatory. For focused work, load the relevant sections; for end-to-end work, read the guide completely.

When to Use This Skill

Use this skill when users want to:

  • Fine-tune object detection models (D-FINE, RT-DETR v2, DETR, YOLOS) on cloud GPUs or local
  • Fine-tune image classification models (timm: MobileNetV3, MobileViT, ResNet, ViT/DINOv3, or any Transformers classifier) on cloud GPUs or local
  • Fine-tune SAM or SAM2 models for segmentation / image matting using bbox or point prompts
  • Train bounding-box detectors on custom datasets
  • Train image classifiers on custom datasets
  • Train segmentation models on custom mask datasets with prompts
  • Run vision training jobs on Hugging Face Jobs infrastructure
  • Ensure trained vision models are permanently saved to the Hub

Prerequisites Checklist

Before starting any training job, verify:

Account & Authentication
  • Hugging Face Account with Pro, Team, or Enterprise plan (Jobs require paid plan)
  • Authenticated login: Check with hf_whoami() (tool) or hf auth whoami (terminal)
  • Token has write permissions
  • MUST pass token in job secrets — see directive #3 below for syntax (MCP tool vs Python API)
Dataset Requirements — Object Detection
  • Dataset must exist on Hub
  • Annotations must use the objects column with bbox, category (and optionally area) sub-fields
  • Bboxes can be in xywh (COCO) or xyxy (Pascal VOC) format — auto-detected and converted
  • Categories can be integers or strings — strings are auto-remapped to integer IDs
  • image_id column is optional — generated automatically if missing
  • ALWAYS validate unknown datasets before GPU training (see Dataset Validation section)
Show full SKILL.md (228 more words)Show less
Dataset Requirements — Image Classification
  • Dataset must exist on Hub
  • Must have an image column (PIL images) and a label column (integer class IDs or strings)
  • The label column can be ClassLabel type (with names) or plain integers/strings — strings are auto-remapped
  • Common column names auto-detected: label, labels, class, fine_label
  • ALWAYS validate unknown datasets before GPU training (see Dataset Validation section)
Dataset Requirements — SAM/SAM2 Segmentation
  • Dataset must exist on Hub
  • Must have an image column (PIL images) and a mask column (binary ground-truth segmentation mask)
  • Must have a prompt — either:
    • A prompt column with JSON containing {"bbox": [x0,y0,x1,y1]} or {"point": [x,y]}
    • OR a dedicated bbox column with [x0,y0,x1,y1] values
    • OR a dedicated point column with [x,y] or [[x,y],...] values
  • Bboxes should be in xyxy format (absolute pixel coordinates)
  • Example dataset: merve/MicroMat-mini (image matting with bbox prompts)
  • ALWAYS validate unknown datasets before GPU training (see Dataset Validation section)
Critical Settings
  • Timeout must exceed expected training time — Default 30min is TOO SHORT. See directive #6 for recommended values.
  • Hub push must be enabled — push_to_hub=True, hub_model_id="username/model-name", token in secrets

Limitations

  • Use this skill only when the task clearly matches its upstream product or API scope.
  • Verify commands, API behavior, pricing, quotas, credentials, and deployment effects against current official documentation before making changes.
  • Do not treat generated examples as a substitute for environment-specific tests, security review, or user approval for destructive or costly actions.

© sickn33, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references) in skills/hugging-face-vision-trainer of sickn33/agentic-awesome-skills.

  • SKILL.md
  • references/detailed-guide.md
  • references/finetune_sam2_trainer.md
  • references/hub_saving.md
  • references/image_classification_training_notebook.md
  • references/object_detection_training_notebook.md
  • references/reliability_principles.md
  • references/timm_trainer.md
  • scripts/dataset_inspector.py
  • scripts/estimate_cost.py
  • scripts/image_classification_training.py
  • scripts/object_detection_training.py
  • scripts/sam_segmentation_training.py

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Hugging Face Vision Trainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hugging Face Vision Trainer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hugging Face Vision Trainer this skillsickn33/agentic-awesome-skills47k1 repos~1.1kAutomated safety check: PassApache-2.0
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0
Huggingface Vision Trainerwaybarrios/opencode-power-pack534—~2.7kAutomated safety check: PassApache-2.0
Hugging Face Transformers Usagedavila7/claude-code-templates33k11 repos~1.2kAutomated safety check: PassMIT
Hugging Face Vision Trainerhenryalouf/ruflow157—~7.4kAutomated safety check: PassMIT
Transformers.jshuggingface/skills11k1 repos~6.2kAutomated safety check: PassApache-2.0

Similar skills

  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface Vision Trainer

    waybarrios/opencode-power-pack

    Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

    534 GitHub stars~2.7k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Transformers Usage

    davila7/claude-code-templates

    Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

    33k GitHub starsUsed in 11 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Train or fine-tune vision models on Hugging Face Jobs for detection, classification, and SAM or SAM2 segmentation.

    157 GitHub stars~7.4k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Transformers.js

    huggingface/skills

    Official

    Runs pre-trained Hugging Face models in JavaScript or TypeScript with Transformers.js, in browsers or Node.js, Bun and Deno, for text, vision, audio and multimodal tasks.

    11k GitHub starsUsed in 1 repo~6.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.

    3.6k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Questions about Hugging Face Vision Trainer

What does Hugging Face Vision Trainer do?

Train object detection, image classification, and SAM or SAM2 segmentation models locally or on Hugging Face Jobs, with dataset validation and results saved to the Hub. Hugging Face Vision Trainer is an agent skill from sickn33/agentic-awesome-skills. Train object detection, image classification, and SAM or SAM2 segmentation models locally or on Hugging Face Jobs, with dataset validation and results saved to the Hub.

When should I use Hugging Face Vision Trainer?

Hugging Face Vision Trainer fits situations like: tasks that involve Computer vision; tasks that involve Model hubs and datasets.

How do I install Hugging Face Vision Trainer in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill hugging-face-vision-trainer -a claude-code`. Or copy the skill folder (skills/hugging-face-vision-trainer in sickn33/agentic-awesome-skills) into .claude/skills/hugging-face-vision-trainer in your project. Claude Code loads it when a task matches its description.

How do I install Hugging Face Vision Trainer in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill hugging-face-vision-trainer -a codex`. Or copy the skill folder (skills/hugging-face-vision-trainer in sickn33/agentic-awesome-skills) into .agents/skills/hugging-face-vision-trainer in your project. Codex loads it when a task matches its description.

Can I use Hugging Face Vision Trainer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill hugging-face-vision-trainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hugging-face-vision-trainer, .gemini/skills/hugging-face-vision-trainer, .github/skills/hugging-face-vision-trainer and .opencode/skills/hugging-face-vision-trainer in your project.

What does Hugging Face Vision Trainer need to run?

Going by SKILL.md and its folder, Hugging Face Vision Trainer needs Python for the scripts in its folder and the command-line tools its instructions call (hf). Our summary lists: Python 3.

Does Hugging Face Vision Trainer access the network?

SKILL.md names 1 domain. As links in the text: hf.co. This is read from the text; nothing was executed.

Is Hugging Face Vision Trainer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Hugging Face Vision Trainer use?

Hugging Face Vision Trainer is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hugging Face Vision Trainer use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 26k tokens, read only when the agent opens those files.

What are the alternatives to Hugging Face Vision Trainer?

Skills that share tags, products or a category with Hugging Face Vision Trainer: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Huggingface Vision Trainer (waybarrios/opencode-power-pack, 534 stars), Hugging Face Transformers Usage (davila7/claude-code-templates, 33k stars) and Hugging Face Vision Trainer (henryalouf/ruflow, 157 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hugging Face Vision Trainer?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.