Hugging Face Transformers Usage
davila7/claude-code-templates
Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
$ npx skills add huggingface/skills --skill huggingface-vision-trainer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install huggingface/skills huggingface-vision-trainer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/huggingface-vision-trainer .claude/skills/huggingface-vision-trainer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "huggingface-vision-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-vision-trainer into .claude/skills/huggingface-vision-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-vision-trainer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/huggingface/skills/tree/main/skills/huggingface-vision-trainerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add huggingface/skills --skill huggingface-vision-trainer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install huggingface/skills huggingface-vision-trainer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/huggingface-vision-trainer .agents/skills/huggingface-vision-trainer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "huggingface-vision-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-vision-trainer into .agents/skills/huggingface-vision-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-vision-trainer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add huggingface/skills --skill huggingface-vision-trainer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install huggingface/skills huggingface-vision-trainer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/huggingface-vision-trainer .cursor/skills/huggingface-vision-trainer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "huggingface-vision-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-vision-trainer into .cursor/skills/huggingface-vision-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-vision-trainer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/huggingface/skills.git --path skills/huggingface-vision-trainer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add huggingface/skills --skill huggingface-vision-trainer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install huggingface/skills huggingface-vision-trainer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/huggingface-vision-trainer .gemini/skills/huggingface-vision-trainer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "huggingface-vision-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-vision-trainer into .gemini/skills/huggingface-vision-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-vision-trainer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install huggingface/skills huggingface-vision-trainerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add huggingface/skills --skill huggingface-vision-trainer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/huggingface-vision-trainer .github/skills/huggingface-vision-trainer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "huggingface-vision-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-vision-trainer into .github/skills/huggingface-vision-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-vision-trainer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add huggingface/skills --skill huggingface-vision-trainer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install huggingface/skills huggingface-vision-trainer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/huggingface-vision-trainer .opencode/skills/huggingface-vision-trainer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "huggingface-vision-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-vision-trainer into .opencode/skills/huggingface-vision-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-vision-trainer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
huggingface-vision-trainerTrains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
The skill covers three kinds of vision training: object detection (D-FINE, RT-DETR v2, DETR, YOLOS), image classification (timm models such as MobileNetV3, MobileViT, ResNet and ViT/DINOv3, or any Transformers classifier) and SAM or SAM2 segmentation with bounding-box or point prompts. Jobs run on Hugging Face Jobs cloud GPUs, so no local GPU is needed, and trained models are saved to the Hub so they persist.
It includes COCO-format dataset preparation, Albumentations augmentation, mAP and mAR evaluation, accuracy metrics, DiceCE loss for segmentation, hardware selection, cost estimation and Trackio monitoring. Helper scripts, dataset_inspector.py, estimate_cost.py and one training script per task, run with uv run. Detection datasets need an objects column with bbox and category fields, and boxes in xywh or xyxy form are detected and converted.
Jobs need a Hugging Face account on a paid plan (Pro, Team or Enterprise), a login whose token has write permission, and that token passed in the job secrets. General Jobs questions go to the hugging-face-jobs skill, and text model training with TRL goes to hugging-face-model-trainer.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c3ff942. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvhfFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.coAlso links to:
hf.coFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Hugging Face Vision Trainer loads about 7.5k tokens when it runs, and up to ~27k if it reads all its reference files. Until then it costs about 201 tokens; SKILL.md has 2,221 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from huggingface/skills at commit c3ff942, republished under its Apache-2.0 licence (© huggingface). 2,221 words, ~7,499 tokens.
.claude/skills/huggingface-vision-trainer/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Train object detection, image classification, and SAM/SAM2 segmentation models on managed cloud GPUs. No local GPU setup required—results are automatically saved to the Hugging Face Hub.
Use this skill when users want to:
hugging-face-jobs — General HF Jobs infrastructure: token authentication, hardware flavors, timeout management, cost estimation, secrets, environment variables, scheduled jobs, and result persistence. Refer to the Jobs skill for any non-training-specific Jobs questions (e.g., "how do secrets work?", "what hardware is available?", "how do I pass tokens?").hugging-face-model-trainer — TRL-based language model training (SFT, DPO, GRPO). Use that skill for text/language model fine-tuning.Helper scripts use PEP 723 inline dependencies. Run them with uv run:
uv run scripts/dataset_inspector.py --dataset username/dataset-name --split train
uv run scripts/estimate_cost.py --helpBefore starting any training job, verify:
hf_whoami() (tool) or hf auth whoami (terminal)objects column with bbox, category (and optionally area) sub-fieldsimage_id column is optional — generated automatically if missingimage column (PIL images) and a label column (integer class IDs or strings)ClassLabel type (with names) or plain integers/strings — strings are auto-remappedlabel, labels, class, fine_labelimage column (PIL images) and a mask column (binary ground-truth segmentation mask)prompt column with JSON containing {"bbox": [x0,y0,x1,y1]} or {"point": [x,y]}bbox column with [x0,y0,x1,y1] valuespoint column with [x,y] or [[x,y],...] valuesmerve/MicroMat-mini (image matting with bbox prompts)push_to_hub=True, hub_model_id="username/model-name", token in secretsValidate dataset format BEFORE launching GPU training to prevent the #1 cause of training failures: format mismatches.
ALWAYS validate for unknown/custom datasets or any dataset you haven't trained with before. Skip for cppe-5 (the default in the training script).
Option 1: Via HF Jobs (recommended — avoids local SSL/dependency issues):
hf_jobs("uv", {
"script": "path/to/dataset_inspector.py",
"script_args": ["--dataset", "username/dataset-name", "--split", "train"]
})Option 2: Locally:
uv run scripts/dataset_inspector.py --dataset username/dataset-name --split trainOption 3: Via HfApi().run_uv_job() (if hf_jobs MCP unavailable):
from huggingface_hub import HfApi
api = HfApi()
api.run_uv_job(
script="scripts/dataset_inspector.py",
script_args=["--dataset", "username/dataset-name", "--split", "train"],
flavor="cpu-basic",
timeout=300,
)✓ READY — Dataset is compatible, use directly✗ NEEDS FORMATTING — Needs preprocessing (mapping code provided in output)The object detection training script (scripts/object_detection_training.py) automatically handles bbox format detection (xyxy→xywh conversion), bbox sanitization, image_id generation, string category→integer remapping, and dataset truncation. No manual preprocessing needed — just ensure the dataset has objects.bbox and objects.category columns.
Copy this checklist and track progress:
Training Progress:
- [ ] Step 1: Verify prerequisites (account, token, dataset)
- [ ] Step 2: Validate dataset format (run dataset_inspector.py)
- [ ] Step 3: Ask user about dataset size and validation split
- [ ] Step 4: Prepare training script (OD: scripts/object_detection_training.py, IC: scripts/image_classification_training.py, SAM: scripts/sam_segmentation_training.py)
- [ ] Step 5: Save script locally, submit job, and report detailsStep 1: Verify prerequisites
Follow the Prerequisites Checklist above.
Step 2: Validate dataset
Run the dataset inspector BEFORE spending GPU time. See "Dataset Validation" section above.
Step 3: Ask user preferences
ALWAYS use the AskUserQuestion tool with option-style format:
AskUserQuestion({
"questions": [
{
"question": "Do you want to run a quick test with a subset of the data first?",
"header": "Dataset Size",
"options": [
{"label": "Quick test run (10% of data)", "description": "Faster, cheaper (~30-60 min, ~$2-5) to validate setup"},
{"label": "Full dataset (Recommended)", "description": "Complete training for best model quality"}
],
"multiSelect": false
},
{
"question": "Do you want to create a validation split from the training data?",
"header": "Split data",
"options": [
{"label": "Yes (Recommended)", "description": "Automatically split 15% of training data for validation"},
{"label": "No", "description": "Use existing validation split from dataset"}
],
"multiSelect": false
},
{
"question": "Which GPU hardware do you want to use?",
"header": "Hardware Flavor",
"options": [
{"label": "t4-small ($0.40/hr)", "description": "1x T4, 16 GB VRAM — sufficient for all OD models under 100M params"},
{"label": "l4x1 ($0.80/hr)", "description": "1x L4, 24 GB VRAM — more headroom for large images or batch sizes"},
{"label": "a10g-large ($1.50/hr)", "description": "1x A10G, 24 GB VRAM — faster training, more CPU/RAM"},
{"label": "a100-large ($2.50/hr)", "description": "1x A100, 80 GB VRAM — fastest, for very large datasets or image sizes"}
],
"multiSelect": false
}
]
})Step 4: Prepare training script
For object detection, use scripts/object_detection_training.py as the production-ready template. For image classification, use scripts/image_classification_training.py. For SAM/SAM2 segmentation, use scripts/sam_segmentation_training.py. All scripts use HfArgumentParser — all configuration is passed via CLI arguments in script_args, NOT by editing Python variables. For timm model details, see references/timm_trainer.md. For SAM2 training details, see references/finetune_sam2_trainer.md.
Step 5: Save script, submit job, and report
submitted_jobs/ in the workspace root (create if needed) with a descriptive name like training_<dataset>_<YYYYMMDD_HHMMSS>.py. Tell the user the path.hf_jobs MCP tool (preferred) or HfApi().run_uv_job() — see directive #1 for both methods. Pass all config via script_args..id attribute), monitoring URL, Trackio dashboard (https://huggingface.co/spaces/{username}/trackio), expected time, and estimated cost.These rules prevent common failures. Follow them exactly.
hf_jobs MCP tool vs Python APIhf_jobs() is an MCP tool, NOT a Python function. Do NOT try to import it from huggingface_hub. Call it as a tool:
hf_jobs("uv", {"script": training_script_content, "flavor": "a10g-large", "timeout": "4h", "secrets": {"HF_TOKEN": "$HF_TOKEN"}})If hf_jobs MCP tool is unavailable, use the Python API directly:
from huggingface_hub import HfApi, get_token
api = HfApi()
job_info = api.run_uv_job(
script="path/to/training_script.py", # file PATH, NOT content
script_args=["--dataset_name", "cppe-5", ...],
flavor="a10g-large",
timeout=14400, # seconds (4 hours)
env={"PYTHONUNBUFFERED": "1"},
secrets={"HF_TOKEN": get_token()}, # MUST use get_token(), NOT "$HF_TOKEN"
)
print(f"Job ID: {job_info.id}")Critical differences between the two methods:
hf_jobs MCP tool | HfApi().run_uv_job() | |
|---|---|---|
script param | Python code string or URL (NOT local paths) | File path to .py file (NOT content) |
| Token in secrets | "$HF_TOKEN" (auto-replaced) | get_token() (actual token value) |
| Timeout format | String ("4h") | Seconds (14400) |
Rules for both methods:
image or command parameters (those belong to run_job(), not run_uv_job())Job config MUST include the token in secrets — syntax depends on submission method (see table above).
Training script requirement: The Transformers Trainer calls create_repo(token=self.args.hub_token) during __init__() when push_to_hub=True. The training script MUST inject HF_TOKEN into training_args.hub_token AFTER parsing args but BEFORE creating the Trainer. The template scripts/object_detection_training.py already includes this:
hf_token = os.environ.get("HF_TOKEN")
if training_args.push_to_hub and not training_args.hub_token:
if hf_token:
training_args.hub_token = hf_tokenIf you write a custom script, you MUST include this token injection before the Trainer(...) call.
login() in custom scripts unless replicating the full pattern from scripts/object_detection_training.pyhub_token=None) — unreliable in Jobshugging-face-jobs skill → Token Usage Guide for full detailsAccess the job identifier using .id (NOT .job_id or .name — these don't exist):
job_info = api.run_uv_job(...) # or hf_jobs("uv", {...})
job_id = job_info.id # Correct -- returns string like "687fb701029421ae5549d998"scripts/object_detection_training.py uses HfArgumentParser — all config is passed via script_args. Boolean arguments have two syntaxes:
bool fields (e.g., push_to_hub, do_train): Use as bare flags (--push_to_hub) or negate with --no_ prefix (--no_remove_unused_columns)Optional[bool] fields (e.g., greater_is_better): MUST pass explicit value (--greater_is_better True). Bare --greater_is_better causes error: expected one argumentRequired flags for object detection:
--no_remove_unused_columns # MUST: preserves image column for pixel_values
--no_eval_do_concat_batches # MUST: images have different numbers of target boxes
--push_to_hub # MUST: environment is ephemeral
--hub_model_id username/model-name
--metric_for_best_model eval_map
--greater_is_better True # MUST pass "True" explicitly (Optional[bool])
--do_train
--do_evalRequired flags for image classification:
--no_remove_unused_columns # MUST: preserves image column for pixel_values
--push_to_hub # MUST: environment is ephemeral
--hub_model_id username/model-name
--metric_for_best_model eval_accuracy
--greater_is_better True # MUST pass "True" explicitly (Optional[bool])
--do_train
--do_evalRequired flags for SAM/SAM2 segmentation:
--remove_unused_columns False # MUST: preserves input_boxes/input_points
--push_to_hub # MUST: environment is ephemeral
--hub_model_id username/model-name
--do_train
--prompt_type bbox # or "point"
--dataloader_pin_memory False # MUST: avoids pin_memory issues with custom collatorDefault 30 min is TOO SHORT for object detection. Set minimum 2-4 hours. Add 30% buffer for model loading, preprocessing, and Hub push.
| Scenario | Timeout |
|---|---|
| Quick test (100-200 images, 5-10 epochs) | 1h |
| Development (500-1K images, 15-20 epochs) | 2-3h |
| Production (1K-5K images, 30 epochs) | 4-6h |
| Large dataset (5K+ images) | 6-12h |
Trackio is always enabled in the object detection training script — it calls trackio.init() and trackio.finish() automatically. No need to pass --report_to trackio. The project name is taken from --output_dir and the run name from --run_name. For image classification, pass --report_to trackio in TrainingArguments.
Dashboard at: https://huggingface.co/spaces/{username}/trackio
| Model | Params | Use case |
|---|---|---|
ustc-community/dfine-small-coco | 10.4M | Best starting point — fast, cheap, SOTA quality |
PekingU/rtdetr_v2_r18vd | 20.2M | Lightweight real-time detector |
ustc-community/dfine-large-coco | 31.4M | Higher accuracy, still efficient |
PekingU/rtdetr_v2_r50vd | 43M | Strong real-time baseline |
ustc-community/dfine-xlarge-obj365 | 63.5M | Best accuracy (pretrained on Objects365) |
PekingU/rtdetr_v2_r101vd | 76M | Largest RT-DETR v2 variant |
Start with ustc-community/dfine-small-coco for fast iteration. Move to D-FINE Large or RT-DETR v2 R50 for better accuracy.
All timm/ models work out of the box via AutoModelForImageClassification (loaded as TimmWrapperForImageClassification). See references/timm_trainer.md for details.
| Model | Params | Use case |
|---|---|---|
timm/mobilenetv3_small_100.lamb_in1k | 2.5M | Ultra-lightweight — mobile/edge, fastest training |
timm/mobilevit_s.cvnets_in1k | 5.6M | Mobile transformer — good accuracy/speed trade-off |
timm/resnet50.a1_in1k | 25.6M | Strong CNN baseline — reliable, well-studied |
timm/vit_base_patch16_dinov3.lvd1689m | 86.6M | Best accuracy — DINOv3 self-supervised ViT |
Start with timm/mobilenetv3_small_100.lamb_in1k for fast iteration. Move to timm/resnet50.a1_in1k or timm/vit_base_patch16_dinov3.lvd1689m for better accuracy.
| Model | Params | Use case |
|---|---|---|
facebook/sam2.1-hiera-tiny | 38.9M | Fastest SAM2 — good for quick experiments |
facebook/sam2.1-hiera-small | 46.0M | Best starting point — good quality/speed balance |
facebook/sam2.1-hiera-base-plus | 80.8M | Higher capacity for complex segmentation |
facebook/sam2.1-hiera-large | 224.4M | Best SAM2 accuracy — requires more VRAM |
facebook/sam-vit-base | 93.7M | Original SAM — ViT-B backbone |
facebook/sam-vit-large | 312.3M | Original SAM — ViT-L backbone |
facebook/sam-vit-huge | 641.1M | Original SAM — ViT-H, best SAM v1 accuracy |
Start with facebook/sam2.1-hiera-small for fast iteration. SAM2 models are generally more efficient than SAM v1 at similar quality. Only the mask decoder is trained by default (vision and prompt encoders are frozen).
All recommended OD and IC models are under 100M params — t4-small (16 GB VRAM, $0.40/hr) is sufficient for all of them. Image classification models are generally smaller and faster than object detection models — t4-small handles even ViT-Base comfortably. For SAM2 models up to hiera-base-plus, t4-small is sufficient since only the mask decoder is trained. For sam2.1-hiera-large or SAM v1 models, use l4x1 or a10g-large. Only upgrade if you hit OOM from large batch sizes — reduce batch size first before switching hardware. Common upgrade path: t4-small → l4x1 ($0.80/hr, 24 GB) → a10g-large ($1.50/hr, 24 GB).
For full hardware flavor list: refer to the hugging-face-jobs skill. For cost estimation: run scripts/estimate_cost.py.
The script_args below are the same for both submission methods. See directive #1 for the critical differences between them.
OD_SCRIPT_ARGS = [
"--model_name_or_path", "ustc-community/dfine-small-coco",
"--dataset_name", "cppe-5",
"--image_square_size", "640",
"--output_dir", "dfine_finetuned",
"--num_train_epochs", "30",
"--per_device_train_batch_size", "8",
"--learning_rate", "5e-5",
"--eval_strategy", "epoch",
"--save_strategy", "epoch",
"--save_total_limit", "2",
"--load_best_model_at_end",
"--metric_for_best_model", "eval_map",
"--greater_is_better", "True",
"--no_remove_unused_columns",
"--no_eval_do_concat_batches",
"--push_to_hub",
"--hub_model_id", "username/model-name",
"--do_train",
"--do_eval",
]from huggingface_hub import HfApi, get_token
api = HfApi()
job_info = api.run_uv_job(
script="scripts/object_detection_training.py",
script_args=OD_SCRIPT_ARGS,
flavor="t4-small",
timeout=14400,
env={"PYTHONUNBUFFERED": "1"},
secrets={"HF_TOKEN": get_token()},
)
print(f"Job ID: {job_info.id}")script_args--model_name_or_path — recommended: "ustc-community/dfine-small-coco" (see model table above)--dataset_name — the Hub dataset ID--image_square_size — 480 (fast iteration) or 800 (better accuracy)--hub_model_id — "username/model-name" for Hub persistence--num_train_epochs — 30 typical for convergence--train_val_split — fraction to split for validation (default 0.15), set if dataset lacks a validation split--max_train_samples — truncate training set (useful for quick test runs, e.g. "785" for ~10% of a 7.8K dataset)--max_eval_samples — truncate evaluation setIC_SCRIPT_ARGS = [
"--model_name_or_path", "timm/mobilenetv3_small_100.lamb_in1k",
"--dataset_name", "ethz/food101",
"--output_dir", "food101_classifier",
"--num_train_epochs", "5",
"--per_device_train_batch_size", "32",
"--per_device_eval_batch_size", "32",
"--learning_rate", "5e-5",
"--eval_strategy", "epoch",
"--save_strategy", "epoch",
"--save_total_limit", "2",
"--load_best_model_at_end",
"--metric_for_best_model", "eval_accuracy",
"--greater_is_better", "True",
"--no_remove_unused_columns",
"--push_to_hub",
"--hub_model_id", "username/food101-classifier",
"--do_train",
"--do_eval",
]from huggingface_hub import HfApi, get_token
api = HfApi()
job_info = api.run_uv_job(
script="scripts/image_classification_training.py",
script_args=IC_SCRIPT_ARGS,
flavor="t4-small",
timeout=7200,
env={"PYTHONUNBUFFERED": "1"},
secrets={"HF_TOKEN": get_token()},
)
print(f"Job ID: {job_info.id}")script_args--model_name_or_path — any timm/ model or Transformers classification model (see model table above)--dataset_name — the Hub dataset ID--image_column_name — column containing PIL images (default: "image")--label_column_name — column containing class labels (default: "label")--hub_model_id — "username/model-name" for Hub persistence--num_train_epochs — 3-5 typical for classification (fewer than OD)--per_device_train_batch_size — 16-64 (classification models use less memory than OD)--train_val_split — fraction to split for validation (default 0.15), set if dataset lacks a validation split--max_train_samples / --max_eval_samples — truncate for quick testsSAM_SCRIPT_ARGS = [
"--model_name_or_path", "facebook/sam2.1-hiera-small",
"--dataset_name", "merve/MicroMat-mini",
"--prompt_type", "bbox",
"--prompt_column_name", "prompt",
"--output_dir", "sam2-finetuned",
"--num_train_epochs", "30",
"--per_device_train_batch_size", "4",
"--learning_rate", "1e-5",
"--logging_steps", "1",
"--save_strategy", "epoch",
"--save_total_limit", "2",
"--remove_unused_columns", "False",
"--dataloader_pin_memory", "False",
"--push_to_hub",
"--hub_model_id", "username/sam2-finetuned",
"--do_train",
"--report_to", "trackio",
]from huggingface_hub import HfApi, get_token
api = HfApi()
job_info = api.run_uv_job(
script="scripts/sam_segmentation_training.py",
script_args=SAM_SCRIPT_ARGS,
flavor="t4-small",
timeout=7200,
env={"PYTHONUNBUFFERED": "1"},
secrets={"HF_TOKEN": get_token()},
)
print(f"Job ID: {job_info.id}")script_args--model_name_or_path — SAM or SAM2 model (see model table above); auto-detects SAM vs SAM2--dataset_name — the Hub dataset ID (e.g., "merve/MicroMat-mini")--prompt_type — "bbox" or "point" — type of prompt in the dataset--prompt_column_name — column with JSON-encoded prompts (default: "prompt")--bbox_column_name — dedicated bbox column (alternative to JSON prompt column)--point_column_name — dedicated point column (alternative to JSON prompt column)--mask_column_name — column with ground-truth masks (default: "mask")--hub_model_id — "username/model-name" for Hub persistence--num_train_epochs — 20-30 typical for SAM fine-tuning--per_device_train_batch_size — 2-4 (SAM models use significant memory)--freeze_vision_encoder / --freeze_prompt_encoder — freeze encoder weights (default: both frozen, only mask decoder trains)--train_val_split — fraction to split for validation (default 0.1)MCP tool (if available):
hf_jobs("ps") # List all jobs
hf_jobs("logs", {"job_id": "your-job-id"}) # View logs
hf_jobs("inspect", {"job_id": "your-job-id"}) # Job detailsPython API fallback:
from huggingface_hub import HfApi
api = HfApi()
api.list_jobs() # List all jobs
api.get_job_logs(job_id="your-job-id") # View logs
api.get_job(job_id="your-job-id") # Job detailsReduce per_device_train_batch_size (try 4, then 2), reduce IMAGE_SIZE, or upgrade hardware.
Run scripts/dataset_inspector.py first. The training script auto-detects xyxy vs xywh, converts string categories to integer IDs, and adds image_id if missing. Ensure objects.bbox contains 4-value coordinate lists in absolute pixels and objects.category contains either integer IDs or string labels.
Verify: (1) job secrets include token (see directive #2), (2) script sets training_args.hub_token BEFORE creating the Trainer, (3) push_to_hub=True is set, (4) correct hub_model_id, (5) token has write permissions.
Increase timeout (see directive #5 table), reduce epochs/dataset, or use checkpoint strategy with hub_strategy="every_save".
The object detection training script handles this gracefully — it falls back to the validation split. Ensure you're using the latest scripts/object_detection_training.py.
torchmetrics.MeanAveragePrecision returns scalar (0-d) tensors for per-class metrics when there's only one class. The template scripts/object_detection_training.py handles this by calling .unsqueeze(0) on these tensors. Ensure you're using the latest template.
Increase epochs (30-50), ensure 500+ images, check per-class mAP for imbalanced classes, try different learning rates (1e-5 to 1e-4), increase image size.
For comprehensive troubleshooting: see references/reliability_principles.md
© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (scripts, references) in skills/huggingface-vision-trainer of huggingface/skills.
Open the folder on GitHubat commit c3ff942
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.
Hugging Face Vision Trainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Hugging Face Vision Trainer this skillhuggingface/skills | 11k | 1 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Transformers Usagedavila7/claude-code-templates | 32k | 11 repos | ~1.2k | Automated safety check: Pass | MIT | |
| bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Runpodericrisco/rsc-harness | 174 | — | ~2.8k | Automated safety check: Pass | MIT | |
| Huggingface Vision Trainerwaybarrios/opencode-power-pack | 533 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Tao Finetune Huggingface ModelNVIDIA/skills | 3.5k | — | ~4.9k | Automated safety check: Notes | Apache-2.0 |
davila7/claude-code-templates
Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.
Orchestra-Research/AI-Research-SKILLs
Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.
ericrisco/rsc-harness
A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…
waybarrios/opencode-power-pack
Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.
NVIDIA/skills
Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
huggingface/skills
Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
huggingface/skills
Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
huggingface/skills
Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.
Works with
Categories
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub. The skill covers three kinds of vision training: object detection (D-FINE, RT-DETR v2, DETR, YOLOS), image classification (timm models such as MobileNetV3, MobileViT, ResNet and ViT/DINOv3, or any Transformers classifier) and SAM or SAM2 segmentation with bounding-box or point prompts. Jobs run on Hugging Face Jobs cloud GPUs, so no local GPU is needed, and trained models are saved to the Hub so they persist.
Hugging Face Vision Trainer fits situations like: fine-tuning an object detector such as D-FINE or RT-DETR v2 on a custom dataset; training an image classifier with a timm model or any Transformers classifier; fine-tuning SAM or SAM2 for segmentation with bbox or point prompts; estimating GPU cost before launching a training job.
Run `npx skills add huggingface/skills --skill huggingface-vision-trainer -a claude-code`. Or copy the skill folder (skills/huggingface-vision-trainer in huggingface/skills) into .claude/skills/huggingface-vision-trainer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add huggingface/skills --skill huggingface-vision-trainer -a codex`. Or copy the skill folder (skills/huggingface-vision-trainer in huggingface/skills) into .agents/skills/huggingface-vision-trainer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill huggingface-vision-trainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huggingface-vision-trainer, .gemini/skills/huggingface-vision-trainer, .github/skills/huggingface-vision-trainer and .opencode/skills/huggingface-vision-trainer in your project.
Going by SKILL.md and its folder, Hugging Face Vision Trainer needs Python for the scripts in its folder, the command-line tools its instructions call (uv and hf) and credentials named HF_TOKEN. Our summary lists: A Hugging Face account on a paid plan (Pro, Team or Enterprise) for Jobs; A Hugging Face token with write permission; uv to run the helper scripts.
SKILL.md names 2 domains. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. As links in the text: hf.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Hugging Face Vision Trainer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Hugging Face Vision Trainer: Hugging Face Transformers Usage (davila7/claude-code-templates, 32k stars), bitsandbytes Model Quantization (Orchestra-Research/AI-Research-SKILLs, 13k stars), Runpod (ericrisco/rsc-harness, 174 stars) and Huggingface Vision Trainer (waybarrios/opencode-power-pack, 533 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,151 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 8, 2026.
Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.