Official agent skill

Tao Finetune Huggingface Model

by NVIDIA in NVIDIA/skills

Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Tao Finetune Huggingface Model

skills CLI
$ npx skills add NVIDIA/skills --skill tao-finetune-huggingface-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-finetune-huggingface-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-finetune-huggingface-model .claude/skills/tao-finetune-huggingface-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-finetune-huggingface-model
GitHub stars
3.6k
Token cost
~4.9k tokens
SKILL.md length
2,072 words
Files
29 (incl. references)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.

  • Works in 6 steps: Inspect & qualify → Hardware audit & NGC image → Research the recipe → …
  • The user wants to fine-tune a HuggingFace model (full
  • SKILL.md covers Dedicated-model routing gate, Inputs, Execution platform and References — fallback safety net, plus 4 more sections
  • Calls docker, python and pytest; needs HF_TOKEN and WANDB_API_KEY

What it does

Tao Finetune Huggingface Model is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches. Use when the user wants to fine-tune a HuggingFace model (full or LoRA), train a vision / VLM / LLM model end-to-end, generate a reproducible HF training pipeline, smoke-test a HuggingFace model locally before scale-up, push a fine-tuned model to the HF Hub with a model card, or emit a self-contained rerun skill for an existing HuggingFace finetune. Supports image…

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 31 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit, NVIDIA GPU (driver ≥ 545, ≥ 24 GB VRAM for ≤3B models), ~40 GB free disk. Optional credentials (read from the…

It sits in AI & LLM Engineering, covering Fine-tuning, Model hubs and datasets and Computer vision. It works with Hugging Face, NVIDIA AI Platform and PyTorch. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user wants to fine-tune a HuggingFace model (full
  • Train a vision / VLM / LLM model end-to-end
  • Generate a reproducible HF training pipeline
  • Smoke-test a HuggingFace model locally before scale-up

Example prompts

  • “/tao-finetune-huggingface-model”

Requirements

  • Python 3
  • Docker
  • A credential in WANDB_API_KEY
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit, NVIDIA GPU (driver ≥ 545, ≥ 24 GB VRAM for ≤3B models), ~40 GB free disk. Optional credentials (read from the session environment) — HF_TOKEN is read only when the model/dataset is gated or `push_to_hub` is on; WANDB_API_KEY and WANDB_PROJECT only when WandB logging is enabled.
  • Pre-approved tools (allowed-tools): Read, Bash, Write

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Inspect & qualify
  2. Hardware audit & NGC image
  3. Research the recipe
  4. Generate project & smoke-test
  5. Train, evaluate, infer
  6. Push & emit rerun skill

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • python
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • apache.org
    • docs.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • WANDB_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit, NVIDIA GPU (driver ≥ 545, ≥ 24 GB VRAM for ≤3B models), ~40 GB free disk. Optional credentials (read from the session environment) — HF_TOKEN is read only when the model/dataset is gated or `push_to_hub` is on; WANDB_API_KEY and WANDB_PROJECT only when WandB logging is enabled.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Finetune Huggingface Model loads about 4.9k tokens when it runs, and up to ~75k if it reads all its reference files. Until then it costs about 251 tokens; SKILL.md has 2,072 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~251
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~75k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 2,072 words, ~4,882 tokens.

Download SKILL.mdSave it as .claude/skills/tao-finetune-huggingface-model/SKILL.md (or your agent's skills folder). This skill also uses 28 other files; get the full folder from GitHub.
name
tao-finetune-huggingface-model
description
Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches. Use when the user wants to fine-tune a HuggingFace model (full or LoRA), train a vision / VLM / LLM model end-to-end, generate a reproducible HF training pipeline, smoke-test a HuggingFace model locally before scale-up, push a fine-tuned model to the HF Hub with a model card, or emit a self-contained rerun skill for an existing HuggingFace finetune. Supports image classification, object detection, semantic / instance / panoptic segmentation, depth estimation, image-text-to-text VLM (SFT / LoRA), and LLM SFT / DPO / GRPO. Six-step workflow: inspect and qualify, hardware and NGC image, research, generate and smoke, train + eval + infer, push and emit rerun skill. Do not use for any Hugging Face model ID claimed by a dedicated `skills/models/*` skill; the model skill and its declared execution environment take precedence.
allowed-tools
Read, Bash, Write
compatibility
Requires docker + nvidia-container-toolkit, NVIDIA GPU (driver ≥ 545, ≥ 24 GB VRAM for ≤3B models), ~40 GB free disk. Optional credentials (read from the session environment) — HF_TOKEN is read only when the model/dataset is gated or `push_to_hub` is on; WANDB_API_KEY and WANDB_PROJECT only when WandB logging is enabled.
license
Apache-2.0
tags
finetuning, huggingface, nvidia-tao, computer-vision, training
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
<!-- Copyright (c) 2026, NVIDIA CORPORATION. All rights reserved. Licensed under the Apache License, Version 2.0; see http://www.apache.org/licenses/LICENSE-2.0 -->

tao-finetune-huggingface-model

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Local NVIDIA GPU fine-tuning for HuggingFace models, grounded in live-fetched documentation with curated references as a fallback safety net. One NGC container, a few focused scripts, one push to HF Hub. Follow the rules in this file; don't improvise.

Dedicated-model routing gate

Before Step 1 or any probe, image selection, package install, venv creation, or training-code generation, resolve model_id against the packaged model-owner registry. Use the absolute skill-bank root from which this file was loaded:

bash
python <bank-root>/scripts/resolve_tao_model.py \
  --skill-bank <bank-root> \
  --model "$MODEL_ID" \
  --format json

The resolver matches model metadata, including huggingface_model_ids, network_arch, skill names, and legacy aliases. Routing is internal: a model ID and task are enough. Never require prompt boilerplate about skills, containers, or checkpoint formats.

  • Exit 0: stop this workflow and follow the owning model skill's environment, action metadata, preflight, and checkpoint preparation.
  • Exit 3: no packaged model skill owns the ID. This is the only result that permits Step 1 of the generic workflow.
  • Any other nonzero exit: ownership discovery is broken or ambiguous. Stop and resolve that error; do not silently fall back to generic Hugging Face training.

Hugging Face hosting never overrides ownership. Do not use this workflow to bypass a matched skill or ask the user to prescribe its internal preparation. For example, nvidia/Cosmos3-Nano routes to tao-finetune-cosmos-reason.

Do not create a host training venv in this workflow. Its default execution path is the NGC container documented below; any venv-based training path requires an explicit user request.

Order of authority (highest first):

  1. User input — explicit model_id, dataset_id, training_method, config.yaml overrides.
  2. Live research — model card, HF repo example, author finetune script, HF task docs, paper; always fetched (Step 3 + references/research-priorities.md).
  3. Curated references (references/*.md) — fallback when live research is silent/ambiguous.
  4. Your training-data memory — last resort; suspect, cross-check against (2)/(3).

Conflict resolution between (2) and (3) and the source-line discrepancy note are in references/research-priorities.md.


Inputs

Required:

  • model_id — HuggingFace model ID, e.g. google/vit-base-patch16-224

Conditional credentials (read from the session environment — exported before launching or sourced from a user-approved env file):

  • HF_TOKEN — only when the model/dataset is gated (read) or push_to_hub is on (write); public + public + push_to_hub: false needs none. Value never read — presence-only via [ -n "$HF_TOKEN" ].
  • WANDB_API_KEY, WANDB_PROJECT — only when WandB is enabled; WANDB_MODE=disabled opts out.

Dataset — exactly one:

  • dataset_id — HuggingFace dataset ID (source: hf)
  • local_dataset_path — local folder or file (source: local); optional local_dataset_format ∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet, csv} (default: auto-detect).
  • (omit) — agent recommends popular datasets (source: recommend)

Optional (have defaults):

  • task_type — auto-detected from config + model card
  • n_train=10000, n_eval=1000, n_epochs=3, lora_r=16
  • output_dir=./output/<model_short_name>
  • hf_model_repo — push target; if unset and HF_TOKEN has write access, auto-derived as <whoami>/<model_short_name>-finetuned.
  • push_to_hub=True — set to False to skip
  • skip_baseline=False — skip zero-shot baseline eval

Optional deliverables (off by default):

yaml
emit_progress_log: false   # output_dir/PROGRESS.md (per-step journal)
emit_report:       false   # reports/report.{pdf,html} with curves & samples
emit_unit_tests:   false   # tests/ with fake-data heterogeneous-batch tests

All values live in output_dir/config.yaml. Never hardcode in Python.


Execution platform

This skill orchestrates what to run; the platform skills own how to run it on a GPU host — read them first.

ConcernAuthoritative skill
GPU host runtime (driver 580, CUDA Toolkit 13.0, NVIDIA Container Toolkit 1.19.0)tao-skill-bank:tao-setup-nvidia-gpu-host
docker run flags, NGC auth, mounts, env passthrough, local/remote Docker job preflight (daemon, GPU smoke)tao-skill-bank:tao-run-on-docker

Default platform: local-docker — build a one-off image (run-<short>:latest) and run it on the local Docker daemon. Ask only when the user explicitly needs a different backend (Brev remote GPU, SLURM/Kubernetes); then run that platform's Preflight first and route the Steps 4–5 docker run commands through it. The GPU-runtime and presence-only credential preflights (values never read), the canonical docker run flag set, discovery of the execution platforms from the installed platform skills (tao-run-on-docker / -slurm / -kubernetes / -brev, plus any external one; on a runtime that surfaces only the core router skills, read skills/platform/tao-run-on-*/SKILL.md frontmatter), and the workflow-specific flags (--entrypoint /bin/bash -lc, PYTORCH_CUDA_ALLOC_CONF, --name hft_train) are in references/workflow-intake-preflight.md.


References — fallback safety net

Consulted only when live research is silent, ambiguous, or unavailable; live docs always win for the specific model and current API. Each step links the references it needs; full catalog in references/detailed-workflow.md.

Always-on: core-rules.md, error-playbook.md, compat-workarounds.md, model-discovery.md, dataset-recommendations.md, dataset-sources.md, dataset-patterns.md, hardware-container.md, research-priorities.md, cv-scripts.md, vlm-scripts.md, docker-runs.md, hub-push.md, pipeline-skill-template.md, deliverables.md. Opt-in (when their flag/need applies): progress-tracking.md, testing.md, reporting.md, workflow-intake-preflight.md, workflow-generate-train.md, workflow-push-rerun.md.

Rule: before falling back, log the live source you tried and why it was insufficient (config.yaml notes:, and PROGRESS.md if enabled). [FETCH LIVE] markers in cv-scripts.md / vlm-scripts.md are a research checklist, not code to inline — refetch the listed URL if a block has no Step 3 finding.


Core rules

Non-negotiable behaviors. Short version (full enumeration — hallucinated-imports list, never-without-approval list, full error-recovery and hardware-sizing tables — in references/core-rules.md, consult before any training-time decision):

  • Your HF-library knowledge is outdated. Fetch live docs (model card, HF repo example, task doc) before writing any ML code — don't generate trainer args / collator / transforms from memory (Step 3).
  • Smoke-test on real data with --max_steps 1 before any full run; no batch launches without a verified smoke.
  • Never silently substitute model_id, dataset_id, or training_method — if what the user asked for doesn't load, stop and ask.
  • Error recovery is minimal-change. OOM → halve batch, double grad_accum, enable gradient checkpointing (no LoRA switch without approval); NaN → reduce LR 10×; flat loss → inspect collator; same error 3× → stop and ask. Don't loop.
  • Dataset columns verified BEFORE the collator — rename in prepare_data.py; restructuring needed → stop and ask.
  • Hardware-sizing thumb (bf16): ≤3B → 24 GB, 7–13B → 80 GB, 30B+ → multi-GPU or LoRA on 1× 80 GB, 70B+ → 8× 80 GB or LoRA. Full finetune won't fit and no LoRA requested → ask before switching.

Workflow — 6 steps

Single pass, sequential; each step has a clear gate before the next begins.

Step 1 — Inspect & qualify

Goal: decide whether to proceed. Probe model + dataset, apply accept/reject, register applicable compat fixes, write the initial config.yaml.

Prerequisites: MODEL_ID, optional DATASET_ID / local_dataset_path, optional HF_TOKEN, OUTPUT_DIR (default ./output/<model_short_name>). Probes run in a CPU-only python:3.12-slim Docker container (bind-mounted .probe/ scratch) so the host needs no virtualenv — Docker must exist first. Docker-presence guard, container env, full probe invocation, and the model/dataset probe scripts are in references/workflow-intake-preflight.md, references/model-discovery.md, and references/dataset-sources.md.

Probe requirements:

  • Model: load AutoConfig, read model-card tags, detect task from architectures + tags + card examples (fallback logging in model-discovery.md).
  • Dataset: for recommended datasets, first present 3-5 choices from dataset-recommendations.md; for local data, bind-mount read-only and use dataset-sources.md format detection.
  • Reject early if the model config fails, the task is out of scope, no recipe source exists, or the dataset cannot load / match the task schema.
  • Evaluate compat-workarounds.md against the model/task; defer hardware-dependent rules to Step 2.

Write the initial config.yaml (model_id, task, dataset_id or local_dataset_path, research_sources: [] filled in Step 3, applicable_workarounds: from Step 1, notes: [] for reference fallbacks, push_to_hub: true default — annotated template in references/workflow-intake-preflight.md). Optionally rm -rf "$OUTPUT_DIR/.probe" once the gate is met.

Gate: config.yaml exists with model, dataset, task, applicable_workarounds; do not proceed if any field is missing.


Show full SKILL.md (943 more words)Show less
Step 2 — Hardware audit & NGC image

Goal: verify Docker + GPU + disk, pick the NGC PyTorch image live, finalize hardware-dependent compat rules.

2a. Audit (hard gate) — three checks (commands in references/workflow-intake-preflight.md):

  1. GPU host runtime — tao-setup-nvidia-gpu-host's setup-nvidia-gpu-host.sh --backend docker --check-only; on fail, ask approval then re-run with --install --yes.
  2. Free-disk soft-warn — override via MIN_DISK_GB (default 100 GB); recommend ≥ 100 GB for NGC base (~20 GB) + HF cache + checkpoints + data.
  3. Conditional credential presence (values never read) — HF_TOKEN only when gated or push_to_hub is on; WANDB_* only when WandB is on.

Do not proceed to Step 4 on a hard-fail — Step 4's docker build pulls a 20+ GB NGC base, and a missing nvidia-container-toolkit only surfaces later as could not select device driver "" with capabilities: [[gpu]]. Record gpu_count, gpu_name, driver_major, vram_gb_per_gpu in config.yaml.

2b. Pick NGC image (live): from the NVIDIA Deep Learning Frameworks support matrix (https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html), PyTorch NGC container section, pick the highest-versioned image where Min driver ≤ detected driver_major and container CUDA ≤ host CUDA Toolkit (match closely so cuDNN / TensorRT line up). Do not reject an image for an aN/bN/rcN PyTorch tag — NGC validates the full image; pick the newest CUDA-aligned one and let compat-workarounds.md handle per-version issues. If the matrix is unreachable, use the fallbacks in references/hardware-container.md; default nvcr.io/nvidia/pytorch:24.09-py3 <!-- unpinned: documented fallback --> (driver ≥ 545; SDPA+GQA bug — if num_key_value_heads < num_attention_heads, set attn_implementation: "eager"). Record ngc_image in config.yaml.

2c. Re-evaluate hardware-dependent compat rules: re-run the compat-workarounds.md walk for entries whose detect needs hw; update applicable_workarounds: in place.

2d. Model-fit check: estimate param_bytes ≈ 2×param_count (bf16); if

60% of vram_gb_per_gpu × 1e9, recommend LoRA in the user-facing summary.

Gate: config.yaml has ngc_image, gpu_count, gpu_name, driver_major, vram_gb_per_gpu; hardware-dependent compat fixes recorded.


Step 3 — Research the recipe

Goal: fetch the live recipe — training-data knowledge of transformers/trl/peft is suspect, so Step 3 is non-negotiable. Walk references/research-priorities.md in priority order (Priority 1 → 6); stop once you have, for the detected task:

  • AutoModel / processor class
  • Train + eval transforms
  • Collator
  • compute_metrics
  • Hyperparameter hints (LR, batch size, epochs, scheduler)

Record findings in meta/recipe.md, append source URLs to config.yaml: research_sources:. A slot with no live finding falls back to the matching scaffold (cv-scripts.md / vlm-scripts.md), logged as "fallback to scaffold — no live source for <slot>" under notes:. Conflict-resolution rules are in references/research-priorities.md.

Gate: every required slot filled, with a source URL or scaffold-fallback note.


Step 4 — Generate project & smoke-test

Goal: write all scripts, build the image, prepare data, run a 1-step smoke on real data (one docker build, two docker runs).

4a. Generate project files in output_dir/: config.yaml, Dockerfile, requirements.txt, prepare_data.py, train.py, run_eval.py, infer.py, optional merge_lora.py, optional tests/, .gitignore. Live Step 3 research is authority; cv-scripts.md / vlm-scripts.md give scaffold shape only. Apply every applicable_workarounds entry as a Dockerfile block, requirement pin, config override, or runtime env var. Hard rules: run_eval.py keeps that exact filename (avoids colliding with the HF evaluate package); every generated .py starts with the NVIDIA Apache-2.0 copyright header and any emitter fails when it is missing; emit_unit_tests: true generates and runs tests per references/testing.md. Script bodies, Dockerfile shape, and the emitter contract are in references/workflow-generate-train.md.

4b. Build, prepare, smoke — docker build -t run-<short>:latest ., then prepare_data and the --smoke --max_steps 1 run (references/docker-runs.md §1-3). Smoke pass criteria (in logs/smoke.log):

  • No exception
  • Loss is finite (not 0.0, not NaN)
  • grad_norm > 0 at step 1

If emit_unit_tests: true, also run pytest tests/ in the container. Any failure → STOP.

4c. Preflight summary — before full training, print and verify: reference URL, dataset columns, Hub target, monitoring target, NGC image, hardware, smoke loss/grad norm.

Gate: project files written, image built, smoke PASSED, preflight has no blank fields.


Step 5 — Train, evaluate, infer

Goal: baseline eval, full training, post-train eval, optional LoRA merge, 5 inference samples (all commands: references/docker-runs.md §4-8).

Sub-stepdocker-runs.mdSkip if
5a. Baseline eval (zero-shot)§4skip_baseline: true
5b. Full training (detached)§5—
5c. LoRA merge§6not VLM+LoRA
5d. Post-train eval§7—
5e. Inference (5 samples)§8—

Multi-GPU: prepend torchrun --nproc_per_node=$gpu_count to python train.py.

While training streams, watch docker logs -f hft_train: loss should drop within 10-20 steps; flat loss (collator/label-masking bug), NaN (LR too high), and OOM all stop the run — recovery in references/core-rules.md. If emit_report: true, run report.py after Step 5e per references/reporting.md.

Gate: all of:

  • checkpoints/final/ (or checkpoints/merged/ for LoRA) exists
  • reports/eval_results.json has a numeric primary metric
  • reports/baseline_results.json exists (unless skipped)
  • reports/inference_samples/ has 5 samples
  • wandb URL shows descending loss

Step 6 — Push & emit rerun skill

Goal: publish the run and make it reproducible without re-research.

Push per references/hub-push.md (weights, model card, eval/baseline JSONs, config.yaml, Dockerfile, requirements.txt, inference samples, reports when emitted) unless push_to_hub: false is explicit. Emit <output_dir>/skills/run-<short>/SKILL.md from references/pipeline-skill-template.md — substitute every placeholder, include full YAML metadata + the NVIDIA copyright HTML comment, and make any emitter fail if those are missing.

Gate (Done criteria): all of:

  • Step 5 gate met
  • HF Hub repo exists at the resolved URL with weights + card + results/ (unless push_to_hub: false)
  • <output_dir>/skills/run-<short>/SKILL.md exists, no <placeholder> left, with metadata + copyright HTML comment per pipeline-skill-template.md

Final message: wandb URL, HF Hub URL, baseline -> fine-tuned primary metric, reports/inference_samples/, and the rerun skill path.


Error playbook

On a known runtime error, consult the symptom → minimal-fix table in references/error-playbook.md (NGC entrypoint, PyTorch/Transformers regressions, numpy ABI, Albumentations bbox, PEFT/checkpointing, LoRA target breadth, CV augmentation gaps, OOM at step 0) before redesigning anything. When a row there fires twice across runs, lift it into compat-workarounds.md with a detect rule — auto-applied in Step 1 before the error can fire.


Communication style

  • Terse. No filler, no restating the request; one-word answers when appropriate.
  • Always include direct Hub and wandb URLs when referencing artifacts.
  • On error: state what went wrong, why, what you changed — no menus.
  • Never present "Option A/B/C" for a request with a clear answer. Act.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 28 other files (references) in skills/tao-finetune-huggingface-model of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • eval.config
  • evals/evals.json
  • references/compat-workarounds.md
  • references/core-rules.md
  • references/cv-scripts.md
  • references/dataset-patterns.md
  • references/dataset-recommendations.md
  • references/dataset-sources.md
  • references/deliverables.md
  • references/detailed-workflow.md
  • references/docker-runs.md
  • references/error-playbook.md
  • references/hardware-container.md
  • references/hub-push.md
  • references/model-discovery.md
  • … and 11 more

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Tao Finetune Huggingface Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Finetune Huggingface Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Finetune Huggingface Model this skillNVIDIA/skills3.6k—~4.9kAutomated safety check: NotesApache-2.0
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0
Hugging Face Transformers Usagedavila7/claude-code-templates33k11 repos~1.2kAutomated safety check: PassMIT
Discover MLrand/cc-polymath181—~574Automated safety check: PassMIT
Huggingface Vision Trainerwaybarrios/opencode-power-pack534—~2.7kAutomated safety check: PassApache-2.0
Dataset Transformationawslabs/agent-plugins9161 repos~3.5kAutomated safety check: PassApache-2.0

Similar skills

  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Transformers Usage

    davila7/claude-code-templates

    Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

    33k GitHub starsUsed in 11 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Discover ML

    rand/cc-polymath

    Automatically discover machine learning and AI skills when working with machine learning, PyTorch, training, inference, RAG, embeddings, fine-tuning, LLM, DSPy, HuggingFace, or diffusion models.

    181 GitHub stars~574 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Huggingface Vision Trainer

    waybarrios/opencode-power-pack

    Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

    534 GitHub stars~2.7k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    916 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing.

    97k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Finetune Huggingface Model

What does Tao Finetune Huggingface Model do?

Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches. Tao Finetune Huggingface Model is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.

When should I use Tao Finetune Huggingface Model?

Tao Finetune Huggingface Model fits situations like: the user wants to fine-tune a HuggingFace model (full; train a vision / VLM / LLM model end-to-end; generate a reproducible HF training pipeline; smoke-test a HuggingFace model locally before scale-up.

How do I install Tao Finetune Huggingface Model in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-finetune-huggingface-model -a claude-code`. Or copy the skill folder (skills/tao-finetune-huggingface-model in NVIDIA/skills) into .claude/skills/tao-finetune-huggingface-model in your project. Claude Code loads it when a task matches its description.

How do I install Tao Finetune Huggingface Model in Codex?

Run `npx skills add NVIDIA/skills --skill tao-finetune-huggingface-model -a codex`. Or copy the skill folder (skills/tao-finetune-huggingface-model in NVIDIA/skills) into .agents/skills/tao-finetune-huggingface-model in your project. Codex loads it when a task matches its description.

Can I use Tao Finetune Huggingface Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-finetune-huggingface-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-finetune-huggingface-model, .gemini/skills/tao-finetune-huggingface-model, .github/skills/tao-finetune-huggingface-model and .opencode/skills/tao-finetune-huggingface-model in your project.

What does Tao Finetune Huggingface Model need to run?

Going by SKILL.md and its folder, Tao Finetune Huggingface Model needs the command-line tools its instructions call (docker, python and pytest) and credentials named HF_TOKEN and WANDB_API_KEY. Our summary lists: Python 3; Docker; A credential in WANDB_API_KEY. Its frontmatter pre-approves these tools: Read, Bash, Write. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit, NVIDIA GPU (driver ≥ 545, ≥ 24 GB VRAM for ≤3B models), ~40 GB free disk. Optional credentials (read from the session environment) — HF_TOKEN is read only when the model/dataset is gated or `push_to_hub` is on; WANDB_API_KEY and WANDB_PROJECT only when WandB logging is enabled..

Does Tao Finetune Huggingface Model access the network?

SKILL.md names 2 domains. As links in the text: apache.org and docs.nvidia.com. This is read from the text; nothing was executed.

Is Tao Finetune Huggingface Model safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Finetune Huggingface Model use?

Tao Finetune Huggingface Model is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Finetune Huggingface Model use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 70k tokens, read only when the agent opens those files.

What are the alternatives to Tao Finetune Huggingface Model?

Skills that share tags, products or a category with Tao Finetune Huggingface Model: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Hugging Face Transformers Usage (davila7/claude-code-templates, 33k stars), Discover ML (rand/cc-polymath, 181 stars) and Huggingface Vision Trainer (waybarrios/opencode-power-pack, 534 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Finetune Huggingface Model?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.