Huggingface LLM Trainer
waybarrios/opencode-power-pack
Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
$ npx skills add huggingface/skills --skill huggingface-llm-trainer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install huggingface/skills huggingface-llm-trainer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/huggingface-llm-trainer .claude/skills/huggingface-llm-trainer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "huggingface-llm-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer into .claude/skills/huggingface-llm-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-llm-trainer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add huggingface/skills --skill huggingface-llm-trainer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install huggingface/skills huggingface-llm-trainer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/huggingface-llm-trainer .agents/skills/huggingface-llm-trainer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "huggingface-llm-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer into .agents/skills/huggingface-llm-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-llm-trainer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add huggingface/skills --skill huggingface-llm-trainer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install huggingface/skills huggingface-llm-trainer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/huggingface-llm-trainer .cursor/skills/huggingface-llm-trainer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "huggingface-llm-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer into .cursor/skills/huggingface-llm-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-llm-trainer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/huggingface/skills.git --path skills/huggingface-llm-trainer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add huggingface/skills --skill huggingface-llm-trainer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install huggingface/skills huggingface-llm-trainer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/huggingface-llm-trainer .gemini/skills/huggingface-llm-trainer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "huggingface-llm-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer into .gemini/skills/huggingface-llm-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-llm-trainer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install huggingface/skills huggingface-llm-trainerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add huggingface/skills --skill huggingface-llm-trainer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/huggingface-llm-trainer .github/skills/huggingface-llm-trainer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "huggingface-llm-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer into .github/skills/huggingface-llm-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-llm-trainer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add huggingface/skills --skill huggingface-llm-trainer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install huggingface/skills huggingface-llm-trainer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/huggingface-llm-trainer .opencode/skills/huggingface-llm-trainer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "huggingface-llm-trainer" agent skill from https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer into .opencode/skills/huggingface-llm-trainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "huggingface-llm-trainer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
huggingface-llm-trainerTrains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
This skill covers training language models on Hugging Face Jobs, so no local GPU is required, with results saved to the Hugging Face Hub. TRL methods include SFT for instruction tuning, DPO for preference alignment, GRPO for online reinforcement learning and reward modeling for RLHF. Unsloth is recommended instead when GPU memory is tight, speed matters, the model is larger than 13B or a vision-language model is involved.
The agent is told to submit work through the hf_jobs MCP tool, passing the training script as a string, and to include Trackio monitoring in every script. Bundled scripts include SFT, DPO, GRPO and Unsloth examples, a dataset inspector, a cost estimator, a benchmark helper and a GGUF converter. Reference notes cover training methods, hardware selection, saving to the Hub, Trackio, reliability, troubleshooting, local training on macOS, and GGUF conversion for local use with Ollama, LM Studio or llama.cpp.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 8 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
hfuvuvxFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.cogithub.comraw.githubusercontent.comgist.githubusercontent.comAlso links to:
hf.codocs.astral.shFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Hugging Face LLM Trainer loads about 7.2k tokens when it runs, and up to ~27k if it reads all its reference files. Until then it costs about 162 tokens; SKILL.md has 2,464 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 2,464 words, ~7,166 tokens.
.claude/skills/huggingface-llm-trainer/SKILL.md (or your agent's skills folder). This skill also uses 18 other files; get the full folder from GitHub.Train language models using TRL (Transformer Reinforcement Learning) on fully managed Hugging Face infrastructure. No local GPU setup required—models train on cloud GPUs and results are automatically saved to the Hugging Face Hub.
TRL provides multiple training methods:
For detailed TRL method documentation:
hf_doc_search("your query", product="trl")
hf_doc_fetch("https://huggingface.co/docs/trl/sft_trainer") # SFT
hf_doc_fetch("https://huggingface.co/docs/trl/dpo_trainer") # DPO
# etc.See also: references/training_methods.md for method overviews and selection guidance
Use this skill when users want to:
Use Unsloth (references/unsloth.md) instead of standard TRL when:
FastVisionModel supportSee references/unsloth.md for complete Unsloth documentation and scripts/unsloth_sft_example.py for a production-ready training script.
When assisting with training jobs:
ALWAYS use hf_jobs() MCP tool - Submit jobs using hf_jobs("uv", {...}), NOT bash trl-jobs commands. The script parameter accepts Python code directly. Do NOT save to local files unless the user explicitly requests it. Pass the script content as a string to hf_jobs(). If user asks to "train a model", "fine-tune", or similar requests, you MUST create the training script AND submit the job immediately using hf_jobs().
Always include Trackio - Every training script should include Trackio for real-time monitoring. Use example scripts in scripts/ as templates.
Provide job details after submission - After submitting, provide job ID, monitoring URL, estimated time, and note that the user can request status checks later.
Use example scripts as templates - Reference scripts/train_sft_example.py, scripts/train_dpo_example.py, etc. as starting points.
Repository scripts use PEP 723 inline dependencies. Run them with uv run:
uv run scripts/estimate_cost.py --help
uv run scripts/dataset_inspector.py --helpBefore starting any training job, verify:
hf_whoami()secrets={"HF_TOKEN": "$HF_TOKEN"} in job config to make token available (the $HF_TOKEN syntax
references your actual token value)datasets.load_dataset()push_to_hub=True, hub_model_id="username/model-name"; Job: secrets={"HF_TOKEN": "$HF_TOKEN"}⚠️ IMPORTANT: Training jobs run asynchronously and can take hours
When user requests training:
scripts/train_sft_example.py as template)hf_jobs() MCP tool with script content inline - don't save to file unless user requestsProvide to user:
Example Response:
✅ Job submitted successfully!
Job ID: abc123xyz
Monitor: https://huggingface.co/jobs/username/abc123xyz
Expected time: ~2 hours
Estimated cost: ~$10
The job is running in the background. Ask me to check status/logs when ready!💡 Tip for Demos: For quick demos on smaller GPUs (t4-small), omit eval_dataset and eval_strategy to save ~40% memory. You'll still see training loss and learning progress.
TRL config classes use max_length (not max_seq_length) to control tokenized sequence length:
# ✅ CORRECT - If you need to set sequence length
SFTConfig(max_length=512) # Truncate sequences to 512 tokens
DPOConfig(max_length=2048) # Longer context (2048 tokens)
# ❌ WRONG - This parameter doesn't exist
SFTConfig(max_seq_length=512) # TypeError!Default behavior: max_length=1024 (truncates from right). This works well for most training.
When to override:
max_length=2048)max_length=512)max_length=None (prevents cutting image tokens)Usually you don't need to set this parameter at all - the examples below use the sensible default.
UV scripts use PEP 723 inline dependencies for clean, self-contained training. This is the primary approach for Claude Code.
hf_jobs("uv", {
"script": """
# /// script
# dependencies = ["trl>=0.12.0", "peft>=0.7.0", "trackio"]
# ///
from datasets import load_dataset
from peft import LoraConfig
from trl import SFTTrainer, SFTConfig
import trackio
dataset = load_dataset("trl-lib/Capybara", split="train")
# Create train/eval split for monitoring
dataset_split = dataset.train_test_split(test_size=0.1, seed=42)
trainer = SFTTrainer(
model="Qwen/Qwen2.5-0.5B",
train_dataset=dataset_split["train"],
eval_dataset=dataset_split["test"],
peft_config=LoraConfig(r=16, lora_alpha=32),
args=SFTConfig(
output_dir="my-model",
push_to_hub=True,
hub_model_id="username/my-model",
num_train_epochs=3,
eval_strategy="steps",
eval_steps=50,
report_to="trackio",
project="meaningful_prject_name", # project name for the training name (trackio)
run_name="meaningful_run_name", # descriptive name for the specific training run (trackio)
)
)
trainer.train()
trainer.push_to_hub()
""",
"flavor": "a10g-large",
"timeout": "2h",
"secrets": {"HF_TOKEN": "$HF_TOKEN"}
})Benefits: Direct MCP tool usage, clean code, dependencies declared inline (PEP 723), no file saving required, full control
When to use: Default choice for all training tasks in Claude Code, custom training logic, any scenario requiring hf_jobs()
⚠️ Important: The script parameter accepts either inline code (as shown above) OR a URL. Local file paths do NOT work.
Why local paths don't work: Jobs run in isolated Docker containers without access to your local filesystem. Scripts must be:
Common mistakes:
# ❌ These will all fail
hf_jobs("uv", {"script": "train.py"})
hf_jobs("uv", {"script": "./scripts/train.py"})
hf_jobs("uv", {"script": "/path/to/train.py"})Correct approaches:
# ✅ Inline code (recommended)
hf_jobs("uv", {"script": "# /// script\n# dependencies = [...]\n# ///\n\n<your code>"})
# ✅ From Hugging Face Hub
hf_jobs("uv", {"script": "https://huggingface.co/user/repo/resolve/main/train.py"})
# ✅ From GitHub
hf_jobs("uv", {"script": "https://raw.githubusercontent.com/user/repo/main/train.py"})
# ✅ From Gist
hf_jobs("uv", {"script": "https://gist.githubusercontent.com/user/id/raw/train.py"})To use local scripts: Upload to HF Hub first:
hf repos create my-training-scripts --type model
hf upload my-training-scripts ./train.py train.py
# Use: https://huggingface.co/USERNAME/my-training-scripts/resolve/main/train.pyTRL provides battle-tested scripts for all methods. Can be run from URLs:
hf_jobs("uv", {
"script": "https://github.com/huggingface/trl/blob/main/trl/scripts/sft.py",
"script_args": [
"--model_name_or_path", "Qwen/Qwen2.5-0.5B",
"--dataset_name", "trl-lib/Capybara",
"--output_dir", "my-model",
"--push_to_hub",
"--hub_model_id", "username/my-model"
],
"flavor": "a10g-large",
"timeout": "2h",
"secrets": {"HF_TOKEN": "$HF_TOKEN"}
})Benefits: No code to write, maintained by TRL team, production-tested When to use: Standard TRL training, quick experiments, don't need custom code Available: Scripts are available from https://github.com/huggingface/trl/tree/main/examples/scripts
The uv-scripts organization provides ready-to-use UV scripts stored as datasets on Hugging Face Hub:
# Discover available UV script collections
dataset_search({"author": "uv-scripts", "sort": "downloads", "limit": 20})
# Explore a specific collection
hub_repo_details(["uv-scripts/classification"], repo_type="dataset", include_readme=True)Popular collections: ocr, classification, synthetic-data, vllm, dataset-creation
When the hf_jobs() MCP tool is unavailable, use the hf jobs CLI directly.
⚠️ CRITICAL: CLI Syntax Rules
# ✅ CORRECT syntax - flags BEFORE script URL
hf jobs uv run --flavor a10g-large --timeout 2h --secrets HF_TOKEN "https://example.com/train.py"
# ❌ WRONG - "run uv" instead of "uv run"
hf jobs run uv "https://example.com/train.py" --flavor a10g-large
# ❌ WRONG - flags AFTER script URL (will be ignored!)
hf jobs uv run "https://example.com/train.py" --flavor a10g-large
# ❌ WRONG - "--secret" instead of "--secrets" (plural)
hf jobs uv run --secret HF_TOKEN "https://example.com/train.py"Key syntax rules:
hf jobs uv run (NOT hf jobs run uv)--flavor, --timeout, --secrets) must come BEFORE the script URL--secrets (plural), not --secretComplete CLI example:
hf jobs uv run \
--flavor a10g-large \
--timeout 2h \
--secrets HF_TOKEN \
"https://huggingface.co/user/repo/resolve/main/train.py"Check job status via CLI:
hf jobs ps # List all jobs
hf jobs logs <job-id> # View logs
hf jobs inspect <job-id> # Job details
hf jobs cancel <job-id> # Cancel a jobThe trl-jobs package provides optimized defaults and one-liner training.
uvx trl-jobs sft \
--model_name Qwen/Qwen2.5-0.5B \
--dataset_name trl-lib/Capybara
Benefits: Pre-configured settings, automatic Trackio integration, automatic Hub push, one-line commands When to use: User working in terminal directly (not Claude Code context), quick local experimentation Repository: https://github.com/huggingface/trl-jobs
⚠️ In Claude Code context, prefer using hf_jobs() MCP tool (Approach 1) when available.
| Model Size | Recommended Hardware | Cost (approx/hr) | Use Case |
|---|---|---|---|
| <1B params | t4-small | ~$0.75 | Demos, quick tests only without eval steps |
| 1-3B params | t4-medium, l4x1 | ~$1.50-2.50 | Development |
| 3-7B params | a10g-small, a10g-large | ~$3.50-5.00 | Production training |
| 7-13B params | a10g-large, a100-large | ~$5-10 | Large models (use LoRA) |
| 13B+ params | a100-large, a10g-largex2 | ~$10-20 | Very large (use LoRA) |
GPU Flavors: cpu-basic/upgrade/performance/xl, t4-small/medium, l4x1/x4, a10g-small/large/largex2/largex4, a100-large, h100/h100x8
Guidelines:
See: references/hardware_guide.md for detailed specifications
⚠️ EPHEMERAL ENVIRONMENT—MUST PUSH TO HUB
The Jobs environment is temporary. All files are deleted when the job ends. If the model isn't pushed to Hub, ALL TRAINING IS LOST.
In training script/config:
SFTConfig(
push_to_hub=True,
hub_model_id="username/model-name", # MUST specify
hub_strategy="every_save", # Optional: push checkpoints
)In job submission:
{
"secrets": {"HF_TOKEN": "$HF_TOKEN"} # Enables authentication
}Before submitting:
push_to_hub=True set in confighub_model_id includes username/repo-namesecrets parameter includes HF_TOKENSee: references/hub_saving.md for detailed troubleshooting
⚠️ DEFAULT: 30 MINUTES—TOO SHORT FOR TRAINING
{
"timeout": "2h" # 2 hours (formats: "90m", "2h", "1.5h", or seconds as integer)
}| Scenario | Recommended | Notes |
|---|---|---|
| Quick demo (50-100 examples) | 10-30 min | Verify setup |
| Development training | 1-2 hours | Small datasets |
| Production (3-7B model) | 4-6 hours | Full datasets |
| Large model with LoRA | 3-6 hours | Depends on dataset |
Always add 20-30% buffer for model/dataset loading, checkpoint saving, Hub push operations, and network delays.
On timeout: Job killed immediately, all unsaved progress lost, must restart from beginning
Identify models to train based on task type or benchmark results.
Use scripts/hf_benchmarks.py to identify top-performing models for specific tasks. This helps the user select a model as the base for training, whilst keeping size and hardware constraints in mind.
# Get help on the benchmarks command:
uv run scripts/hf_benchmarks.py --help# Search for benchmarks containing whose name contains the text `ocr`
uv run scripts/hf_benchmarks.py search --query ocr
# Get the ranked leaderboard for the allenai/olmOCR-bench benchmark
uv run scripts/hf_benchmarks.py leaderboard allenai/olmOCR-benchOffer to estimate cost when planning jobs with known parameters. Use scripts/estimate_cost.py:
uv run scripts/estimate_cost.py \
--model meta-llama/Llama-2-7b-hf \
--dataset trl-lib/Capybara \
--hardware a10g-large \
--dataset-size 16000 \
--epochs 3Output includes estimated time, cost, recommended timeout (with buffer), and optimization suggestions.
When to offer: User planning a job, asks about cost/time, choosing hardware, job will run >1 hour or cost >$5
Production-ready templates with all best practices:
Load these scripts for correctly:
scripts/train_sft_example.py - Complete SFT training with Trackio, LoRA, checkpointsscripts/train_dpo_example.py - DPO training for preference learningscripts/train_grpo_example.py - GRPO training for online RLThese scripts demonstrate proper Hub saving, Trackio integration, checkpoint management, and optimized parameters. Pass their content inline to hf_jobs() or use as templates for custom scripts.
Trackio provides real-time metrics visualization. See references/trackio_guide.md for complete setup guide.
Key points:
trackio to dependenciesreport_to="trackio" and run_name="meaningful_name"Use sensible defaults unless user specifies otherwise. When generating training scripts with Trackio:
Default Configuration:
{username}/trackio (use "trackio" as default space name)User overrides: If user requests specific trackio configuration (custom space, run naming, grouping, or additional config), apply their preferences instead of defaults.
This is useful for managing multiple jobs with the same configuration or keeping training scripts portable.
See references/trackio_guide.md for complete documentation including grouping runs for experiments.
# List all jobs
hf_jobs("ps")
# Inspect specific job
hf_jobs("inspect", {"job_id": "your-job-id"})
# View logs
hf_jobs("logs", {"job_id": "your-job-id"})Remember: Wait for user to request status checks. Avoid polling repeatedly.
Validate dataset format BEFORE launching GPU training to prevent the #1 cause of training failures: format mismatches.
prompt, chosen, rejected)ALWAYS validate for:
Skip validation for known TRL datasets:
trl-lib/ultrachat_200k, trl-lib/Capybara, HuggingFaceH4/ultrachat_200k, etc.hf_jobs("uv", {
"script": "https://huggingface.co/datasets/mcp-tools/skills/raw/main/dataset_inspector.py",
"script_args": ["--dataset", "username/dataset-name", "--split", "train"]
})The script is fast, and will usually complete synchronously.
The output shows compatibility for each training method:
✓ READY - Dataset is compatible, use directly✗ NEEDS MAPPING - Compatible but needs preprocessing (mapping code provided)✗ INCOMPATIBLE - Cannot be used for this methodWhen mapping is needed, the output includes a "MAPPING CODE" section with copy-paste ready Python code.
# 1. Inspect dataset (costs ~$0.01, <1 min on CPU)
hf_jobs("uv", {
"script": "https://huggingface.co/datasets/mcp-tools/skills/raw/main/dataset_inspector.py",
"script_args": ["--dataset", "argilla/distilabel-math-preference-dpo", "--split", "train"]
})
# 2. Check output markers:
# ✓ READY → proceed with training
# ✗ NEEDS MAPPING → apply mapping code below
# ✗ INCOMPATIBLE → choose different method/dataset
# 3. If mapping needed, apply before training:
def format_for_dpo(example):
return {
'prompt': example['instruction'],
'chosen': example['chosen_response'],
'rejected': example['rejected_response'],
}
dataset = dataset.map(format_for_dpo, remove_columns=dataset.column_names)
# 4. Launch training job with confidenceMost DPO datasets use non-standard column names. Example:
Dataset has: instruction, chosen_response, rejected_response
DPO expects: prompt, chosen, rejectedThe validator detects this and provides exact mapping code to fix it.
After training, convert models to GGUF format for use with llama.cpp, Ollama, LM Studio, and other local inference tools.
What is GGUF:
When to convert:
See: references/gguf_conversion.md for complete conversion guide, including production-ready conversion script, quantization options, hardware requirements, usage examples, and troubleshooting.
Quick conversion:
hf_jobs("uv", {
"script": "<see references/gguf_conversion.md for complete script>",
"flavor": "a10g-large",
"timeout": "45m",
"secrets": {"HF_TOKEN": "$HF_TOKEN"},
"env": {
"ADAPTER_MODEL": "username/my-finetuned-model",
"BASE_MODEL": "Qwen/Qwen2.5-0.5B",
"OUTPUT_REPO": "username/my-model-gguf"
}
})See references/training_patterns.md for detailed examples including:
Fix (try in order):
per_device_train_batch_size=1, increase gradient_accumulation_steps=8. Effective batch size is per_device_train_batch_size x gradient_accumulation_steps. For best performance keep effective batch size close to 128. gradient_checkpointing=TrueFix:
uv run https://huggingface.co/datasets/mcp-tools/skills/raw/main/dataset_inspector.py \
--dataset name --split trainFix:
hf_jobs("logs", {"job_id": "..."})"timeout": "3h" (add 30% to estimated time)num_train_epochs, use smaller dataset, enable max_stepssave_strategy="steps", save_steps=500, hub_strategy="every_save"Note: Default 30min is insufficient for real training. Minimum 1-2 hours.
Fix:
secrets={"HF_TOKEN": "$HF_TOKEN"}push_to_hub=True, hub_model_id="username/model-name"mcp__huggingface__hf_whoami()hub_private_repo=True)Fix: Add to PEP 723 header:
# /// script
# dependencies = ["trl>=0.12.0", "peft>=0.7.0", "trackio", "missing-package"]
# ///Common issues:
mcp__huggingface__hf_whoami(), token permissions, secrets parameterSee: references/troubleshooting.md for complete troubleshooting guide
references/training_methods.md - Overview of SFT, DPO, GRPO, KTO, PPO, Reward Modelingreferences/training_patterns.md - Common training patterns and examplesreferences/unsloth.md - Unsloth for fast VLM training (~2x speed, 60% less VRAM)references/gguf_conversion.md - Complete GGUF conversion guidereferences/trackio_guide.md - Trackio monitoring setupreferences/hardware_guide.md - Hardware specs and selectionreferences/hub_saving.md - Hub authentication troubleshootingreferences/troubleshooting.md - Common issues and solutionsreferences/local_training_macos.md - Local training on macOSscripts/train_sft_example.py - Production SFT templatescripts/train_dpo_example.py - Production DPO templatescripts/train_grpo_example.py - Production GRPO templatescripts/unsloth_sft_example.py - Unsloth text LLM training template (faster, less VRAM)scripts/estimate_cost.py - Estimate time and cost (offer when appropriate)scripts/convert_to_gguf.py - Complete GGUF conversion scriptscripts/hf_benchmarks.py - Search for benchmark results and leaderboards by task, alias or free text.uv run or hf_jobs)script parameter accepts Python code directly; no file saving required unless user requestsscripts/estimate_cost.pyhf_jobs("uv", {...}) with inline scripts; TRL maintained scripts for standard training; avoid bash trl-jobs commands in Claude Code© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 18 other files (scripts, references) in skills/huggingface-llm-trainer of huggingface/skills.
Open the folder on GitHubat commit ca0325b
We found 13 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.
Hugging Face LLM Trainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Hugging Face LLM Trainer this skillhuggingface/skills | 11k | 3 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| Huggingface LLM Trainerwaybarrios/opencode-power-pack | 533 | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| Open Weightsericrisco/rsc-harness | 167 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Resolvealexziskind1/model-shelf | 130 | — | ~792 | Automated safety check: Pass | MIT | |
| Generate Openenv Envadithya-s-k/FineEnvs | 443 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Xybrid Initxybrid-ai/xybrid | 466 | — | ~3k | Automated safety check: Pass | Apache-2.0 |
waybarrios/opencode-power-pack
Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.
ericrisco/rsc-harness
A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
adithya-s-k/FineEnvs
Builds an OpenEnv (Hugging Face) variant of an RL environment.
xybrid-ai/xybrid
Generate model metadata for an ML model so it works with xybrid.
sickn33/agentic-awesome-skills
Train or fine-tune language and vision models using TRL (Transformer Reinforcement Learning) or Unsloth with Hugging Face Jobs infrastructure.
huggingface/skills
Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
huggingface/skills
Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
huggingface/skills
Indexes research papers on the Hugging Face Hub from arXiv, links them to models and datasets, claims authorship and generates markdown research articles from templates.
Categories
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF. This skill covers training language models on Hugging Face Jobs, so no local GPU is required, with results saved to the Hugging Face Hub. TRL methods include SFT for instruction tuning, DPO for preference alignment, GRPO for online reinforcement learning and reward modeling for RLHF.
Hugging Face LLM Trainer fits situations like: fine-tuning a language model on cloud GPUs without local setup; running SFT, DPO or GRPO training with TRL on Hugging Face Jobs; converting a trained model to GGUF for local use; estimating training cost and choosing hardware before launching a job.
Run `npx skills add huggingface/skills --skill huggingface-llm-trainer -a claude-code`. Or copy the skill folder (skills/huggingface-llm-trainer in huggingface/skills) into .claude/skills/huggingface-llm-trainer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add huggingface/skills --skill huggingface-llm-trainer -a codex`. Or copy the skill folder (skills/huggingface-llm-trainer in huggingface/skills) into .agents/skills/huggingface-llm-trainer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill huggingface-llm-trainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huggingface-llm-trainer, .gemini/skills/huggingface-llm-trainer, .github/skills/huggingface-llm-trainer and .opencode/skills/huggingface-llm-trainer in your project.
Going by SKILL.md and its folder, Hugging Face LLM Trainer needs Python for the scripts in its folder, the command-line tools its instructions call (hf, uv and uvx) and credentials named HF_TOKEN. Our summary lists: Hugging Face Hub authentication; Access to Hugging Face Jobs for cloud GPU training; The hf_jobs MCP tool.
SKILL.md names 6 domains. In commands or code: huggingface.co, github.com, raw.githubusercontent.com and gist.githubusercontent.com; the agent is likely to contact these when it follows the instructions. As links in the text: hf.co and docs.astral.sh. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Hugging Face LLM Trainer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Hugging Face LLM Trainer: Huggingface LLM Trainer (waybarrios/opencode-power-pack, 533 stars), Open Weights (ericrisco/rsc-harness, 167 stars), Resolve (alexziskind1/model-shelf, 130 stars) and Generate Openenv Env (adithya-s-k/FineEnvs, 443 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,148 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.
Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.