Hugging Face LLM Trainer
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
A skill your agent uses when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then…
$ npx skills add ericrisco/rsc-harness --skill finetuning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness finetuning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/finetuning .claude/skills/finetuning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "finetuning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/finetuning into .claude/skills/finetuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "finetuning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/finetuningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill finetuning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness finetuning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/finetuning .agents/skills/finetuning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "finetuning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/finetuning into .agents/skills/finetuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "finetuning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill finetuning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness finetuning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/finetuning .cursor/skills/finetuning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "finetuning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/finetuning into .cursor/skills/finetuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "finetuning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/finetuning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill finetuning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness finetuning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/finetuning .gemini/skills/finetuning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "finetuning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/finetuning into .gemini/skills/finetuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "finetuning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness finetuningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill finetuning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/finetuning .github/skills/finetuning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "finetuning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/finetuning into .github/skills/finetuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "finetuning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill finetuning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness finetuning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/finetuning .opencode/skills/finetuning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "finetuning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/finetuning into .opencode/skills/finetuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "finetuning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
finetuningA skill your agent uses when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then…
Finetuning is an agent skill from ericrisco/rsc-harness. Use when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then preference optimization (DPO/ORPO/KTO/GRPO), and for fine-tune vs prompt vs RAG. NOT adding facts to a model (that is rag); NOT the single-GPU Unsloth backend or GGUF export (that is unsloth).
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/hyperparameters-and-eval.md`).
It sits in AI & LLM Engineering, covering Fine-tuning. It works with llama.cpp. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
arxiv.orghuggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Finetuning loads about 3.8k tokens when it runs, and up to ~7.1k if it reads all its reference files. Until then it costs about 94 tokens; SKILL.md has 1,556 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,556 words, ~3,760 tokens.
.claude/skills/finetuning/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.You own the discipline of adapting an open-weight model: deciding whether to fine-tune at all,
then running SFT and (optionally) preference optimization with trl + peft, backend-agnostic.
You are judged by whether the tuned model reliably produces the target form/behavior on a
held-out set — not by train loss, and not by vibes.
The one sentence that routes half of all "should I fine-tune?" questions correctly:
fine-tuning teaches form and behavior; RAG supplies facts. If the ask is "know our latest
prices / docs / tickets," that is retrieval (../rag/SKILL.md), not training. If the ask is "sound
like us, always emit this JSON, follow this reasoning pattern," that is here.
Fine-tuning is the last lever, not the first. Exhaust the cheaper, reversible options first; each row below is a real off-ramp.
| If the goal is… | Do this first | Fine-tune only when… |
|---|---|---|
| The model should know current/company facts | RAG (../rag/SKILL.md) — retrieve + ground | never for facts; facts go stale, weights don't update |
| One-off format/tone, small volume | Prompt + few-shot (prompt-engineering) | the prompt is huge, brittle, or you pay for it every call |
| Behavior depends on a long document | Longer context / put it in the prompt | context won't fit, or per-call token cost is the bottleneck |
| Consistent form/behavior at scale, latency/cost sensitive | — | prompting plateaus AND you have (or can build) good examples |
| A capability the base model just can't do | — | you have a reward signal or demonstration data for it |
Route out explicitly. Facts / freshness / citations → ../rag/SKILL.md. Squeezing a prompt before
spending money → prompt-engineering. Picking which base model (size/license/task) → open-weights.
Building the JSONL/preference corpus → training-data (LLM corpora, NOT tabular cleaning — that is
data-cleaning). A fast single-GPU run + GGUF export → ../unsloth/SKILL.md (same LoRA/QLoRA
concepts, one optimized implementation; this skill stays backend-agnostic). Downloading the base or
pushing the adapter/merged model → huggingface. Serving the result → ../vllm/SKILL.md.
The cheapest fine-tune is the one you didn't need. Prompt + RAG solves most "make it behave" asks at zero training cost and updates instantly. Fine-tune when that ceiling is real, measured, and you can afford to re-run it every time the base model or data changes.
trl consolidated into a v1.x line (v1.0 landed ~2026; docs at author time referenced
~v1.8). Every method has a Trainer + a Config dataclass that inherits
transformers.TrainingArguments (SFTTrainer/SFTConfig, DPOTrainer/DPOConfig, …). Confirm
the current major before pinning: pip show trl / the TRL docs.transformers is on a v5.x line; peft, bitsandbytes, accelerate, datasets
round out the stack. Do not freeze a pin as "the version" — say "current major is ~X, verify."trl.experimental.* (e.g. from trl.experimental.orpo import ORPOTrainer
at author time). Import paths churn — check the method's doc page before copying an import.open-weights.Three options on one memory↔quality axis. Default to QLoRA unless you have a proven reason not to.
| Method | What trains | Rough VRAM (7–8B) | Use when |
|---|---|---|---|
| Full FT | every weight, fp16/bf16 | very high (needs multi-GPU / offload) | you have the hardware and a large, high-quality corpus and adapters underfit |
| LoRA | small low-rank adapter matrices; base frozen (fp16) | high | base fits in fp16 and you want adapter portability + speed |
| QLoRA | LoRA adapters over a 4-bit NF4 frozen base | lowest — single consumer GPU for 7–13B | the default; fine-tune big models on one GPU with ~no quality loss |
QLoRA (Dettmers et al., arXiv:2305.14314): load the base in 4-bit NF4 with double quantization, keep it frozen, and train LoRA adapters in bf16 on top. It made single-GPU fine-tuning of large models practical at near-full-FT quality.
import torch
from transformers import BitsAndBytesConfig
from peft import LoraConfig
from trl import SFTTrainer, SFTConfig
# 4-bit NF4 base (QLoRA). Verify arg names against current bitsandbytes/transformers.
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16,
)
# LoRA over ALL linear layers — the safe target for QLoRA (PEFT quantization guide).
peft_config = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
bias="none", task_type="CAUSAL_LM",
target_modules="all-linear", # or explicit ["q_proj","k_proj","v_proj","o_proj",...]
)
trainer = SFTTrainer(
model="Qwen/Qwen2.5-7B-Instruct", # instruct base → chat template already present
args=SFTConfig(output_dir="out", max_length=2048, packing=True,
learning_rate=2e-4, num_train_epochs=2, bf16=True),
train_dataset=dataset, # conversational: TRL applies the chat template
peft_config=peft_config,
quantization_config=bnb, # SFTTrainer + peft_config + this == QLoRA
)
trainer.train()
trainer.save_model("out") # saves the ADAPTER, not a merged modelpeft_config + quantization_config on SFTTrainer is the current one-liner for QLoRA — no manual
get_peft_model / prepare_model_for_kbit_training wiring needed. save_model writes the small
adapter; merge to a standalone model only when serving requires it (see references).
SFT → preference optimization is the standard post-training arc. SFT teaches the model the
format and gives it the behavior by imitation. Preference optimization then sharpens which of
several plausible outputs is better. Most projects need only SFT; add a preference stage when "the
outputs are fine but I want the good one preferred" is the remaining gap.
| Method | Data it needs | Stage | Pick it when |
|---|---|---|---|
| SFT | demonstrations (chat/messages or prompt→completion) | base of everything | always first (except ORPO) |
| DPO (2305.18290) | paired chosen/rejected | after SFT | you have pairwise preferences; the workhorse aligner |
| ORPO (2403.07691) | paired preferences | replaces SFT+DPO (single stage, ref-free) | you want one pass from a base model and have pairs |
| KTO (2402.01306) | unpaired binary good/bad labels | after SFT | you have thumbs-up/down, not matched pairs |
| GRPO (2402.03300; DeepSeek-R1 2501.12948) | a reward function (verifier), no pairs | after SFT | correctness is checkable (math/code/format) → RL for reasoning |
Rule of thumb: have pairs → DPO (or ORPO to fuse the two stages); have only up/down votes → KTO; can score an answer programmatically → GRPO. Preference optimization uses a tiny learning rate.
# DPO after SFT — dataset has prompt / chosen / rejected columns.
from trl import DPOTrainer, DPOConfig
trainer = DPOTrainer(
model="out", # your SFT checkpoint (or SFT+adapter)
args=DPOConfig(output_dir="dpo-out", beta=0.1, # beta = KL strength to the ref model
learning_rate=5e-7, max_length=1024,
precompute_ref_log_probs=True), # saves memory; ref model auto-created
train_dataset=pref_dataset,
peft_config=peft_config, # LoRA works for preference stages too
)
trainer.train()# GRPO — no preference pairs, a reward FUNCTION that returns a score per completion.
from trl import GRPOTrainer, GRPOConfig
def format_reward(completions, **kwargs): # signature: gets completions (+ dataset cols via kwargs)
return [1.0 if "\\boxed{" in c[0]["content"] else 0.0 for c in completions]
trainer = GRPOTrainer(
model="out",
reward_funcs=[format_reward], # one or many; GRPOConfig.reward_weights to combine
args=GRPOConfig(output_dir="grpo-out", num_generations=8, # group size per prompt
beta=0.04, learning_rate=1e-6, use_vllm=True), # vLLM speeds rollouts
train_dataset=prompts_dataset,
)
trainer.train()Full runnable SFT→DPO and GRPO scripts, ORPO/KTO variants, dataset schemas, and adapter-merge steps
are in references/methods.md.
More rows is not the win. LIMA (arXiv:2305.11206) got strong
instruction-following from ~1,000 carefully curated examples — "less is more for alignment."
A thousand clean, on-distribution, correctly-templated examples beat 100k scraped noisy ones, which
actively teach the model bad form. Building and validating that corpus (JSONL messages, preference
pairs, dedup, contamination checks) is training-data — bring it here already clean.
r and lora_alpha: r = adapter rank (capacity); effective scaling = lora_alpha / r.
Common heuristic alpha ≈ 2·r (e.g. r=16→alpha=32) so scaling ≈ 2; then adjust LR, not both.
Start r=8–16 for style/format, higher (32–64+) for harder behavior. (Newer "LoRA-without-regret"
guidance favors target_modules="all-linear" + higher rank + tuned LR — verify current advice.)target_modules: "all-linear" is the safe default. Targeting too few / wrong-named modules
is a top silent failure — the run "succeeds," loss barely moves, the adapter learned ~nothing.SFTConfig default). Preference optimization is far lower — DPO ~5e-7, GRPO ~1e-6.gradient_accumulation_steps when VRAM
caps per_device_train_batch_size. Enable packing=True + gradient_checkpointing to fit more.warmup_ratio (~0.03–0.1) stabilizes the early, high-gradient steps.Fine-tuning on a narrow task can degrade general ability the base model had. Three mitigations, cheapest first: use LoRA/QLoRA (base weights frozen — inherently gentler than full FT); keep the LR low and epochs few; and replay — mix a slice of general instruction data into your task data so the model doesn't forget how to be a general assistant. If a tuned model suddenly "got dumber" at everything else, this is the usual cause.
A vibe-check is not an eval. Before training, split off a held-out set the model never sees, and define a concrete task metric (exact-match / JSON-valid rate / rubric score / a task-specific score). Judge the run on that, plus eval loss.
training-data.)agent-eval. Bring your task metric here.The full tuning + forgetting + evaluation playbook is in references/hyperparameters-and-eval.md.
| Anti-pattern | Why it breaks | Do instead |
|---|---|---|
| Fine-tune to add facts / fresh knowledge | Weights memorize poorly and go stale; hallucinations | ../rag/SKILL.md — retrieve + ground |
| Fine-tune before trying prompt + few-shot | Slow, costly, irreversible for a prompt-solvable ask | prompt-engineering first |
Wrong / too-few target_modules | Adapter has no capacity where it matters → learns ~nothing | "all-linear" (or correct proj names) |
| Judge success by train loss | Falls even while the model overfits | Held-out eval set + task metric + eval loss |
| Crank epochs "to learn it better" | Overfits, forgets, memorizes noise | 1–3 epochs; stop when eval loss turns up |
| Train with a wrong/absent chat template | Inference emits garbage / never stops | Match train template to serve; align eos_token |
| A few dozen examples for full FT | Not enough signal; unstable | Curate ~hundreds–thousands (LIMA) or use LoRA |
| Fine-tune a model you can't legally deploy | License blocks your use case | Check the model card first → open-weights |
open-weights); read the actual model card.training-data); dedup vs eval.eos_token aligned.target_modules="all-linear" (or verified names); alpha ≈ 2·r; LoRA LR ~1e-4–2e-4.trl/peft/transformers current majors and import paths at author time.© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/finetuning of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
Finetuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Finetuning this skillericrisco/rsc-harness | 156 | — | ~3.8k | Automated safety check: Pass | MIT | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 3 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| Gemma Trainergoogle-gemma/gemma-skills | 1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Unsloth Finetuningsickn33/agentic-awesome-skills | 47k | 1 repos | ~4.1k | Automated safety check: Pass | Apache-2.0 | |
| Quantized Exportwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Wan Flf Videoartokun/comfyui-mcp | 793 | — | ~5.1k | Automated safety check: Pass | MIT |
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
google-gemma/gemma-skills
Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g.
sickn33/agentic-awesome-skills
Fine-tune and post-train LLMs with Unsloth Core on a single consumer GPU: VRAM sizing, LoRA/QLoRA, GRPO/DPO, chat-template correctness, and GGUF export.
wshobson/agents
Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.
artokun/comfyui-mcp
Build WAN 2.2 First-Last-Frame video workflows. An agent skill from artokun/comfyui-mcp.
AnastasiyaW/codex-claude-code-config
Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then…. Finetuning is an agent skill from ericrisco/rsc-harness. Use when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then preference optimization (DPO/ORPO/KTO/GRPO), and for fine-tune vs prompt vs RAG.
Finetuning fits situations like: adapting an open-weight model to a target form; behavior — tone; reasoning pattern — via LoRA/QLoRA; full fine-tuning with TRL SFTTrainer.
Run `npx skills add ericrisco/rsc-harness --skill finetuning -a claude-code`. Or copy the skill folder (skills/finetuning in ericrisco/rsc-harness) into .claude/skills/finetuning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill finetuning -a codex`. Or copy the skill folder (skills/finetuning in ericrisco/rsc-harness) into .agents/skills/finetuning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill finetuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/finetuning, .gemini/skills/finetuning, .github/skills/finetuning and .opencode/skills/finetuning in your project.
Going by SKILL.md and its folder, Finetuning needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: arxiv.org and huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Finetuning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Finetuning: Hugging Face LLM Trainer (huggingface/skills, 11k stars), Gemma Trainer (google-gemma/gemma-skills, 1k stars), Unsloth Finetuning (sickn33/agentic-awesome-skills, 47k stars) and Quantized Export (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.