Agent skill

Finetuning Method Selection

by wshobson in wshobson/agents

Decide whether to fine-tune at all, and route to the right method (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) and base model.

MITAuto-check passedAI & LLM Engineering

Install Finetuning Method Selection

skills CLI
$ npx skills add wshobson/agents --skill finetuning-method-selection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents finetuning-method-selection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/finetuning-method-selection .claude/skills/finetuning-method-selection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
finetuning-method-selection
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
987 words
Files
3 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Decide whether to fine-tune at all, and route to the right method (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) and base model.

  • Starting any fine-tuning effort
  • SKILL.md covers When to Use This Skill, Quick Reference, Off-Ramps First and Method Router, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Unsure whether RAG

What it does

Finetuning Method Selection is an agent skill from wshobson/agents. Decide whether to fine-tune at all, and route to the right method (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) and base model. Use when starting any fine-tuning effort, when unsure whether RAG or prompting would suffice, or when choosing between preference-optimization and reinforcement methods.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/memory-math.md` and `references/model-catalog.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Starting any fine-tuning effort
  • Unsure whether RAG
  • Prompting would suffice
  • Choosing between preference-optimization and reinforcement methods

Example prompts

  • “/finetuning-method-selection”

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Finetuning Method Selection loads about 2k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 987 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 987 words, ~1,958 tokens.

Download SKILL.mdSave it as .claude/skills/finetuning-method-selection/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
finetuning-method-selection
description
Decide whether to fine-tune at all, and route to the right method (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) and base model. Use when starting any fine-tuning effort, when unsure whether RAG or prompting would suffice, or when choosing between preference-optimization and reinforcement methods.

Fine-Tuning Method Selection

This is the router skill for the fine-tuning lifecycle: it decides whether fine-tuning is the right tool at all, and if so, which method and which base-model size class. Every other skill in this plugin assumes this routing already happened — start here before opening lora-qlora-recipes, preference-optimization, or grpo-rlvr-training.

When to Use This Skill

  • Starting any fine-tuning effort, before a framework or base model has been chosen.
  • Unsure whether RAG or prompt engineering would solve the problem more cheaply than training.
  • Choosing between preference optimization (DPO family) and a reinforcement method (GRPO/RLVR) for the same underlying task.
  • Sizing a candidate model/method combination before committing to a run.

Quick Reference

SituationRoute
Facts change often (prices, docs, news)RAG, not fine-tuning
Desired behavior still being figured outPrompt engineering
Stable domain knowledge, ≥500MB textCPT then SFT — see Off-Ramps First
Have input/output demonstrationsSFT — see lora-qlora-recipes
Have preference pairs or thumbs-up/downDPO/ORPO/KTO — see preference-optimization
Have a verifiable pass/fail signalGRPO+RLVR — see grpo-rlvr-training
No eval harness yetStop — see eval-harness-first

Off-Ramps First

Most requests that sound like "fine-tune this" are served better and cheaper elsewhere. Check these off-ramps before opening a training run:

  • Knowledge-bound and volatile (the gap is facts that change — prices, docs, current events): route to RAG, not fine-tuning. A fine-tuned model bakes in a snapshot; volatile facts go stale immediately.
  • Behavior-bound and shifting (the desired behavior is still being figured out, or changes per request): route to prompt engineering. Fine-tuning locks in a behavior; don't lock in one that hasn't stabilized yet.
  • Stable, dense domain knowledge: this is where continued pretraining (CPT) enters, sized by how much domain text exists:
Domain text volumeRoute
<10MBRAG only
10MB–500MBRAG + fine-tune
500MB–10GBCPT, then SFT
>10GBCPT required

CPT learning rate ≈ 10% of the pretraining LR. CPT is guidance-only in this plugin — sizing and LR guidance live here, but this plugin does not execute a CPT run.

Method Router

Once the off-ramps are ruled out, this is the full decision tree (verbatim from the research this plugin is built on):

New FACTS?  volatile → RAG | stable+dense → CPT (LR ~10% of pretrain) → SFT
New BEHAVIOR? shifting → prompt-engineering | stable:
  demos → SFT (LoRA/QLoRA, all-linear, α=2r)
  preference pairs → DPO (SimPO if length-bias, ORPO if memory-bound)
  unpaired 👍/👎 → KTO
  verifiable success → RLVR + GRPO (DAPO/GSPO/Dr.GRPO per failure mode)
Deploy: FP8 (Hopper+) | NVFP4 (Blackwell scale) | AWQ (older) | GGUF+imatrix (edge)
BEFORE ANY OF THIS: the eval harness must exist first.

Read the tree top-down: answer "new facts or new behavior," then follow the branch that matches the data shape in hand (demos, preference pairs, thumbs up/down, or verifiable success/failure). The data shape picks the method — not the other way around.

Worked Routing Examples
  • "Users want the assistant to follow our support macros exactly." Behavior is stable and demonstrable from transcripts → demos → SFT.
  • "We have pairs of good/bad responses from reviewer thumbs-up/down, unpaired." → unpaired signal → KTO, not DPO (DPO needs paired preferences).
  • "The model can already solve some of these math problems and we can grade correctness automatically." → verifiable success signal → GRPO+RLVR, and only after confirming the model succeeds at least sometimes (see Key Routing Facts below).
  • "We want the model to know this week's pricing page." → volatile facts → RAG, no training run at all.

Key Routing Facts

  • Loss-function choice is low-leverage. A 240-H100-run study found method choice worth ~1 percentage point versus ~50 points for model scale, and zero of 20 DPO variants beat vanilla DPO. Don't spend a routing decision agonizing over DPO-variant selection — spend it on getting the data shape and scale right.
  • DPO is for taste, GRPO+RLVR is for reasoning. Preference pairs that encode a subjective judgment (tone, style, "which answer is better") route to DPO. Tasks with a verifiable pass/fail signal (math, code, tool calls) route to GRPO+RLVR instead.
  • RL is not the fix for a model that never succeeds. GRPO and other RL methods sharpen an existing capability — they don't teach one from zero. If the model doesn't yet understand the task or output format, run SFT first; only bring in RL once the model succeeds at least sometimes.
Show full SKILL.md (371 more words)Show less
Common Routing Mistakes
  • Reaching for fine-tuning to fix facts that change weekly — that's a RAG problem, and fine-tuning will just go stale faster than the source data does.
  • Picking a DPO variant before checking whether the actual bottleneck is data quality or model scale — variant choice is the ~1pp lever, not the ~50pp one.
  • Starting an RL run on a model that fails every rollout — route to SFT first so RL has something to sharpen.
  • Treating CPT as the default for "the model doesn't know our domain" — check the data volume thresholds first; under 500MB, RAG or RAG+fine-tune iterates faster than a CPT run.

Model Selection

Base-model choice is size-class first, family second, and it goes stale fast — so it lives in exactly one place: references/model-catalog.md. That file is the only place in this plugin (and in the DGX Spark ops plugin) that names a base model family. Neither this skill nor references/memory-math.md names one; both describe models by size class only (for example, "8B-class LoRA," not a model name).

The catalog is dated on purpose — model rankings turn over quarterly. It carries a "last verified" date and a refresh checklist. Before trusting a row, check that date; if stale, work the refresh checklist in the catalog before recommending a model from it.

Precedence when the catalog and a method skill disagree: the catalog's per-row Notes column states hardware/size-class feasibility, not a method recommendation — lora-qlora-recipes's LoRA vs QLoRA vs Full FT table (routed by task shape) governs the actual method choice.

Memory Feasibility

Before committing to a method, size it: total memory ≈ params × dtype bytes + optimizer state + gradients + activations. Work each term for the chosen dtype and method (full fine-tune, LoRA, or QLoRA) — worked worksheets and size-class examples live in references/memory-math.md.

On DGX Spark specifically, unified-memory behavior breaks the naive estimate (transient load peaks, nvidia-smi underreporting, thermal throttling on long runs). Once the dgx-spark-ops plugin is installed, defer Spark-specific feasibility calls to its spark-memory-thermal-ops skill rather than re-deriving them here.

Once this skill has picked a method, hand off to the skill that executes it:

  • lora-qlora-recipes — SFT via LoRA/QLoRA
  • preference-optimization — DPO, ORPO, KTO
  • grpo-rlvr-training — GRPO with verifiable rewards

No method is selected before the eval harness exists — see eval-harness-first.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/llm-finetuning/skills/finetuning-method-selection of wshobson/agents.

  • SKILL.md
  • references/memory-math.md
  • references/model-catalog.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Finetuning Method Selection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Finetuning Method Selection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Finetuning Method Selection this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Dataset Evaluationawslabs/agent-plugins9122 repos~1.3kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    912 GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 13 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 12 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 12 repos~814 tokens
    Auto-check passed
  • Portfolio Risk Metrics

    wshobson/agents

    Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.

    40k GitHub starsUsed in 12 repos~502 tokens
    Auto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 11 repos~1.3k tokens
    Auto-check passed

Questions about Finetuning Method Selection

What does Finetuning Method Selection do?

Decide whether to fine-tune at all, and route to the right method (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) and base model. Finetuning Method Selection is an agent skill from wshobson/agents. Decide whether to fine-tune at all, and route to the right method (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) and base model.

When should I use Finetuning Method Selection?

Finetuning Method Selection fits situations like: starting any fine-tuning effort; unsure whether RAG; prompting would suffice; choosing between preference-optimization and reinforcement methods.

How do I install Finetuning Method Selection in Claude Code?

Run `npx skills add wshobson/agents --skill finetuning-method-selection -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/finetuning-method-selection in wshobson/agents) into .claude/skills/finetuning-method-selection in your project. Claude Code loads it when a task matches its description.

How do I install Finetuning Method Selection in Codex?

Run `npx skills add wshobson/agents --skill finetuning-method-selection -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/finetuning-method-selection in wshobson/agents) into .agents/skills/finetuning-method-selection in your project. Codex loads it when a task matches its description.

Can I use Finetuning Method Selection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill finetuning-method-selection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/finetuning-method-selection, .gemini/skills/finetuning-method-selection, .github/skills/finetuning-method-selection and .opencode/skills/finetuning-method-selection in your project.

What does Finetuning Method Selection need to run?

SKILL.md names no scripts, command-line tools or credentials: Finetuning Method Selection is instructions for the agent only.

Does Finetuning Method Selection access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Finetuning Method Selection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Finetuning Method Selection use?

Finetuning Method Selection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Finetuning Method Selection use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Finetuning Method Selection?

Skills that share tags, products or a category with Finetuning Method Selection: Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Sentence-Transformers Training Router (huggingface/skills, 11k stars) and Dataset Evaluation (awslabs/agent-plugins, 912 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Finetuning Method Selection?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,254 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.