Agent skill

Preference Optimization

by wshobson in wshobson/agents

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO.

MITAuto-check passedAI & LLM Engineering

Install Preference Optimization

skills CLI
$ npx skills add wshobson/agents --skill preference-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents preference-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/preference-optimization .claude/skills/preference-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
preference-optimization
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
1,004 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO.

  • Works in 4 steps: Sample completions from the current policy → Score or rank the completions (reward… → Run a DPO pass using the current… → …
  • Preference pairs
  • SKILL.md covers Method Selection, The Low-Leverage Truth, Production Pattern: Iterative… and Pair Construction, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Preference Optimization is an agent skill from wshobson/agents. Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/method-configs.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Preference pairs
  • Thumbs-up/down feedback exist
  • Choosing between preference-optimization methods
  • A DPO run needs hyperparameters

Example prompts

  • “/preference-optimization”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Sample completions from the current policy
  2. Score or rank the completions (reward model,
  3. Run a DPO pass using the current checkpoint as
  4. The resulting checkpoint becomes both the new

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Preference Optimization loads about 2k tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 1,004 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 1,004 words, ~1,960 tokens.

Download SKILL.mdSave it as .claude/skills/preference-optimization/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
preference-optimization
description
Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

Preference Optimization

This skill assumes finetuning-method-selection already routed here because the data shape is preference pairs or unpaired thumbs-up/down feedback, not demonstrations (that's lora-qlora-recipes) or a verifiable reward signal (that's grpo-rlvr-training). What follows is method selection among the DPO family, the evidence for how much that selection actually matters, the production training pattern, and how to build the pairs in the first place.

Input: a routing decision (preference optimization) plus preference pairs or unpaired feedback, usually from an SFT checkpoint. Output format: a validated method choice plus a config — the kwarg values in references/method-configs.md, not free-form advice — that llm-finetuning-training-engineer consumes directly.

Method Selection

Data shapeMethodKey parameters
Preference pairs, default caseDPOβ=0.1, LR 5e-7–1e-6, 1–2 epochs
Memory-bound or no SFT checkpointORPOreference-free, fused SFT+preference in one loss
Unpaired thumbs-up/downKTObinary label per example, no pairing needed
Length bias observed, sweep budget availableSimPOreference-free; see sweep grid below
  • DPO is the safe default. Use β=0.1 and a learning rate of 5e-7 to 1e-6 for 1–2 epochs. This LR is lower than the SFT LR that produced the checkpoint being aligned — porting an SFT- scale LR into a DPO run is the most common misconfiguration here, not an edge case.
  • ORPO routes in when memory is the constraint, or when there's no separate SFT checkpoint to start from — it's reference-free and fuses the SFT and preference objectives into one loss, skipping the separate SFT pass and the reference-model memory cost DPO carries.
  • KTO routes in when feedback is unpaired binary signal (thumbs-up/down) rather than matched preference pairs — don't force unpaired feedback into synthetic pairs to use DPO instead.
  • SimPO fixes DPO's length bias but only pays off with disciplined sweeping — its published gains are a ceiling reported under a tuned sweep, not a baseline any single config will reproduce. Route here only when there's sweep budget; use DPO instead if there isn't.
  • Classic RLHF (reward model + PPO) is retired outside frontier labs. Don't reach for it in a production pipeline — every method above is cheaper and better-supported for the same data shapes.
Worked Examples
  • "We have an SFT checkpoint and clean paired preference data, no length-bias complaints yet." → default case → DPO at β=0.1.
  • "Reviewers click thumbs-up/down per response; nothing is paired." → unpaired signal → KTO, not DPO — don't synthesize pairs to force DPO onto unpaired data.
  • "GPU budget doesn't cover a separate SFT pass plus a DPO reference model." → memory-bound, no separate checkpoint → ORPO.
  • "DPO output favors longer answers regardless of quality, and there's time to run a sweep." → length bias plus sweep budget → SimPO. Skip it if the sweep budget isn't actually there.

The Low-Leverage Truth

A 2026 240-H100-run study (arXiv 2603.19335) is the load-bearing evidence behind the table above: loss-function choice is worth roughly 1 percentage point of leverage, model scale is worth roughly 50. Zero of 20 DPO variants tested beat vanilla DPO. Rankings also invert with scale — a variant that wins in a small pilot can lose at deployment size.

Two practical consequences:

  • Don't spend a routing decision agonizing over DPO-variant bake-offs. The table above is sufficient; deeper variant selection is low-leverage compared to data quality and scale.
  • Validate at deployment scale before trusting a ranking. A method comparison run on a small pilot model doesn't transfer to the production size class — re-check the winner once scale changes.

This is also why the Method Selection table above is deliberately short: it encodes the ~1pp lever, not a ranking of DPO variants that the same study shows doesn't hold up across scale. Treat any variant-selection advice that isn't in that table — including advice that claims a specific variant "wins" — as unproven until it's been validated at the target deployment size.

Show full SKILL.md (389 more words)Show less

Production Pattern: Iterative On-Policy DPO

A single offline DPO pass on a static preference dataset is a starting point, not the production pattern. The policy drifts away from the distribution the pairs were sampled from as training proceeds, and a static dataset goes stale against that drift. Production pipelines run DPO iteratively and on-policy instead:

  1. Sample completions from the current policy checkpoint.
  2. Score or rank the completions (reward model, judge, or task grader).
  3. Run a DPO pass using the current checkpoint as the reference model.
  4. The resulting checkpoint becomes both the new policy and the new reference for the next round.

Repeat. Each round's reference model is the prior round's output, not a fixed initial checkpoint — that's what keeps the preference signal on-policy instead of scoring against an increasingly stale distribution.

A single-pass DPO run is still a reasonable first iteration — it just isn't the whole pipeline. Plan for at least one more round once the first checkpoint exists, rather than treating pass one as the finished artifact.

Pair Construction

Build DPO/ORPO pairs from same-task passing-vs-failing trajectories — two attempts at the same underlying task, not unrelated best-and-worst examples pulled from different tasks. Within that trajectory set, select the rejected member at μ−2σ of the reward distribution, never the minimum. Naive best-vs-worst pair construction (max reward vs. absolute minimum) degrades as scale increases; the μ−2σ selection is more robust to the same scale sensitivity the low-leverage study surfaced above.

sorted_by_reward = sort(trajectories, key=reward)
chosen   = sorted_by_reward[-1]                # highest reward
mu, sigma = mean(rewards), stdev(rewards)
rejected = closest(sorted_by_reward, mu - 2 * sigma)
# NOT sorted_by_reward[0] — the absolute minimum
# is the naive best-vs-worst construction that
# degrades as scale increases.

For the mechanics of turning graded traces into these pairs — including rejection sampling and judge-scored delta selection — see trace-to-training-data.

References

Complete TRL config blocks per method — DPOConfig, ORPOConfig, KTOConfig, and the SimPO sweep grid — plus Unsloth wrappers and a catastrophic-forgetting note live in references/method-configs.md. Those configs use the same current-TRL API conventions established in lora-qlora-recipes's references/unsloth-trl-mapping.md (processing_class, not tokenizer=).

references/method-configs.md also carries the catastrophic-forgetting note: a too-high learning rate is the usual cause when a preference-tuned checkpoint loses general capability, and the fix is almost always to drop the LR toward the low end of the range in the Method Selection table above before reaching for any other remediation.

Related skills: finetuning-method-selection routes here once preference pairs or unpaired feedback exist; lora-qlora-recipes produces the SFT checkpoint DPO/KTO/SimPO align (ORPO's fused path can skip it); trace-to-training-data converts passing/failing trajectories into the pairs this skill's Pair Construction section consumes.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/llm-finetuning/skills/preference-optimization of wshobson/agents.

  • SKILL.md
  • references/method-configs.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Preference Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Preference Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Preference Optimization this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Explore Codelllllllama/RigorPilot-Skills4971 repos~648Automated safety check: PassMIT
Repo DevelopmentVectorSpaceLab/AREX-Skill328—~746Automated safety check: PassApache-2.0
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Fix Art IssuesOpenPipe/ART11k—~840Automated safety check: NotesApache-2.0
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Explore Code

    lllllllama/RigorPilot-Skills

    Rigor Improve implementation leaf skill for auditable candidate implementation in deep learning research repositories.

    497 GitHub starsUsed in 1 repo~648 tokens
    AI & LLM EngineeringAuto-check passed
  • Repo Development

    VectorSpaceLab/AREX-Skill

    A skill your agent uses when modifying PEFT itself, preparing a PEFT pull request, adding a new PEFT method, selecting contributor tests, or checking PEFT contribution/style/backward-compatibility…

    328 GitHub stars~746 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Fix Art Issues

    OpenPipe/ART

    Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

    11k GitHub stars~840 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Safactory Workflows

    AI45Lab/SAfactory

    Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.

    236 GitHub stars~1.8k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 13 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 12 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 12 repos~814 tokens
    Auto-check passed
  • Portfolio Risk Metrics

    wshobson/agents

    Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.

    40k GitHub starsUsed in 12 repos~502 tokens
    Auto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 11 repos~1.3k tokens
    Auto-check passed

Questions about Preference Optimization

What does Preference Optimization do?

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Preference Optimization is an agent skill from wshobson/agents. Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO.

When should I use Preference Optimization?

Preference Optimization fits situations like: preference pairs; thumbs-up/down feedback exist; choosing between preference-optimization methods; A DPO run needs hyperparameters.

How do I install Preference Optimization in Claude Code?

Run `npx skills add wshobson/agents --skill preference-optimization -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/preference-optimization in wshobson/agents) into .claude/skills/preference-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Preference Optimization in Codex?

Run `npx skills add wshobson/agents --skill preference-optimization -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/preference-optimization in wshobson/agents) into .agents/skills/preference-optimization in your project. Codex loads it when a task matches its description.

Can I use Preference Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill preference-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/preference-optimization, .gemini/skills/preference-optimization, .github/skills/preference-optimization and .opencode/skills/preference-optimization in your project.

What does Preference Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: Preference Optimization is instructions for the agent only.

Does Preference Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Preference Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Preference Optimization use?

Preference Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Preference Optimization use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Preference Optimization?

Skills that share tags, products or a category with Preference Optimization: Explore Code (lllllllama/RigorPilot-Skills, 497 stars), Repo Development (VectorSpaceLab/AREX-Skill, 328 stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and Fix Art Issues (OpenPipe/ART, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Preference Optimization?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,254 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.