Agent skill

Finetuning

by evo-hq in evo-hq/evo

This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Finetuning

skills CLI
$ npx skills add evo-hq/evo --skill finetuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install evo-hq/evo finetuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/evo-hq/evo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/evo/skills/finetuning .claude/skills/finetuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
finetuning
GitHub stars
1.5k
Token cost
~4.5k tokens
SKILL.md length
2,079 words
Files
9 (incl. references)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe…

  • Works in 5 steps: Periodic checkpoint every N steps (e.g.… → Mini-eval after each checkpoint on a… → Early-stop on regression: track best… → …
  • Mentions fine-tuning
  • SKILL.md covers Pick the technique by reward…, Research the literature before…, Before committing the budget:… and Long training: checkpoint,…, plus 9 more sections
  • Needs WANDB_API_KEY and HF_TOKEN

What it does

Finetuning is an agent skill from evo-hq/evo. This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe, reward design, or weight updates. Decision tree by reward shape, smoke-run gate, three failure diagnostics, five false-progress patterns. Provider recipes and I/O contract in references/.

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `references/diagnostics.md`, `references/false-progress.md` and `references/glue.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents. The licence is Apache-2.0.

When your agent uses it

  • Mentions fine-tuning
  • Training recipe

Example prompts

  • “/finetuning”

Requirements

  • A credential in WANDB_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Periodic checkpoint every N steps (e.g. every 0.25 epoch, or every 200 steps — whichever is faster).
  2. Mini-eval after each checkpoint on a small held-out subset (5–10 items, not the full held-out — that's reserved for the final committed…
  3. Early-stop on regression: track best mid-eval score; stop if it hasn't improved in patience checkpoints (typically 2). Don't burn 60 more…
  4. Save the BEST checkpoint, not the last. Early-stop means the current model is probably past its peak; the checkpoint you commit should be…
  5. Log every mid-eval score to your tracker (see ## Stream training metrics live). The user watching the live dashboard sees the trajectory…

What it can do on your machine

Read from SKILL.md and the folder at commit c70c04b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • WANDB_API_KEY
    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Finetuning loads about 4.5k tokens when it runs, and up to ~9k if it reads all its reference files. Until then it costs about 98 tokens; SKILL.md has 2,079 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from evo-hq/evo at commit c70c04b, republished under its Apache-2.0 licence (© evo-hq). 2,079 words, ~4,492 tokens.

Download SKILL.mdSave it as .claude/skills/finetuning/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
finetuning
description
This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe, reward design, or weight updates. Decision tree by reward shape, smoke-run gate, three failure diagnostics, five false-progress patterns. Provider recipes and I/O contract in references/.
evo_version
0.8.0

Finetuning

Priors, not rules. Only firm guardrails: held-out eval you never train on, no leakage, trust evo's recorded numbers over the run's self-report. Override anything else against the gate.

Pick the technique by reward shape

Decide on the reward first, technique second. Choosing the comfortable technique over the matching one is the most common failure.

Reward shapeTechnique
Verifiable (exact match, unit tests, parser-decidable)RL (GRPO / RLOO / PPO) — reward includes format, so the model learns to emit verifier-acceptable shape
Preference pairs (chosen vs rejected)DPO / KTO / ORPO — cheaper than full RL, no rollouts
Demonstrations only (curated traces, chat data)SFT — install format/tone/capability the base lacks
Have a scorer + want SFT stabilityRFT — sample, filter by reward, SFT on survivors

"SFT-then-RL" is not a law. For a competent base model on a verifiable benchmark, RL-from-base often beats SFT-then-RL end-to-end.

Research the literature before the first commit

The decision tree above is the structural prior. The empirical answer for this model on this benchmark usually has a recent paper, blog, or HF Space recipe behind it -- and what beats baseline on a 4B base model in 2026 is not what the agent's pre-training data captures. Before picking the technique for exp_0001 (the first experiment after baseline), invoke evo:ideator with a literature brief:

Task(
    subagent_type="evo:ideator",
    prompt="brief=literature\n"
           "model_family=<e.g. Qwen3-4B-Base, Llama-3.1-8B-Base>\n"
           "benchmark=<name + URL/paper if known>\n"
           "objective=<one line: what beats baseline looks like>\n"
           "constraints=<budget, data sources allowed, gated models forbidden, etc>"
)

The ideator returns ranked proposals with references (arXiv, HF Hub, GitHub, blogs). Read them before picking from the reward-shape table. A paper showing GRPO-from-base works on <model_family> for a similar verifiable benchmark beats applying the table cold.

Run this once before exp_0001, and again whenever the optimize loop hits a plateau (the "stuck across distinct techniques" diagnostic below). Not every subsequent experiment needs a literature pass -- the table + diagnostics carry the rest.

Before committing the budget: smoke-run

Run the full pipeline on ~10 examples for ~1 minute. Must produce: a checkpoint the benchmark can load AND a non-zero eval on a held-out item. If not, the recipe is broken — fix it, don't scale it. dtype mismatch, tokenizer/template drift, OOM at this batch size, empty artifacts dir despite falling loss — all surface on 10 examples. Running longer doesn't surface them differently, just more expensively.

Long training: checkpoint, mid-eval, early-stop in-script

Training for an hour and getting one number at the end is the wrong granularity for evo's tree search. By the time you know the recipe failed, you've spent the budget. Build the verification into the training script, not around it.

Pattern for any training run expected to exceed ~30 min wall-clock:

  1. Periodic checkpoint every N steps (e.g. every 0.25 epoch, or every 200 steps — whichever is faster).
  2. Mini-eval after each checkpoint on a small held-out subset (5–10 items, not the full held-out — that's reserved for the final committed score). Same scorer as the real eval; the model just sees fewer items.
  3. Early-stop on regression: track best mid-eval score; stop if it hasn't improved in patience checkpoints (typically 2). Don't burn 60 more minutes once the trajectory has flattened or reverted.
  4. Save the BEST checkpoint, not the last. Early-stop means the current model is probably past its peak; the checkpoint you commit should be the one that scored highest mid-training, not whatever the trainer happened to leave behind.
  5. Log every mid-eval score to your tracker (see ## Stream training metrics live). The user watching the live dashboard sees the trajectory build up step-by-step instead of staring at the loss curve hoping it transfers.

HuggingFace TRL: implement as a TrainerCallback on on_step_end — save checkpoint, run the mini-eval via vLLM or HF transformers, compare to best_score, set control.should_training_stop = True on stall. Pattern is one ~30-line class.

Keep vLLM warm across mid-evals when you can (one serve process, reload adapter between checkpoints) — cold-starting vLLM every 200 steps adds 5 min of overhead per checkpoint.

Use a tighter mini-eval subset than the full held-out. The mini-eval is a signal, not the score that gets committed. If the mini-eval scores ≥ baseline on its subset, run the full held-out as the eval-gate scoring pass at the end. If it doesn't, early-stop.

This is Pattern B from the design tradeoff with multi-node staging (Pattern A — break the training into multiple committed evo nodes, each a stage). Pattern B keeps the experiment as one evo node with the verification logic inside the script; it's simpler to write and avoids per-stage vLLM spin-up, at the cost of less tree-search introspection. Multi-stage as separate nodes is preferable when you want the orchestrator to be able to branch alternative continuations from any mid-training checkpoint.

Cap retries at training scale

evo run allows up to max_attempts=3 retries per experiment by default. That budget was designed for second-scale benchmarks where retrying after an edit-bug fix is free. At training scale (~hours per attempt), it's the wrong tradeoff — by attempt 2 you've spent more compute than just trying a fresh hypothesis would cost.

For training-heavy workspaces, set the cap to 1 once at init:

bash
evo config set max-attempts 1

One attempt, one shot. Regression → evo discard → new branch from parent with a different hypothesis. This pairs with the in-script early-stop above: each attempt is single-shot, but its internal verification keeps it from burning the budget on a clearly-failing trajectory.

The "fix-and-rerun" retry pattern still applies for sub-minute benchmarks; leave the default max_attempts=3 there.

Four diagnostics

Stuck at 0 on a verifiable benchmark after 2+ SFT runs. Technique class is wrong, not the recipe. Pivot to RL with the verifier as reward; SFT loss can be healthy while the model emits unparseable output.

Base scores below random before any training (knowledge-heavy benchmark). Model lacks the knowledge, not the format. Post-training shapes existing knowledge; it does not install new knowledge. Right axis: continued pre-training on a domain corpus, distillation from a stronger model that has the knowledge, or retrieval-augmented inference.

delta <= 0 across several committed train moves. Method exhausted on this target. Try a different method, change the data, or improve the harness instead of the weights.

Stuck at the same non-zero score across 3+ experiments spanning distinct techniques. When 3+ committed experiments — across structurally different techniques (e.g. SFT, GRPO, RFT) — all land at the same non-zero score, the bottleneck is not the training method. The most common cause is a train↔verifier objective mismatch: the model has learned to emit answers in one format, but the verifier expects a different one. Examples: training data uses \boxed{X} but the verifier prompt requests ANSWER: X (or vice versa); training uses one chat template, eval uses another; training optimizes step-by-step CoT but the verifier wants the answer alone.

Diagnostic action: spot-check 3 training examples and 3 eval-prompt examples side by side. If a perfect-score training example would NOT pass the verifier (or vice versa), the objective is mismatched. Realign the training data format to the verifier's expected output, OR change the eval prompt (if rules allow). Do NOT try a fourth training-technique variant before doing this spot-check.

What never counts as progress

Five patterns produce a number going up without the model improving. See references/false-progress.md for examples + detection.

  1. Training on the held-out set — direct or transitive (public instruction datasets sometimes contain eval-derived items).
  2. Embedding eval items in "synthetic" data, even renamed or paraphrased.
  3. Generating training data conditioned on per-eval-item failure logs.
  4. Submitting a checkpoint you didn't train (off-the-shelf instruct model; parent's checkpoint unchanged).
  5. Training a different objective than the verifier scores.

The verifier should catch these. List is here so the train move doesn't produce them.

Show full SKILL.md (849 more words)Show less

Surviving session compaction

Write the dataset URL, method choice, user-imposed constraints, and hyperparameters you converged on to methodlog.md in the experiment worktree. One line each. Re-read after any context reset, before the next train move. Prevents silent dataset swaps between experiments and re-running ablations.

Numbers that matter (in order)

  1. A reward you trust — verifiable beats a learned reward (which gets hacked).
  2. A held-out eval you never train on.
  3. On-policy freshness for RL — train on current policy's samples, not stale ones.
  4. LoRA LR ~10x full-FT; rank 32 is a fine default. LoRA ~ full-FT for RL and small-data SFT; lags on large SFT.

Method/provider-specific numbers (LR, KL, group size) live in the recipe under references/.

Stream training metrics live

A long training run is observability-blind until the experiment commits — without a live tracker, nobody can tell if loss is converging, if the GPU is idle, or if the recipe is silently broken. They get one number at the end. Wire a tracker into the training script by default.

Detection prior — apply when the corresponding env var is set, skip otherwise. Don't install a tracker the user didn't opt into:

Env varTrackerTRL one-liner
WANDB_API_KEYwandbSFTConfig(report_to="wandb")
TRACKIO_SPACE_IDtrackio (wandb-compatible OSS, logs to a public HF Space)SFTConfig(report_to="trackio")
MLFLOW_TRACKING_URImlflowSFTConfig(report_to="mlflow")
(none set)nonetrain without a tracker; don't invent one

For custom training loops, use tracker.init(project=..., name=f"exp_{exp_id}") + tracker.log({"loss": ..., "step": ...}) — concrete patterns in references/observability.md.

Use EVO_EXPERIMENT_ID as the run name so each experiment shows up as its own line in the tracker dashboard. The same env detection applies to HuggingFace datasets / Hub uploads: if HF_TOKEN is set, treat gated datasets and private Hub pushes as available.

Warm-start from a parent / prior checkpoint

When the orchestrator branches an experiment from a committed or preserved checkpoint with evo new --from-artifact <exp[:label]>, evo exposes that artifact's path to your recipe as EVO_SEED_ARTIFACT (and, for back-compat, the same value as EVO_PARENT_POLICY). Warm-start from it rather than re-training from base — re-training from base every time burns the budget on duplicated work and stops the tree from accumulating capability across generations. To make a run reusable this way you must DECLARE your checkpoint as an artifact: write it to EVO_CHECKPOINT_DIR and name it in the benchmark result's artifacts field (full contract in references/glue.md). Only declared artifacts are preserved on discard and seedable via --from-artifact.

Concrete pattern:

python
seed = os.environ.get("EVO_SEED_ARTIFACT") or os.environ.get("EVO_PARENT_POLICY")
if seed and os.path.exists(seed):
    print(f"warm-starting from {seed}")
    model = AutoModelForCausalLM.from_pretrained(seed, ...)
else:
    print("no seed; loading base")
    model = AutoModelForCausalLM.from_pretrained(BASE_MODEL, ...)

Override only when the brief explicitly asks for a fresh-from-base ablation. The full I/O contract is in references/glue.md.

Configure for training, not inference. Put the whole training computation on the accelerator you're training on, and don't enable inference-oriented conveniences for a training run. Auto device-mapping / model-sharding / CPU-offload exist to fit oversized models for inference by spreading or offloading layers; inside a training step they either break the backward pass or silently fall back to slower memory — so training still "runs" but crawls, with no error to surface the problem (the most dangerous case: it looks like it's working). Shard only when the model genuinely doesn't fit one device, and then use the framework's training parallelism path, not an inference placement shortcut. Same logic for other inference-mode defaults that leak into training (eval-mode quantization, kv-cache, dropout off). (Concrete instance — HuggingFace: load with device_map={"": 0} / .to("cuda"), never device_map="auto", which errors with a meta-device gradient mismatch or offloads to CPU at a large slowdown; for real multi-GPU use accelerate/FSDP/DDP.)

Cache expensive intermediates

LoRA adapters, filtered/curated datasets, tokenized datasets, computed embeddings, generated rollouts -- expensive to produce, large, and gitignored. They don't ride the experiment branch. They also don't have to be rebuilt per experiment.

Write expensive artifacts to a stable, workspace-level path; check for them first, compute only on miss. Subsequent experiments (siblings, descendants, or re-runs of the same experiment after a worktree clean) read the same path.

Convention: under .evo/cache/, sibling to run_<NNNN>/. Already gitignored (via .evo/ in the workspace's git excludes). Survives across runs -- it's not nested inside any run_<id>/, so evo new/evo run/evo reset don't touch it.

Pattern:

python
import os
from pathlib import Path
# walk up from cwd to find the workspace root (the dir that has .evo/)
def _workspace_root() -> Path:
    p = Path.cwd().resolve()
    for d in [p, *p.parents]:
        if (d / ".evo").is_dir():
            return d
    raise RuntimeError("not inside an evo workspace")

cache = _workspace_root() / ".evo" / "cache" / "datasets"
cache.mkdir(parents=True, exist_ok=True)
# Cache key embeds every input that changes the artifact: dataset name,
# filter recipe version, tokenizer, max length, etc. Different recipe ->
# different key, so a sibling experiment with a different filter keeps
# its own cache without trampling yours.
key = cache / "numina-cot-r1-filter-v2-qwen3-tok-3072.arrow"
if key.exists():
    ds = datasets.Dataset.load_from_disk(str(key))
else:
    ds = build_and_filter_dataset()
    ds.save_to_disk(str(key))

High-value caches (not exhaustive): curated/tokenized training corpora (tokenization is the slow part on millions of rows); LoRA adapters produced by prior experiments that a sibling might warm-start from (the parent path is already handled by EVO_PARENT_POLICY above; this is for sibling-reachable named adapters); computed embeddings, retrieval indexes, precomputed eval-time generations.

Don't duplicate the HuggingFace Hub cache (~/.cache/huggingface/). That handles from_pretrained downloads automatically and is user-level, already shared across all experiments.

Anti-pattern: writing the artifact inside the experiment's worktree (<worktree>/some_cache/). Worktrees are gitignored for these files, the artifact doesn't propagate to descendants via the git tree, and a worktree clean / gc removes it. Use the workspace-level .evo/cache/ instead.

A first-class named registry (evo asset put/get/list/use) for these is tracked in issue #55. The path convention above is the lightweight version anyone can adopt today.

References

Pull via Read tool when the trigger applies. Tree organized by category -- core contracts first, then provider-specific recipes under rl/, sft/, serving/.

finetuning/references/
│
├── glue.md             writing train.py -- I/O contract evo expects.
│                       Read FIRST when starting any training code.
├── trace-schema.md     TrainingTrace JSON shape (per-step train trace fields)
├── diagnostics.md      held_out_score / delta / reward_saturation /
│                       generalization_gap -- read when interpreting a result
├── false-progress.md   the five patterns + how to detect them.
│                       Read when a score improves implausibly fast or
│                       breaks the smoke gate.
├── observability.md    wandb / trackio / mlflow wiring -- env-driven detection,
│                       TRL report_to options, custom-loop patterns.
│                       Read when writing a training script.
│
├── rl/                 RL framework recipes (rollouts + reward + policy update)
│   └── art.md          ART (Algorithm-Refined Training)
│
├── sft/                SFT framework recipes
│   └── tinker.md       Tinker SFT runner
│
└── serving/            Eval-time inference framework references
    └── vllm.md         vLLM serving config + LoRA-multi (load multiple
                        adapters in one server -- saves cold-start per experiment)

Cross-skill references also worth pulling during finetuning work:

  • discover/references/sdk_python.py / sdk_node.js -- wiring per-task instrumentation in the benchmark
  • discover/references/inline_instrumentation.py -- inline fallback when SDK can't be used (copy as-is)
  • references/evo-wait.md -- waiting for training / eval without burning context

© evo-hq, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in plugins/evo/skills/finetuning of evo-hq/evo.

  • SKILL.md
  • references/diagnostics.md
  • references/false-progress.md
  • references/glue.md
  • references/observability.md
  • references/rl/art.md
  • references/serving/vllm.md
  • references/sft/tinker.md
  • references/trace-schema.md

Open the folder on GitHubat commit c70c04b

Compare with similar skills

Finetuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Finetuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Finetuning this skillevo-hq/evo1.5k—~4.5kAutomated safety check: PassApache-2.0
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Dataset Evaluationawslabs/agent-plugins9152 repos~1.3kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from evo-hq/evo

  • Discover

    evo-hq/evo

    Initialize evo for the current repository by exploring the codebase, proposing unexplored optimization dimensions, constructing the benchmark inside a baseline worktree, and running the first…

    1.5k GitHub stars~12k tokensUpdated 3 days ago
    Auto-check: notes
  • Report

    evo-hq/evo

    Read-only evo run reporting. An agent skill from evo-hq/evo.

    1.5k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Infra Setup

    evo-hq/evo

    Non-user-invocable provider/setup reference for evo backend switching, prerequisite checks, and auth/install guidance.

    1.5k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Ship

    evo-hq/evo

    Land the winning experiment from an evo run as a clean, mergeable change -- open a PR when the repo has a remote, otherwise merge into the working branch.

    1.5k GitHub stars~1.7k tokensUpdated 3 days ago
    Auto-check passed
  • Subagent

    evo-hq/evo

    Protocol that evo optimization subagents follow when dispatched from /optimize.

    1.5k GitHub starsUsed in 1 repo~7.1k tokens
    Auto-check: notes
  • Optimize

    evo-hq/evo

    Drive structured autoresearch iteration after evo:discover and the baseline commit.

    1.5k GitHub stars~13k tokensUpdated 3 days ago
    Auto-check passed

Questions about Finetuning

What does Finetuning do?

This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe…. Finetuning is an agent skill from evo-hq/evo. This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe, reward design, or weight updates.

When should I use Finetuning?

Finetuning fits situations like: mentions fine-tuning; training recipe.

How do I install Finetuning in Claude Code?

Run `npx skills add evo-hq/evo --skill finetuning -a claude-code`. Or copy the skill folder (plugins/evo/skills/finetuning in evo-hq/evo) into .claude/skills/finetuning in your project. Claude Code loads it when a task matches its description.

How do I install Finetuning in Codex?

Run `npx skills add evo-hq/evo --skill finetuning -a codex`. Or copy the skill folder (plugins/evo/skills/finetuning in evo-hq/evo) into .agents/skills/finetuning in your project. Codex loads it when a task matches its description.

Can I use Finetuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add evo-hq/evo --skill finetuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/finetuning, .gemini/skills/finetuning, .github/skills/finetuning and .opencode/skills/finetuning in your project.

What does Finetuning need to run?

Going by SKILL.md and its folder, Finetuning needs credentials named WANDB_API_KEY and HF_TOKEN. Our summary lists: A credential in WANDB_API_KEY.

Does Finetuning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Finetuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Finetuning use?

Finetuning is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Finetuning use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Finetuning?

Skills that share tags, products or a category with Finetuning: Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Sentence-Transformers Training Router (huggingface/skills, 11k stars) and Dataset Evaluation (awslabs/agent-plugins, 915 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Finetuning?

evo-hq (a GitHub organization) maintains it in evo-hq/evo, which has 1,466 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 5, 2026.

Source: evo-hq/evo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.