Agent skill

Finetuning Model Onboarding

by overmind-core in overmind-core/overmind

Rules for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior for an existing one — engine-agnostic customization via family hooks instead of if/else in…

AGPL-3.0Auto-check passedAI & LLM Engineering

Install Finetuning Model Onboarding

skills CLI
$ npx skills add overmind-core/overmind --skill finetuning-model-onboarding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install overmind-core/overmind finetuning-model-onboarding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/overmind-core/overmind.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/finetuning-model-onboarding .claude/skills/finetuning-model-onboarding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
finetuning-model-onboarding
GitHub stars
597
Token cost
~3.2k tokens
SKILL.md length
1,418 words
Files
1
Skills in repo
20
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Rules for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior for an existing one — engine-agnostic customization via family hooks instead of if/else in…

  • Works in 7 steps: Never branch on model/family in the… → Chat template must be… → If a model fails on its assigned… → …
  • Onboarding a model to Modal+Unsloth finetuning
  • SKILL.md covers 1. Never branch on…, 2. Chat template must be…, 3. If a model fails on its… and 4. Deriving…, plus 4 more sections
  • Calls modal

What it does

Finetuning Model Onboarding is an agent skill from overmind-core/overmind. Rules for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior for an existing one — engine-agnostic customization via family hooks instead of if/else in the shared script, TRL-compatible chat templates ({% generation %} / pretoktrl), H100-before-H200 GPU probing, populating realmaxcontextlength/validatedcontextlength in models.json, and setting finetuning cost/pricing for a new model. Use when onboarding a model to Modal+Unsloth finetuning, adding a ModelFamily…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Fine-tuning. It works with Qwen. The repository describes itself as: The platform for continuously improving AI agents. The licence is AGPL-3.0.

When your agent uses it

  • Onboarding a model to Modal+Unsloth finetuning
  • Adding a ModelFamily
  • Pretoktrl falls back with a training-compatible chat template error
  • A model fails on its assigned GPU/context length

Example prompts

  • “Use the finetuning-model-onboarding skill to rule for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior…”
  • “/finetuning-model-onboarding”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Never branch on model/family in the shared script
  2. Chat template must be training-compatible (pretok Path A)
  3. If a model fails on its assigned environment, re-derive the GPU, don't patch around it
  4. Deriving real_max_context_length and validated_context_length
  5. Everything model-specific goes in models.json
  6. Cost: usually nothing to add, one marketing floor to set
  7. New ModelFamily needs a frontend icon

What it can do on your machine

Read from SKILL.md and the folder at commit 3dec73c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • modal

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Finetuning Model Onboarding loads about 3.2k tokens when it runs. Until then it costs about 176 tokens; SKILL.md has 1,418 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~176
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from overmind-core/overmind at commit 3dec73c, republished under its AGPL-3.0 licence (© overmind-core). 1,418 words, ~3,151 tokens.

Download SKILL.mdSave it as .claude/skills/finetuning-model-onboarding/SKILL.md (or your agent's skills folder).
name
finetuning-model-onboarding
description
Rules for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior for an existing one — engine-agnostic customization via family hooks instead of if/else in the shared script, TRL-compatible chat templates ({% generation %} / pretok_trl), H100-before-H200 GPU probing, populating real_max_context_length/validated_context_length in models.json, and setting finetuning cost/pricing for a new model. Use when onboarding a model to Modal+Unsloth finetuning, adding a ModelFamily, pretok_trl falls back with a training-compatible chat template error, a model fails on its assigned GPU/context length, or pricing a new model's training cost.

Finetuning model onboarding

All finetuning runs on Modal with Unsloth (overbae/services/sft_assets/engine_unsloth.py). Every model and family in this pipeline goes through the same engine — the rules below exist to keep it that way.

1. Never branch on model/family in the shared script

engine_unsloth.py and overbae/services/finetuning_runner.py/finetuning_policy.py are shared by every model. If a new model needs different behavior:

  • Do not add if model_id == "..." or if family == "..." branches to the engine, runner, or policy.
  • Do add/extend a family hooks module under overbae/services/sft_assets/families/<family>.py, subclassing DefaultHooks (families/__init__.py) and overriding only the hook methods you need (env_overrides, load_kwargs, post_load, device_map, sft_config_overrides, peft_lora_dropout, peft_gradient_checkpointing, fast_model_cls). Export a single hooks = XxxHooks().
  • Register the family in modal_shared/modelfam/families.py as a FamilySpec (patterns, images, hooks_module="families.<family>"), then add it to the ordered _FAMILIES tuple in modal_shared/modelfam/registry.py — specific patterns before general ones (e.g. qwen35/qwen_coder before qwen_mm before qwen).
  • The engine calls hooks uniformly (load_hooks(_spec.hooks_module) then _hooks.<method>(...)) — it never knows which family it's running. That's the contract to preserve.

Reference existing hooks files for scale: qwen35.py (two-line SDPA override) up to gpt_oss.py (monkeypatches + adapter base-model swap) — match the size of the override to the size of the actual quirk, don't build more than the model needs.

2. Chat template must be training-compatible (pretok Path A)

Unsloth labels come from overbae/services/sft_assets/pretok.py. Path A is TRL get_training_chat_template + return_assistant_tokens_mask. That only works if the live tokenizer template already has {% generation %} / {% endgeneration %} markers, or TRL can exact-match it to a bundled original, or we swap in a hand-patched twin.

If the log shows:

diag: pretok_trl unsupported (ValueError('The chat template is not training-compatible (missing prefix-preservation or `{% generation %}` markers) and patching is not supported for this template. ...')); using multi_header fallback

the job may still train, but assistant-only masks are the weaker Path B. Onboarding is not done until Path A succeeds (pretok paths: contains trl_training_template or native_generation_markers, not only multi_header_fallback).

How to check (tokenizer issue — LoRA is enough; skip a second Full run):

  1. Probe the catalog hf_model_id, not the catalog id. Unsloth repos often ship a different chat_template.jinja than the upstream id; TRL and our patches match the live string after from_pretrained.
  2. Tiny Modal LoRA (MAX_STEPS=2, short MAX_LENGTH) is enough. Read runs/<run_id>/train_stdout.log for pretok paths: / pretok_trl unsupported.
  3. One size is not enough when Unsloth templates differ inside a family (Gemma 4 E2B vs 12B, Qwen3-0.6B vs 8B, Qwen3.5-0.8B vs 4B). Hash the Hub template for each hf_model_id you enable.

How to fix — do not branch in engine_unsloth.py / pretok.py:

  1. Save the live Hub template as a base jinja under overbae/services/sft_assets/ (existing dirs: llama_templates/, qwen_templates/, gemma_templates/, …).
  2. Copy it to *_training.jinja and wrap only assistant-generated spans in {% generation %} / {% endgeneration %}. {% generation %} is a real Jinja block: it cannot open inside {% if %} and close after {% endif %}, and {% if %}/{% else %}/{% endif %} must sit wholly inside or wholly outside it. The training file must render byte-identical text to the base; markers are invisible in the rendered string. Qwen3's twins omit the empty <think> block the base inserts on a final assistant turn: that block is not prefix-preserving, and serving keeps thinking off. Compile every twin (see test_every_training_template_compiles).
  3. Register (base, training) in KNOWN_TEMPLATE_PATCHES in overbae/services/sft_assets/training_chat_template.py. pretok.py and engine_unsloth.py both call patch_known_training_template — that is the only dispatch table.
  4. Add a pair assertion in tests/test_sft_training_chat_template.py (or rely on test_every_patch_pair_exists_and_training_has_markers).
  5. sft_assets is baked into the train image (add_local_dir). A pretok/jinja change does nothing until modal deploy of overbae/modal/modal_sft_worker.py. Re-run the LoRA probe after deploy.

3. If a model fails on its assigned environment, re-derive the GPU, don't patch around it

GPU selection for training lives in finetuning_runner.py's _GPU_TABLE / _GPU_TABLE_FULL (params_b → gpu_type/count) plus the _LONG_CONTEXT_H200 bump and any family-specific clamp (e.g. clamp_gemma4_training_gpu — Gemma4 forced onto 1×H200 because Unsloth's device_map="balanced" multi-GPU split is broken for it).

If a new model OOMs or errors on its assigned GPU:

  1. Check whether the failure is architectural (needs a real device_map/env override → family hook, §1) or capacity (needs a different GPU/context ceiling → §4).
  2. If it's a one-off model/family quirk, add a clamp or override scoped to that family only (see the Gemma4 precedent), not a new general rule in the shared table.
  3. Re-run the context-length probe (§4) rather than guessing a new ceiling.

4. Deriving real_max_context_length and validated_context_length

For every new model, before it's usable for finetuning:

  1. Find the real max context length. This is the HF max_position_embeddings (no RoPE/YaRN scaling) — it goes in finetuning.context_length in models.json, distinct from the top-level context_length (published inference window, which may be YaRN-extended).
  2. Probe each enabled training type (full, lora) independently. Use scripts/calibrate_activation_budget.py (--experiment e3 context sweep, spawns real Modal sft_unsloth jobs and reads peak VRAM) to find the largest context that actually trains without OOM.
  3. Try H100 first, then H200. TRAINING_GPU_VRAM_GB = {"H100": 80.0, "H200": 141.0} (finetuning_policy.py) — H100 has less VRAM, so it's the cheaper GPU and must be tried first. Only fall back to H200 if the model can't reach its real max context length on H100.
  4. Record the result per training type in models.json under finetuning.training_type.<full|lora>:
    • context_length: the largest context that trained successfully (equal to finetuning.context_length if the full ceiling was reached, lower otherwise).
    • validated_context_length: true once probed — this flag means "this number came from an actual training run on H100/H200," not an assumption.
  5. Set finetuning.real_max_context_length to the max across validated training types, and keep every training_type.*.context_length <= real_max_context_length.

tests/test_modelfam.py::test_baseten_real_max_context_length enforces this shape for every backend: "baseten" entry — run it before considering onboarding done.

Show full SKILL.md (553 more words)Show less

5. Everything model-specific goes in models.json

Model configuration — context lengths, batch sizes, training type enablement, GPU/VRAM-relevant architecture fields (hidden_size, num_attn_layers, num_kv_heads, head_dim, fp8_supported), pricing, disabled state — belongs in overbae/modal/models.json, not scattered across Python as constants or conditionals. The file's own "comment" field documents each field; read it before adding a new one. If a field doesn't exist yet and is genuinely per-model data (not behavior), add it to the schema there rather than hardcoding it in a script.

6. Cost: usually nothing to add, one marketing floor to set

Actual training cost is computed, not stored per model — overbae/services/finetuning_pricing.py derives it from fields already in models.json:

  • Baseten/Modal: estimate_training_cost() picks GPU count from _BASETEN_GPU_COUNT_TIERS keyed on total_params_b, estimates duration from FLOPs (6·N·D full / 4·N·D LoRA at 35% assumed MFU), and bills GPU-count × minutes × the fixed H100 per-minute rate.
  • Together: training_price_per_million() looks up a $/1M-token rate from _TOGETHER_SFT_TIERS, again keyed on total_params_b.

So as long as the new model's total_params_b is set correctly in models.json, cost estimation works automatically — don't add a new pricing branch or per-model rate to finetuning_pricing.py. The only case it returns None is a >100B-param Together model, which needs an individually negotiated rate (not in this catalog today).

The one thing to add by hand is the marketing "from" floor: models.json's pricing.train_from_usd (paired with pricing.run_from_usd_per_1m_output — model_library.py's _pricing() drops the whole block from the API response unless both are set). This is a display-only number for the model library UI, not read by the cost estimator. Set it by calling estimate_training_run()/estimate_training_cost() for a small representative dataset on the new model and rounding to a customer-facing number consistent with similarly-sized peer models already in the catalog (e.g. dense ~1-4B models cluster around 0.5–0.7, larger dense/MoE tiers step up to 1.0–2.5, frontier-scale up to 4.5–5.0).

7. New ModelFamily needs a frontend icon

If the family has no existing icon mapping, it falls through to a generic simpleicons/placeholder fallback (see model-provider.ts's SIMPLEICONS_SLUG_FIXES, and model-provider-chip.tsx's ProviderLogo fallback chain). To add one:

  1. Add a ProviderId variant and PROVIDER_BY_SLUG entry in frontend/src/components/model-provider.ts (or a PROVIDER_ALIASES entry if it should map onto an existing provider, e.g. Meta/NVIDIA).
  2. If a dedicated icon should render (not the CDN fallback), import the @lobehub/icons component and add it to PROVIDER_ICONS in frontend/src/components/model-provider-chip.tsx.
  3. Extend inferProviderFromModelId() if the model id doesn't carry an explicit provider prefix.

Checklist for onboarding a new model

  • Family resolves correctly (modal_shared/modelfam/registry.py pattern order) — add a FamilySpec only if genuinely new, otherwise reuse
  • Family quirks live in a hooks module, not in engine_unsloth.py/finetuning_runner.py/finetuning_policy.py
  • Chat template is pretok Path A on the hf_model_id tokenizer (log pretok paths: is trl_training_template or native_generation_markers). If TRL cannot auto-patch, add a byte-identical {% generation %} twin and register it in training_chat_template.py KNOWN_TEMPLATE_PATCHES; LoRA-only is enough to verify. Redeploy the SFT worker after jinja/pretok edits
  • finetuning.context_length set to the real HF max position embeddings
  • Context probed per training type, H100 before H200, via calibrate_activation_budget.py
  • training_type.{full,lora}.context_length + validated_context_length set from actual probe results
  • real_max_context_length set and consistent with validated training types
  • tests/test_modelfam.py passes, including test_baseten_real_max_context_length
  • total_params_b set correctly (drives auto-computed training cost — no manual rate needed)
  • pricing.train_from_usd + pricing.run_from_usd_per_1m_output set for the model library display
  • Frontend icon mapped if the family is new
  • Benchmark artifact (overbae/services/benchmarks/data/benchmark_results.json) refreshed by a maintainer once the model is in models.json, so its benchmark results feed model recommendations; the sync runs outside this repo

© overmind-core, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/finetuning-model-onboarding of overmind-core/overmind.

Open the folder on GitHubat commit 3dec73c

Compare with similar skills

Finetuning Model Onboarding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Finetuning Model Onboarding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Finetuning Model Onboarding this skillovermind-core/overmind597—~3.2kAutomated safety check: PassAGPL-3.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
Train SftOpenPipe/ART11k—~2.9kAutomated safety check: PassApache-2.0
slime RL Post-TrainingOrchestra-Research/AI-Research-SKILLs13k5 repos~2.8kAutomated safety check: PassMIT
Qwen21sorryhyun/anima_lora125—~1.9kAutomated safety check: NotesMIT
LlamafactoryPrism-Shadow/penguin-harness2.5k—~855Automated safety check: PassApache-2.0

Similar skills

  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Train Sft

    OpenPipe/ART

    SFT training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • slime RL Post-Training

    Orchestra-Research/AI-Research-SKILLs

    Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

    13k GitHub starsUsed in 5 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen21

    sorryhyun/anima_lora

    Qwen-Image-2.1 LoRA line (NOT Anima) — running cache/train through the daemon, make gui-qwen, the CacheRequest/TrainRequest flag surface and how to add a field, model-dir resolution, cache layout…

    125 GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Llamafactory

    Prism-Shadow/penguin-harness

    Fine-tune LLMs with LlamaFactory — register datasets, train via YAML configs, merge LoRA adapters and serve the result.

    2.5k GitHub stars~855 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Flux Txt2img

    artokun/comfyui-mcp

    Build Flux txt2img workflows with Flux.1 Dev (SRPO), Flux 2 Klein 9B, Turbo LoRAs, FluxGuidance, and DualCLIPLoader patterns

    795 GitHub stars~3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from overmind-core/overmind

All 20 skills in this repo
  • API Endpoints

    overmind-core/overmind

    End-to-end workflow for adding or changing a backend API endpoint — which module the serializer and view belong in, URL registration, OpenAPI client regeneration, and typed consumption from the…

    597 GitHub stars~830 tokensUpdated today
    Auto-check: notes
  • Frontend Design

    overmind-core/overmind

    Overmind Console design system — semantic tokens, shared primitives, geometry and icons, the border-contrast floor, the duplicated table implementations, and the verification scripts.

    597 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • MCP

    overmind-core/overmind

    End-to-end workflow for adding or changing Overmind MCP tools, resources, prompts, authentication, or result contracts — server layers, catalog registration, MCP-impact classification, and required…

    597 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • PR Etiquette

    overmind-core/overmind

    How to open a complete pull request on overmind-core/overmind — the CI gates, the cross-cutting surfaces a change must carry with it (MCP, blast radius, the docs repo), gh pr edit being broken here…

    597 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Seed Demo Data

    overmind-core/overmind

    Run or modify the seeddemo management command (the one-project Support Copilot demo) without breaking the beat-safety invariants that keep celery workers from re-driving seeded rows.

    597 GitHub stars~973 tokensUpdated today
    Auto-check passed
  • Overmind Agent

    overmind-core/overmind

    Inspect an Overmind project's agent map, capabilities, behaviour contracts, repository provenance and evaluation coverage.

    597 GitHub stars~797 tokensUpdated today
    Auto-check passed

Works with

Questions about Finetuning Model Onboarding

What does Finetuning Model Onboarding do?

Rules for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior for an existing one — engine-agnostic customization via family hooks instead of if/else in…. Finetuning Model Onboarding is an agent skill from overmind-core/overmind.json, and setting finetuning cost/pricing for a new model.

When should I use Finetuning Model Onboarding?

Finetuning Model Onboarding fits situations like: onboarding a model to Modal+Unsloth finetuning; adding a ModelFamily; pretoktrl falls back with a training-compatible chat template error; A model fails on its assigned GPU/context length.

How do I install Finetuning Model Onboarding in Claude Code?

Run `npx skills add overmind-core/overmind --skill finetuning-model-onboarding -a claude-code`. Or copy the skill folder (.agents/skills/finetuning-model-onboarding in overmind-core/overmind) into .claude/skills/finetuning-model-onboarding in your project. Claude Code loads it when a task matches its description.

How do I install Finetuning Model Onboarding in Codex?

Run `npx skills add overmind-core/overmind --skill finetuning-model-onboarding -a codex`. Or copy the skill folder (.agents/skills/finetuning-model-onboarding in overmind-core/overmind) into .agents/skills/finetuning-model-onboarding in your project. Codex loads it when a task matches its description.

Can I use Finetuning Model Onboarding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add overmind-core/overmind --skill finetuning-model-onboarding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/finetuning-model-onboarding, .gemini/skills/finetuning-model-onboarding, .github/skills/finetuning-model-onboarding and .opencode/skills/finetuning-model-onboarding in your project.

What does Finetuning Model Onboarding need to run?

Going by SKILL.md and its folder, Finetuning Model Onboarding needs the command-line tools its instructions call (modal). Our summary lists: Python 3.

Does Finetuning Model Onboarding access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Finetuning Model Onboarding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Finetuning Model Onboarding use?

Finetuning Model Onboarding is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Finetuning Model Onboarding use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Finetuning Model Onboarding?

Skills that share tags, products or a category with Finetuning Model Onboarding: Train Rl (OpenPipe/ART, 11k stars), Train Sft (OpenPipe/ART, 11k stars), slime RL Post-Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Qwen21 (sorryhyun/anima_lora, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Finetuning Model Onboarding?

overmind-core (a GitHub organization) maintains it in overmind-core/overmind, which has 597 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 8, 2026.

Source: overmind-core/overmind on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.