Agent skill

Omh Model Finetuning

by rlaope in rlaope/oh-my-hermes

[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and…

MITAuto-check passedAI & LLM Engineering

Install Omh Model Finetuning

skills CLI
$ npx skills add rlaope/oh-my-hermes --skill omh-model-finetuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rlaope/oh-my-hermes omh-model-finetuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rlaope/oh-my-hermes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/omh-model-finetuning .claude/skills/omh-model-finetuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
omh-model-finetuning
GitHub stars
3.2k
Token cost
~2.4k tokens
SKILL.md length
1,247 words
Files
2 (incl. references)
Skills in repo
143
Repo updated
First seen
Licence
MIT

At a glance

[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and…

  • The user says: model-finetuning
  • SKILL.md covers Why This Exists, First Steps, Do Not Use When and Examples, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Model finetuning

What it does

Omh Model Finetuning is an agent skill from rlaope/oh-my-hermes. [omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval. Use when the user says: model-finetuning, model finetuning, model fine-tuning, fine-tune a model, fine-tune the model, fine tune a model, fine tune the model, fine-tune.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/finetuning-method.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: All in one plugin for Hermes Agent ⚚ the coding intelligence, a long-term memory system and model optimized workflow packages. The licence is MIT.

When your agent uses it

  • The user says: model-finetuning
  • Model finetuning
  • Model fine-tuning
  • Fine-tune a model

Example prompts

  • “/omh-model-finetuning”

What it can do on your machine

Read from SKILL.md and the folder at commit 41de9dc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Omh Model Finetuning loads about 2.4k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 1,247 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rlaope/oh-my-hermes at commit 41de9dc, republished under its MIT licence (© rlaope). 1,247 words, ~2,390 tokens.

Download SKILL.mdSave it as .claude/skills/omh-model-finetuning/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
omh-model-finetuning
description
[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval. Use when the user says: model-finetuning, model finetuning, model fine-tuning, fine-tune a model, fine-tune the model, fine tune a model, fine tune the model, fine-tune.

Model Finetuning

This is a Hermes-native model-finetuning workflow skill.

Why This Exists

model-finetuning exists because producing a model had no owner: model-optimization onboards a model into OMH, inference-serving serves one that exists, and llm-app-dev builds on top of one, while an SFT or DPO question reached workflow-learning with no baseline comparison and no way to answer that training is not needed.

First Steps

  • Ask what the untuned model and the best prompt score on the held-out examples before discussing any method.
  • Ask what shape the data has -- demonstrations, preference pairs, or answers a program can check.

Do Not Use When

  • The ask is onboarding a new model generation into OMH's routing, calibration, or pricing; use model-optimization.
  • The ask is serving an existing model behind an endpoint or benchmarking that endpoint; use inference-serving.
  • The ask is building an application on top of a hosted model -- RAG, structured output, prompt versions; use llm-app-dev.
  • The ask is learning from an OMH run, a missed route, or a skill improvement candidate; use workflow-learning.

Examples

Good example:

  • Prompt: should we fine-tune a model for our support replies or is a better prompt enough
  • Expected behavior: Ask for the held-out examples and the current prompt's score, try few-shot and retrieval against them, and return do_not_finetune if one closes the gap; otherwise pick SFT from the reply demonstrations and gate promotion on beating the untuned baseline.
  • Why: Most prompt-shaped gaps close without training, and training first hides that the cheaper fix was enough.

Bad example:

  • Prompt: the fine-tuned checkpoint got 0.82 on our eval, ship it
  • Expected behavior: Refuse to promote on a standalone score: run the same held-out eval on the untuned baseline and compare before promotion.
  • Why: A score with no baseline cannot show the training helped at all.

Completion Checklist

  • The fine-tune decision is stated, and do_not_finetune was considered first.
  • The method is chosen from the data's shape and its failure mode is named.
  • The held-out split was drawn before training and checked for overlap.
  • Promotion cites the same eval observed on the untuned baseline and the candidate.
  • OMH ran nothing, and every score cites observed output or is marked unverified.

Recovery Notes

  • If no held-out eval exists, building one is the first step; say so before any training plan.
  • If the candidate does not beat the baseline, keep the baseline serving and report the gap rather than retraining blindly.

Workflow Lane

  • Current lane: Research and company ops (product-docs, source-finder, web-research, research, model-optimization, inference-serving, model-finetuning, research-brief, +20 more) - research, signals, ops, and briefings.
  • If intent belongs to another lane, hand back to oh-my-hermes or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: omh-routing/references/skill-common-rail.md.

Use When

Use when someone wants to fine-tune a model on their own data, or is deciding whether to: supervised fine-tuning (SFT), preference tuning (DPO), reinforcement learning from verifiable rewards (RLVR), or a LoRA adapter. The output is a decision on whether to train at all, the method chosen from the data available, a training data plan with a held-out split, a comparison against the untuned baseline, and a checkpoint promotion gate; OMH trains nothing and runs no eval.

Strong routing signals: `model-finetuning`, `model finetuning`, `model fine-tuning`, `fine-tune a model`, `fine-tune the model`, `fine tune a model`, `fine tune the model`, `fine-tune`, `fine tune`, `fine-tuning`, `fine tuning`, `fine-tuned model`, `fine-tuned checkpoint`, `finetune`, `finetuning`, `sft`, `supervised fine-tuning`, `dpo`, `direct preference optimization`, `rlvr`, `verifiable rewards`, `lora`, `qlora`, `lora adapter`, `preference data`, `untuned baseline`, `held-out eval`

Catalog Metadata

Category: planning Phase: model-finetuning Hermes role: planner Quality tier: baseline-comparison-gated Reasoning demand: standard

Quality bar:

  • Measure the untuned model and the cheaper fixes on the held-out eval before proposing any training.
  • Load references/finetuning-method.md for the decision ladder, the method table, the data checklist, and the promotion procedure instead of recalling them.
  • Choose the method from the data that exists, not from the method that is fashionable.
  • Compare every candidate against the untuned baseline on the same eval, never against its own previous run alone.
  • Keep prepared, trained, evaluated, and promoted as separate states for every checkpoint.

Handoff policy:

Keep the fine-tune decision, the method choice, the data plan, the baseline comparison, and the promotion gate in Hermes. Losses, eval scores, and comparisons are recorded only from executor, operator, or wrapper observed output; OMH never launches a training run, calls a model, or runs an eval.

Required inputs:

  • the task the model fails at, and the held-out examples that show the failure
  • what prompting, few-shot examples, or retrieval were already tried, and what they scored
  • the data available: demonstrations, ranked or paired preferences, or answers a program can check
  • the base model, its license, and the compute or provider the operator will train on
  • observed eval scores for the untuned baseline and every candidate checkpoint
Show full SKILL.md (456 more words)Show less

Expected outputs:

  • finetune_decision/v1
  • training_method_choice/v1
  • training_data_plan/v1
  • baseline_comparison/v1
  • checkpoint_promotion_gate/v1

Artifact expectations:

  • finetune_decision/v1 names the measured gap on the held-out eval and what prompting, few-shot, and retrieval scored against it, and returns do_not_finetune when one of them closes the gap -- a complete outcome, reached before any training step
  • training_method_choice/v1 picks SFT, DPO, or RLVR from the shape of the data -- demonstrations, preference pairs, or a verifiable reward -- and full weights or an adapter, and names the failure mode of the method chosen
  • training_data_plan/v1 names each source and its license, the dedupe and filtering, and a held-out split drawn before training and checked for overlap with the training set
  • baseline_comparison/v1 runs the same held-out eval on the untuned baseline and each candidate, per metric, plus a regression check on general capability the task does not cover
  • checkpoint_promotion_gate/v1 promotes a checkpoint only when it beats the untuned baseline by a stated margin with no regression past a stated tolerance, and otherwise keeps the baseline serving

Safety rules:

  • Decide whether to fine-tune before any training step; do_not_finetune is a complete answer when prompting or retrieval closes the measured gap.
  • Never promote a checkpoint on a standalone score; promotion needs the same held-out eval observed on the untuned baseline and on the candidate.
  • Draw the held-out split before training and keep it out of the training data; an eval the model trained on proves nothing.
  • Do not train on data whose license or consent does not permit it.
  • OMH trains nothing, calls no model, and runs no eval; every loss, score, and comparison comes from observed output or is marked unverified.

Runtime Evidence

Preferred harness for this skill: research.

sh
omh runtime record --skill model-finetuning --harness research --status started

Record observed delegation results; otherwise return not_available or not_observed. Prepared OMH routing is not execution, review, CI, merge-readiness, or merge evidence.

  • Treat wrapper memory/context summaries as advisory local context, not proof of opaque Hermes memory reads or changes. Preserve workflow intent and stop conditions; verify before claiming completion. Reply in the user's own words and the host's own voice: its SOUL.md persona owns reply language, tone, speech level, and sentence endings, progress updates included (where it sets no language, use the one the user wrote in), and OMH shapes structure and content only; OMH's record terms (surface, lane, wrapper, handoff, evidence boundary, not_observed) stay in records and tool calls, never in the sentence the user reads unless they ask about one; and when a stop condition or a decision the user owns ends the turn, offer the next action as a question rather than declaring what will not be done.

Use Hermes-native subagent/delegation features when available: native subagents -> Hermes delegation when available, otherwise sequential lanes.

Shared product, compatibility, topology, memory, harness, and execution rules: omh-routing/references/skill-common-rail.md. Load it when applicable; otherwise name an unavailable capability.

© rlaope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/omh-model-finetuning of rlaope/oh-my-hermes.

  • SKILL.md
  • references/finetuning-method.md

Open the folder on GitHubat commit 41de9dc

Compare with similar skills

Omh Model Finetuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Omh Model Finetuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Omh Model Finetuning this skillrlaope/oh-my-hermes3.2k—~2.4kAutomated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.0
Dataset Evaluationawslabs/agent-plugins9151 repos~1.3kAutomated safety check: PassApache-2.0
Train SftOpenPipe/ART11k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Sft

    OpenPipe/ART

    SFT training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Fine Tuning With Trl

    Orchestra-Research/AI-Research-SKILLs

    Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.

    13k GitHub starsUsed in 6 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from rlaope/oh-my-hermes

All 143 skills in this repo
  • Omh Accessibility Audit

    rlaope/oh-my-hermes

    [omh] Screen-reader or keyboard accessibility gaps: prepare WCAG, keyboard, focus, screen-reader, target-size, and reflow evidence gates for UI surfaces.

    3.2k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Omh Agent Evaluation

    rlaope/oh-my-hermes

    [omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics.

    3.2k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Omh Agent Instructions

    rlaope/oh-my-hermes

    [omh] Agent instruction file for a repo -- AGENTS.md, CLAUDE.md, a Cursor rule: write or update what an agent cannot derive from the code, inside a marked region, with every command verified or…

    3.2k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Omh Agent Ops Review

    rlaope/oh-my-hermes

    [omh] AI agent progress for managers: help managers inspect AI-agent progress, blockers, quality gates, and throughput levers.

    3.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Omh AI Slop Cleaner

    rlaope/oh-my-hermes

    [omh] Messy or AI-generated code to clean up: delete AI-generated slop, dead code, and duplication while observable behavior stays identical.

    3.2k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Omh App Debugging

    rlaope/oh-my-hermes

    [omh] Application code misbehaves -- a wrong value, a flaky test, a lost update: reproduce it first, form competing hypotheses, discriminate them with the cheapest observation, and only then fix the…

    3.2k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed

Questions about Omh Model Finetuning

What does Omh Model Finetuning do?

[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and…. Omh Model Finetuning is an agent skill from rlaope/oh-my-hermes. [omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval.

When should I use Omh Model Finetuning?

Omh Model Finetuning fits situations like: the user says: model-finetuning; model finetuning; model fine-tuning; fine-tune a model.

How do I install Omh Model Finetuning in Claude Code?

Run `npx skills add rlaope/oh-my-hermes --skill omh-model-finetuning -a claude-code`. Or copy the skill folder (skills/omh-model-finetuning in rlaope/oh-my-hermes) into .claude/skills/omh-model-finetuning in your project. Claude Code loads it when a task matches its description.

How do I install Omh Model Finetuning in Codex?

Run `npx skills add rlaope/oh-my-hermes --skill omh-model-finetuning -a codex`. Or copy the skill folder (skills/omh-model-finetuning in rlaope/oh-my-hermes) into .agents/skills/omh-model-finetuning in your project. Codex loads it when a task matches its description.

Can I use Omh Model Finetuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rlaope/oh-my-hermes --skill omh-model-finetuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/omh-model-finetuning, .gemini/skills/omh-model-finetuning, .github/skills/omh-model-finetuning and .opencode/skills/omh-model-finetuning in your project.

What does Omh Model Finetuning need to run?

SKILL.md names no scripts, command-line tools or credentials: Omh Model Finetuning is instructions for the agent only.

Does Omh Model Finetuning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Omh Model Finetuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Omh Model Finetuning use?

Omh Model Finetuning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Omh Model Finetuning use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 708 tokens, read only when the agent opens those files.

What are the alternatives to Omh Model Finetuning?

Skills that share tags, products or a category with Omh Model Finetuning: Sentence-Transformers Training Router (huggingface/skills, 11k stars), Train Rl (OpenPipe/ART, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Dataset Evaluation (awslabs/agent-plugins, 915 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Omh Model Finetuning?

rlaope (a GitHub user) maintains it in rlaope/oh-my-hermes, which has 3,233 GitHub stars. The repository holds 143 skills in this directory. The repository was last updated on October 8, 2026.

Source: rlaope/oh-my-hermes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.