Sentence-Transformers Training Router
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters.
$ npx skills add wshobson/agents --skill lora-qlora-recipes -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents lora-qlora-recipes --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/lora-qlora-recipes .claude/skills/lora-qlora-recipes && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "lora-qlora-recipes" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/lora-qlora-recipes into .claude/skills/lora-qlora-recipes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lora-qlora-recipes", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/lora-qlora-recipesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill lora-qlora-recipes -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents lora-qlora-recipes --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/llm-finetuning/skills/lora-qlora-recipes .agents/skills/lora-qlora-recipes && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "lora-qlora-recipes" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/lora-qlora-recipes into .agents/skills/lora-qlora-recipes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lora-qlora-recipes", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill lora-qlora-recipes -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents lora-qlora-recipes --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/llm-finetuning/skills/lora-qlora-recipes .cursor/skills/lora-qlora-recipes && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "lora-qlora-recipes" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/lora-qlora-recipes into .cursor/skills/lora-qlora-recipes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lora-qlora-recipes", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/llm-finetuning/skills/lora-qlora-recipes--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill lora-qlora-recipes -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents lora-qlora-recipes --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/llm-finetuning/skills/lora-qlora-recipes .gemini/skills/lora-qlora-recipes && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "lora-qlora-recipes" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/lora-qlora-recipes into .gemini/skills/lora-qlora-recipes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lora-qlora-recipes", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents lora-qlora-recipesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill lora-qlora-recipes -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/llm-finetuning/skills/lora-qlora-recipes .github/skills/lora-qlora-recipes && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "lora-qlora-recipes" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/lora-qlora-recipes into .github/skills/lora-qlora-recipes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lora-qlora-recipes", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill lora-qlora-recipes -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents lora-qlora-recipes --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/llm-finetuning/skills/lora-qlora-recipes .opencode/skills/lora-qlora-recipes && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "lora-qlora-recipes" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/lora-qlora-recipes into .opencode/skills/lora-qlora-recipes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lora-qlora-recipes", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
lora-qlora-recipesConfigure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters.
Lora Qlora Recipes is an agent skill from wshobson/agents. Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters. Use when writing or reviewing a LoRA/QLoRA training configuration, choosing rank/alpha/target modules, or deciding between LoRA, QLoRA, and full fine-tuning.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/hyperparameters.md` and `references/unsloth-trl-mapping.md`).
It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Lora Qlora Recipes loads about 1.9k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 925 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 925 words, ~1,907 tokens.
.claude/skills/lora-qlora-recipes/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.This skill assumes the routing decision already
happened — finetuning-method-selection should
have already pointed here because the data shape
is demonstrations (SFT), not preference pairs or
a verifiable reward signal. What follows is the
current best-practice recipe for configuring the
adapter itself: which modules to target, how to
size rank and alpha, what learning rate to use,
and when QLoRA buys real headroom versus when it
just adds risk. Dataset preparation and quality
checks are a separate concern — see
dataset-curation.
Input: a routing decision (SFT via LoRA/
QLoRA) plus a target size class.
Output format: a validated adapter config —
the kwarg values below, not free-form advice —
that llm-finetuning-training-engineer consumes
directly when it generates a runnable script.
The reference recipe is "LoRA Without Regret" (Thinking Machines/Schulman, 2025-09), now the settled convention for LoRA/QLoRA SFT.
Target all-linear modules, not just attention:
target_modules = [
"q_proj", "k_proj", "v_proj", "o_proj", # attention
"gate_proj", "up_proj", "down_proj", # MLP — matters most
]The MLP layers (gate_proj, up_proj,
down_proj) matter most — attention-only
targeting was the older, weaker convention.
Dropping modules to save memory is a Failure
Mode below, not a valid optimization.
lora_alpha = 2 * r is the settled
convention (NeurIPS 2025 "intruder dimensions"
result). Don't hand-tune alpha independently of
rank — derive it from rank every time.references/hyperparameters.md.Rank is task-shaped, not a single global default:
| Task | Rank |
|---|---|
| RL (GRPO/RLVR adapters) | 1–32 |
| General default | 16–32 |
| SFT at scale | up to ~256 |
Higher rank isn't automatically better — it raises capacity to memorize as fast as it raises capacity to generalize. Start at the row matching the task, and only move up a row if the lower rank measurably underfits on held-out eval, not as a default hedge.
Keep effective batch size under 32. This recipe was validated at that scale — pushing effective batch higher is an untested extrapolation, not a free throughput win.
Unsloth is the reference implementation this
plugin assumes as the default fast path — except
for messages-shaped conversational SFT with
assistant_only_loss=True, where Unsloth
2026.7.x's compiled trainer has no messages-shaped
path at all and the plain-TRL escape hatch
(references/unsloth-trl-mapping.md) is the
default for that combination, not a rare-regression
fallback. Its out-of-the-box defaults, and why
each one is set that way:
lora_dropout=0 — the optimized kernel
path assumes zero dropout; setting a nonzero
value forfeits the fused-kernel speedup.bias="none" — bias terms add adapter
parameters for negligible quality gain at this
rank range.use_gradient_checkpointing="unsloth" —
Unsloth's checkpointing variant, not vanilla HF
checkpointing; saves roughly 30% VRAM over
no checkpointing.optim="adamw_8bit" — 8-bit AdamW cuts
optimizer-state memory with negligible quality
impact at LoRA/QLoRA adapter scale.random_state fixed — pins LoRA
initialization for reproducibility across runs;
treat it like any other seed, not a tunable.These show up together on the get_peft_model
call:
model = FastLanguageModel.get_peft_model(
model,
r=32,
target_modules=target_modules,
lora_alpha=64, # 2 * r
lora_dropout=0,
bias="none",
use_gradient_checkpointing="unsloth",
random_state=3407,
)Exact kwarg names and their plain-TRL/PEFT
equivalents, plus a full worked config including
SFTConfig: references/unsloth-trl-mapping.md
and references/hyperparameters.md.
| Situation | Default choice |
|---|---|
| Adapting behavior on demonstrations | LoRA |
| Base model doesn't fit in bf16 at target rank | QLoRA |
| Injecting dense new domain knowledge | Full FT (see finetuning-method-selection) |
| Unsure which one | LoRA — upgrade to QLoRA only if memory forces it |
dgx-spark-ops plugin's
spark-memory-thermal-ops skill covers the
full OOM remediation ladder (bf16 LoRA is the
next thing to try, not a further QLoRA
shrink).fp16 divergence on non-BF16 GPUs. Training
in fp16 on hardware that doesn't have solid
BF16 support is a known source of loss spikes
and silent divergence. Force bf16=True
wherever the hardware supports it; don't fall
back to fp16 as if it were equivalent. Check
hardware support before picking a dtype:
python -c "import torch; print(torch.cuda.is_bf16_supported())"Rank too high on a small dataset overfits. A rank picked for "SFT at scale" (up to ~256) on a dataset that doesn't have scale behind it memorizes rather than generalizes. Match rank to the Rank by Task table above, not to the largest number available.
Removing target modules to save memory costs
quality for negligible savings. The adapter
parameters on gate_proj/up_proj/down_proj
are a small fraction of total model size — cutting
them barely moves memory but measurably hurts
quality. If memory is tight, move to QLoRA or
reduce rank/batch/pack length before trimming
target modules.
All three failure modes share a pattern: they look like a training-loop bug (loss spikes, plateaus, memorization) but are actually a config choice that contradicts the reference recipe above. Check configuration against this skill before debugging the training loop itself.
references/hyperparameters.md — full rank/
alpha/LR tables by task type, rsLoRA notes,
batch/packing interactions, and a complete
worked Unsloth config block.references/unsloth-trl-mapping.md — every
Unsloth kwarg mapped to its TRL/PEFT
equivalent, current TRL API notes, and the
escape-hatch rule for when to drop back to
plain TRL.Related skills: finetuning-method-selection
routes here; dataset-curation covers the data
side this skill doesn't; llm-finetuning-training-engineer
is the downstream consumer of the config this
skill produces.
© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in plugins/llm-finetuning/skills/lora-qlora-recipes of wshobson/agents.
Open the folder on GitHubat commit 46891e7
Lora Qlora Recipes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Lora Qlora Recipes this skillwshobson/agents | 40k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Sentence-Transformers Training Routerhuggingface/skills | 11k | 1 repos | ~2.6k | Automated safety check: Pass | Apache-2.0 | |
| Train RlOpenPipe/ART | 11k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~830 | Automated safety check: Pass | Apache-2.0 | |
| Dataset Evaluationawslabs/agent-plugins | 916 | 1 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Train SftOpenPipe/ART | 11k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
OpenPipe/ART
RL training reference for the ART framework. An agent skill from OpenPipe/ART.
R6410418/Jackrong-llm-finetuning-guide
Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.
awslabs/agent-plugins
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).
OpenPipe/ART
SFT training reference for the ART framework. An agent skill from OpenPipe/ART.
Orchestra-Research/AI-Research-SKILLs
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
Categories
Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters. Lora Qlora Recipes is an agent skill from wshobson/agents. Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters.
Lora Qlora Recipes fits situations like: reviewing a LoRA/QLoRA training configuration; choosing rank/alpha/target modules; deciding between LoRA; full fine-tuning.
Run `npx skills add wshobson/agents --skill lora-qlora-recipes -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/lora-qlora-recipes in wshobson/agents) into .claude/skills/lora-qlora-recipes in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill lora-qlora-recipes -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/lora-qlora-recipes in wshobson/agents) into .agents/skills/lora-qlora-recipes in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill lora-qlora-recipes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lora-qlora-recipes, .gemini/skills/lora-qlora-recipes, .github/skills/lora-qlora-recipes and .opencode/skills/lora-qlora-recipes in your project.
Going by SKILL.md and its folder, Lora Qlora Recipes needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Lora Qlora Recipes is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Lora Qlora Recipes: Sentence-Transformers Training Router (huggingface/skills, 11k stars), Train Rl (OpenPipe/ART, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Dataset Evaluation (awslabs/agent-plugins, 916 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.