Agent skill

Lora Qlora Recipes

by wshobson in wshobson/agents

Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters.

MITAuto-check passedAI & LLM Engineering

Install Lora Qlora Recipes

skills CLI
$ npx skills add wshobson/agents --skill lora-qlora-recipes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents lora-qlora-recipes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/lora-qlora-recipes .claude/skills/lora-qlora-recipes && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
lora-qlora-recipes
GitHub stars
40k
Token cost
~1.9k tokens
SKILL.md length
925 words
Files
3 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters.

  • Reviewing a LoRA/QLoRA training configuration
  • SKILL.md covers The Reference Recipe, Unsloth Defaults, LoRA vs QLoRA vs Full FT and Failure Modes, plus 1 more section
  • Calls python
  • Choosing rank/alpha/target modules

What it does

Lora Qlora Recipes is an agent skill from wshobson/agents. Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters. Use when writing or reviewing a LoRA/QLoRA training configuration, choosing rank/alpha/target modules, or deciding between LoRA, QLoRA, and full fine-tuning.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/hyperparameters.md` and `references/unsloth-trl-mapping.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Reviewing a LoRA/QLoRA training configuration
  • Choosing rank/alpha/target modules
  • Deciding between LoRA
  • Full fine-tuning

Example prompts

  • “/lora-qlora-recipes”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Lora Qlora Recipes loads about 1.9k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 925 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 925 words, ~1,907 tokens.

Download SKILL.mdSave it as .claude/skills/lora-qlora-recipes/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
lora-qlora-recipes
description
Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters. Use when writing or reviewing a LoRA/QLoRA training configuration, choosing rank/alpha/target modules, or deciding between LoRA, QLoRA, and full fine-tuning.

LoRA & QLoRA Recipes

This skill assumes the routing decision already happened — finetuning-method-selection should have already pointed here because the data shape is demonstrations (SFT), not preference pairs or a verifiable reward signal. What follows is the current best-practice recipe for configuring the adapter itself: which modules to target, how to size rank and alpha, what learning rate to use, and when QLoRA buys real headroom versus when it just adds risk. Dataset preparation and quality checks are a separate concern — see dataset-curation.

Input: a routing decision (SFT via LoRA/ QLoRA) plus a target size class. Output format: a validated adapter config — the kwarg values below, not free-form advice — that llm-finetuning-training-engineer consumes directly when it generates a runnable script.

The Reference Recipe

The reference recipe is "LoRA Without Regret" (Thinking Machines/Schulman, 2025-09), now the settled convention for LoRA/QLoRA SFT.

Target Modules

Target all-linear modules, not just attention:

python
target_modules = [
    "q_proj", "k_proj", "v_proj", "o_proj",   # attention
    "gate_proj", "up_proj", "down_proj",      # MLP — matters most
]

The MLP layers (gate_proj, up_proj, down_proj) matter most — attention-only targeting was the older, weaker convention. Dropping modules to save memory is a Failure Mode below, not a valid optimization.

Alpha and Learning Rate
  • lora_alpha = 2 * r is the settled convention (NeurIPS 2025 "intruder dimensions" result). Don't hand-tune alpha independently of rank — derive it from rank every time.
  • LoRA learning rate ≈ 10x the equivalent full-fine-tune LR. For QLoRA specifically, 2e-4 is the standard starting point. Full hyperparameter tables and worked examples: references/hyperparameters.md.
Rank by Task

Rank is task-shaped, not a single global default:

TaskRank
RL (GRPO/RLVR adapters)1–32
General default16–32
SFT at scaleup to ~256

Higher rank isn't automatically better — it raises capacity to memorize as fast as it raises capacity to generalize. Start at the row matching the task, and only move up a row if the lower rank measurably underfits on held-out eval, not as a default hedge.

Effective Batch Size

Keep effective batch size under 32. This recipe was validated at that scale — pushing effective batch higher is an untested extrapolation, not a free throughput win.

Unsloth Defaults

Unsloth is the reference implementation this plugin assumes as the default fast path — except for messages-shaped conversational SFT with assistant_only_loss=True, where Unsloth 2026.7.x's compiled trainer has no messages-shaped path at all and the plain-TRL escape hatch (references/unsloth-trl-mapping.md) is the default for that combination, not a rare-regression fallback. Its out-of-the-box defaults, and why each one is set that way:

  • lora_dropout=0 — the optimized kernel path assumes zero dropout; setting a nonzero value forfeits the fused-kernel speedup.
  • bias="none" — bias terms add adapter parameters for negligible quality gain at this rank range.
  • use_gradient_checkpointing="unsloth" — Unsloth's checkpointing variant, not vanilla HF checkpointing; saves roughly 30% VRAM over no checkpointing.
  • optim="adamw_8bit" — 8-bit AdamW cuts optimizer-state memory with negligible quality impact at LoRA/QLoRA adapter scale.
  • random_state fixed — pins LoRA initialization for reproducibility across runs; treat it like any other seed, not a tunable.

These show up together on the get_peft_model call:

python
model = FastLanguageModel.get_peft_model(
    model,
    r=32,
    target_modules=target_modules,
    lora_alpha=64,               # 2 * r
    lora_dropout=0,
    bias="none",
    use_gradient_checkpointing="unsloth",
    random_state=3407,
)

Exact kwarg names and their plain-TRL/PEFT equivalents, plus a full worked config including SFTConfig: references/unsloth-trl-mapping.md and references/hyperparameters.md.

Show full SKILL.md (442 more words)Show less

LoRA vs QLoRA vs Full FT

SituationDefault choice
Adapting behavior on demonstrationsLoRA
Base model doesn't fit in bf16 at target rankQLoRA
Injecting dense new domain knowledgeFull FT (see finetuning-method-selection)
Unsure which oneLoRA — upgrade to QLoRA only if memory forces it
  • QLoRA = NF4-quantized frozen base weights + BF16 adapters. This is what makes a 65B-class model trainable on 48GB — the quantized base is the memory win, not the adapter itself.
  • Full fine-tuning is not a default. Reserve it for dense knowledge injection where the goal is changing what the model knows at the weight level, not adapting a behavior. For everything else in this skill's scope, LoRA or QLoRA is the starting assumption.
  • On DGX Spark, QLoRA can OOM before an equivalent bf16 LoRA run would, even though QLoRA's steady-state footprint is smaller — bitsandbytes dequantization buffers are transient CUDA-side allocations that spike during load. A QLoRA OOM is not proof the model doesn't fit; the dgx-spark-ops plugin's spark-memory-thermal-ops skill covers the full OOM remediation ladder (bf16 LoRA is the next thing to try, not a further QLoRA shrink).

Failure Modes

  • fp16 divergence on non-BF16 GPUs. Training in fp16 on hardware that doesn't have solid BF16 support is a known source of loss spikes and silent divergence. Force bf16=True wherever the hardware supports it; don't fall back to fp16 as if it were equivalent. Check hardware support before picking a dtype:

    bash
    python -c "import torch; print(torch.cuda.is_bf16_supported())"
  • Rank too high on a small dataset overfits. A rank picked for "SFT at scale" (up to ~256) on a dataset that doesn't have scale behind it memorizes rather than generalizes. Match rank to the Rank by Task table above, not to the largest number available.

  • Removing target modules to save memory costs quality for negligible savings. The adapter parameters on gate_proj/up_proj/down_proj are a small fraction of total model size — cutting them barely moves memory but measurably hurts quality. If memory is tight, move to QLoRA or reduce rank/batch/pack length before trimming target modules.

All three failure modes share a pattern: they look like a training-loop bug (loss spikes, plateaus, memorization) but are actually a config choice that contradicts the reference recipe above. Check configuration against this skill before debugging the training loop itself.

References

  • references/hyperparameters.md — full rank/ alpha/LR tables by task type, rsLoRA notes, batch/packing interactions, and a complete worked Unsloth config block.
  • references/unsloth-trl-mapping.md — every Unsloth kwarg mapped to its TRL/PEFT equivalent, current TRL API notes, and the escape-hatch rule for when to drop back to plain TRL.

Related skills: finetuning-method-selection routes here; dataset-curation covers the data side this skill doesn't; llm-finetuning-training-engineer is the downstream consumer of the config this skill produces.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/llm-finetuning/skills/lora-qlora-recipes of wshobson/agents.

  • SKILL.md
  • references/hyperparameters.md
  • references/unsloth-trl-mapping.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Lora Qlora Recipes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Lora Qlora Recipes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Lora Qlora Recipes this skillwshobson/agents40k—~1.9kAutomated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.0
Dataset Evaluationawslabs/agent-plugins9161 repos~1.3kAutomated safety check: PassApache-2.0
Train SftOpenPipe/ART11k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    916 GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Sft

    OpenPipe/ART

    SFT training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Fine Tuning With Trl

    Orchestra-Research/AI-Research-SKILLs

    Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.

    13k GitHub starsUsed in 6 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 6 days ago
    Auto-check passed

Questions about Lora Qlora Recipes

What does Lora Qlora Recipes do?

Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters. Lora Qlora Recipes is an agent skill from wshobson/agents. Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters.

When should I use Lora Qlora Recipes?

Lora Qlora Recipes fits situations like: reviewing a LoRA/QLoRA training configuration; choosing rank/alpha/target modules; deciding between LoRA; full fine-tuning.

How do I install Lora Qlora Recipes in Claude Code?

Run `npx skills add wshobson/agents --skill lora-qlora-recipes -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/lora-qlora-recipes in wshobson/agents) into .claude/skills/lora-qlora-recipes in your project. Claude Code loads it when a task matches its description.

How do I install Lora Qlora Recipes in Codex?

Run `npx skills add wshobson/agents --skill lora-qlora-recipes -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/lora-qlora-recipes in wshobson/agents) into .agents/skills/lora-qlora-recipes in your project. Codex loads it when a task matches its description.

Can I use Lora Qlora Recipes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill lora-qlora-recipes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lora-qlora-recipes, .gemini/skills/lora-qlora-recipes, .github/skills/lora-qlora-recipes and .opencode/skills/lora-qlora-recipes in your project.

What does Lora Qlora Recipes need to run?

Going by SKILL.md and its folder, Lora Qlora Recipes needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Lora Qlora Recipes access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Lora Qlora Recipes safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Lora Qlora Recipes use?

Lora Qlora Recipes is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Lora Qlora Recipes use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.

What are the alternatives to Lora Qlora Recipes?

Skills that share tags, products or a category with Lora Qlora Recipes: Sentence-Transformers Training Router (huggingface/skills, 11k stars), Train Rl (OpenPipe/ART, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Dataset Evaluation (awslabs/agent-plugins, 916 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Lora Qlora Recipes?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.