Agent skill

Fine-Tuning Expert

by Jeffallan in Jeffallan/claude-skills

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

MITAuto-check passedAI & LLM Engineering

Install Fine-Tuning Expert

skills CLI
$ npx skills add Jeffallan/claude-skills --skill fine-tuning-expert -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Jeffallan/claude-skills fine-tuning-expert --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/fine-tuning-expert .claude/skills/fine-tuning-expert && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fine-tuning-expert
GitHub stars
12k
Token cost
~1.7k tokens
SKILL.md length
333 words
Files
6 (incl. references)
Skills in repo
58
Repo updated
First seen
Licence
MIT

At a glance

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

  • Works in 5 steps: Dataset preparation — Validate and… → Method selection — Choose PEFT technique… → Training — Configure hyperparameters,… → …
  • Adapting a foundation model to a specific task with LoRA or QLoRA adapters
  • SKILL.md covers Core Workflow, Reference Guide, Minimal Working Example — LoRA… and Constraints, plus 1 more section
  • Calls python

What it does

The workflow has five stages: validate and format the training data, pick a parameter-efficient method, train while watching loss curves, evaluate against the base model, and deploy. Method advice is LoRA for most tasks, QLoRA with 4-bit quantization when GPU memory is tight, and a full fine-tune only for small models. Each stage has a checkpoint, such as fixing every dataset error before training starts and treating a plateauing or rising validation loss as a sign of overfitting.

Reference files cover LoRA and PEFT, dataset preparation, hyperparameters such as learning rates and batch sizes, evaluation metrics, and deployment steps like merging adapter weights and serving. Evaluation gathers perplexity, task metrics such as BLEU and ROUGE, and latency figures. The worked example uses the datasets, transformers and peft libraries, with a QLoRA variant built on BitsAndBytesConfig and a snippet that merges the adapter into the base model.

When your agent uses it

  • Adapting a foundation model to a specific task with LoRA or QLoRA adapters
  • Preparing and checking a JSONL training dataset before a run
  • Choosing learning rates, batch sizes and schedulers for a fine-tuning job
  • Comparing a tuned model with its base model on held-out data
  • Merging adapter weights and quantizing a tuned model for serving

Example prompts

  • “Set up a QLoRA run on our support transcripts that fits on one consumer GPU.”
  • “Check data.jsonl for formatting errors before I start training.”
  • “Our validation loss rises after the second epoch, so help me work out whether it is overfitting.”
  • “Merge the LoRA adapter into the base model and measure inference throughput.”

Requirements

  • Python with Hugging Face `transformers`, `peft` and `datasets`

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Dataset preparation — Validate and format data; run quality checks before training starts
  2. Method selection — Choose PEFT technique based on GPU memory and task requirements
  3. Training — Configure hyperparameters, monitor loss curves, checkpoint regularly
  4. Evaluation — Benchmark against the base model; test on held-out set and edge cases
  5. Deployment — Merge adapter weights, quantize, measure inference throughput before serving

What it can do on your machine

Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • synergetic.solutions
    • jeffallan.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fine-Tuning Expert loads about 1.7k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 135 tokens; SKILL.md has 333 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 333 words, ~1,695 tokens.

Download SKILL.mdSave it as .claude/skills/fine-tuning-expert/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
fine-tuning-expert
description
Use when fine-tuning LLMs, training custom models, or adapting foundation models for specific tasks. Invoke for configuring LoRA/QLoRA adapters, preparing JSONL training datasets, setting hyperparameters for fine-tuning runs, adapter training, transfer learning, finetuning with Hugging Face PEFT, OpenAI fine-tuning, instruction tuning, RLHF, DPO, or quantizing and deploying fine-tuned models. Trigger terms include: LoRA, QLoRA, PEFT, finetuning, fine-tuning, adapter tuning, LLM training, model training, custom model.
license
MIT
metadata.author
https://github.com/Jeffallan
metadata.company
https://synergetic.solutions
metadata.version
1.1.0
metadata.domain
data-ml
metadata.triggers
fine-tuning, fine tuning, finetuning, LoRA, QLoRA, PEFT, adapter tuning, transfer learning, model training, custom model, LLM training, instruction tuning…
metadata.role
expert
metadata.scope
implementation
metadata.output-format
code
metadata.related-skills
devops-engineer

Fine-Tuning Expert

Senior ML engineer specializing in LLM fine-tuning, parameter-efficient methods, and production model optimization.

Core Workflow

  1. Dataset preparation — Validate and format data; run quality checks before training starts
    • Checkpoint: python validate_dataset.py --input data.jsonl — fix all errors before proceeding
  2. Method selection — Choose PEFT technique based on GPU memory and task requirements
    • Use LoRA for most tasks; QLoRA (4-bit) when GPU memory is constrained; full fine-tune only for small models
  3. Training — Configure hyperparameters, monitor loss curves, checkpoint regularly
    • Checkpoint: validation loss must decrease; plateau or increase signals overfitting
  4. Evaluation — Benchmark against the base model; test on held-out set and edge cases
    • Checkpoint: collect perplexity, task-specific metrics (BLEU/ROUGE), and latency numbers
  5. Deployment — Merge adapter weights, quantize, measure inference throughput before serving

Reference Guide

Load detailed guidance based on context:

TopicReferenceLoad When
LoRA/PEFTreferences/lora-peft.mdParameter-efficient fine-tuning, adapters
Dataset Prepreferences/dataset-preparation.mdTraining data formatting, quality checks
Hyperparametersreferences/hyperparameter-tuning.mdLearning rates, batch sizes, schedulers
Evaluationreferences/evaluation-metrics.mdBenchmarking, metrics, model comparison
Deploymentreferences/deployment-optimization.mdModel merging, quantization, serving

Minimal Working Example — LoRA Fine-Tuning with Hugging Face PEFT

python
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM, TrainingArguments
from peft import LoraConfig, get_peft_model, TaskType
from trl import SFTTrainer
import torch

# 1. Load base model and tokenizer
model_id = "meta-llama/Llama-3-8B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

# 2. Configure LoRA adapter
lora_config = LoraConfig(
    task_type=TaskType.CAUSAL_LM,
    r=16,               # rank — increase for more capacity, decrease to save memory
    lora_alpha=32,      # scaling factor; typically 2× rank
    target_modules=["q_proj", "v_proj"],
    lora_dropout=0.05,
    bias="none",
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()  # verify: should be ~0.1–1% of total params

# 3. Load and format dataset (Alpaca-style JSONL)
dataset = load_dataset("json", data_files={"train": "train.jsonl", "test": "test.jsonl"})

def format_prompt(example):
    return {"text": f"### Instruction:\n{example['instruction']}\n\n### Response:\n{example['output']}"}

dataset = dataset.map(format_prompt)

# 4. Training arguments
training_args = TrainingArguments(
    output_dir="./checkpoints",
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,     # effective batch size = 16
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    warmup_ratio=0.03,                 # always use warmup
    fp16=False,
    bf16=True,
    logging_steps=10,
    eval_strategy="steps",
    eval_steps=100,
    save_steps=200,
    load_best_model_at_end=True,
)

# 5. Train
trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=dataset["train"],
    eval_dataset=dataset["test"],
    dataset_text_field="text",
    max_seq_length=2048,
)
trainer.train()

# 6. Save adapter weights only
model.save_pretrained("./lora-adapter")
tokenizer.save_pretrained("./lora-adapter")

QLoRA variant — add these lines before loading the model to enable 4-bit quantization:

python
from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=bnb_config, device_map="auto")

Merge adapter into base model for deployment:

python
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16)
merged = PeftModel.from_pretrained(base, "./lora-adapter").merge_and_unload()
merged.save_pretrained("./merged-model")

Constraints

MUST DO
  • Validate dataset quality before training
  • Use parameter-efficient methods for large models (>7B)
  • Monitor training/validation loss curves
  • Document hyperparameters and training config
  • Version datasets and model checkpoints
  • Always include a learning rate warmup
MUST NOT DO
  • Skip data quality validation
  • Overfit on small datasets — use regularisation (dropout, weight decay) and early stopping
  • Merge incompatible adapters (mismatched rank, base model, or target modules)
  • Deploy without evaluation against a held-out set and latency benchmark

Output Templates

When implementing fine-tuning, always provide:

  1. Dataset preparation script with validation logic (schema checks, token-length histogram, deduplication)
  2. Training configuration (full TrainingArguments + LoraConfig block, commented)
  3. Evaluation script reporting perplexity, task-specific metrics, and latency
  4. Brief design rationale — why this PEFT method, rank, and learning rate were chosen for this task

Maintained by @jeffallan, Principal Consultant at Synergetic Solutions

Documentation

© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/fine-tuning-expert of Jeffallan/claude-skills.

  • SKILL.md
  • references/dataset-preparation.md
  • references/deployment-optimization.md
  • references/evaluation-metrics.md
  • references/hyperparameter-tuning.md
  • references/lora-peft.md

Open the folder on GitHubat commit 1be15d8

Compare with similar skills

Fine-Tuning Expert next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fine-Tuning Expert compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fine-Tuning Expert this skillJeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT
Vllmericrisco/rsc-harness180—~3.6kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs13k2 repos~2.5kAutomated safety check: PassMIT
Open Weightsericrisco/rsc-harness180—~4.1kAutomated safety check: PassMIT
Dataset Transformationawslabs/agent-plugins9161 repos~3.5kAutomated safety check: PassApache-2.0

Similar skills

  • Vllm

    ericrisco/rsc-harness

    A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…

    180 GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • bitsandbytes Model Quantization

    Orchestra-Research/AI-Research-SKILLs

    Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

    13k GitHub starsUsed in 2 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Open Weights

    ericrisco/rsc-harness

    A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

    180 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    916 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Mesh API

    mr-tbot/mesh-api

    Interact with a Meshtastic LoRa mesh network through MESH-API — list nodes, read messages, send texts, and check connection status.

    180 GitHub stars~1.8k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from Jeffallan/claude-skills

All 58 skills in this repo
  • API Designer

    Jeffallan/claude-skills

    Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.

    12k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • CLI Developer

    Jeffallan/claude-skills

    Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.

    12k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Kubernetes Specialist

    Jeffallan/claude-skills

    Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Laravel Specialist

    Jeffallan/claude-skills

    Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Apache Spark Engineer

    Jeffallan/claude-skills

    Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Questions about Fine-Tuning Expert

What does Fine-Tuning Expert do?

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment. The workflow has five stages: validate and format the training data, pick a parameter-efficient method, train while watching loss curves, evaluate against the base model, and deploy. Method advice is LoRA for most tasks, QLoRA with 4-bit quantization when GPU memory is tight, and a full fine-tune only for small models.

When should I use Fine-Tuning Expert?

Fine-Tuning Expert fits situations like: adapting a foundation model to a specific task with LoRA or QLoRA adapters; preparing and checking a JSONL training dataset before a run; choosing learning rates, batch sizes and schedulers for a fine-tuning job; comparing a tuned model with its base model on held-out data.

How do I install Fine-Tuning Expert in Claude Code?

Run `npx skills add Jeffallan/claude-skills --skill fine-tuning-expert -a claude-code`. Or copy the skill folder (skills/fine-tuning-expert in Jeffallan/claude-skills) into .claude/skills/fine-tuning-expert in your project. Claude Code loads it when a task matches its description.

How do I install Fine-Tuning Expert in Codex?

Run `npx skills add Jeffallan/claude-skills --skill fine-tuning-expert -a codex`. Or copy the skill folder (skills/fine-tuning-expert in Jeffallan/claude-skills) into .agents/skills/fine-tuning-expert in your project. Codex loads it when a task matches its description.

Can I use Fine-Tuning Expert in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill fine-tuning-expert -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fine-tuning-expert, .gemini/skills/fine-tuning-expert, .github/skills/fine-tuning-expert and .opencode/skills/fine-tuning-expert in your project.

What does Fine-Tuning Expert need to run?

Going by SKILL.md and its folder, Fine-Tuning Expert needs the command-line tools its instructions call (python). Our summary lists: Python with Hugging Face `transformers`, `peft` and `datasets`.

Does Fine-Tuning Expert access the network?

SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.

Is Fine-Tuning Expert safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fine-Tuning Expert use?

Fine-Tuning Expert is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fine-Tuning Expert use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.

What are the alternatives to Fine-Tuning Expert?

Skills that share tags, products or a category with Fine-Tuning Expert: Vllm (ericrisco/rsc-harness, 180 stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), bitsandbytes Model Quantization (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Open Weights (ericrisco/rsc-harness, 180 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fine-Tuning Expert?

Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,802 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.

Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.