Agent skill

LLM Fine Tuning

by sickn33 in sickn33/agentic-awesome-skills

Set up infrastructure for fine-tuning LLMs with QLoRA, LoRA, and full fine-tuning using Hugging Face TRL, Axolotl, and distributed training with DeepSpeed or FSDP.

MITAuto-check passedAI & LLM Engineering

Install LLM Fine Tuning

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill llm-fine-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills llm-fine-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-fine-tuning .claude/skills/llm-fine-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-fine-tuning
GitHub stars
47k
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
302 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Set up infrastructure for fine-tuning LLMs with QLoRA, LoRA, and full fine-tuning using Hugging Face TRL, Axolotl, and distributed training with DeepSpeed or FSDP.

  • Tasks that involve Fine-tuning
  • SKILL.md covers When to Use This Skill, Prerequisites, Quick Start: QLoRA Fine-Tuning and Axolotl (Production…, plus 8 more sections
  • Calls pip and python; needs HF_TOKEN and HUGGING_FACE_HUB_TOKEN
  • Tasks that involve Deep learning

What it does

LLM Fine Tuning is an agent skill from sickn33/agentic-awesome-skills. Set up infrastructure for fine-tuning LLMs with QLoRA, LoRA, and full fine-tuning using Hugging Face TRL, Axolotl, and distributed training with DeepSpeed or FSDP.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

It sits in AI & LLM Engineering, covering Fine-tuning and Deep learning. It works with Hugging Face. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Fine-tuning
  • Tasks that involve Deep learning

Example prompts

  • “/llm-fine-tuning”

Requirements

  • Python 3
  • A credential in HUGGING_FACE_HUB_TOKEN
  • A credential in WANDB_API_KEY
  • Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • HUGGING_FACE_HUB_TOKEN
    • WANDB_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

LLM Fine Tuning loads about 2.3k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 302 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 302 words, ~2,338 tokens.

Download SKILL.mdSave it as .claude/skills/llm-fine-tuning/SKILL.md (or your agent's skills folder).
name
llm-fine-tuning
description
Set up infrastructure for fine-tuning LLMs with QLoRA, LoRA, and full fine-tuning using Hugging Face TRL, Axolotl, and distributed training with DeepSpeed or FSDP.
compatibility
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

LLM Fine-Tuning Infrastructure

Train and fine-tune open-source LLMs efficiently — from LoRA on a single GPU to distributed full fine-tuning across multi-node clusters.

When to Use This Skill

Use this skill when:

  • Fine-tuning an LLM on domain-specific data (legal, medical, code, support)
  • Running QLoRA to fine-tune 70B models on consumer GPUs
  • Setting up distributed training with DeepSpeed or FSDP
  • Exporting fine-tuned adapters for production serving
  • Implementing RLHF, DPO, or instruction tuning pipelines

Prerequisites

  • NVIDIA GPU(s) with 24GB+ VRAM (RTX 4090 / A100 / H100)
  • CUDA 12.1+ and nvidia-smi working
  • Python 3.10+ with pip
  • Hugging Face account and HF_TOKEN for gated models
  • 500GB+ disk for model weights and training data

Quick Start: QLoRA Fine-Tuning

bash
pip install transformers datasets trl peft bitsandbytes accelerate

python - <<'EOF'
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model
from trl import SFTTrainer, SFTConfig
import torch

model_id = "meta-llama/Llama-3.1-8B-Instruct"

# 4-bit quantization (QLoRA)
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

model = AutoModelForCausalLM.from_pretrained(
    model_id, quantization_config=bnb_config, device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_id)

# LoRA configuration
peft_config = LoraConfig(
    r=16,                    # rank
    lora_alpha=32,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)

dataset = load_dataset("your-org/your-dataset", split="train")

trainer = SFTTrainer(
    model=model,
    args=SFTConfig(
        output_dir="./output",
        num_train_epochs=3,
        per_device_train_batch_size=2,
        gradient_accumulation_steps=8,
        learning_rate=2e-4,
        bf16=True,
        logging_steps=10,
        save_strategy="epoch",
        report_to="wandb",
    ),
    train_dataset=dataset,
    peft_config=peft_config,
    processing_class=tokenizer,
)
trainer.train()
trainer.save_model("./fine-tuned-model")
EOF

Axolotl (Production Fine-Tuning Framework)

yaml
# config.yaml — Axolotl QLoRA config for Llama 3.1
base_model: meta-llama/Llama-3.1-8B-Instruct
model_type: LlamaForCausalLM
tokenizer_type: PreTrainedTokenizerFast

load_in_4bit: true
adapter: qlora
lora_r: 32
lora_alpha: 64
lora_dropout: 0.05
lora_target_modules:
  - q_proj
  - k_proj
  - v_proj
  - o_proj
  - gate_proj
  - up_proj
  - down_proj

datasets:
  - path: your-org/your-dataset
    type: alpaca              # or sharegpt, chat_template, etc.

dataset_prepared_path: ./prepared-data
val_set_size: 0.05
output_dir: ./output

sequence_len: 4096
sample_packing: true         # pack multiple short samples for efficiency

micro_batch_size: 2
gradient_accumulation_steps: 8
num_epochs: 3
learning_rate: 2e-4
optimizer: adamw_bnb_8bit
lr_scheduler: cosine
warmup_ratio: 0.05

bf16: true
flash_attention: true

logging_steps: 10
eval_steps: 100
save_steps: 200
wandb_project: my-fine-tune
bash
# Run with Axolotl
pip install axolotl[flash-attn,deepspeed]
accelerate launch -m axolotl.cli.train config.yaml

Distributed Training with DeepSpeed

json
// deepspeed_zero3.json — ZeRO Stage 3 (split optimizer + gradients + params)
{
  "zero_optimization": {
    "stage": 3,
    "offload_optimizer": {"device": "cpu", "pin_memory": true},
    "offload_param": {"device": "cpu", "pin_memory": true},
    "overlap_comm": true,
    "contiguous_gradients": true,
    "sub_group_size": 1e9,
    "reduce_bucket_size": "auto",
    "stage3_prefetch_bucket_size": "auto",
    "stage3_param_persistence_threshold": "auto",
    "stage3_max_live_parameters": 1e9,
    "stage3_max_reuse_distance": 1e9,
    "gather_16bit_weights_on_model_save": true
  },
  "bf16": {"enabled": true},
  "gradient_clipping": 1.0,
  "train_batch_size": "auto",
  "train_micro_batch_size_per_gpu": "auto"
}
bash
# Launch 4-GPU DeepSpeed training
deepspeed --num_gpus=4 train.py \
  --deepspeed deepspeed_zero3.json \
  --model_name meta-llama/Llama-3.1-70B-Instruct \
  --output_dir ./output

DPO / RLHF Alignment

python
from trl import DPOTrainer, DPOConfig
from datasets import load_dataset

# Dataset format: {"prompt": ..., "chosen": ..., "rejected": ...}
dataset = load_dataset("your-org/preference-data")

trainer = DPOTrainer(
    model=model,
    ref_model=None,           # None = implicit reference with peft
    args=DPOConfig(
        output_dir="./dpo-output",
        beta=0.1,             # KL divergence weight
        num_train_epochs=1,
        per_device_train_batch_size=1,
        gradient_accumulation_steps=16,
        learning_rate=5e-7,
        bf16=True,
    ),
    train_dataset=dataset["train"],
    peft_config=peft_config,
    processing_class=tokenizer,
)
trainer.train()

Merging LoRA Adapters for Deployment

python
from peft import PeftModel
from transformers import AutoModelForCausalLM

# Load base model in full precision
base_model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B-Instruct",
    torch_dtype=torch.bfloat16,
    device_map="cpu",
)

# Load and merge LoRA adapter
model = PeftModel.from_pretrained(base_model, "./fine-tuned-model")
merged_model = model.merge_and_unload()

# Save merged model (ready for vLLM serving)
merged_model.save_pretrained("./merged-model", safe_serialization=True)
tokenizer.save_pretrained("./merged-model")

# Push to Hugging Face Hub
merged_model.push_to_hub("your-org/your-fine-tuned-model")

Kubernetes Training Job

yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: llm-fine-tune
spec:
  template:
    spec:
      restartPolicy: OnFailure
      nodeSelector:
        nvidia.com/gpu.product: A100-SXM4-80GB
      containers:
      - name: trainer
        image: nvcr.io/nvidia/pytorch:24.05-py3
        command: ["accelerate", "launch", "-m", "axolotl.cli.train", "/config/config.yaml"]
        resources:
          limits:
            nvidia.com/gpu: "4"
            memory: "320Gi"
          requests:
            nvidia.com/gpu: "4"
        volumeMounts:
        - name: config
          mountPath: /config
        - name: model-cache
          mountPath: /root/.cache/huggingface
        - name: output
          mountPath: /output
        env:
        - name: HUGGING_FACE_HUB_TOKEN
          valueFrom:
            secretKeyRef:
              name: hf-token
              key: token
        - name: WANDB_API_KEY
          valueFrom:
            secretKeyRef:
              name: wandb-token
              key: key
      volumes:
      - name: config
        configMap:
          name: axolotl-config
      - name: model-cache
        persistentVolumeClaim:
          claimName: model-cache-pvc
      - name: output
        persistentVolumeClaim:
          claimName: training-output-pvc

Common Issues

IssueCauseFix
CUDA out of memoryBatch too largeReduce micro_batch_size; increase gradient_accumulation_steps
Training loss NaNLearning rate too highLower LR to 1e-4 or 5e-5; add warmup
Slow trainingNo Flash AttentionInstall flash-attn; enable flash_attention: true
Poor fine-tune qualityBad data formattingValidate dataset format; check sample_packing compatibility
Adapter merge errorsMixed quantizationMerge in bf16 on CPU, not in 4-bit

Best Practices

  • Use Flash Attention 2 — it's 2–4× faster and uses less memory.
  • Monitor training loss/eval loss via W&B or MLflow; overfit = more dropout or less data.
  • Validate with a held-out eval set (5–10%); MMLU or custom evals for quality gates.
  • Start with LoRA r=16 before increasing — higher rank = more parameters, diminishing returns.
  • Use sample_packing in Axolotl to maximize GPU utilization on short sequences.
  • vllm-server (vllm-server) - Serve fine-tuned models
  • gpu-server-management (gpu-server-management) - GPU setup
  • llm-inference-scaling (llm-inference-scaling) - Deploy at scale
  • ai-pipeline-orchestration (ai-pipeline-orchestration) - Training pipelines

Limitations

  • Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
  • Docs-only import: upstream scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/llm-fine-tuning of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLM Fine Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Fine Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Fine Tuning this skillsickn33/agentic-awesome-skills47k1 repos~2.3kAutomated safety check: PassMIT
LLM Fine TuningBagelHole/DevOps-Security-Agent-Skills1.2k—~2.2kAutomated safety check: PassMIT
AI ML Skillswentorai/research-plugins2981 repos~993Automated safety check: PassMIT
Cosmos3 Post TrainingNVIDIA/cosmos-framework560—~2.7kAutomated safety check: PassCustom licence
Discover MLrand/cc-polymath181—~574Automated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • LLM Fine Tuning

    BagelHole/DevOps-Security-Agent-Skills

    Set up infrastructure for fine-tuning LLMs with QLoRA, LoRA, and full fine-tuning using Hugging Face TRL, Axolotl, and distributed training with DeepSpeed or FSDP.

    1.2k GitHub stars~2.2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • AI ML Skills

    wentorai/research-plugins

    27 ai & machine learning skills. An agent skill from wentorai/research-plugins.

    298 GitHub starsUsed in 1 repo~993 tokens
    AI & LLM EngineeringAuto-check passed
  • Cosmos3 Post Training

    NVIDIA/cosmos-framework

    Official

    Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired…

    560 GitHub stars~2.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Discover ML

    rand/cc-polymath

    Automatically discover machine learning and AI skills when working with machine learning, PyTorch, training, inference, RAG, embeddings, fine-tuning, LLM, DSPy, HuggingFace, or diffusion models.

    181 GitHub stars~574 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Questions about LLM Fine Tuning

What does LLM Fine Tuning do?

Set up infrastructure for fine-tuning LLMs with QLoRA, LoRA, and full fine-tuning using Hugging Face TRL, Axolotl, and distributed training with DeepSpeed or FSDP. LLM Fine Tuning is an agent skill from sickn33/agentic-awesome-skills. Set up infrastructure for fine-tuning LLMs with QLoRA, LoRA, and full fine-tuning using Hugging Face TRL, Axolotl, and distributed training with DeepSpeed or FSDP.

When should I use LLM Fine Tuning?

LLM Fine Tuning fits situations like: tasks that involve Fine-tuning; tasks that involve Deep learning.

How do I install LLM Fine Tuning in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill llm-fine-tuning -a claude-code`. Or copy the skill folder (skills/llm-fine-tuning in sickn33/agentic-awesome-skills) into .claude/skills/llm-fine-tuning in your project. Claude Code loads it when a task matches its description.

How do I install LLM Fine Tuning in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill llm-fine-tuning -a codex`. Or copy the skill folder (skills/llm-fine-tuning in sickn33/agentic-awesome-skills) into .agents/skills/llm-fine-tuning in your project. Codex loads it when a task matches its description.

Can I use LLM Fine Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill llm-fine-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-fine-tuning, .gemini/skills/llm-fine-tuning, .github/skills/llm-fine-tuning and .opencode/skills/llm-fine-tuning in your project.

What does LLM Fine Tuning need to run?

Going by SKILL.md and its folder, LLM Fine Tuning needs the command-line tools its instructions call (pip and python) and credentials named HF_TOKEN, HUGGING_FACE_HUB_TOKEN and WANDB_API_KEY. Our summary lists: Python 3; A credential in HUGGING_FACE_HUB_TOKEN; A credential in WANDB_API_KEY. Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..

Does LLM Fine Tuning access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is LLM Fine Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Fine Tuning use?

LLM Fine Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Fine Tuning use?

About 2.3k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Fine Tuning?

Skills that share tags, products or a category with LLM Fine Tuning: LLM Fine Tuning (BagelHole/DevOps-Security-Agent-Skills, 1.2k stars), AI ML Skills (wentorai/research-plugins, 298 stars), Cosmos3 Post Training (NVIDIA/cosmos-framework, 560 stars) and Discover ML (rand/cc-polymath, 181 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Fine Tuning?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.