Agent skill

Gemma Trainer

by google-gemma in google-gemma/gemma-skills

Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Gemma Trainer

skills CLI
$ npx skills add google-gemma/gemma-skills --skill gemma-trainer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-gemma/gemma-skills gemma-trainer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-gemma/gemma-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gemma-trainer .claude/skills/gemma-trainer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemma-trainer
GitHub stars
1k
Token cost
~1.9k tokens
SKILL.md length
894 words
Files
6 (incl. assets)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g.

  • Works in 6 steps: Core Principles: Local Fine-Tuning Setup → Choosing the Right Training Method → Dataset Preparation & Validation → …
  • This skill when the user wants to train
  • SKILL.md covers 1. Core Principles: Local…, 2. Choosing the Right Training…, 3. Dataset Preparation &… and 4. Fine-Tuning Workflows, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Gemma Trainer is an agent skill from google-gemma/gemma-skills. Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g. SFT, DPO, RLHF, Reward Modeling) on local hardware. Covers TRL, Unsloth, dataset preparation, validation, and GGUF/LiteRT conversion.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including assets (for example `assets/dataset_prep.py`, `assets/distill_dataset.py` and `assets/dpo_train.py`).

It sits in AI & LLM Engineering, covering Fine-tuning. It works with llama.cpp. The repository describes itself as: Skills for the Gemma and model/agent interactions. The licence is Apache-2.0.

When your agent uses it

  • This skill when the user wants to train
  • Adapt Gemma models (e.g

Example prompts

  • “/gemma-trainer”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Core Principles: Local Fine-Tuning Setup
  2. Choosing the Right Training Method
  3. Dataset Preparation & Validation
  4. Fine-Tuning Workflows
  5. Multimodal Fine-Tuning (Vision & Audio)
  6. Post-Training & Deployment Utilities

What it can do on your machine

Read from SKILL.md and the folder at commit f86bcc6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • unsloth.ai
    • github.com
    • developers.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemma Trainer loads about 1.9k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 894 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google-gemma/gemma-skills at commit f86bcc6, republished under its Apache-2.0 licence (© google-gemma). 894 words, ~1,929 tokens.

Download SKILL.mdSave it as .claude/skills/gemma-trainer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
gemma-trainer
description
Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g. SFT, DPO, RLHF, Reward Modeling) on local hardware. Covers TRL, Unsloth, dataset preparation, validation, and GGUF/LiteRT conversion.

Gemma Training and Fine-Tuning Skill

1. Core Principles: Local Fine-Tuning Setup

When training locally, memory efficiency and execution speed are huge. Always guide the user to follow these best practices:

  • Prioritize Unsloth: For local single-GPU training, always recommend Unsloth. It supports Gemma 4 natively, uses up to 70% less memory, and is up to 2x faster than standard Hugging Face PEFT training.
  • Fall Back to TRL: For multi-GPU environments (using DDP/FSDP) or when Unsloth is unavailable, use Hugging Face TRL (SFTTrainer, DPOTrainer) coupled with PEFT and bitsandbytes (for QLoRA).
  • Always use QLoRA (4-bit Quantization): Crucial for fitting Gemma models (like Gemma 4 12B/31B) into consumer VRAM.
  • Manage Context Window & Max Length: Although Gemma 4 supports up to a 256K context window, recommend training with a context window of 2048 to 8192 tokens locally to prevent Out-Of-Memory (OOM) errors.

2. Choosing the Right Training Method

Help the user choose the correct workflow based on their goal:

  • Supervised Fine-Tuning (SFT): Teaching new domains, specialized task instructions, or custom output structures.
    • Prerequisites: Raw text, instruction pairs, or chat logs.
    • Output: Adaptor trained on prompt/completion pairs.
  • Direct Preference Optimization (DPO): Aligning model style, behavior, tone, or safety with human preferences.
    • Prerequisites: A previously SFT-trained Gemma model and preferred pairwise datasets.
    • Output: Aligning model weights directly without a separate reward head.
  • Reward Modeling (RM): Training a scoring system to evaluate response quality.
    • Prerequisites: Binary preference pairwise datasets.
    • Output: A classification-style reward head on top of Gemma.

3. Dataset Preparation & Validation

Formatting issues are the #1 cause of poor training runs. Ensure you validate files using the utility script [assets/dataset_prep.py].

Gemma Chat Prompt Format

Ensure the dataset matches Gemma's official chat template:

<|turn>system
Your instruction here<turn|>
<|turn>user
Your query here<turn|>
<|turn>model
Your response here<turn|>

To avoid formatting drift, use the tokenizer's apply_chat_template during dataset tokenization.

Format Specifications
  • SFT (Supervised Fine-Tuning): Format as a list of conversation turns.
    json
    {
      "messages": [
        {"role": "user", "content": "Tell me a joke."},
        {"role": "model", "content": "Why did the computer go to the doctor? It had a virus!"}
      ]
    }
  • DPO (Direct Preference Optimization): Requires pairwise samples containing a prompt, a chosen (better) response, and a rejected (worse) response.
    json
    {
      "prompt": "Write a python function to compute factorial.",
      "chosen": "def factorial(n):\n    return 1 if n <= 1 else n * factorial(n - 1)",
      "rejected": "factorial is computed using recursion or loops. Just import math."
    }
  • Reward Modeling: Format identically to DPO datasets. The RewardTrainer evaluates the pair and learns to output a higher logit score for chosen than for rejected.
Dataset Distillation & Synthesis (Teacher-Student)

Local knowledge distillation allows you to train small, lightweight student models (such as Gemma 4 E2B) using high-quality dataset outputs generated by larger, highly capable teacher models (such as Gemma 4 31B or Gemma 4 26B A4B).

Use the [assets/distill_dataset.py] utility script to generate fine-tuning datasets on your local machine.

4. Fine-Tuning Workflows

Supervised Fine-Tuning (SFT)

Use the [assets/sft_train.py] asset to launch a local QLoRA fine-tuning session.

  • LoRA Hyperparameters:
    • Rank (r): 16 or 32 (Higher rank captures complex behaviors but consumes more memory).
    • Alpha (lora_alpha): 32 or 64 (Rule of thumb: lora_alpha = 2 * r).
    • Dropout (lora_dropout): 0.05 or 0.1 (Forcing the model to learn more robust features rather than relying on specific paths).
    • Target Modules: Use PEFT's Gemma 4 defaults scope to the LM layers.
    • Learning Rate: 2e-4 for QLoRA; 2e-5 for full fine-tuning.
Show full SKILL.md (357 more words)Show less
Direct Preference Optimization (DPO)

Use the [assets/dpo_train.py] template to execute alignment.

  • Rules for DPO:
    • Always perform SFT on the base model using your instruction format before running DPO. Running DPO directly on an out-of-domain base model usually fails or degrades output formatting.
    • Set beta (DPO temperature parameter) to 0.1. Values between 0.1 and 0.5 control how strictly the model adheres to the reference policy.
Reward Modeling (RM)

Use the [assets/reward_train.py] template to train an evaluation model.

  • Initializes the model with a sequence classification head (AutoModelForSequenceClassification with num_labels=1).
  • Trains the single scalar reward value to distinguish preferred responses.

5. Multimodal Fine-Tuning (Vision & Audio)

Gemma 4 models are natively multimodal. To fine-tune them on images or audio:

Vision SFT
  • Use Gemma 4 E2B/E4B/12B/26B/31B models.
  • Use standard Hugging Face SFTTrainer with a custom visual data collator.
  • Prepare your dataset containing local image paths or PIL images, alongside corresponding conversation instructions:
    json
    {
      "messages": [
        {"role": "user", "content": [
            {"type": "image", "url": "path/to/image.png"},
            {"type": "text", "text": "Describe this image."}
        ]},
        {"role": "assistant", "content": [
            {"type": "text", "text": "An abstract oil painting with vibrant warm gradients."}
        ]}
      ]
    }
    
Audio SFT
  • Use Gemma 4 E2B/E4B/12B models.
  • Feed raw audio arrays (sampled at 16kHz) through the model processor to produce input_features.
  • Maintain conversational formatting, replacing image type with audio type in the message format.
    json
    {
      "messages": [
        {"role": "user", "content": [
            {"type": "text", "text": "Describe this audio."},
            {"type": "audio", "url": "path/to/audio.wav"}
        ]},
        {"role": "assistant", "content": [
            {"type": "text", "text": "This is an audio file of a bird chirping."}
        ]}
      ]
    }

6. Post-Training & Deployment Utilities

GGUF Conversion

Once your LoRA training is finished, you can convert your model to GGUF format.

If you trained your model using Unsloth, you can export directly to GGUF natively. This automatically handles merging and quantization.

Fetch Saving to GGUF for the best practice.

Option 2: Manual Conversion with llama.cpp

If you did not use Unsloth, you can convert your merged Hugging Face model directory manually using llama.cpp.

On-Device Deployment with LiteRT-LM (.litertlm)

LiteRT-LM is optimized for running models like Gemma 4 E2B and Gemma 4 E4B on mobile, web, and IoT hardware with hardware acceleration (CPU, GPU, NPU). Fetch LiteRT-LM guide for the best practice.

© google-gemma, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (assets) in skills/gemma-trainer of google-gemma/gemma-skills.

  • SKILL.md
  • assets/dataset_prep.py
  • assets/distill_dataset.py
  • assets/dpo_train.py
  • assets/reward_train.py
  • assets/sft_train.py

Open the folder on GitHubat commit f86bcc6

Compare with similar skills

Gemma Trainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemma Trainer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemma Trainer this skillgoogle-gemma/gemma-skills1k—~1.9kAutomated safety check: PassApache-2.0
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Unsloth Finetuningsickn33/agentic-awesome-skills47k1 repos~4.1kAutomated safety check: PassApache-2.0
Quantized Exportwshobson/agents40k—~2kAutomated safety check: PassMIT
Wan Flf Videoartokun/comfyui-mcp795—~5.1kAutomated safety check: PassMIT
ML Research LabAnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMIT

Similar skills

  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Unsloth Finetuning

    sickn33/agentic-awesome-skills

    Fine-tune and post-train LLMs with Unsloth Core on a single consumer GPU: VRAM sizing, LoRA/QLoRA, GRPO/DPO, chat-template correctness, and GGUF export.

    47k GitHub starsUsed in 1 repo~4.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Quantized Export

    wshobson/agents

    Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

    40k GitHub stars~2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Wan Flf Video

    artokun/comfyui-mcp

    Build WAN 2.2 First-Last-Frame video workflows. An agent skill from artokun/comfyui-mcp.

    795 GitHub stars~5.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • ML Research Lab

    AnastasiyaW/codex-claude-code-config

    Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

    154 GitHub stars~794 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Finetuning

    ericrisco/rsc-harness

    A skill your agent uses when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then…

    167 GitHub stars~3.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from google-gemma/gemma-skills

  • Gemma Dev

    google-gemma/gemma-skills

    Trigger this skill when building applications with Gemma or for general knowledge inquiries related to Gemma models (e.g.

    1k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Gemma Trainer

What does Gemma Trainer do?

Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g. Gemma Trainer is an agent skill from google-gemma/gemma-skills.g.

When should I use Gemma Trainer?

Gemma Trainer fits situations like: this skill when the user wants to train; adapt Gemma models (e.g.

How do I install Gemma Trainer in Claude Code?

Run `npx skills add google-gemma/gemma-skills --skill gemma-trainer -a claude-code`. Or copy the skill folder (skills/gemma-trainer in google-gemma/gemma-skills) into .claude/skills/gemma-trainer in your project. Claude Code loads it when a task matches its description.

How do I install Gemma Trainer in Codex?

Run `npx skills add google-gemma/gemma-skills --skill gemma-trainer -a codex`. Or copy the skill folder (skills/gemma-trainer in google-gemma/gemma-skills) into .agents/skills/gemma-trainer in your project. Codex loads it when a task matches its description.

Can I use Gemma Trainer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-gemma/gemma-skills --skill gemma-trainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemma-trainer, .gemini/skills/gemma-trainer, .github/skills/gemma-trainer and .opencode/skills/gemma-trainer in your project.

What does Gemma Trainer need to run?

Going by SKILL.md and its folder, Gemma Trainer needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Gemma Trainer access the network?

SKILL.md names 3 domains. As links in the text: unsloth.ai, github.com and developers.google.com. This is read from the text; nothing was executed.

Is Gemma Trainer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gemma Trainer use?

Gemma Trainer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemma Trainer use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gemma Trainer?

Skills that share tags, products or a category with Gemma Trainer: Hugging Face LLM Trainer (huggingface/skills, 11k stars), Unsloth Finetuning (sickn33/agentic-awesome-skills, 47k stars), Quantized Export (wshobson/agents, 40k stars) and Wan Flf Video (artokun/comfyui-mcp, 795 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemma Trainer?

google-gemma (a GitHub organization) maintains it in google-gemma/gemma-skills, which has 1,004 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 6, 2026.

Source: google-gemma/gemma-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.