Official agent skill

TRL Post-Training

by huggingface in huggingface/skills

Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install TRL Post-Training

skills CLI
$ npx skills add huggingface/skills --skill trl-training -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huggingface/skills trl-training --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/trl-training .claude/skills/trl-training && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
trl-training
GitHub stars
11k
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
301 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA.

  • Writing a TRL script for supervised fine-tuning or preference tuning
  • SKILL.md covers Dataset formats, SFT: the fields that matter, GRPO: online RL and CLI
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Debugging dataset format or reward function errors in GRPO training

What it does

Each TRL method pairs a `*Trainer` class with a `*Config` dataclass, and the skill maps them to the data they expect: `SFTTrainer` for language-modeling or prompt-completion data, `DPOTrainer` for chosen and rejected pairs, `GRPOTrainer` for prompts plus reward functions, `KTOTrainer` for unpaired labels, `RewardTrainer` for a scalar reward model and `DistillationTrainer` for on-policy distillation with a teacher. Less stable trainers sit in `trl.experimental`.

Practical rules follow. Pass the model as a string and put loading options in `model_init_kwargs`, let the chat template apply itself to conversational datasets, and give a `LoraConfig` through `peft_config` for LoRA. SFT settings such as `max_length` and `packing` are explained. GRPO reward functions receive keyword arguments including `prompts` and `completions`, should accept extra keyword arguments for dataset columns they ignore, and return one float per completion.

When your agent uses it

  • Writing a TRL script for supervised fine-tuning or preference tuning
  • Debugging dataset format or reward function errors in GRPO training
  • Choosing between DPO, KTO and reward-model training for the data you have

Example prompts

  • “Write an SFT script with TRL for my chat dataset and add LoRA.”
  • “My GRPO reward function crashes on extra dataset columns, so fix it.”
  • “Set up DPO training with TRL on a chosen and rejected preference dataset.”

Requirements

  • Python with `trl`, `transformers` and `datasets` installed

What it can do on your machine

Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

TRL Post-Training loads about 1.1k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 301 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 301 words, ~1,078 tokens.

Download SKILL.mdSave it as .claude/skills/trl-training/SKILL.md (or your agent's skills folder).
name
trl-training
description
Post-train LLMs with TRL (Transformers Reinforcement Learning) — SFT, DPO, GRPO, KTO, and reward-model training. Use when writing or debugging training code with the TRL Python API or the trl CLI.
license
Apache-2.0
metadata.author
huggingface
metadata.documentation
https://huggingface.co/docs/trl

TRL

Each method pairs a *Trainer class with a *Config dataclass. Configs extend transformers.TrainingArguments, so all of its arguments work in any trainer config.

TrainerDataset type
SFTTrainerlanguage modeling or prompt-completion
DPOTrainerpreference (chosen/rejected pairs)
GRPOTrainerprompt-only + reward function(s)
DistillationTrainerprompt-only + a teacher model (on-policy distillation)
KTOTrainerunpaired preference (per-sample bool label)
RewardTrainerpreference (chosen/rejected pairs); trains a scalar reward model, not a policy

Many more trainers (OnlineDPO, ORPO, CPO, GKD, …) live in trl.experimental with unstable APIs: https://huggingface.co/docs/trl/experimental_overview

python
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer

trainer = SFTTrainer(
    model="Qwen/Qwen2.5-0.5B",  # model ID or a PreTrainedModel instance
    args=SFTConfig(output_dir="Qwen2.5-0.5B-SFT"),
    train_dataset=load_dataset("trl-lib/Capybara", split="train"),
)
trainer.train()

Pass model as a string and route loading kwargs through model_init_kwargs (e.g. {"dtype": "bfloat16", "attn_implementation": "kernels-community/flash-attn2"}) instead of calling from_pretrained yourself. The tokenizer/processor is inferred from the model; pass processing_class only when it differs. For LoRA, pass peft_config=LoraConfig(...).

Dataset formats

Conversational: {"messages": [{"role": ..., "content": ...}]} (language modeling) or {"prompt": [...], "completion": [...]}. The chat template is applied automatically — never apply it yourself. Extra columns are allowed; GRPO forwards them to reward functions. Reference: https://huggingface.co/docs/trl/dataset_formats

SFT: the fields that matter

python
SFTConfig(
    max_length=1024,        # truncation length; None disables truncation
    packing=True,           # pack sequences into max_length blocks: fewer pad tokens, higher throughput
    padding_free=True,      # flatten batch, no padding; requires FlashAttention; implied by packing
    use_liger_kernel=True,  # fused Liger kernels, reduces peak memory
    assistant_only_loss=True,  # loss only on assistant turns (conversational datasets)
)

GRPO: online RL

python
def reward_len(completions, **kwargs):
    return [-abs(20 - len(c[0]["content"])) for c in completions]

trainer = GRPOTrainer(
    model="Qwen/Qwen2.5-0.5B-Instruct",
    reward_funcs=reward_len,  # or a list; rewards are summed
    args=GRPOConfig(output_dir="Qwen2.5-0.5B-GRPO", max_completion_length=512),
    train_dataset=load_dataset("trl-lib/DeepMath-103K", split="train"),
)

Reward functions are called with keyword arguments prompts, completions, completion_ids, trainer_state, plus every extra dataset column — accept **kwargs for the ones you ignore. Return list[float], one reward per completion. With conversational data, completions is a list of message lists, not strings.

The generation batch is per_device_train_batch_size × num_processes × steps_per_generation (or set generation_batch_size directly) and must be divisible by num_generations (default 8). Generation is the usual bottleneck — enable vLLM with use_vllm=True: vllm_mode="colocate" shares the training GPUs (size with vllm_gpu_memory_utilization); vllm_mode="server" uses a separate trl vllm-serve --model <model_id>.

AsyncGRPOTrainer (trl.experimental.async_grpo) implements the same algorithm with generation decoupled from training: a background worker streams completions from a vLLM server while the training loop consumes them, so the two overlap instead of alternating.

CLI

Flags mirror the config fields: trl sft --model_name_or_path Qwen/Qwen2.5-0.5B --dataset_name trl-lib/Capybara. YAML via --config; distributed presets via --accelerate_config zero3 (Python scripts: accelerate launch train.py).

© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/trl-training of huggingface/skills.

Open the folder on GitHubat commit ca0325b

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

TRL Post-Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

TRL Post-Training compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
TRL Post-Training this skillhuggingface/skills11k1 repos~1.1kAutomated safety check: PassApache-2.0
verl RL TrainingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.4kAutomated safety check: PassMIT
bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs13k3 repos~2.5kAutomated safety check: PassMIT
Hugging Face Transformers Usagedavila7/claude-code-templates32k12 repos~1.2kAutomated safety check: PassMIT
Huggingface LLM Trainerwaybarrios/opencode-power-pack533—~3kAutomated safety check: PassApache-2.0
SimPO Preference TrainingOrchestra-Research/AI-Research-SKILLs13k5 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • verl RL Training

    Orchestra-Research/AI-Research-SKILLs

    Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends.

    13k GitHub starsUsed in 3 repos~2.4k tokens
    AI & LLM EngineeringAuto-check passed
  • bitsandbytes Model Quantization

    Orchestra-Research/AI-Research-SKILLs

    Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

    13k GitHub starsUsed in 3 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Transformers Usage

    davila7/claude-code-templates

    Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

    32k GitHub starsUsed in 12 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface LLM Trainer

    waybarrios/opencode-power-pack

    Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

    533 GitHub stars~3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • SimPO Preference Training

    Orchestra-Research/AI-Research-SKILLs

    Walks through aligning language models with SimPO, a reference-free preference optimization method, using accelerate configs for Mistral 7B, Llama 3 8B and math-focused models.

    13k GitHub starsUsed in 5 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • torchforge RL Training

    Orchestra-Research/AI-Research-SKILLs

    Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs.

    13k GitHub starsUsed in 3 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed

More from huggingface/skills

All 25 skills in this repo
  • Official

    Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

    11k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    Auto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed

Questions about TRL Post-Training

What does TRL Post-Training do?

Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA. Each TRL method pairs a `*Trainer` class with a `*Config` dataclass, and the skill maps them to the data they expect: `SFTTrainer` for language-modeling or prompt-completion data, `DPOTrainer` for chosen and rejected pairs, `GRPOTrainer` for prompts plus reward functions, `KTOTrainer` for unpaired labels, `RewardTrainer` for a scalar reward model and `DistillationTrainer` for on-policy distillation with a teacher.experimental`.

When should I use TRL Post-Training?

TRL Post-Training fits situations like: writing a TRL script for supervised fine-tuning or preference tuning; debugging dataset format or reward function errors in GRPO training; choosing between DPO, KTO and reward-model training for the data you have.

How do I install TRL Post-Training in Claude Code?

Run `npx skills add huggingface/skills --skill trl-training -a claude-code`. Or copy the skill folder (skills/trl-training in huggingface/skills) into .claude/skills/trl-training in your project. Claude Code loads it when a task matches its description.

How do I install TRL Post-Training in Codex?

Run `npx skills add huggingface/skills --skill trl-training -a codex`. Or copy the skill folder (skills/trl-training in huggingface/skills) into .agents/skills/trl-training in your project. Codex loads it when a task matches its description.

Can I use TRL Post-Training in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill trl-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/trl-training, .gemini/skills/trl-training, .github/skills/trl-training and .opencode/skills/trl-training in your project.

What does TRL Post-Training need to run?

SKILL.md names no scripts, command-line tools or credentials: TRL Post-Training is instructions for the agent only. Our summary lists: Python with `trl`, `transformers` and `datasets` installed.

Does TRL Post-Training access the network?

SKILL.md names 1 domain. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is TRL Post-Training safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does TRL Post-Training use?

TRL Post-Training is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does TRL Post-Training use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to TRL Post-Training?

Skills that share tags, products or a category with TRL Post-Training: verl RL Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), bitsandbytes Model Quantization (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Transformers Usage (davila7/claude-code-templates, 32k stars) and Huggingface LLM Trainer (waybarrios/opencode-power-pack, 533 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains TRL Post-Training?

huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,148 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.

Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.