Agent skill

Rl Post Training

by benchflow-ai in benchflow-ai/skillsbench

Diagnostic guide for RL-based post-training of language models (GRPO, PPO, REINFORCE, DPO).

Apache-2.0Auto-check passedAI & LLM Engineering

Install Rl Post Training

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill rl-post-training -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench rl-post-training --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/debug-trl-grpo/environment/skills/rl-post-training .claude/skills/rl-post-training && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rl-post-training
GitHub stars
1.8k
Token cost
~1.6k tokens
SKILL.md length
787 words
Files
4 (incl. scripts, references)
Skills in repo
189
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnostic guide for RL-based post-training of language models (GRPO, PPO, REINFORCE, DPO).

  • Works in 5 steps: Verify Reward Signal → Verify Advantage Computation → Verify Log-Probability Computation → …
  • Tasks that involve Fine-tuning
  • SKILL.md covers Core Concepts, Diagnostic Methodology, Fixing, Not Rewriting and Key Mathematical Invariants, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Rl Post Training is an agent skill from benchflow-ai/skillsbench. Diagnostic guide for RL-based post-training of language models (GRPO, PPO, REINFORCE, DPO). Use proactively when debugging a training pipeline that shows no improvement, loss anomalies, reward stagnation, NaN gradients, or other unexpected behavior during reinforcement learning fine-tuning. Work through all pipeline stages — reward, advantages, log-probs, loss, generation/decoding — before concluding the diagnosis is complete; stopping after one or two fixes is a common failure mode. Covers log-probability math…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/common-pitfalls.md`, `references/diagnostic-workflow.md` and `scripts/verify_pipeline.py`).

It sits in AI & LLM Engineering, covering Fine-tuning and Reinforcement learning. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Fine-tuning
  • Tasks that involve Reinforcement learning

Example prompts

  • “/rl-post-training”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Verify Reward Signal
  2. Verify Advantage Computation
  3. Verify Log-Probability Computation
  4. Verify Loss Computation
  5. Verify Generation and Decoding

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rl Post Training loads about 1.6k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 159 tokens; SKILL.md has 787 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 787 words, ~1,644 tokens.

Download SKILL.mdSave it as .claude/skills/rl-post-training/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
rl-post-training
description
Diagnostic guide for RL-based post-training of language models (GRPO, PPO, REINFORCE, DPO). Use proactively when debugging a training pipeline that shows no improvement, loss anomalies, reward stagnation, NaN gradients, or other unexpected behavior during reinforcement learning fine-tuning. Work through all pipeline stages — reward, advantages, log-probs, loss, generation/decoding — before concluding the diagnosis is complete; stopping after one or two fixes is a common failure mode. Covers log-probability math, advantage estimation, numerical stability, reward processing, and generation/decoding pipeline issues.

RL Post-Training — Concepts & Diagnostic Guide

Core Concepts

RL post-training optimizes a language model's policy using reward signals. The standard pipeline:

prompt → generate completions → score with reward → compute advantages → policy gradient update

Each stage has distinct failure modes. When a model "shows no improvement," the bug could be anywhere in this pipeline.

Diagnostic Methodology

When RL training produces no improvement, work through these stages in order. Each stage depends on the previous one being correct.

Stage 1: Verify Reward Signal
  • Are rewards non-constant? If all rewards are identical, there is no learning signal.
  • Do rewards correlate with completion quality? Spot-check decoded completions against their scores.
  • Is the reward function being called on the correct text? Check that decoding/stripping preserves the content the reward function needs to evaluate.
Stage 2: Verify Advantage Computation
  • Are advantages non-zero when rewards vary? If they collapse to ~0, the policy gradient vanishes.
  • Check the magnitude and dtype of every numerical-stability constant in the advantage path (additive epsilons, clipping bounds). Compare each to what the math requires.
  • Check the group size G. G ≤ 2 makes std either undefined or extremely noisy.
Stage 3: Verify Log-Probability Computation
  • Verify bounds: log-probs of valid tokens must be non-positive.
  • Compare your implementation against F.log_softmax on a small deterministic input — a numerical match rules out sign errors, wrong gathering axis, and off-by-one subtraction.
  • On a near-one-hot input, confirm the dominant token's log-prob is close to 0 (not close to the min).
Stage 4: Verify Loss Computation
  • Is the loss changing across steps? Flat loss suggests zero gradients upstream.
  • Log the KL term and the policy-gradient term separately. Either dominating the other is diagnostic.
  • Check the fraction of clipped samples. Near-100% clipping means the clip range is starving the signal.
Stage 5: Verify Generation and Decoding
  • Are padding and decoder artefacts stripped from the text the reward function sees?
  • Does every completion shape your model can emit survive the decoding path with non-empty output where a human would expect non-empty output? Include degenerate cases (no formatting markers, only a prefix, only a suffix) in the round-trip test.
  • Print or log a sample of the actual strings handed to the reward function — mismatches between "what the model generated" and "what the reward saw" are often visible at a glance.

Fixing, Not Rewriting

A bug in a branched function is a bug in one branch, not a verdict on the whole function. Before editing:

  • Enumerate the input shapes the function handles today and the output each branch produces. The branches exist because callers rely on them.
  • Identify which input-output pairs violate the intended contract. Those are the only branches you need to change.
  • If your diff collapses a multi-branch function to a one-liner, you have almost certainly broken a contract a different caller depends on. Re-read the call sites before committing.

The same principle applies to epsilon values, sign conventions, and clipping bounds — if a constant looks wrong, replace it with a correct constant; don't remove the surrounding numerical-stability logic.

Show full SKILL.md (303 more words)Show less

Key Mathematical Invariants

These invariants should always hold. If any is violated, there is a bug.

InvariantWhat it meansHow to check
log_prob <= 0Log of a probability is non-positiveassert (log_probs <= 1e-6).all()
log-probs match F.log_softmaxManual implementation equals the referencetorch.allclose(manual, F.log_softmax(...).gather(...))
sum(softmax(logits)) == 1Probabilities sum to 1assert torch.allclose(softmax.sum(-1), ones)
0 < epsilon << 1Additive epsilons exist for numerical stability onlyassert 0 < epsilon < 1
advantages != 0 when rewards varyNon-constant rewards yield non-zero advantagesassert advantages.abs().max() > 0.1
Decoding preserves contentEvery generation shape the model emits survives the decode path with the content a human reader would expectRound-trip representative samples; confirm non-empty where non-empty is expected

Common Pitfall Categories

Detailed pitfall catalog with examples is in references/common-pitfalls.md. Summary:

  1. Sign errors in log-space — log-softmax, KL divergence, DPO log-ratio
  2. Numerical stability constants out of range — additive epsilons, clip bounds, temperature
  3. String processing in decoding — cleanup that blanks shapes the rules didn't anticipate
  4. Reward / decoding format mismatch — reward function sees text with the shape it needs removed
  5. Reference model drift — reference model not frozen, or gradients flowing through it
  6. Gradient flow breakage — detached tensors, in-place ops, .item()/numpy round-trips in the loss path
  7. Configuration misuse — SFT-scale LR, KL coefficient dominating, clip range too tight, G too small

Available References

FileContentsWhen to load
references/common-pitfalls.mdCatalog of pitfall categories with symptoms and detection strategiesWhen you have a suspicious area and want to match it against known pattern shapes
references/diagnostic-workflow.mdStage-by-stage diagnostic procedures with verification snippetsWhen you need guidance on which invariant to check in which order
scripts/verify_pipeline.pyRunnable diagnostic that prints log-prob / advantage / decoding values on small fixed inputs so you can compare them against the invariants aboveBefore declaring a fix done: run this and inspect the output, paying attention to lines marked ?

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in tasks/debug-trl-grpo/environment/skills/rl-post-training of benchflow-ai/skillsbench.

  • SKILL.md
  • references/common-pitfalls.md
  • references/diagnostic-workflow.md
  • scripts/verify_pipeline.py

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Rl Post Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rl Post Training compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rl Post Training this skillbenchflow-ai/skillsbench1.8k—~1.6kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.0
Fine Tuning With TrlOrchestra-Research/AI-Research-SKILLs13k6 repos~2.9kAutomated safety check: PassMIT
Optim AgentOptim-Agent/optim-agent801—~1.3kAutomated safety check: PassMIT
Safactory WorkflowsAI45Lab/SAfactory236—~1.8kAutomated safety check: PassNone

Similar skills

  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Fine Tuning With Trl

    Orchestra-Research/AI-Research-SKILLs

    Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.

    13k GitHub starsUsed in 6 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Optim Agent

    Optim-Agent/optim-agent

    A skill your agent uses when the user wants to optimize configurable system parameters against a measurable scalar objective, especially for model training, inference, quantitative strategies…

    801 GitHub stars~1.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Safactory Workflows

    AI45Lab/SAfactory

    Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.

    236 GitHub stars~1.8k tokensUpdated 15 days ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed

More from benchflow-ai/skillsbench

All 189 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Rl Post Training

What does Rl Post Training do?

Diagnostic guide for RL-based post-training of language models (GRPO, PPO, REINFORCE, DPO). Rl Post Training is an agent skill from benchflow-ai/skillsbench. Diagnostic guide for RL-based post-training of language models (GRPO, PPO, REINFORCE, DPO).

When should I use Rl Post Training?

Rl Post Training fits situations like: tasks that involve Fine-tuning; tasks that involve Reinforcement learning.

How do I install Rl Post Training in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill rl-post-training -a claude-code`. Or copy the skill folder (tasks/debug-trl-grpo/environment/skills/rl-post-training in benchflow-ai/skillsbench) into .claude/skills/rl-post-training in your project. Claude Code loads it when a task matches its description.

How do I install Rl Post Training in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill rl-post-training -a codex`. Or copy the skill folder (tasks/debug-trl-grpo/environment/skills/rl-post-training in benchflow-ai/skillsbench) into .agents/skills/rl-post-training in your project. Codex loads it when a task matches its description.

Can I use Rl Post Training in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill rl-post-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rl-post-training, .gemini/skills/rl-post-training, .github/skills/rl-post-training and .opencode/skills/rl-post-training in your project.

What does Rl Post Training need to run?

Going by SKILL.md and its folder, Rl Post Training needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Rl Post Training access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rl Post Training safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Rl Post Training use?

Rl Post Training is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rl Post Training use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.6k tokens, read only when the agent opens those files.

What are the alternatives to Rl Post Training?

Skills that share tags, products or a category with Rl Post Training: Train Rl (OpenPipe/ART, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Fine Tuning With Trl (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Optim Agent (Optim-Agent/optim-agent, 801 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rl Post Training?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,834 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.