Agent skill

Train Rl

by OpenPipe in OpenPipe/ART

RL training reference for the ART framework. An agent skill from OpenPipe/ART.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Train Rl

skills CLI
$ npx skills add OpenPipe/ART --skill train-rl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenPipe/ART train-rl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenPipe/ART.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/train-rl .claude/skills/train-rl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
train-rl
GitHub stars
11k
Token cost
~2.4k tokens
SKILL.md length
1,055 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

RL training reference for the ART framework. An agent skill from OpenPipe/ART.

  • Works in 6 steps: Replayability → Task Shape → Reward Choice → …
  • The user asks to create
  • SKILL.md covers Blocking Decisions, Required Questions, 1. Replayability and 2. Task Shape, plus 7 more sections
  • Needs OPENAI_API_KEY and WANDB_API_KEY

What it does

Train Rl is an agent skill from OpenPipe/ART. RL training reference for the ART framework. Use when the user asks to create, write, or help with an RL training script, reinforcement learning, GRPO, reward functions, RULER scoring, rollout functions, or anything related to RL fine-tuning.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Fine-tuning and Reinforcement learning. It works with Qwen. The repository describes itself as: Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama… The licence is Apache-2.0.

When your agent uses it

  • The user asks to create
  • Help with an RL training script
  • Reinforcement learning
  • Reward functions

Example prompts

  • “/train-rl”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY
  • A credential in WANDB_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Replayability
  2. Task Shape
  3. Reward Choice
  4. Data Splits and Validation
  5. Base Parameters
  6. Hyperparameters

What it can do on your machine

Read from SKILL.md and the folder at commit 12162f2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • WANDB_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Train Rl loads about 2.4k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 1,055 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenPipe/ART at commit 12162f2, republished under its Apache-2.0 licence (© OpenPipe). 1,055 words, ~2,423 tokens.

Download SKILL.mdSave it as .claude/skills/train-rl/SKILL.md (or your agent's skills folder).
name
train-rl
description
RL training reference for the ART framework. Use when the user asks to create, write, or help with an RL training script, reinforcement learning, GRPO, reward functions, RULER scoring, rollout functions, or anything related to RL fine-tuning.

ART RL Workflow

Use this skill when the user wants an RL script or help adapting an existing ART agent for RL.

Keep the process simple:

  1. Inspect the existing environment or agent first.
  2. Ask one question at a time until the blocking decisions are resolved.
  3. Generate a runnable script that reuses the real rollout/tool logic instead of approximating it.

This skill is an interactive wizard. Do not write the script immediately.

Rules:

  • Ask one question at a time.
  • Wait for the user's answer before asking the next question.
  • Use the repo to make recommendations, but still ask the user to confirm every required choice.
  • Do not skip required questions just because the code suggests a likely answer.
  • Do not generate the final script until all required questions below have been answered.

Blocking Decisions

You must resolve these before writing the final script:

  1. Is the environment replayable?
  2. Is this multi-turn RL or single-turn training on static examples?
  3. Is reward programmatic, RULER, or custom?
  4. Is the backend ServerlessBackend or LocalBackend?

Required Questions

You must collect answers for all of these, one at a time, before generating code:

  1. Replayability confirmation
  2. What behavior the agent should optimize
  3. Where training scenarios come from
  4. If multi-turn: how turns work and when the episode ends
  5. Reward choice
  6. If RULER: judge model
  7. Training split or source
  8. Validation split or source
  9. Base model
  10. Project name
  11. Run name
  12. Backend
  13. Hyperparameters: use defaults or customize
  14. Iteration mode: fixed-dataset epoch loop or manual/open-ended loop

If the repo already makes an answer likely, present that as a recommendation and ask the user to confirm or correct it. That still counts as a question and still requires a user response.

1. Replayability

Inspect the repo before asking.

  • Recommend multi-turn RL when episodes can be recreated and rolled out repeatedly from the same initial state.
  • Recommend single-turn static training when the task depends on live humans, mutable production systems, or other unreproducible state.

If replayability is clear, say so and ask for confirmation. Example:

This looks replayable because each episode starts from fixed local state and the tools only read from it, so I recommend multi-turn RL. Please confirm there is no hidden live dependency.

If it is not clear, ask whether the task has a replayable environment or only logged/static scenarios.

2. Task Shape

Ask only for the details needed to implement the rollout, but do not skip the required task questions.

Always gather:

  • What the agent is supposed to accomplish.
  • Where training scenarios come from.

For multi-turn tasks, also gather:

  • What the observation is on each turn.
  • What actions or tools the agent can take.
  • When the episode ends.

When adapting an existing agent:

  • Preserve its prompt and tool behavior by default.
  • Reuse its real tool execution path and message schema.
  • Prefer storing structured terminal outputs such as final_answer directly on the trajectory when useful.
  • If the existing environment already has a natural typed terminal object or evaluation artifact, keep it in the trajectory or structured logs instead of reducing everything to free-form text.

3. Reward Choice

Start from RULER as the default.

Use this rule:

  • Choose programmatic reward only when correctness is robustly checkable with code.
  • Choose RULER for open-text answers or tool-use behavior where exact matching is brittle.
  • Choose custom only when the task genuinely mixes multiple reward sources.

Explain RULER briefly once:

  • RULER is an LLM judge that compares trajectories within a group and scores which ones are better.

If the user chooses programmatic reward:

  • Put the score in trajectory.reward.
  • Keep extra signals in trajectory.metrics.
  • Do not invent weak heuristic rewards.
  • If the environment already exposes robust auxiliary signals such as correctness, source overlap, pass/fail, completion rate, or tool-error rate, keep logging them even when they are not the main reward.

If the user chooses RULER:

  • Prefer ruler_score_group(...) with the default rubric.
  • Recommend OPENAI_API_KEY validation at startup.
  • Recommend openai/gpt-5.4 as the default judge model.
  • Ask which judge model to use before generating the script.
Show full SKILL.md (380 more words)Show less

4. Data Splits and Validation

For fixed datasets:

  • Prefer a held-out validation split.
  • Prefer capped periodic validation during training.
  • Do not run a full held-out pass at step 0 unless the user asks for it.

For validation:

  • Log validation groups at a concrete training step.
  • Make sure the metric used for checkpoint cleanup is actually logged.
  • With the default await model.delete_checkpoints(), validation must produce val/reward.

5. Base Parameters

Ask for, explicitly and separately:

  • Base model
  • Project name
  • Run name
  • Backend

Do not present a single "recommended starting point" model by default. Offer all allowed base models:

  • OpenPipe/Qwen3-14B-Instruct
  • Qwen/Qwen3-30B-A3B-Instruct-2507
  • meta-llama/Llama-3.1-8B-Instruct

Environment requirements:

  • ServerlessBackend: require WANDB_API_KEY
  • RULER: require OPENAI_API_KEY

6. Hyperparameters

Ask whether to use these starting defaults or customize them:

  • Learning rate: 1e-5
  • Rollouts per group: 4
  • Groups per step: 2

Iteration defaults:

  • For fixed datasets, prefer iterate_dataset(..., initial_step=await model.get_step()).
  • For non-fixed/open-ended generation, use a manual step loop.

Implementation Guardrails

These are the main ART-specific rules that matter in practice:

  • Reuse the real environment/agent entrypoints when they already exist.
  • Prefer building Trajectory.messages_and_choices directly for multi-turn tool use.
  • Use backend.train(model, trajectory_groups, ...) plus await model.log(...).
  • Call await backend.close() before exit.
  • Catch recoverable inference or tool errors after the trajectory has started and return a partial trajectory with numeric or bool metrics.
  • Pass art.TrajectoryGroup(...) awaitables directly into art.gather_trajectory_groups(...). Do not await them early.
  • If using RULER rescoring, prefer after_each=lambda group: ruler_score_group(...).
  • Preserve group.exceptions if you rebuild groups after rollout.
  • Default max_exceptions to scale with the active batch size, typically args.rollouts_per_group * len(batch.items) for training and the analogous validation batch size. Do not hard-code a small fixed value unless the user explicitly wants that.
  • For fixed datasets, make resume behavior explicit with initial_step=await model.get_step().
  • Delete old checkpoints by default after successful training and validation logging unless the user wants to keep all of them.

Minimal Training Pattern

Use this as the default pattern for fixed datasets with RULER:

python
from art.rewards import ruler_score_group
from art.utils.iterate_dataset import iterate_dataset


async def rollout(model: TrainableModel, scenario: Scenario) -> art.Trajectory:
    ...


for batch in iterate_dataset(
    train_scenarios,
    groups_per_step=args.groups_per_step,
    num_epochs=args.num_epochs,
    initial_step=await model.get_step(),
):
    train_groups = await art.gather_trajectory_groups(
        [
            art.TrajectoryGroup(
                (rollout(model, scenario) for _ in range(args.rollouts_per_group)),
                metadata={"scenario_id": scenario.id},
            )
            for scenario in batch.items
        ],
        after_each=lambda group: ruler_score_group(
            group,
            judge_model=args.judge_model,
        ),
        max_exceptions=args.rollouts_per_group * len(batch.items),
    )

    train_result = await backend.train(
        model,
        train_groups,
        learning_rate=args.learning_rate,
    )

    await model.log(
        train_groups,
        metrics=train_result.metrics,
        step=train_result.step,
        split="train",
    )

    if should_validate(train_result.step):
        val_groups = await art.gather_trajectory_groups(
            [
                art.TrajectoryGroup(
                    (rollout(model, scenario) for _ in range(args.rollouts_per_group)),
                    metadata={"scenario_id": scenario.id},
                )
                for scenario in validation_scenarios
            ],
            after_each=lambda group: ruler_score_group(
                group,
                judge_model=args.judge_model,
            ),
            max_exceptions=args.rollouts_per_group * len(validation_scenarios),
        )
        await model.log(
            val_groups,
            metrics={"reward": ...},
            step=train_result.step,
            split="val",
        )
        await model.delete_checkpoints()

Final Script Requirements

Every generated script should:

  • Validate required environment variables early.
  • Follow the repo's existing env-loading convention.
  • Use the selected backend.
  • Print the final inference model name and a short usage example.
  • Close the backend cleanly.

If you fail to find enough information from the repo, say what is missing and ask the next single blocking question. Do not fabricate environment behavior, reward logic, or dataset structure.

© OpenPipe, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/train-rl of OpenPipe/ART.

Open the folder on GitHubat commit 12162f2

Compare with similar skills

Train Rl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Train Rl compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Train Rl this skillOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
slime RL Post-TrainingOrchestra-Research/AI-Research-SKILLs13k5 repos~2.8kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.0
Fine Tuning With TrlOrchestra-Research/AI-Research-SKILLs13k7 repos~2.9kAutomated safety check: PassMIT
Optim AgentOptim-Agent/optim-agent800—~1.3kAutomated safety check: PassMIT

Similar skills

  • slime RL Post-Training

    Orchestra-Research/AI-Research-SKILLs

    Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

    13k GitHub starsUsed in 5 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Fine Tuning With Trl

    Orchestra-Research/AI-Research-SKILLs

    Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.

    13k GitHub starsUsed in 7 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Optim Agent

    Optim-Agent/optim-agent

    A skill your agent uses when the user wants to optimize configurable system parameters against a measurable scalar objective, especially for model training, inference, quantitative strategies…

    800 GitHub stars~1.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Safactory Workflows

    AI45Lab/SAfactory

    Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.

    236 GitHub stars~1.8k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed

More from OpenPipe/ART

  • Fix Art Issues

    OpenPipe/ART

    Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

    11k GitHub stars~840 tokensUpdated 2 days ago
    Auto-check: notes
  • Train Sft

    OpenPipe/ART

    SFT training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Train Rl

What does Train Rl do?

RL training reference for the ART framework. An agent skill from OpenPipe/ART. Train Rl is an agent skill from OpenPipe/ART. RL training reference for the ART framework.

When should I use Train Rl?

Train Rl fits situations like: the user asks to create; help with an RL training script; reinforcement learning; reward functions.

How do I install Train Rl in Claude Code?

Run `npx skills add OpenPipe/ART --skill train-rl -a claude-code`. Or copy the skill folder (.agents/skills/train-rl in OpenPipe/ART) into .claude/skills/train-rl in your project. Claude Code loads it when a task matches its description.

How do I install Train Rl in Codex?

Run `npx skills add OpenPipe/ART --skill train-rl -a codex`. Or copy the skill folder (.agents/skills/train-rl in OpenPipe/ART) into .agents/skills/train-rl in your project. Codex loads it when a task matches its description.

Can I use Train Rl in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenPipe/ART --skill train-rl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/train-rl, .gemini/skills/train-rl, .github/skills/train-rl and .opencode/skills/train-rl in your project.

What does Train Rl need to run?

Going by SKILL.md and its folder, Train Rl needs credentials named OPENAI_API_KEY and WANDB_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in WANDB_API_KEY.

Does Train Rl access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Train Rl safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Train Rl use?

Train Rl is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Train Rl use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Train Rl?

Skills that share tags, products or a category with Train Rl: slime RL Post-Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Fine Tuning With Trl (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Train Rl?

OpenPipe (a GitHub organization) maintains it in OpenPipe/ART, which has 10,786 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 5, 2026.

Source: OpenPipe/ART on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.