slime RL Post-Training
Orchestra-Research/AI-Research-SKILLs
Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.
RL training reference for the ART framework. An agent skill from OpenPipe/ART.
$ npx skills add OpenPipe/ART --skill train-rl -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install OpenPipe/ART train-rl --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/OpenPipe/ART.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/train-rl .claude/skills/train-rl && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "train-rl" agent skill from https://github.com/OpenPipe/ART/tree/main/.agents/skills/train-rl into .claude/skills/train-rl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-rl", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/OpenPipe/ART/tree/main/.agents/skills/train-rlType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add OpenPipe/ART --skill train-rl -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install OpenPipe/ART train-rl --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenPipe/ART.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/train-rl .agents/skills/train-rl && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "train-rl" agent skill from https://github.com/OpenPipe/ART/tree/main/.agents/skills/train-rl into .agents/skills/train-rl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-rl", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OpenPipe/ART --skill train-rl -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install OpenPipe/ART train-rl --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenPipe/ART.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/train-rl .cursor/skills/train-rl && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "train-rl" agent skill from https://github.com/OpenPipe/ART/tree/main/.agents/skills/train-rl into .cursor/skills/train-rl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-rl", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/OpenPipe/ART.git --path .agents/skills/train-rl--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add OpenPipe/ART --skill train-rl -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install OpenPipe/ART train-rl --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenPipe/ART.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/train-rl .gemini/skills/train-rl && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "train-rl" agent skill from https://github.com/OpenPipe/ART/tree/main/.agents/skills/train-rl into .gemini/skills/train-rl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-rl", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install OpenPipe/ART train-rlInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add OpenPipe/ART --skill train-rl -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/OpenPipe/ART.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/train-rl .github/skills/train-rl && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "train-rl" agent skill from https://github.com/OpenPipe/ART/tree/main/.agents/skills/train-rl into .github/skills/train-rl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-rl", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OpenPipe/ART --skill train-rl -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install OpenPipe/ART train-rl --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenPipe/ART.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/train-rl .opencode/skills/train-rl && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "train-rl" agent skill from https://github.com/OpenPipe/ART/tree/main/.agents/skills/train-rl into .opencode/skills/train-rl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "train-rl", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
train-rlRL training reference for the ART framework. An agent skill from OpenPipe/ART.
Train Rl is an agent skill from OpenPipe/ART. RL training reference for the ART framework. Use when the user asks to create, write, or help with an RL training script, reinforcement learning, GRPO, reward functions, RULER scoring, rollout functions, or anything related to RL fine-tuning.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Fine-tuning and Reinforcement learning. It works with Qwen. The repository describes itself as: Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama… The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 12162f2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYWANDB_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Train Rl loads about 2.4k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 1,055 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from OpenPipe/ART at commit 12162f2, republished under its Apache-2.0 licence (© OpenPipe). 1,055 words, ~2,423 tokens.
.claude/skills/train-rl/SKILL.md (or your agent's skills folder).Use this skill when the user wants an RL script or help adapting an existing ART agent for RL.
Keep the process simple:
This skill is an interactive wizard. Do not write the script immediately.
Rules:
You must resolve these before writing the final script:
ServerlessBackend or LocalBackend?You must collect answers for all of these, one at a time, before generating code:
If the repo already makes an answer likely, present that as a recommendation and ask the user to confirm or correct it. That still counts as a question and still requires a user response.
Inspect the repo before asking.
multi-turn RL when episodes can be recreated and rolled out repeatedly from the same initial state.single-turn static training when the task depends on live humans, mutable production systems, or other unreproducible state.If replayability is clear, say so and ask for confirmation. Example:
This looks replayable because each episode starts from fixed local state and the tools only read from it, so I recommend multi-turn RL. Please confirm there is no hidden live dependency.
If it is not clear, ask whether the task has a replayable environment or only logged/static scenarios.
Ask only for the details needed to implement the rollout, but do not skip the required task questions.
Always gather:
For multi-turn tasks, also gather:
When adapting an existing agent:
final_answer directly on the trajectory when useful.Start from RULER as the default.
Use this rule:
programmatic reward only when correctness is robustly checkable with code.RULER for open-text answers or tool-use behavior where exact matching is brittle.custom only when the task genuinely mixes multiple reward sources.Explain RULER briefly once:
RULER is an LLM judge that compares trajectories within a group and scores which ones are better.If the user chooses programmatic reward:
trajectory.reward.trajectory.metrics.If the user chooses RULER:
ruler_score_group(...) with the default rubric.OPENAI_API_KEY validation at startup.openai/gpt-5.4 as the default judge model.For fixed datasets:
0 unless the user asks for it.For validation:
await model.delete_checkpoints(), validation must produce val/reward.Ask for, explicitly and separately:
Do not present a single "recommended starting point" model by default. Offer all allowed base models:
OpenPipe/Qwen3-14B-InstructQwen/Qwen3-30B-A3B-Instruct-2507meta-llama/Llama-3.1-8B-InstructEnvironment requirements:
ServerlessBackend: require WANDB_API_KEYRULER: require OPENAI_API_KEYAsk whether to use these starting defaults or customize them:
1e-542Iteration defaults:
iterate_dataset(..., initial_step=await model.get_step()).These are the main ART-specific rules that matter in practice:
Trajectory.messages_and_choices directly for multi-turn tool use.backend.train(model, trajectory_groups, ...) plus await model.log(...).await backend.close() before exit.art.TrajectoryGroup(...) awaitables directly into art.gather_trajectory_groups(...). Do not await them early.after_each=lambda group: ruler_score_group(...).group.exceptions if you rebuild groups after rollout.max_exceptions to scale with the active batch size, typically args.rollouts_per_group * len(batch.items) for training and the analogous validation batch size. Do not hard-code a small fixed value unless the user explicitly wants that.initial_step=await model.get_step().Use this as the default pattern for fixed datasets with RULER:
from art.rewards import ruler_score_group
from art.utils.iterate_dataset import iterate_dataset
async def rollout(model: TrainableModel, scenario: Scenario) -> art.Trajectory:
...
for batch in iterate_dataset(
train_scenarios,
groups_per_step=args.groups_per_step,
num_epochs=args.num_epochs,
initial_step=await model.get_step(),
):
train_groups = await art.gather_trajectory_groups(
[
art.TrajectoryGroup(
(rollout(model, scenario) for _ in range(args.rollouts_per_group)),
metadata={"scenario_id": scenario.id},
)
for scenario in batch.items
],
after_each=lambda group: ruler_score_group(
group,
judge_model=args.judge_model,
),
max_exceptions=args.rollouts_per_group * len(batch.items),
)
train_result = await backend.train(
model,
train_groups,
learning_rate=args.learning_rate,
)
await model.log(
train_groups,
metrics=train_result.metrics,
step=train_result.step,
split="train",
)
if should_validate(train_result.step):
val_groups = await art.gather_trajectory_groups(
[
art.TrajectoryGroup(
(rollout(model, scenario) for _ in range(args.rollouts_per_group)),
metadata={"scenario_id": scenario.id},
)
for scenario in validation_scenarios
],
after_each=lambda group: ruler_score_group(
group,
judge_model=args.judge_model,
),
max_exceptions=args.rollouts_per_group * len(validation_scenarios),
)
await model.log(
val_groups,
metrics={"reward": ...},
step=train_result.step,
split="val",
)
await model.delete_checkpoints()Every generated script should:
If you fail to find enough information from the repo, say what is missing and ask the next single blocking question. Do not fabricate environment behavior, reward logic, or dataset structure.
© OpenPipe, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/train-rl of OpenPipe/ART.
Open the folder on GitHubat commit 12162f2
Train Rl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Train Rl this skillOpenPipe/ART | 11k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| slime RL Post-TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 3 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~830 | Automated safety check: Pass | Apache-2.0 | |
| Fine Tuning With TrlOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Optim AgentOptim-Agent/optim-agent | 800 | — | ~1.3k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
R6410418/Jackrong-llm-finetuning-guide
Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.
Orchestra-Research/AI-Research-SKILLs
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.
Optim-Agent/optim-agent
A skill your agent uses when the user wants to optimize configurable system parameters against a measurable scalar objective, especially for model training, inference, quantitative strategies…
AI45Lab/SAfactory
Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.
OpenPipe/ART
Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.
OpenPipe/ART
SFT training reference for the ART framework. An agent skill from OpenPipe/ART.
Works with
Categories
RL training reference for the ART framework. An agent skill from OpenPipe/ART. Train Rl is an agent skill from OpenPipe/ART. RL training reference for the ART framework.
Train Rl fits situations like: the user asks to create; help with an RL training script; reinforcement learning; reward functions.
Run `npx skills add OpenPipe/ART --skill train-rl -a claude-code`. Or copy the skill folder (.agents/skills/train-rl in OpenPipe/ART) into .claude/skills/train-rl in your project. Claude Code loads it when a task matches its description.
Run `npx skills add OpenPipe/ART --skill train-rl -a codex`. Or copy the skill folder (.agents/skills/train-rl in OpenPipe/ART) into .agents/skills/train-rl in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenPipe/ART --skill train-rl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/train-rl, .gemini/skills/train-rl, .github/skills/train-rl and .opencode/skills/train-rl in your project.
Going by SKILL.md and its folder, Train Rl needs credentials named OPENAI_API_KEY and WANDB_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in WANDB_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Train Rl is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Train Rl: slime RL Post-Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Fine Tuning With Trl (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
OpenPipe (a GitHub organization) maintains it in OpenPipe/ART, which has 10,786 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 5, 2026.
Source: OpenPipe/ART on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.