Topic · AI & LLM Engineering

Best reinforcement learning skills for Claude Code, Codex and other agents.

Skills that train and evaluate reinforcement-learning agents.
skills
67
official
4

Reinforcement learning skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Reinforcement learning skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

huggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.06 days ago
2

Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

OpenPipe/ART11k—~840Automated safety check: NotesApache-2.02 days ago
3

RL training reference for the ART framework. An agent skill from OpenPipe/ART.

OpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.02 days ago
4

Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

R6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.02 mo ago
5

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.

Orchestra-Research/AI-Research-SKILLs13k7 repos~2.9kAutomated safety check: PassMIT3 mo ago
6

A skill your agent uses when the user wants to optimize configurable system parameters against a measurable scalar objective, especially for model training, inference, quantitative strategies…

Optim-Agent/optim-agent800—~1.3kAutomated safety check: PassMIT1 mo ago
7

Builds a Verifiers (PrimeIntellect) variant of an RL environment.

adithya-s-k/FineEnvs4211 repo~2.3kAutomated safety check: PassApache-2.0yesterday
8

Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.

AI45Lab/SAfactory236—~1.8kAutomated safety check: PassNo licence13 days ago
9

Router for adding a diffusion or omni pipeline to verl-omni.

verl-project/verl-omni1.2k—~1kAutomated safety check: PassApache-2.0today
10

Guide for adding a new reward scorer to verl-omni and wiring it into a run.

verl-project/verl-omni1.2k—~648Automated safety check: PassApache-2.0today
11

Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

waybarrios/opencode-power-pack533—~3kAutomated safety check: PassApache-2.0yesterday
12

Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs.

adithya-s-k/FineEnvs421—~2.1kAutomated safety check: PassApache-2.0yesterday
13

Run, configure, retry, and validate AReno SFT, DPO, GSPO, GRPO, PPO, and agentic training.

inclusionAI/AReno323—~782Automated safety check: PassApache-2.013 days ago
14

Walks through aligning language models with SimPO, a reference-free preference optimization method, using accelerate configs for Mistral 7B, Llama 3 8B and math-focused models.

Orchestra-Research/AI-Research-SKILLs13k5 repos~1.5kAutomated safety check: PassMIT3 mo ago
15

Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.8kAutomated safety check: PassMIT3 mo ago
16

Builds an OpenEnv (Hugging Face) variant of an RL environment.

adithya-s-k/FineEnvs421—~2.4kAutomated safety check: PassApache-2.0yesterday
17

Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package.

adithya-s-k/FineEnvs421—~2.3kAutomated safety check: NotesApache-2.0yesterday
18

Add or modify an AReno algorithm, trainer, loss, advantage calculation, role model, or algorithm-specific configuration.

inclusionAI/AReno323—~501Automated safety check: PassApache-2.013 days ago
19

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.2kAutomated safety check: PassMIT3 mo ago
20

High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.1kAutomated safety check: NotesMIT3 mo ago
21

Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.5kAutomated safety check: PassMIT3 mo ago
22

Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.4kAutomated safety check: PassMIT3 mo ago
23

This skill should be used when working with reinforcement learning tasks including high-performance RL training, custom environment development, vectorized parallel simulation, multi-agent systems…

davila7/claude-code-templates32k9 repos~3.4kAutomated safety check: PassMITtoday
24

A skill your agent uses for reinforcement learning tasks including training RL agents (PPO, SAC, DQN, TD3, DDPG, A2C, etc.), creating custom Gym environments, implementing callbacks for monitoring…

davila7/claude-code-templates32k9 repos~2.4kAutomated safety check: PassMITtoday
25

Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym.

adithya-s-k/FineEnvs421—~2.3kAutomated safety check: NotesApache-2.0yesterday
26

Explains how to train a model to be harmless with self-critique, revision and AI-generated preference feedback, with Hugging Face and TRL code for each stage.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2kAutomated safety check: PassMIT3 mo ago
27

Guides GRPO reinforcement-learning fine-tuning of language models with TRL, centered on designing reward functions for formats, verifiable tasks and reasoning.

Orchestra-Research/AI-Research-SKILLs13k4 repos~4.3kAutomated safety check: PassMIT3 mo ago
28

Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths.

LegoX/Lego-RL108—~4.2kAutomated safety check: PassApache-2.02 days ago
29

Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA.

huggingface/skills11k1 repo~1.1kAutomated safety check: PassApache-2.06 days ago
30

Configure article metadata via MDX frontmatter. An agent skill from adithya-s-k/FineEnvs.

adithya-s-k/FineEnvs421—~1.1kAutomated safety check: PassApache-2.0yesterday
31

Create self-contained D3 HTML embed charts for the research article template.

adithya-s-k/FineEnvs421—~1.1kAutomated safety check: PassApache-2.0yesterday
32

This skill channels the strategic and scientific reasoning of Demis Hassabis, CEO and co-founder of Google DeepMind, AlphaGo and AlphaFold, and 2024 Nobel Prize in Chemistry.

K-Dense-AI/mimeo282—~2kAutomated safety check: PassMIT1 mo ago
33

Deploy the article to a Hugging Face Space. An agent skill from adithya-s-k/FineEnvs.

adithya-s-k/FineEnvs421—~864Automated safety check: PassApache-2.0yesterday
34

Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge.

agentscope-ai/OpenJudge8671 repo~1.6kAutomated safety check: PassApache-2.026 days ago
35

Build, run, debug, and operate applications on AWS Lambda MicroVMs — Firecracker-isolated, snapshot-resumable serverless compute environments that run inside a container with up to 8-hour lifetimes.

awslabs/agent-plugins9121 repo~4.1kAutomated safety check: PassApache-2.02 days ago
36

Guide for using SLIME (LLM post-training framework for RL Scaling).

yzlnew/infra-skills149—~3.2kAutomated safety check: PassNo licence3 mo ago
37

Version-aware guidance for PufferLib reinforcement-learning environments, vectorization, policies, PuffeRL training, evaluation, and safe checkpoint review.

K-Dense-AI/scientific-agent-skills48k1 repo~3.9kAutomated safety check: NotesMIT2 days ago
38

Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint…

K-Dense-AI/scientific-agent-skills48k1 repo~3.7kAutomated safety check: NotesMIT2 days ago
39

Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning).

waybarrios/opencode-power-pack5333 repos~2.1kAutomated safety check: PassApache-2.0yesterday
40

Best practices for reinforcement learning policy optimization.

aiming-lab/AutoResearchClaw15k—~329Automated safety check: PassMIT1 mo ago
41

Build, test, and debug Hermes Agent RL environments for Atropos training.

Tommy-yw/RunbookHermes546—~3.3kAutomated safety check: PassMIT4 mo ago
42

AI-powered LinkedIn content suite: generate posts, carousels, newsletters, and 30-day calendars with niche-specific SEO rules and a reinforcement-learning personal memory system.

sickn33/agentic-awesome-skills47k1 repo~4.7kAutomated safety check: PassMITyesterday
43

Applies the reasoning style of Ilya Sutskever (deep learning pioneer, co-founder of OpenAI and Safe Superintelligence Inc.) to problems involving AI architecture, scaling laws, alignment, and…

K-Dense-AI/mimeo282—~1.7kAutomated safety check: PassMIT1 mo ago
44

Create and train AI learning plugins with AgentDB's 9 reinforcement learning algorithms.

aiskillstore/marketplace4307 repos~3kAutomated safety check: PassNo licencetoday
45

Reinforcement learning for trade execution optimization including order splitting, adaptive timing, and impact minimization

agiprolabs/claude-trading-skills410—~2kAutomated safety check: PassMIT1 mo ago
46

Knowledge base from "Spinning Up in Deep RL" by Joshua Achiam (OpenAI, MIT-licensed).

alirezarezvani/claude-skills28k—~2.8kAutomated safety check: PassMIT1 mo ago
47

Fine-tune LLMs using the Tinker API. An agent skill from sundial-org/skills.

sundial-org/skills152—~1.2kAutomated safety check: PassNo licence2 mo ago
48

Applies the reasoning, principles, and frameworks of Jürgen Schmidhuber (LSTM co-inventor and deep learning pioneer).

K-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT1 mo ago

Questions, answered from the data.

What is the best reinforcement learning skill?

Hugging Face LLM Trainer (official) from huggingface/skills ranks first of the 67 reinforcement learning skills listed here, with the highest score: its repository has 11k GitHub stars, 3 other GitHub owners carry a copy, its SKILL.md loads about 7.2k tokens and it passes the automated safety check with no findings. Next come Fix Art Issues and Train Rl.

Which reinforcement learning skills are official?

4 of the 67 reinforcement learning skills are official, published by the vendor's own GitHub organization: Hugging Face LLM Trainer, TRL Post-Training, AWS Lambda Microvms and Nemo Rl Brev Etiquette.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.