Topic · AI & LLM Engineering
Best reinforcement learning skills for Claude Code, Codex and other agents.
- skills
- 67
- official
- 4
Reinforcement learning skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF. | huggingface/ | 11k | 3 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 2 | Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART. | OpenPipe/ | 11k | — | ~840 | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 3 | 3.Train Rl RL training reference for the ART framework. An agent skill from OpenPipe/ART. | OpenPipe/ | 11k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 4 | Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO. | R6410418/ | 1.7k | — | ~830 | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 5 | Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. | Orchestra-Research/ | 13k | 7 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | A skill your agent uses when the user wants to optimize configurable system parameters against a measurable scalar objective, especially for model training, inference, quantitative strategies… | Optim-Agent/ | 800 | — | ~1.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 7 | Builds a Verifiers (PrimeIntellect) variant of an RL environment. | adithya-s-k/ | 421 | 1 repo | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training. | AI45Lab/ | 236 | — | ~1.8k | Automated safety check: Pass | No licence | 13 days ago |
| 9 | Router for adding a diffusion or omni pipeline to verl-omni. | verl-project/ | 1.2k | — | ~1k | Automated safety check: Pass | Apache-2.0 | today |
| 10 | Guide for adding a new reward scorer to verl-omni and wiring it into a run. | verl-project/ | 1.2k | — | ~648 | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion. | waybarrios/ | 533 | — | ~3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 12 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 421 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 13 | Run, configure, retry, and validate AReno SFT, DPO, GSPO, GRPO, PPO, and agentic training. | inclusionAI/ | 323 | — | ~782 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 14 | Walks through aligning language models with SimPO, a reference-free preference optimization method, using accelerate configs for Mistral 7B, Llama 3 8B and math-focused models. | Orchestra-Research/ | 13k | 5 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 15 | Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models. | Orchestra-Research/ | 13k | 5 repos | ~2.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 16 | Builds an OpenEnv (Hugging Face) variant of an RL environment. | adithya-s-k/ | 421 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 17 | Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package. | adithya-s-k/ | 421 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 18 | Add or modify an AReno algorithm, trainer, loss, advantage calculation, role model, or algorithm-specific configuration. | inclusionAI/ | 323 | — | ~501 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 19 | Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. | Orchestra-Research/ | 13k | 3 repos | ~2.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 20 | High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 3 repos | ~2.1k | Automated safety check: Notes | MIT | 3 mo ago |
| 21 | Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs. | Orchestra-Research/ | 13k | 3 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 22 | Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends. | Orchestra-Research/ | 13k | 3 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 23 | 23.Pufferlib This skill should be used when working with reinforcement learning tasks including high-performance RL training, custom environment development, vectorized parallel simulation, multi-agent systems… | davila7/ | 32k | 9 repos | ~3.4k | Automated safety check: Pass | MIT | today |
| 24 | A skill your agent uses for reinforcement learning tasks including training RL agents (PPO, SAC, DQN, TD3, DDPG, A2C, etc.), creating custom Gym environments, implementing callbacks for monitoring… | davila7/ | 32k | 9 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 25 | Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym. | adithya-s-k/ | 421 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 26 | Explains how to train a model to be harmless with self-critique, revision and AI-generated preference feedback, with Hugging Face and TRL code for each stage. | Orchestra-Research/ | 13k | 3 repos | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 27 | Guides GRPO reinforcement-learning fine-tuning of language models with TRL, centered on designing reward functions for formats, verifiable tasks and reasoning. | Orchestra-Research/ | 13k | 4 repos | ~4.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 28 | 28.Dashboard Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths. | LegoX/ | 108 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 29 | Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA. | huggingface/ | 11k | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 30 | Configure article metadata via MDX frontmatter. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 421 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 31 | Create self-contained D3 HTML embed charts for the research article template. | adithya-s-k/ | 421 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 32 | This skill channels the strategic and scientific reasoning of Demis Hassabis, CEO and co-founder of Google DeepMind, AlphaGo and AlphaFold, and 2024 Nobel Prize in Chemistry. | K-Dense-AI/ | 282 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 33 | Deploy the article to a Hugging Face Space. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 421 | — | ~864 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 34 | 34.02 Rl Reward Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge. | agentscope-ai/ | 867 | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | 26 days ago |
| 35 | Build, run, debug, and operate applications on AWS Lambda MicroVMs — Firecracker-isolated, snapshot-resumable serverless compute environments that run inside a container with up to 8-hour lifetimes. | awslabs/ | 912 | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 36 | 36.Slime User Guide for using SLIME (LLM post-training framework for RL Scaling). | yzlnew/ | 149 | — | ~3.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 37 | 37.Pufferlib Version-aware guidance for PufferLib reinforcement-learning environments, vectorization, policies, PuffeRL training, evaluation, and safe checkpoint review. | K-Dense-AI/ | 48k | 1 repo | ~3.9k | Automated safety check: Notes | MIT | 2 days ago |
| 38 | Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint… | K-Dense-AI/ | 48k | 1 repo | ~3.7k | Automated safety check: Notes | MIT | 2 days ago |
| 39 | 39.Trl Training Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning). | waybarrios/ | 533 | 3 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 40 | Best practices for reinforcement learning policy optimization. | aiming-lab/ | 15k | — | ~329 | Automated safety check: Pass | MIT | 1 mo ago |
| 41 | Build, test, and debug Hermes Agent RL environments for Atropos training. | Tommy-yw/ | 546 | — | ~3.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 42 | AI-powered LinkedIn content suite: generate posts, carousels, newsletters, and 30-day calendars with niche-specific SEO rules and a reinforcement-learning personal memory system. | sickn33/ | 47k | 1 repo | ~4.7k | Automated safety check: Pass | MIT | yesterday |
| 43 | Applies the reasoning style of Ilya Sutskever (deep learning pioneer, co-founder of OpenAI and Safe Superintelligence Inc.) to problems involving AI architecture, scaling laws, alignment, and… | K-Dense-AI/ | 282 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 44 | Create and train AI learning plugins with AgentDB's 9 reinforcement learning algorithms. | aiskillstore/ | 430 | 7 repos | ~3k | Automated safety check: Pass | No licence | today |
| 45 | 45.Rl Execution Reinforcement learning for trade execution optimization including order splitting, adaptive timing, and impact minimization | agiprolabs/ | 410 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 46 | Knowledge base from "Spinning Up in Deep RL" by Joshua Achiam (OpenAI, MIT-licensed). | alirezarezvani/ | 28k | — | ~2.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 47 | 47.Tinker Fine-tune LLMs using the Tinker API. An agent skill from sundial-org/skills. | sundial-org/ | 152 | — | ~1.2k | Automated safety check: Pass | No licence | 2 mo ago |
| 48 | Applies the reasoning, principles, and frameworks of Jürgen Schmidhuber (LSTM co-inventor and deep learning pioneer). | K-Dense-AI/ | 282 | — | ~1.9k | Automated safety check: Pass | MIT | 1 mo ago |
Questions, answered from the data.
What is the best reinforcement learning skill?
Hugging Face LLM Trainer (official) from huggingface/skills ranks first of the 67 reinforcement learning skills listed here, with the highest score: its repository has 11k GitHub stars, 3 other GitHub owners carry a copy, its SKILL.md loads about 7.2k tokens and it passes the automated safety check with no findings. Next come Fix Art Issues and Train Rl.
Which reinforcement learning skills are official?
4 of the 67 reinforcement learning skills are official, published by the vendor's own GitHub organization: Hugging Face LLM Trainer, TRL Post-Training, AWS Lambda Microvms and Nemo Rl Brev Etiquette.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- AI interpretability23