Search
Reinforcement learning
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART. | OpenPipe/ | 11k | — | ~840 | Automated safety check: Notes | Apache-2.0 | today |
| 2 | 2.Train Rl RL training reference for the ART framework. An agent skill from OpenPipe/ART. | OpenPipe/ | 11k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO. | R6410418/ | 1.7k | — | ~830 | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 4 | Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. | Orchestra-Research/ | 13k | 6 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 5 | A skill your agent uses when the user wants to optimize configurable system parameters against a measurable scalar objective, especially for model training, inference, quantitative strategies… | Optim-Agent/ | 801 | — | ~1.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 6 | Builds a Verifiers (PrimeIntellect) variant of an RL environment. | adithya-s-k/ | 456 | 1 repo | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 7 | Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training. | AI45Lab/ | 236 | — | ~1.8k | Automated safety check: Pass | No licence | 15 days ago |
| 8 | Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF. | huggingface/ | 11k | 1 repo | ~7.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | Router for adding a diffusion or omni pipeline to verl-omni. | verl-project/ | 1.2k | — | ~1k | Automated safety check: Pass | Apache-2.0 | today |
| 10 | Guide for adding a new reward scorer to verl-omni and wiring it into a run. | verl-project/ | 1.2k | — | ~648 | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion. | waybarrios/ | 533 | — | ~3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 12 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 456 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 13 | Run, configure, retry, and validate AReno SFT, DPO, GSPO, GRPO, PPO, and agentic training. | inclusionAI/ | 323 | — | ~782 | Automated safety check: Pass | Apache-2.0 | today |
| 14 | Builds an OpenEnv (Hugging Face) variant of an RL environment. | adithya-s-k/ | 456 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 15 | Walks through aligning language models with SimPO, a reference-free preference optimization method, using accelerate configs for Mistral 7B, Llama 3 8B and math-focused models. | Orchestra-Research/ | 13k | 4 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 16 | Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models. | Orchestra-Research/ | 13k | 4 repos | ~2.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package. | adithya-s-k/ | 456 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 18 | Add or modify an AReno algorithm, trainer, loss, advantage calculation, role model, or algorithm-specific configuration. | inclusionAI/ | 323 | — | ~501 | Automated safety check: Pass | Apache-2.0 | today |
| 19 | Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym. | adithya-s-k/ | 456 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 20 | 20.Pufferlib This skill should be used when working with reinforcement learning tasks including high-performance RL training, custom environment development, vectorized parallel simulation, multi-agent systems… | davila7/ | 32k | 8 repos | ~3.4k | Automated safety check: Pass | MIT | today |
| 21 | A skill your agent uses for reinforcement learning tasks including training RL agents (PPO, SAC, DQN, TD3, DDPG, A2C, etc.), creating custom Gym environments, implementing callbacks for monitoring… | davila7/ | 32k | 8 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 22 | Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. | Orchestra-Research/ | 13k | 2 repos | ~2.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 23 | High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 2 repos | ~2.1k | Automated safety check: Notes | MIT | 3 mo ago |
| 24 | Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs. | Orchestra-Research/ | 13k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 25 | Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 26 | Explains how to train a model to be harmless with self-critique, revision and AI-generated preference feedback, with Hugging Face and TRL code for each stage. | Orchestra-Research/ | 13k | 2 repos | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 27 | 27.Dashboard Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths. | LegoX/ | 111 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 28 | Guides GRPO reinforcement-learning fine-tuning of language models with TRL, centered on designing reward functions for formats, verifiable tasks and reasoning. | Orchestra-Research/ | 13k | 3 repos | ~4.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 29 | Configure article metadata via MDX frontmatter. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 456 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 30 | Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA. | huggingface/ | 11k | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 31 | Create self-contained D3 HTML embed charts for the research article template. | adithya-s-k/ | 456 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 32 | Deploy the article to a Hugging Face Space. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 456 | — | ~864 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 33 | This skill channels the strategic and scientific reasoning of Demis Hassabis, CEO and co-founder of Google DeepMind, AlphaGo and AlphaFold, and 2024 Nobel Prize in Chemistry. | K-Dense-AI/ | 282 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 34 | Build, run, debug, and operate applications on AWS Lambda MicroVMs — Firecracker-isolated, snapshot-resumable serverless compute environments that run inside a container with up to 8-hour lifetimes. | awslabs/ | 915 | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | today |
| 35 | 35.Slime User Guide for using SLIME (LLM post-training framework for RL Scaling). | yzlnew/ | 149 | — | ~3.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 36 | 36.Pufferlib Version-aware guidance for PufferLib reinforcement-learning environments, vectorization, policies, PuffeRL training, evaluation, and safe checkpoint review. | K-Dense-AI/ | 48k | 1 repo | ~3.9k | Automated safety check: Notes | MIT | 4 days ago |
| 37 | Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint… | K-Dense-AI/ | 48k | 1 repo | ~3.7k | Automated safety check: Notes | MIT | 4 days ago |
| 38 | Best practices for reinforcement learning policy optimization. | aiming-lab/ | 15k | — | ~329 | Automated safety check: Pass | MIT | 1 mo ago |
| 39 | 39.Trl Training Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning). | waybarrios/ | 533 | 2 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 40 | Build, test, and debug Hermes Agent RL environments for Atropos training. | Tommy-yw/ | 546 | — | ~3.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 41 | AI-powered LinkedIn content suite: generate posts, carousels, newsletters, and 30-day calendars with niche-specific SEO rules and a reinforcement-learning personal memory system. | sickn33/ | 47k | 1 repo | ~4.7k | Automated safety check: Pass | MIT | today |
| 42 | Applies the reasoning style of Ilya Sutskever (deep learning pioneer, co-founder of OpenAI and Safe Superintelligence Inc.) to problems involving AI architecture, scaling laws, alignment, and… | K-Dense-AI/ | 282 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 43 | 43.02 Rl Reward Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge. | agentscope-ai/ | 870 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 28 days ago |
| 44 | Create and train AI learning plugins with AgentDB's 9 reinforcement learning algorithms. | aiskillstore/ | 430 | 6 repos | ~3k | Automated safety check: Pass | No licence | today |
| 45 | 45.Rl Execution Reinforcement learning for trade execution optimization including order splitting, adaptive timing, and impact minimization | agiprolabs/ | 410 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 46 | Knowledge base from "Spinning Up in Deep RL" by Joshua Achiam (OpenAI, MIT-licensed). | alirezarezvani/ | 28k | — | ~2.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 47 | 47.Tinker Fine-tune LLMs using the Tinker API. An agent skill from sundial-org/skills. | sundial-org/ | 152 | — | ~1.2k | Automated safety check: Pass | No licence | 2 mo ago |
| 48 | Applies the reasoning, principles, and frameworks of Jürgen Schmidhuber (LSTM co-inventor and deep learning pioneer). | K-Dense-AI/ | 282 | — | ~1.9k | Automated safety check: Pass | MIT | 1 mo ago |