Search
Reinforcement learning
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Applies the reasoning of Pieter Abbeel, robotics and reinforcement learning expert, UC Berkeley professor, and co-founder of Covariant. | K-Dense-AI/ | 282 | — | ~1.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 50 | Reach for this skill whenever you are discussing reinforcement learning, agentic AI systems, AI alignment, continual learning, or the philosophical limits of large language models. | K-Dense-AI/ | 282 | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 51 | Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'. | K-Dense-AI/ | 282 | — | ~1.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 52 | 52.Grpo Reference for the GRPO (Group Relative Policy Optimization) algorithm. | benchflow-ai/ | 1.8k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 53 | Diagnostic guide for RL-based post-training of language models (GRPO, PPO, REINFORCE, DPO). | benchflow-ai/ | 1.8k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 54 | Build practice environments and evaluation gyms where agents can try, fail, and learn from feedback. | LearnPrompt/ | 110 | — | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 55 | Brev instance operating guidance for NeMo-RL agents working in /home/ubuntu/RL with limited workspace disk, a larger /ephemeral volume, and optional /home/ubuntu/RL/.env secrets. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 56 | Generates A/B test plans and optimization checklists for your App Store product page — icon, screenshots, and app previews. | gustavscirulis/ | 116 | 1 repo | ~1.9k | Automated safety check: Pass | Unknown | 5 mo ago |
| 57 | A skill your agent uses when implementing staking contracts, reward distribution systems, or yield farming. | ccashwell/ | 131 | — | ~1.7k | Automated safety check: Pass | MIT | 11 days ago |
| 58 | 58.Trl Reference for the TRL (Transformer Reinforcement Learning) library codebase. | benchflow-ai/ | 1.8k | — | ~989 | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 59 | Reinforcement learning fundamentals, algorithms, and research | wentorai/ | 298 | 1 repo | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 60 | Vectorized multi-agent reinforcement learning simulator. An agent skill from wentorai/research-plugins. | wentorai/ | 298 | 1 repo | ~960 | Automated safety check: Pass | MIT | 3 mo ago |
| 61 | Train and optimize AI agents using Microsoft's Agent Lightning framework with reinforcement learning. | coco-research/ | 513 | — | ~2k | Automated safety check: Pass | Unknown | yesterday |
| 62 | Your AI research and engineering brain trust. An agent skill from coco-research/coco. | coco-research/ | 513 | — | ~3.7k | Automated safety check: Pass | Unknown | yesterday |
| 63 | Build and review OpenRLHF supervised/preference training plans for SFT, reward models, DPO, IPO, and cDPO. | VectorSpaceLab/ | 331 | — | ~1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 64 | 64.Torchrl Use TorchRL for TensorDict-first reinforcement-learning environments, collectors, replay buffers, modules, objectives, LLM/RLHF/VLA workflows, services, rendering, and maintainer-safe repository… | VectorSpaceLab/ | 331 | — | ~1.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 65 | Generates an executable, end-to-end VERL reinforcement learning quickstart runbook for Ascend/NPU (docker image, dataset preprocessing, model setup, mainppo training, and examples/run.sh flow). | ascend-ai-coding/ | 174 | — | ~592 | Automated safety check: Pass | No licence | yesterday |
| 66 | A skill your agent uses when positioning an AAMAS submission against multiagent, game-theory, and reinforcement-learning literature spread across AAMAS, AAAI, IJCAI, NeurIPS, ICML, EC, and JAAMAS… | brycewang-stanford/ | 1.2k | — | ~970 | Automated safety check: Pass | MIT | 14 days ago |