Search
Python · Reinforcement learning
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Builds a Verifiers (PrimeIntellect) variant of an RL environment. | adithya-s-k/ | 456 | 1 repo | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 2 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 456 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | Walks through aligning language models with SimPO, a reference-free preference optimization method, using accelerate configs for Mistral 7B, Llama 3 8B and math-focused models. | Orchestra-Research/ | 13k | 4 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package. | adithya-s-k/ | 456 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 5 | Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs. | Orchestra-Research/ | 13k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA. | huggingface/ | 11k | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | Version-aware guidance for PufferLib reinforcement-learning environments, vectorization, policies, PuffeRL training, evaluation, and safe checkpoint review. | K-Dense-AI/ | 48k | 1 repo | ~3.9k | Automated safety check: Notes | MIT | 4 days ago |