Search
Docker · Reinforcement learning
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Builds a Verifiers (PrimeIntellect) variant of an RL environment. | adithya-s-k/ | 461 | 1 repo | ~2.3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 2 | Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training. | AI45Lab/ | 236 | — | ~1.8k | Automated safety check: Pass | No licence | 17 days ago |
| 3 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 461 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 4 | Builds an OpenEnv (Hugging Face) variant of an RL environment. | adithya-s-k/ | 461 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 5 | Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models. | Orchestra-Research/ | 13k | 4 repos | ~2.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package. | adithya-s-k/ | 461 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 7 | Deploy the article to a Hugging Face Space. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 461 | — | ~864 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 8 | Generates an executable, end-to-end VERL reinforcement learning quickstart runbook for Ascend/NPU (docker image, dataset preprocessing, model setup, mainppo training, and examples/run.sh flow). | ascend-ai-coding/ | 174 | — | ~592 | Automated safety check: Pass | No licence | yesterday |