Search
Qwen · Reinforcement learning
4 skills found.
Category:
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART. | OpenPipe/ | 11k | — | ~840 | Automated safety check: Notes | Apache-2.0 | today |
| 2 | 2.Train Rl RL training reference for the ART framework. An agent skill from OpenPipe/ART. | OpenPipe/ | 11k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models. | Orchestra-Research/ | 13k | 4 repos | ~2.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. | Orchestra-Research/ | 13k | 2 repos | ~2.2k | Automated safety check: Pass | MIT | 3 mo ago |