Search

AI & LLM Engineering · pandas · For researchers

1 skill found.

Skills

Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago