Search

AI & LLM Engineering · DeepSeek · For researchers

2 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Walks through aligning language models with SimPO, a reference-free preference optimization method, using accelerate configs for Mistral 7B, Llama 3 8B and math-focused models.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.5kAutomated safety check: PassMIT3 mo ago
2

Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.8kAutomated safety check: PassMIT3 mo ago