Search

Python · Reinforcement learning

8 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Builds a Verifiers (PrimeIntellect) variant of an RL environment.

adithya-s-k/FineEnvs4561 repo~2.3kAutomated safety check: PassApache-2.0yesterday
2

Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs.

adithya-s-k/FineEnvs456—~2.1kAutomated safety check: PassApache-2.0yesterday
3

Walks through aligning language models with SimPO, a reference-free preference optimization method, using accelerate configs for Mistral 7B, Llama 3 8B and math-focused models.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.5kAutomated safety check: PassMIT3 mo ago
4

Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package.

adithya-s-k/FineEnvs456—~2.3kAutomated safety check: NotesApache-2.0yesterday
5

Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.5kAutomated safety check: PassMIT3 mo ago
6

Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago
7

Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA.

huggingface/skills11k1 repo~1.1kAutomated safety check: PassApache-2.0yesterday
8

Version-aware guidance for PufferLib reinforcement-learning environments, vectorization, policies, PuffeRL training, evaluation, and safe checkpoint review.

K-Dense-AI/scientific-agent-skills48k1 repo~3.9kAutomated safety check: NotesMIT4 days ago