Search

Reinforcement learning

66 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

OpenPipe/ART11k—~840Automated safety check: NotesApache-2.0today
2

RL training reference for the ART framework. An agent skill from OpenPipe/ART.

OpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0today
3

Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

R6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.03 mo ago
4

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.

Orchestra-Research/AI-Research-SKILLs13k6 repos~2.9kAutomated safety check: PassMIT3 mo ago
5

A skill your agent uses when the user wants to optimize configurable system parameters against a measurable scalar objective, especially for model training, inference, quantitative strategies…

Optim-Agent/optim-agent801—~1.3kAutomated safety check: PassMIT1 mo ago
6

Builds a Verifiers (PrimeIntellect) variant of an RL environment.

adithya-s-k/FineEnvs4561 repo~2.3kAutomated safety check: PassApache-2.0yesterday
7

Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.

AI45Lab/SAfactory236—~1.8kAutomated safety check: PassNo licence15 days ago
8

Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

huggingface/skills11k1 repo~7.2kAutomated safety check: PassApache-2.0yesterday
9

Router for adding a diffusion or omni pipeline to verl-omni.

verl-project/verl-omni1.2k—~1kAutomated safety check: PassApache-2.0today
10

Guide for adding a new reward scorer to verl-omni and wiring it into a run.

verl-project/verl-omni1.2k—~648Automated safety check: PassApache-2.0today
11

Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

waybarrios/opencode-power-pack533—~3kAutomated safety check: PassApache-2.03 days ago
12

Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs.

adithya-s-k/FineEnvs456—~2.1kAutomated safety check: PassApache-2.0yesterday
13

Run, configure, retry, and validate AReno SFT, DPO, GSPO, GRPO, PPO, and agentic training.

inclusionAI/AReno323—~782Automated safety check: PassApache-2.0today
14

Builds an OpenEnv (Hugging Face) variant of an RL environment.

adithya-s-k/FineEnvs456—~2.4kAutomated safety check: PassApache-2.0yesterday
15

Walks through aligning language models with SimPO, a reference-free preference optimization method, using accelerate configs for Mistral 7B, Llama 3 8B and math-focused models.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.5kAutomated safety check: PassMIT3 mo ago
16

Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.8kAutomated safety check: PassMIT3 mo ago
17

Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package.

adithya-s-k/FineEnvs456—~2.3kAutomated safety check: NotesApache-2.0yesterday
18

Add or modify an AReno algorithm, trainer, loss, advantage calculation, role model, or algorithm-specific configuration.

inclusionAI/AReno323—~501Automated safety check: PassApache-2.0today
19

Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym.

adithya-s-k/FineEnvs456—~2.3kAutomated safety check: NotesApache-2.0yesterday
20

This skill should be used when working with reinforcement learning tasks including high-performance RL training, custom environment development, vectorized parallel simulation, multi-agent systems…

davila7/claude-code-templates32k8 repos~3.4kAutomated safety check: PassMITtoday
21

A skill your agent uses for reinforcement learning tasks including training RL agents (PPO, SAC, DQN, TD3, DDPG, A2C, etc.), creating custom Gym environments, implementing callbacks for monitoring…

davila7/claude-code-templates32k8 repos~2.4kAutomated safety check: PassMITtoday
22

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.2kAutomated safety check: PassMIT3 mo ago
23

High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.1kAutomated safety check: NotesMIT3 mo ago
24

Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.5kAutomated safety check: PassMIT3 mo ago
25

Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago
26

Explains how to train a model to be harmless with self-critique, revision and AI-generated preference feedback, with Hugging Face and TRL code for each stage.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2kAutomated safety check: PassMIT3 mo ago
27

Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths.

LegoX/Lego-RL111—~4.2kAutomated safety check: PassApache-2.0yesterday
28

Guides GRPO reinforcement-learning fine-tuning of language models with TRL, centered on designing reward functions for formats, verifiable tasks and reasoning.

Orchestra-Research/AI-Research-SKILLs13k3 repos~4.3kAutomated safety check: PassMIT3 mo ago
29

Configure article metadata via MDX frontmatter. An agent skill from adithya-s-k/FineEnvs.

adithya-s-k/FineEnvs456—~1.1kAutomated safety check: PassApache-2.0yesterday
30

Reference for post-training language models with TRL: which trainer and dataset format to use for SFT, DPO, GRPO, KTO and reward models, and how to add LoRA.

huggingface/skills11k1 repo~1.1kAutomated safety check: PassApache-2.0yesterday
31

Create self-contained D3 HTML embed charts for the research article template.

adithya-s-k/FineEnvs456—~1.1kAutomated safety check: PassApache-2.0yesterday
32

Deploy the article to a Hugging Face Space. An agent skill from adithya-s-k/FineEnvs.

adithya-s-k/FineEnvs456—~864Automated safety check: PassApache-2.0yesterday
33

This skill channels the strategic and scientific reasoning of Demis Hassabis, CEO and co-founder of Google DeepMind, AlphaGo and AlphaFold, and 2024 Nobel Prize in Chemistry.

K-Dense-AI/mimeo282—~2kAutomated safety check: PassMIT1 mo ago
34

Build, run, debug, and operate applications on AWS Lambda MicroVMs — Firecracker-isolated, snapshot-resumable serverless compute environments that run inside a container with up to 8-hour lifetimes.

awslabs/agent-plugins9151 repo~4.1kAutomated safety check: PassApache-2.0today
35

Guide for using SLIME (LLM post-training framework for RL Scaling).

yzlnew/infra-skills149—~3.2kAutomated safety check: PassNo licence3 mo ago
36

Version-aware guidance for PufferLib reinforcement-learning environments, vectorization, policies, PuffeRL training, evaluation, and safe checkpoint review.

K-Dense-AI/scientific-agent-skills48k1 repo~3.9kAutomated safety check: NotesMIT4 days ago
37

Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint…

K-Dense-AI/scientific-agent-skills48k1 repo~3.7kAutomated safety check: NotesMIT4 days ago
38

Best practices for reinforcement learning policy optimization.

aiming-lab/AutoResearchClaw15k—~329Automated safety check: PassMIT1 mo ago
39

Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning).

waybarrios/opencode-power-pack5332 repos~2.1kAutomated safety check: PassApache-2.03 days ago
40

Build, test, and debug Hermes Agent RL environments for Atropos training.

Tommy-yw/RunbookHermes546—~3.3kAutomated safety check: PassMIT4 mo ago
41

AI-powered LinkedIn content suite: generate posts, carousels, newsletters, and 30-day calendars with niche-specific SEO rules and a reinforcement-learning personal memory system.

sickn33/agentic-awesome-skills47k1 repo~4.7kAutomated safety check: PassMITtoday
42

Applies the reasoning style of Ilya Sutskever (deep learning pioneer, co-founder of OpenAI and Safe Superintelligence Inc.) to problems involving AI architecture, scaling laws, alignment, and…

K-Dense-AI/mimeo282—~1.7kAutomated safety check: PassMIT1 mo ago
43

Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge.

agentscope-ai/OpenJudge870—~1.6kAutomated safety check: PassApache-2.028 days ago
44

Create and train AI learning plugins with AgentDB's 9 reinforcement learning algorithms.

aiskillstore/marketplace4306 repos~3kAutomated safety check: PassNo licencetoday
45

Reinforcement learning for trade execution optimization including order splitting, adaptive timing, and impact minimization

agiprolabs/claude-trading-skills410—~2kAutomated safety check: PassMIT1 mo ago
46

Knowledge base from "Spinning Up in Deep RL" by Joshua Achiam (OpenAI, MIT-licensed).

alirezarezvani/claude-skills28k—~2.8kAutomated safety check: PassMIT1 mo ago
47

Fine-tune LLMs using the Tinker API. An agent skill from sundial-org/skills.

sundial-org/skills152—~1.2kAutomated safety check: PassNo licence2 mo ago
48

Applies the reasoning, principles, and frameworks of Jürgen Schmidhuber (LSTM co-inventor and deep learning pioneer).

K-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT1 mo ago