Agent skill

slime RL Post-Training

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

MITAuto-check passedAI & LLM Engineering

Install slime RL Post-Training

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill slime-rl-training -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs slime-rl-training --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/06-post-training/slime .claude/skills/slime-rl-training && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
slime-rl-training
GitHub stars
13k
Used in
4 other repos
Token cost
~2.8k tokens
SKILL.md length
462 words
Files
3 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

  • Works in 6 steps: Prepare Data → Configure Model → Launch Training → …
  • Running GRPO training on GLM, Qwen3, DeepSeek V3 or Llama 3 models
  • SKILL.md covers When to Use slime, Key Features, Architecture Overview and Installation, plus 7 more sections
  • Calls python, pip and docker; reaches github.com

What it does

slime is a post-training framework from Tsinghua's THUDM team that joins Megatron-LM for training with SGLang for fast rollout generation, and a data buffer handles prompt management and sample storage. The skill gives an architecture overview, Docker and from-source installation, and a quick start for GRPO training that sources a model script such as qwen3-4B.sh. Supported parallelism includes tensor, pipeline, data and sequence modes.

Its first workflow is standard GRPO training for reasoning models: confirm the prerequisites (a Docker environment or Megatron-LM plus SGLang, a Hugging Face or Megatron checkpoint, and JSONL data), prepare prompt and label records in plain or chat format, then pick a pre-configured script from scripts/models. It points to miles, verl and torchforge as alternatives, and references cover the API and troubleshooting. The excerpt is cut off before the later workflow steps.

When your agent uses it

  • Running GRPO training on GLM, Qwen3, DeepSeek V3 or Llama 3 models
  • Combining Megatron-LM training with SGLang for rollout generation
  • Building a custom data generation workflow with a flexible data buffer
  • Setting up a Docker-based reinforcement-learning post-training run

Example prompts

  • “Set up slime in Docker and run the GRPO quick start with the Qwen3 4B model script.”
  • “Convert our math question and answer pairs into the JSONL format slime expects.”
  • “Explain how slime splits work between Megatron-LM training and SGLang rollouts.”
  • “Should we use slime, verl or torchforge for RL on a Llama 3 model?”

Requirements

  • Docker with GPU access, or Megatron-LM and SGLang installed
  • A Hugging Face or Megatron model checkpoint
  • Training data in JSONL format

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Prepare Data
  2. Configure Model
  3. Launch Training
  4. Monitor Training
  5. Define Custom Generate Function
  6. Launch with Custom Function

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip
    • docker
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • thudm.github.io
    • lmsys.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

slime RL Post-Training loads about 2.8k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 60 tokens; SKILL.md has 462 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 462 words, ~2,768 tokens.

Download SKILL.mdSave it as .claude/skills/slime-rl-training/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
slime-rl-training
description
Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Reinforcement Learning, Megatron-LM, SGLang, GRPO, Post-Training, GLM
dependencies
sglang-router>=0.2.3, ray, torch>=2.0.0, transformers>=4.40.0

slime: LLM Post-Training Framework for RL Scaling

slime is an LLM post-training framework from Tsinghua's THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation.

When to Use slime

Choose slime when you need:

  • Megatron-LM native training with SGLang inference
  • Custom data generation workflows with flexible data buffers
  • Training GLM, Qwen3, DeepSeek V3, or Llama 3 models
  • Research-grade framework with production backing (Z.ai)

Consider alternatives when:

  • You need enterprise-grade stability features → use miles
  • You want flexible backend swapping → use verl
  • You need PyTorch-native abstractions → use torchforge

Key Features

  • Training: Megatron-LM with full parallelism support (TP, PP, DP, SP)
  • Rollout: SGLang-based high-throughput generation with router
  • Data Buffer: Flexible prompt management and sample storage
  • Models: GLM-4.x, Qwen3, DeepSeek V3/R1, Llama 3

Architecture Overview

┌─────────────────────────────────────────────────────────┐
│                    Data Buffer                          │
│ - Prompt initialization and management                  │
│ - Custom data generation and filtering                  │
│ - Rollout sample storage                                │
└─────────────┬───────────────────────────┬───────────────┘
              │                           │
┌─────────────▼───────────┐ ┌─────────────▼───────────────┐
│ Training (Megatron-LM)  │ │ Rollout (SGLang + Router)   │
│ - Actor model training  │ │ - Response generation       │
│ - Critic (optional)     │ │ - Reward/verifier output    │
│ - Weight sync to rollout│ │ - Multi-turn support        │
└─────────────────────────┘ └─────────────────────────────┘

Installation

bash
# Recommended: Docker
docker pull slimerl/slime:latest
docker run --rm --gpus all --ipc=host --shm-size=16g \
  -it slimerl/slime:latest /bin/bash

# Inside container
cd /root/slime && pip install -e . --no-deps
From Source
bash
git clone https://github.com/THUDM/slime.git
cd slime
pip install -r requirements.txt
pip install -e .

Quick Start: GRPO Training

bash
# Source model configuration
source scripts/models/qwen3-4B.sh

# Launch training
python train.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 4 \
    --rollout-num-gpus 4 \
    --advantage-estimator grpo \
    --use-kl-loss --kl-loss-coef 0.001 \
    --rollout-batch-size 32 \
    --n-samples-per-prompt 8 \
    --global-batch-size 256 \
    --num-rollout 3000 \
    --prompt-data /path/to/data.jsonl \
    ${MODEL_ARGS[@]} ${CKPT_ARGS[@]}

Workflow 1: Standard GRPO Training

Use this workflow for training reasoning models with group-relative advantages.

Prerequisites Checklist
  • Docker environment or Megatron-LM + SGLang installed
  • Model checkpoint (HuggingFace or Megatron format)
  • Training data in JSONL format
Step 1: Prepare Data
python
# data.jsonl format
{"prompt": "What is 2 + 2?", "label": "4"}
{"prompt": "Solve: 3x = 12", "label": "x = 4"}

Or with chat format:

python
{
    "prompt": [
        {"role": "system", "content": "You are a math tutor."},
        {"role": "user", "content": "What is 15 + 27?"}
    ],
    "label": "42"
}
Step 2: Configure Model

Choose a pre-configured model script:

bash
# List available models
ls scripts/models/
# glm4-9B.sh, qwen3-4B.sh, qwen3-30B-A3B.sh, deepseek-v3.sh, llama3-8B.sh, ...

# Source your model
source scripts/models/qwen3-4B.sh
Step 3: Launch Training
bash
python train.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 8 \
    --rollout-num-gpus 8 \
    --advantage-estimator grpo \
    --use-kl-loss \
    --kl-loss-coef 0.001 \
    --prompt-data /path/to/train.jsonl \
    --input-key prompt \
    --label-key label \
    --apply-chat-template \
    --rollout-batch-size 32 \
    --n-samples-per-prompt 8 \
    --global-batch-size 256 \
    --num-rollout 3000 \
    --save-interval 100 \
    --eval-interval 50 \
    ${MODEL_ARGS[@]}
Step 4: Monitor Training
  • Check TensorBoard: tensorboard --logdir outputs/
  • Verify reward curves are increasing
  • Monitor GPU utilization across nodes

Workflow 2: Asynchronous Training

Use async mode for higher throughput by overlapping rollout and training.

When to Use Async
  • Large models with long generation times
  • High GPU idle time in synchronous mode
  • Sufficient memory for buffering
Launch Async Training
bash
python train_async.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 8 \
    --rollout-num-gpus 8 \
    --advantage-estimator grpo \
    --async-buffer-size 4 \
    --prompt-data /path/to/train.jsonl \
    ${MODEL_ARGS[@]}
Async-Specific Parameters
bash
--async-buffer-size 4        # Number of rollouts to buffer
--update-weights-interval 2  # Sync weights every N rollouts

Workflow 3: Multi-Turn Agentic Training

Use this workflow for training agents with tool use or multi-step reasoning.

Prerequisites
  • Custom generate function for multi-turn logic
  • Tool/environment interface
Show full SKILL.md (185 more words)Show less
Step 1: Define Custom Generate Function
python
# custom_generate.py
async def custom_generate(args, samples, evaluation=False):
    """Multi-turn generation with tool calling."""
    for sample in samples:
        conversation = sample.prompt

        for turn in range(args.max_turns):
            # Generate response
            response = await generate_single(conversation)

            # Check for tool call
            tool_call = extract_tool_call(response)
            if tool_call:
                tool_result = execute_tool(tool_call)
                conversation.append({"role": "assistant", "content": response})
                conversation.append({"role": "tool", "content": tool_result})
            else:
                break

        sample.response = response
        sample.reward = compute_reward(sample)

    return samples
Step 2: Launch with Custom Function
bash
python train.py \
    --custom-generate-function-path custom_generate.py \
    --max-turns 5 \
    --prompt-data /path/to/agent_data.jsonl \
    ${MODEL_ARGS[@]}

See examples/search-r1/ for a complete multi-turn search example.


Configuration Reference

Three Argument Categories

slime uses three types of arguments:

1. Megatron Arguments (passed directly):

bash
--tensor-model-parallel-size 2
--pipeline-model-parallel-size 1
--num-layers 32
--hidden-size 4096

2. SGLang Arguments (prefixed with --sglang-):

bash
--sglang-mem-fraction-static 0.8
--sglang-context-length 8192
--sglang-log-level INFO

3. slime Arguments:

bash
# Resource allocation
--actor-num-nodes 1
--actor-num-gpus-per-node 8
--rollout-num-gpus 8
--colocate  # Share GPUs between training/inference

# Data
--prompt-data /path/to/data.jsonl
--input-key prompt
--label-key label

# Training loop
--num-rollout 3000
--rollout-batch-size 32
--n-samples-per-prompt 8
--global-batch-size 256

# Algorithm
--advantage-estimator grpo  # or: gspo, ppo, reinforce_plus_plus
--use-kl-loss
--kl-loss-coef 0.001
Key Constraints
rollout_batch_size × n_samples_per_prompt = global_batch_size × num_steps_per_rollout

Example: 32 × 8 = 256 × 1


Data Buffer System

slime's data buffer enables flexible data management:

Basic Data Source
python
class RolloutDataSource:
    def get_samples(self, num_samples):
        """Fetch prompts from dataset."""
        return self.dataset.sample(num_samples)

    def add_samples(self, samples):
        """Called after generation (no-op by default)."""
        pass
Buffered Data Source (Off-Policy)
python
class RolloutDataSourceWithBuffer(RolloutDataSource):
    def __init__(self):
        self.buffer = []

    def add_samples(self, samples):
        """Store generated samples for reuse."""
        self.buffer.extend(samples)

    def buffer_filter(self, args, buffer, num_samples):
        """Custom selection logic (prioritized, stratified, etc.)."""
        return select_best(buffer, num_samples)

Common Issues and Solutions

Issue: SGLang Engine Crash

Symptoms: Inference engine dies mid-training

Solutions:

bash
# Enable fault tolerance
--use-fault-tolerance

# Increase memory allocation
--sglang-mem-fraction-static 0.85

# Reduce batch size
--rollout-batch-size 16
Issue: Weight Sync Timeout

Symptoms: Training hangs after rollout

Solutions:

bash
# Increase sync interval
--update-weights-interval 5

# Use colocated mode (no network transfer)
--colocate
Issue: OOM During Training

Symptoms: CUDA OOM in backward pass

Solutions:

bash
# Enable gradient checkpointing
--recompute-activations

# Reduce micro-batch size
--micro-batch-size 1

# Enable sequence parallelism
--sequence-parallel
Issue: Slow Data Loading

Symptoms: GPU idle during data fetch

Solutions:

bash
# Increase data workers
--num-data-workers 4

# Use streaming dataset
--streaming-data

Supported Models

Model FamilyConfigurations
GLMGLM-4.5, GLM-4.6, GLM-4.7, GLM-Z1-9B
QwenQwen3 (4B, 8B, 30B-A3B), Qwen3-MoE, Qwen2.5
DeepSeekV3, V3.1, R1
LlamaLlama 3 (8B, 70B)
OthersKimi K2, Moonlight-16B

Each model has pre-configured scripts in scripts/models/.


Advanced Topics

Co-location Mode

Share GPUs between training and inference to reduce memory:

bash
python train.py \
    --colocate \
    --actor-num-gpus-per-node 8 \
    --sglang-mem-fraction-static 0.4 \
    ${MODEL_ARGS[@]}
Custom Reward Model
python
# custom_rm.py
class CustomRewardModel:
    def __init__(self, model_path):
        self.model = load_model(model_path)

    def compute_reward(self, prompts, responses):
        inputs = self.tokenize(prompts, responses)
        scores = self.model(inputs)
        return scores.tolist()
bash
--custom-rm-path custom_rm.py
Evaluation Multi-Task
bash
--eval-prompt-data aime /path/to/aime.jsonl \
--eval-prompt-data gsm8k /path/to/gsm8k.jsonl \
--n-samples-per-eval-prompt 16

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in 06-post-training/slime of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/api-reference.md
  • references/troubleshooting.md

Open the folder on GitHubat commit 773a529

Used in 4 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

slime RL Post-Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

slime RL Post-Training compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
slime RL Post-Training this skillOrchestra-Research/AI-Research-SKILLs13k4 repos~2.8kAutomated safety check: PassMIT
Model Architecture Diagram FinderBBuf/AI-Infra-Auto-Driven-SKILLS938—~1.2kAutomated safety check: PassNone
Weave Router Local Testingweave-os/router5.6k—~3.1kAutomated safety check: NotesApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
Safactory WorkflowsAI45Lab/SAfactory236—~1.8kAutomated safety check: PassNone
Fix Art IssuesOpenPipe/ART11k—~840Automated safety check: NotesApache-2.0

Similar skills

  • Model Architecture Diagram Finder

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.

    938 GitHub stars~1.2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Stands up the Weave model router in Docker Compose and drives it with claude -p against a real or mocked upstream to reproduce and verify routing and streaming behavior.

    5.6k GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Safactory Workflows

    AI45Lab/SAfactory

    Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.

    236 GitHub stars~1.8k tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • Fix Art Issues

    OpenPipe/ART

    Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

    11k GitHub stars~840 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    408 GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about slime RL Post-Training

What does slime RL Post-Training do?

Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models. slime is a post-training framework from Tsinghua's THUDM team that joins Megatron-LM for training with SGLang for fast rollout generation, and a data buffer handles prompt management and sample storage.sh.

When should I use slime RL Post-Training?

slime RL Post-Training fits situations like: running GRPO training on GLM, Qwen3, DeepSeek V3 or Llama 3 models; combining Megatron-LM training with SGLang for rollout generation; building a custom data generation workflow with a flexible data buffer; setting up a Docker-based reinforcement-learning post-training run.

How do I install slime RL Post-Training in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill slime-rl-training -a claude-code`. Or copy the skill folder (06-post-training/slime in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/slime-rl-training in your project. Claude Code loads it when a task matches its description.

How do I install slime RL Post-Training in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill slime-rl-training -a codex`. Or copy the skill folder (06-post-training/slime in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/slime-rl-training in your project. Codex loads it when a task matches its description.

Can I use slime RL Post-Training in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill slime-rl-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/slime-rl-training, .gemini/skills/slime-rl-training, .github/skills/slime-rl-training and .opencode/skills/slime-rl-training in your project.

What does slime RL Post-Training need to run?

Going by SKILL.md and its folder, slime RL Post-Training needs the command-line tools its instructions call (python, pip, docker and git). Our summary lists: Docker with GPU access, or Megatron-LM and SGLang installed; A Hugging Face or Megatron model checkpoint; Training data in JSONL format.

Does slime RL Post-Training access the network?

SKILL.md names 3 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: thudm.github.io and lmsys.org. This is read from the text; nothing was executed.

Is slime RL Post-Training safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does slime RL Post-Training use?

slime RL Post-Training is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does slime RL Post-Training use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.8k tokens, read only when the agent opens those files.

What are the alternatives to slime RL Post-Training?

Skills that share tags, products or a category with slime RL Post-Training: Model Architecture Diagram Finder (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Weave Router Local Testing (weave-os/router, 5.6k stars), Train Rl (OpenPipe/ART, 11k stars) and Safactory Workflows (AI45Lab/SAfactory, 236 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains slime RL Post-Training?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.