RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent.

MITAuto-check passedDevOps & Cloud

Install Slime

skills CLI
$ npx skills add Luciole-Studio/Misaka-Agent --skill slime -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Luciole-Studio/Misaka-Agent slime --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/slime .claude/skills/slime && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
slime
GitHub stars
139
Used in
1 other repo
Token cost
~2.7k tokens
SKILL.md length
462 words
Files
3 (incl. references)
Skills in repo
76
Repo updated
First seen
Licence
MIT

At a glance

RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent.

  • Works in 6 steps: Prepare Data → Configure Model → Launch Training → …
  • Tasks that involve MLOps
  • SKILL.md covers When to Use slime, Key Features, Architecture Overview and Installation, plus 7 more sections
  • Calls python, pip and docker; reaches github.com

What it does

Slime is an agent skill from Luciole-Studio/Misaka-Agent. RL post-training for LLMs with Megatron and SGLang.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/api-reference.md` and `references/troubleshooting.md`).

It sits in DevOps & Cloud, covering MLOps. It works with NVIDIA AI Platform, SGLang and Zhipu GLM. The repository describes itself as: A multi-agent research system for the humanities and social sciences. The licence is MIT.

When your agent uses it

  • Tasks that involve MLOps

Example prompts

  • “/slime”

Requirements

  • Python 3
  • Docker

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Prepare Data
  2. Configure Model
  3. Launch Training
  4. Monitor Training
  5. Define Custom Generate Function
  6. Launch with Custom Function

What it can do on your machine

Read from SKILL.md and the folder at commit b94464a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip
    • docker
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • thudm.github.io
    • lmsys.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Slime loads about 2.7k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 14 tokens; SKILL.md has 462 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~14
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Luciole-Studio/Misaka-Agent at commit b94464a, republished under its MIT licence (© Luciole-Studio). 462 words, ~2,735 tokens.

Download SKILL.mdSave it as .claude/skills/slime/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
slime
description
RL post-training for LLMs with Megatron and SGLang.
version
1.0.0
author
Orchestra Research
license
MIT
dependencies
sglang-router>=0.2.3, ray, torch>=2.0.0, transformers>=4.40.0
platforms
linux, macos

slime: LLM Post-Training Framework for RL Scaling

slime is an LLM post-training framework from Tsinghua's THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation.

When to Use slime

Choose slime when you need:

  • Megatron-LM native training with SGLang inference
  • Custom data generation workflows with flexible data buffers
  • Training GLM, Qwen3, DeepSeek V3, or Llama 3 models
  • Research-grade framework with production backing (Z.ai)

Consider alternatives when:

  • You need enterprise-grade stability features → use miles
  • You want flexible backend swapping → use verl
  • You need PyTorch-native abstractions → use torchforge

Key Features

  • Training: Megatron-LM with full parallelism support (TP, PP, DP, SP)
  • Rollout: SGLang-based high-throughput generation with router
  • Data Buffer: Flexible prompt management and sample storage
  • Models: GLM-4.x, Qwen3, DeepSeek V3/R1, Llama 3

Architecture Overview

┌─────────────────────────────────────────────────────────┐
│                    Data Buffer                          │
│ - Prompt initialization and management                  │
│ - Custom data generation and filtering                  │
│ - Rollout sample storage                                │
└─────────────┬───────────────────────────┬───────────────┘
              │                           │
┌─────────────▼───────────┐ ┌─────────────▼───────────────┐
│ Training (Megatron-LM)  │ │ Rollout (SGLang + Router)   │
│ - Actor model training  │ │ - Response generation       │
│ - Critic (optional)     │ │ - Reward/verifier output    │
│ - Weight sync to rollout│ │ - Multi-turn support        │
└─────────────────────────┘ └─────────────────────────────┘

Installation

bash
# Recommended: Docker
docker pull slimerl/slime:latest
docker run --rm --gpus all --ipc=host --shm-size=16g \
  -it slimerl/slime:latest /bin/bash

# Inside container
cd /root/slime && pip install -e . --no-deps
From Source
bash
git clone https://github.com/THUDM/slime.git
cd slime
pip install -r requirements.txt
pip install -e .

Quick Start: GRPO Training

bash
# Source model configuration
source scripts/models/qwen3-4B.sh

# Launch training
python train.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 4 \
    --rollout-num-gpus 4 \
    --advantage-estimator grpo \
    --use-kl-loss --kl-loss-coef 0.001 \
    --rollout-batch-size 32 \
    --n-samples-per-prompt 8 \
    --global-batch-size 256 \
    --num-rollout 3000 \
    --prompt-data /path/to/data.jsonl \
    ${MODEL_ARGS[@]} ${CKPT_ARGS[@]}

Workflow 1: Standard GRPO Training

Use this workflow for training reasoning models with group-relative advantages.

Prerequisites Checklist
  • Docker environment or Megatron-LM + SGLang installed
  • Model checkpoint (HuggingFace or Megatron format)
  • Training data in JSONL format
Step 1: Prepare Data
python
# data.jsonl format
{"prompt": "What is 2 + 2?", "label": "4"}
{"prompt": "Solve: 3x = 12", "label": "x = 4"}

Or with chat format:

python
{
    "prompt": [
        {"role": "system", "content": "You are a math tutor."},
        {"role": "user", "content": "What is 15 + 27?"}
    ],
    "label": "42"
}
Step 2: Configure Model

Choose a pre-configured model script:

bash
# List available models
ls scripts/models/
# glm4-9B.sh, qwen3-4B.sh, qwen3-30B-A3B.sh, deepseek-v3.sh, llama3-8B.sh, ...

# Source your model
source scripts/models/qwen3-4B.sh
Step 3: Launch Training
bash
python train.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 8 \
    --rollout-num-gpus 8 \
    --advantage-estimator grpo \
    --use-kl-loss \
    --kl-loss-coef 0.001 \
    --prompt-data /path/to/train.jsonl \
    --input-key prompt \
    --label-key label \
    --apply-chat-template \
    --rollout-batch-size 32 \
    --n-samples-per-prompt 8 \
    --global-batch-size 256 \
    --num-rollout 3000 \
    --save-interval 100 \
    --eval-interval 50 \
    ${MODEL_ARGS[@]}
Step 4: Monitor Training
  • Check TensorBoard: tensorboard --logdir outputs/
  • Verify reward curves are increasing
  • Monitor GPU utilization across nodes

Workflow 2: Asynchronous Training

Use async mode for higher throughput by overlapping rollout and training.

When to Use Async
  • Large models with long generation times
  • High GPU idle time in synchronous mode
  • Sufficient memory for buffering
Launch Async Training
bash
python train_async.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 8 \
    --rollout-num-gpus 8 \
    --advantage-estimator grpo \
    --async-buffer-size 4 \
    --prompt-data /path/to/train.jsonl \
    ${MODEL_ARGS[@]}
Async-Specific Parameters
bash
--async-buffer-size 4        # Number of rollouts to buffer
--update-weights-interval 2  # Sync weights every N rollouts

Workflow 3: Multi-Turn Agentic Training

Use this workflow for training agents with tool use or multi-step reasoning.

Prerequisites
  • Custom generate function for multi-turn logic
  • Tool/environment interface
Show full SKILL.md (185 more words)Show less
Step 1: Define Custom Generate Function
python
# custom_generate.py
async def custom_generate(args, samples, evaluation=False):
    """Multi-turn generation with tool calling."""
    for sample in samples:
        conversation = sample.prompt

        for turn in range(args.max_turns):
            # Generate response
            response = await generate_single(conversation)

            # Check for tool call
            tool_call = extract_tool_call(response)
            if tool_call:
                tool_result = execute_tool(tool_call)
                conversation.append({"role": "assistant", "content": response})
                conversation.append({"role": "tool", "content": tool_result})
            else:
                break

        sample.response = response
        sample.reward = compute_reward(sample)

    return samples
Step 2: Launch with Custom Function
bash
python train.py \
    --custom-generate-function-path custom_generate.py \
    --max-turns 5 \
    --prompt-data /path/to/agent_data.jsonl \
    ${MODEL_ARGS[@]}

See examples/search-r1/ for a complete multi-turn search example.


Configuration Reference

Three Argument Categories

slime uses three types of arguments:

1. Megatron Arguments (passed directly):

bash
--tensor-model-parallel-size 2
--pipeline-model-parallel-size 1
--num-layers 32
--hidden-size 4096

2. SGLang Arguments (prefixed with --sglang-):

bash
--sglang-mem-fraction-static 0.8
--sglang-context-length 8192
--sglang-log-level INFO

3. slime Arguments:

bash
# Resource allocation
--actor-num-nodes 1
--actor-num-gpus-per-node 8
--rollout-num-gpus 8
--colocate  # Share GPUs between training/inference

# Data
--prompt-data /path/to/data.jsonl
--input-key prompt
--label-key label

# Training loop
--num-rollout 3000
--rollout-batch-size 32
--n-samples-per-prompt 8
--global-batch-size 256

# Algorithm
--advantage-estimator grpo  # or: gspo, ppo, reinforce_plus_plus
--use-kl-loss
--kl-loss-coef 0.001
Key Constraints
rollout_batch_size × n_samples_per_prompt = global_batch_size × num_steps_per_rollout

Example: 32 × 8 = 256 × 1


Data Buffer System

slime's data buffer enables flexible data management:

Basic Data Source
python
class RolloutDataSource:
    def get_samples(self, num_samples):
        """Fetch prompts from dataset."""
        return self.dataset.sample(num_samples)

    def add_samples(self, samples):
        """Called after generation (no-op by default)."""
        pass
Buffered Data Source (Off-Policy)
python
class RolloutDataSourceWithBuffer(RolloutDataSource):
    def __init__(self):
        self.buffer = []

    def add_samples(self, samples):
        """Store generated samples for reuse."""
        self.buffer.extend(samples)

    def buffer_filter(self, args, buffer, num_samples):
        """Custom selection logic (prioritized, stratified, etc.)."""
        return select_best(buffer, num_samples)

Common Issues and Solutions

Issue: SGLang Engine Crash

Symptoms: Inference engine dies mid-training

Solutions:

bash
# Enable fault tolerance
--use-fault-tolerance

# Increase memory allocation
--sglang-mem-fraction-static 0.85

# Reduce batch size
--rollout-batch-size 16
Issue: Weight Sync Timeout

Symptoms: Training hangs after rollout

Solutions:

bash
# Increase sync interval
--update-weights-interval 5

# Use colocated mode (no network transfer)
--colocate
Issue: OOM During Training

Symptoms: CUDA OOM in backward pass

Solutions:

bash
# Enable gradient checkpointing
--recompute-activations

# Reduce micro-batch size
--micro-batch-size 1

# Enable sequence parallelism
--sequence-parallel
Issue: Slow Data Loading

Symptoms: GPU idle during data fetch

Solutions:

bash
# Increase data workers
--num-data-workers 4

# Use streaming dataset
--streaming-data

Supported Models

Model FamilyConfigurations
GLMGLM-4.5, GLM-4.6, GLM-4.7, GLM-Z1-9B
QwenQwen3 (4B, 8B, 30B-A3B), Qwen3-MoE, Qwen2.5
DeepSeekV3, V3.1, R1
LlamaLlama 3 (8B, 70B)
OthersKimi K2, Moonlight-16B

Each model has pre-configured scripts in scripts/models/.


Advanced Topics

Co-location Mode

Share GPUs between training and inference to reduce memory:

bash
python train.py \
    --colocate \
    --actor-num-gpus-per-node 8 \
    --sglang-mem-fraction-static 0.4 \
    ${MODEL_ARGS[@]}
Custom Reward Model
python
# custom_rm.py
class CustomRewardModel:
    def __init__(self, model_path):
        self.model = load_model(model_path)

    def compute_reward(self, prompts, responses):
        inputs = self.tokenize(prompts, responses)
        scores = self.model(inputs)
        return scores.tolist()
bash
--custom-rm-path custom_rm.py
Evaluation Multi-Task
bash
--eval-prompt-data aime /path/to/aime.jsonl \
--eval-prompt-data gsm8k /path/to/gsm8k.jsonl \
--n-samples-per-eval-prompt 16

Resources

© Luciole-Studio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in misaka/core/skills/assets/optional/mlops/slime of Luciole-Studio/Misaka-Agent.

  • SKILL.md
  • references/api-reference.md
  • references/troubleshooting.md

Open the folder on GitHubat commit b94464a

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Luciole-Studio/Misaka-Agent, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Slime next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Slime compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Slime this skillLuciole-Studio/Misaka-Agent1391 repos~2.7kAutomated safety check: PassMIT
Upgrade Depsareal-project/AReaL5.8k—~6kAutomated safety check: PassApache-2.0
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0

Similar skills

  • Upgrade Deps

    areal-project/AReaL

    Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed

More from Luciole-Studio/Misaka-Agent

All 76 skills in this repo
  • Kanban Video Orchestrator

    Luciole-Studio/Misaka-Agent

    Plan and run multi-agent video production pipelines. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check: notes
  • Ast Grep

    Luciole-Studio/Misaka-Agent

    AST-aware structural code search and rewrite via ast-grep. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Drug Discovery

    Luciole-Studio/Misaka-Agent

    Drug discovery: ChEMBL search, drug-likeness, interactions. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Fitness Nutrition

    Luciole-Studio/Misaka-Agent

    Workout planning, macros, and body metrics via wger/USDA. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Hyperframes

    Luciole-Studio/Misaka-Agent

    Render MP4/WebM videos from HTML compositions. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Osint Investigation

    Luciole-Studio/Misaka-Agent

    Follow the money via public records and sanctions data. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check passed

Questions about Slime

What does Slime do?

RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent. Slime is an agent skill from Luciole-Studio/Misaka-Agent. RL post-training for LLMs with Megatron and SGLang.

When should I use Slime?

Slime fits situations like: tasks that involve MLOps.

How do I install Slime in Claude Code?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill slime -a claude-code`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/slime in Luciole-Studio/Misaka-Agent) into .claude/skills/slime in your project. Claude Code loads it when a task matches its description.

How do I install Slime in Codex?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill slime -a codex`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/slime in Luciole-Studio/Misaka-Agent) into .agents/skills/slime in your project. Codex loads it when a task matches its description.

Can I use Slime in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Luciole-Studio/Misaka-Agent --skill slime -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/slime, .gemini/skills/slime, .github/skills/slime and .opencode/skills/slime in your project.

What does Slime need to run?

Going by SKILL.md and its folder, Slime needs the command-line tools its instructions call (python, pip, docker and git). Our summary lists: Python 3; Docker.

Does Slime access the network?

SKILL.md names 3 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: thudm.github.io and lmsys.org. This is read from the text; nothing was executed.

Is Slime safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Slime use?

Slime is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Slime use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.8k tokens, read only when the agent opens those files.

What are the alternatives to Slime?

Skills that share tags, products or a category with Slime: Upgrade Deps (areal-project/AReaL, 5.8k stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars) and Dstack Prototyping (dstackai/dstack, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Slime?

Luciole-Studio (a GitHub organization) maintains it in Luciole-Studio/Misaka-Agent, which has 139 GitHub stars. The repository holds 76 skills in this directory. The repository was last updated on October 8, 2026.

Source: Luciole-Studio/Misaka-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.