Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

MITAuto-check passedAI & LLM Engineering

Install Miles Rl Training

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill miles-rl-training -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs miles-rl-training --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/06-post-training/miles .claude/skills/miles-rl-training && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
miles-rl-training
GitHub stars
13k
Used in
2 other repos
Token cost
~2.2k tokens
SKILL.md length
697 words
Files
3 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

  • Works in 4 steps: Environment Setup → Configure Training → Enable Speculative Decoding → …
  • Training large MoE models with FP8/INT4
  • SKILL.md covers When to Use miles, Key Features, Installation and Quick Start, plus 6 more sections
  • Calls python, docker and pip; reaches github.com

What it does

Miles Rl Training is an agent skill from Orchestra-Research/AI-Research-SKILLs. Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/api-reference.md` and `references/troubleshooting.md`).

It sits in AI & LLM Engineering, covering Reinforcement learning and LLM inference and serving. It works with DeepSeek and Qwen. The repository describes itself as: Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent… The licence is MIT.

When your agent uses it

  • Training large MoE models with FP8/INT4
  • Needing train-inference alignment
  • Requiring speculative RL for maximum throughput

Example prompts

  • “Use the miles-rl-training skill to provide guidance for enterprise-grade RL training using miles, a production-ready fork of slime”
  • “/miles-rl-training”

Requirements

  • Python 3
  • Docker

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Environment Setup
  2. Configure Training
  3. Enable Speculative Decoding
  4. Enable Online MTP Training (Optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • docker
    • pip
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • lmsys.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Miles Rl Training loads about 2.2k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 697 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 697 words, ~2,224 tokens.

Download SKILL.mdSave it as .claude/skills/miles-rl-training/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
miles-rl-training
description
Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Reinforcement Learning, MoE, FP8, INT4, Enterprise, SGLang, Megatron-LM
dependencies
sglang-router>=0.2.3, ray, torch>=2.0.0, transformers>=4.40.0

miles: Enterprise-Grade RL for Large-Scale Model Training

miles is a high-performance, enterprise-ready RL framework optimized for large-scale model post-training. Built as a production fork of slime, it addresses critical challenges in MoE training stability, low-precision training, and train-inference alignment.

When to Use miles

Choose miles when you need:

  • Training 1TB+ MoE models (DeepSeek V3, Qwen3-MoE)
  • FP8 or INT4 quantization-aware training
  • Bit-wise identical train-inference alignment
  • Speculative RL for maximum throughput
  • Production stability with enterprise support

Consider alternatives when:

  • You want the research-grade original → use slime
  • You need flexible backend swapping → use verl
  • You want PyTorch-native abstractions → use torchforge

Key Features

Low-Precision Training
  • Unified FP8: End-to-end FP8 for both inference and training
  • INT4 QAT: 1TB models on single-machine VRAM (H200)
  • Rollout Routing Replay (R3): Bit-wise expert alignment for MoE
Performance Optimizations
  • Speculative RL: 25%+ rollout speedup with online SFT draft models
  • Zero-Copy Weight Sync: CUDA IPC zero-copy mapping
  • Partial Rollout: Recycle half-finished trajectories
Train-Inference Alignment
  • TIS/MIS: Truncated/Masked Importance Sampling for off-policy correction
  • Kernel-level optimization: FlashAttention-3, DeepGEMM integration

Installation

bash
# Recommended: Docker
docker pull radixark/miles:latest
docker run --rm --gpus all --ipc=host --shm-size=16g \
  -it radixark/miles:latest /bin/bash

# From source
git clone https://github.com/radixark/miles.git
cd miles
pip install -r requirements.txt
pip install -e .

Quick Start

miles inherits slime's configuration system. Basic training:

bash
python train.py \
    --advantage-estimator grpo \
    --model-name qwen3-30b-a3b \
    --hf-checkpoint /path/to/qwen3-30b-a3b-hf \
    --rollout-batch-size 512 \
    --n-samples-per-prompt 8

Workflow 1: Large MoE Training

Use this workflow for training large MoE models like DeepSeek V3 or Qwen3-MoE.

Prerequisites Checklist
  • H100/H200 GPUs with FP8 support
  • MoE model (DeepSeek V3, Qwen3-MoE)
  • Docker environment with miles
Step 1: Environment Setup
bash
# FP8 block scaling (recommended for stability)
export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1
export CUDA_DEVICE_MAX_CONNECTIONS=1
Step 2: Configure Training
bash
python train.py \
    --actor-num-gpus-per-node 8 \
    --rollout-num-gpus 8 \
    --hf-checkpoint /path/to/deepseek-v3 \
    --advantage-estimator grpo \
    --tensor-model-parallel-size 8 \
    --expert-model-parallel-size 4 \
    --prompt-data /path/to/data.jsonl \
    --num-rollout 3000
Verification Checklist
  • Model loads without errors
  • Routing decisions are consistent
  • No NaN/Inf in loss values

Workflow 2: Speculative RL Training

Use this workflow for maximum rollout throughput with EAGLE speculative decoding.

How Speculative RL Works
  1. Small draft model generates candidate tokens
  2. Target model verifies in parallel
  3. Draft model updated via online SFT to track policy
Step 1: Enable Speculative Decoding

miles supports EAGLE speculative decoding via SGLang:

bash
python train.py \
    --actor-num-gpus-per-node 8 \
    --hf-checkpoint /path/to/target-model \
    --sglang-speculative-algorithm EAGLE \
    --sglang-speculative-num-steps 3 \
    --sglang-speculative-eagle-topk 1 \
    --sglang-speculative-num-draft-tokens 4 \
    --sglang-speculative-draft-model-path /path/to/draft-model \
    --advantage-estimator grpo \
    --prompt-data /path/to/data.jsonl
Step 2: Enable Online MTP Training (Optional)

For online SFT of draft model during training:

bash
--mtp-num-layers 1 \
--enable-mtp-training \
--mtp-loss-scaling-factor 0.2

Note: Online MTP training requires a torch dist checkpoint with MTP weights. Add --mtp-num-layers 1 during checkpoint conversion from HuggingFace.

Expected Speedup
  • Standard rollout: Baseline
  • Speculative RL: 25-40% faster rollout
  • With partial rollout: Additional 10-15% throughput

Configuration Reference

miles inherits all slime arguments. See slime API Reference for the complete list.

Cluster Resources (from slime)
bash
--actor-num-nodes 1
--actor-num-gpus-per-node 8
--rollout-num-gpus 8
--rollout-num-gpus-per-engine 2
--colocate
Megatron Parallelism (from slime)
bash
--tensor-model-parallel-size 8
--pipeline-model-parallel-size 2
--expert-model-parallel-size 4    # MoE expert parallelism
Speculative Decoding (miles-specific)
bash
--sglang-speculative-algorithm EAGLE
--sglang-speculative-num-steps 3
--sglang-speculative-eagle-topk 1
--sglang-speculative-num-draft-tokens 4
--sglang-enable-draft-weights-cpu-backup
--sglang-speculative-draft-model-path /your/draft/model/path
Online MTP Training (miles-specific)
bash
--mtp-num-layers 1
--enable-mtp-training
--mtp-loss-scaling-factor 0.2

Key Features (Conceptual)

The following features are documented in miles but specific CLI flags may vary. Consult the miles repository for latest configuration.

Unified FP8 Pipeline

End-to-end FP8 sampling and training that eliminates quantization-induced discrepancy causing RL collapse in MoE models.

Show full SKILL.md (288 more words)Show less
Rollout Routing Replay (R3)

Records expert routing decisions during SGLang inference and replays them during Megatron training for bit-wise expert alignment.

How R3 Works:

  1. During SGLang inference, expert routing decisions are recorded
  2. Routing decisions stored in sample.rollout_routed_experts
  3. During Megatron training, routing is replayed instead of recomputed
  4. Ensures identical expert selection between train and inference
INT4 Quantization-Aware Training

Enables single-machine deployment of 1TB+ models (e.g., on H200).

Memory Savings with INT4:

Model SizeBF16 VRAMINT4 VRAMReduction
70B140GB45GB3.1x
235B470GB150GB3.1x
671B1.3TB420GB3.1x
Train-Inference Alignment

miles achieves "exactly 0 KL divergence" between training and inference through:

  • Flash Attention 3
  • DeepGEMM
  • Batch-invariant kernels from Thinking Machines Lab
  • torch.compile integration

Sample Data Structure

miles uses the same Sample dataclass as slime with the rollout_routed_experts field for MoE routing replay:

python
@dataclass
class Sample:
    prompt: str | list[dict]
    tokens: list[int]
    response: str
    reward: float | dict
    loss_mask: list[int]
    status: Status
    metadata: dict
    rollout_log_probs: list[float]
    rollout_routed_experts: list[list[int]]  # MoE routing for R3

See slime API Reference for the complete Sample definition.


Common Issues and Solutions

Issue: FP8 Training Collapse

Symptoms: Loss explodes, NaN values

Solutions:

  • Use block scaling: export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1
  • Reduce learning rate: --lr 5e-7
  • Ensure MoE routing is consistent between train/inference
Issue: Speculative Draft Drift

Symptoms: Low acceptance rate over time

Solutions:

  • Enable online MTP training to keep draft model aligned
  • Reduce speculative steps: --sglang-speculative-num-steps 2
  • Use CPU backup: --sglang-enable-draft-weights-cpu-backup
Issue: Train-Inference Mismatch

Symptoms: Policy divergence, reward collapse

Solutions:

  • Use TIS for off-policy correction: --use-tis --tis-threshold 0.9
  • Verify log probs match between SGLang and Megatron
  • Enable R3 for MoE models

Supported Models

FamilyModelsMoE Support
DeepSeekR1, V3, V3.2Full
Qwen2, 2.5, 3 (including MoE)Full
Llama3, 3.1, 3.3, 4Dense only
Gemma2, 3, 3NDense only
GLM4.5, 4.6, 4.7Dense only
MiniMaxM2, M2.1Full

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in 06-post-training/miles of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/api-reference.md
  • references/troubleshooting.md

Open the folder on GitHubat commit 773a529

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Miles Rl Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Miles Rl Training compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Miles Rl Training this skillOrchestra-Research/AI-Research-SKILLs13k2 repos~2.2kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Serving LLMs On Instinctamd/skills408—~4kAutomated safety check: NotesMIT
LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~3.9kAutomated safety check: PassNone
Update Ollama Cloud Modelsheypinchy/pinchy182—~3.9kAutomated safety check: NotesAGPL-3.0
Quark Torch File2file Quantizationamd/Quark182—~2.3kAutomated safety check: PassMIT

Similar skills

  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    408 GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    938 GitHub stars~3.9k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.

    182 GitHub stars~3.9k tokensUpdated 19 days ago
    AI & LLM EngineeringAuto-check: notes
  • Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.

    182 GitHub stars~2.3k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Inspect a target model and prepare metadata for Quark PTQ planning.

    182 GitHub stars~2k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Works with

Questions about Miles Rl Training

What does Miles Rl Training do?

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Miles Rl Training is an agent skill from Orchestra-Research/AI-Research-SKILLs. Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

When should I use Miles Rl Training?

Miles Rl Training fits situations like: training large MoE models with FP8/INT4; needing train-inference alignment; requiring speculative RL for maximum throughput.

How do I install Miles Rl Training in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill miles-rl-training -a claude-code`. Or copy the skill folder (06-post-training/miles in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/miles-rl-training in your project. Claude Code loads it when a task matches its description.

How do I install Miles Rl Training in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill miles-rl-training -a codex`. Or copy the skill folder (06-post-training/miles in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/miles-rl-training in your project. Codex loads it when a task matches its description.

Can I use Miles Rl Training in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill miles-rl-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/miles-rl-training, .gemini/skills/miles-rl-training, .github/skills/miles-rl-training and .opencode/skills/miles-rl-training in your project.

What does Miles Rl Training need to run?

Going by SKILL.md and its folder, Miles Rl Training needs the command-line tools its instructions call (python, docker, pip and git). Our summary lists: Python 3; Docker.

Does Miles Rl Training access the network?

SKILL.md names 2 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: lmsys.org. This is read from the text; nothing was executed.

Is Miles Rl Training safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Miles Rl Training use?

Miles Rl Training is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Miles Rl Training use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Miles Rl Training?

Skills that share tags, products or a category with Miles Rl Training: Add Model (guoqingbao/xinfer, 334 stars), Serving LLMs On Instinct (amd/skills, 408 stars), LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars) and Update Ollama Cloud Models (heypinchy/pinchy, 182 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Miles Rl Training?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.