Agent skill

Rlm

by brainqub3 in brainqub3/RLM

Recursive Language Model (RLM) loop for processing a context that is too large to read into the conversation directly.

MITAuto-check: notes

Install Rlm

skills CLI
$ npx skills add brainqub3/RLM --skill rlm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brainqub3/RLM rlm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brainqub3/RLM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/rlm .claude/skills/rlm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rlm
GitHub stars
393
Token cost
~3k tokens
SKILL.md length
1,239 words
Files
20 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Recursive Language Model (RLM) loop for processing a context that is too large to read into the conversation directly.

  • Works in 4 steps: Initialise — load the context, read only… → Probe — understand the format with… → Decompose + sub-query — write code that… → …
  • Classifying every item
  • SKILL.md covers Mental model, When to use this, Inputs ($ARGUMENTS) and The loop (Algorithm 1), plus 4 more sections
  • Runs Python scripts from its folder; calls python and claude

What it does

Rlm is an agent skill from brainqub3/RLM. Recursive Language Model (RLM) loop for processing a context that is too large to read into the conversation directly. Loads the context as a variable in a persistent Python REPL and answers the query by writing code that probes, chunks, and programmatically sub-queries a cheap LLM (llmquery) over slices of it, then aggregates. Use this WHENEVER the user points you at a big context file/log/transcript/codebase/scraped corpus (anything from ~50K chars up to millions) and asks a question that needs most of the…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 25 other files, including scripts (for example `eval/README.md`, `eval/_cache/build_eval.py` and `eval/_cache/pyarrow_fetch.py`).

It works with Python. The repository describes itself as: Claude code setup as an RLM scaffhold Implemented by Brainqub3. The licence is MIT.

When your agent uses it

  • Classifying every item
  • Multi-hop lookup
  • Summarising the whole thing -- ESPECIALLY when the answer depends on almost every line and a single retrieval/grep wont do
  • It even if the user doesnt say RLM: phrases like this file is huge

Example prompts

  • “depends on almost every line”
  • “t do. Trigger it even if the user doesn”
  • “: phrases like”
  • “/rlm”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit, Grep, Glob

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Initialise — load the context, read only its metadata
  2. Probe — understand the format with small, cheap code
  3. Decompose + sub-query — write code that calls the LLM over chunks
  4. Aggregate + answer — compute the final answer, set it in the REPL

What it can do on your machine

Read from SKILL.md and the folder at commit 0039c00. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rlm loads about 3k tokens when it runs. Until then it costs about 249 tokens; SKILL.md has 1,239 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~249
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit, Grep, Glob

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from brainqub3/RLM at commit 0039c00, republished under its MIT licence (© brainqub3). 1,239 words, ~2,995 tokens.

Download SKILL.mdSave it as .claude/skills/rlm/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.
name
rlm
description
Recursive Language Model (RLM) loop for processing a context that is too large to read into the conversation directly. Loads the context as a variable in a persistent Python REPL and answers the query by writing code that probes, chunks, and programmatically sub-queries a cheap LLM (`llm_query`) over slices of it, then aggregates. Use this WHENEVER the user points you at a big context file/log/transcript/codebase/scraped corpus (anything from ~50K chars up to millions) and asks a question that needs most of the content -- counting, aggregating, classifying every item, multi-hop lookup, or summarising the whole thing -- ESPECIALLY when the answer "depends on almost every line" and a single retrieval/grep won't do. Trigger it even if the user doesn't say "RLM": phrases like "this file is huge", "go through the whole log", "how many X across all of these", "label every row", or "it won't fit in context" are all signals to use this skill. Prefer it over dumping the file into chat.
allowed-tools
Bash, Read, Write, Edit, Grep, Glob

rlm — Recursive Language Model loop

A faithful instantiation of Recursive Language Models (Zhang, Kraska, Khattab; arXiv:2512.24601), Algorithm 1, on Claude Code's primitives. The paper's insight: an arbitrarily long prompt should not be fed into a model's context window at all. It should live in an environment the model interacts with programmatically, recursively calling a model over slices of it. That's what this skill does.

Mental model

You (the main Claude Code conversation) are the root model. You do not read the big context into this conversation. Instead:

  • The context lives as a context variable inside a persistent Python REPL (scripts/rlm_repl.py). You only ever see metadata about it (length, a short prefix) and the truncated stdout of code you run — never the whole thing. This is the one rule that lets the context be far larger than any window.
  • You answer by writing REPL code that probes the context, decomposes it, and calls a cheap sub-LM over the pieces:
    • llm_query(prompt) / llm_query_map(prompts) — a single / a parallel batch of plain sub-LM calls (a nested headless claude -p, tools off). This is the leaf: it reads a bounded chunk in its own window and returns text.
    • rlm_query(context, query) — a full recursive RLM over a sub-context, for sub-tasks that are themselves too big for one leaf call (depth > 1). Falls back to llm_query at the depth limit.
  • You build intermediate results into REPL variables/buffers, then return the answer by setting it in the REPL: FINAL(answer) or FINAL_VAR(varname).

The division of labour that makes this work: the LLM does the semantics (classify this question, extract this fact, summarise this section); your Python does the bookkeeping (loop over every chunk, count, aggregate, format). Do not ask the LLM to count or do arithmetic over the whole corpus, and do not try to do the semantics yourself in Python with keyword heuristics — that is exactly the failure mode the paper's ablations show. Split the work along that seam.

When to use this

Use it when the context won't fit comfortably in the conversation and the task needs broad access to it: aggregation/counting over every item, labelling every row, multi-hop questions across a corpus, whole-document summarisation, or "the answer depends on almost every line". For a one-off needle lookup in a file you can just grep, you don't need this.

Inputs ($ARGUMENTS)

  • context=<path> (required): path to the large context file.
  • query=<question> (required): what to answer about it.
  • Optional: sub_model=<alias> (default haiku), max_workers=<int> (default 8), max_depth=<int> (default 1; >1 enables recursive rlm_query).

If the user didn't supply them, ask for the context file path and the query.

Set optional knobs via environment before running, e.g.: export RLM_SUB_MODEL=haiku RLM_MAX_WORKERS=8 RLM_MAX_DEPTH=1.

The loop (Algorithm 1)

Run these via the Bash tool. State persists between calls in .claude/rlm_state/state.pkl. By default, init also creates a standalone audit replay package under .claude/rlm_runs/<run_id>/; every exec saves the submitted Python as steps/step_XXXX.py.

1. Initialise — load the context, read only its metadata
bash
python .claude/skills/rlm/scripts/rlm_repl.py init <context_path>

This prints the context's type, char/line/token estimate, and a short prefix. Do not read the context file with the Read tool — that defeats the purpose. It also prints the audit replay package path. Use --no-audit only when you do not want standalone step scripts.

2. Probe — understand the format with small, cheap code

Look at the shape of the data before deciding a strategy. Print small slices and structure, not the bulk:

bash
python .claude/skills/rlm/scripts/rlm_repl.py exec <<'PY'
print(peek(0, 1500))                 # head
lines = [l for l in content.splitlines() if l.strip()]
print("lines:", len(lines))
print("sample:", lines[1] if len(lines) > 1 else "")
PY

Ask: Is it line-oriented? JSON objects? Markdown sections? Logs with timestamps? The format dictates the chunking.

3. Decompose + sub-query — write code that calls the LLM over chunks

This is the core. Chunk the context, build one prompt per chunk, and fan the semantic work out to the sub-LM with llm_query_map (parallel). Keep the chunks fat (a leaf can hold a large slice — batch to minimise call count) but small enough that the sub-LM stays accurate. Accumulate results in a variable; let Python do the aggregation.

bash
python .claude/skills/rlm/scripts/rlm_repl.py exec <<'PY'
# Example shape for an aggregation task: derive records from the actual format,
# ask leaf LMs for semantic labels, then count/aggregate in Python.
records = [line.strip() for line in content.splitlines() if line.strip()]

# Fill these from the user's query and what you observed while probing. Do not
# assume the file's delimiter, item marker, or labels before inspecting it.
question = "What should be classified or extracted for each record?"
categories = ["category_a", "category_b", "category_c"]

def build(batch, start):
    body = "\n".join(f"{start+i}: {record}" for i, record in enumerate(batch))
    return (
        f"{question}\n"
        f"Use exactly one of these categories: {', '.join(categories)}.\n"
        "Output exactly one line per record as 'N: <category>'. No extra text.\n\n"
        + body
    )

BATCH = 50
prompts = [build(records[s:s+BATCH], s) for s in range(0, len(records), BATCH)]
outs = llm_query_map(prompts)          # parallel sub-LM calls; order preserved

import re
from collections import Counter
labels = {}
for out in outs:
    for ln in out.splitlines():
        m = re.match(r"\s*(\d+)\s*[:.\)]\s*(.+)", ln)
        if m:
            labels[int(m.group(1))] = m.group(2).strip().strip("*[]").lower()
missing = [i for i in range(len(records)) if i not in labels]
counts = Counter(labels.values())
print("classified:", len(labels), "/", len(records), "missing:", len(missing))
print("counts:", dict(counts))
PY

Because the REPL is persistent, items, labels, and counts survive into your next exec. Inspect, sanity-check, and re-run pieces as needed. Save durable intermediate text with add_buffer(...) (it lives in the buffers list).

4. Aggregate + answer — compute the final answer, set it in the REPL

Do the final arithmetic/formatting in Python, then set the answer. The answer must be a REPL variable or literal — not just something you say in chat (so it can be arbitrarily long and is captured verbatim):

bash
python .claude/skills/rlm/scripts/rlm_repl.py exec <<'PY'
top = counts.most_common(1)[0][0]
answer = f"Label: {top}"
FINAL_VAR("answer")          # or: FINAL(f"Label: {top}")
PY
python .claude/skills/rlm/scripts/rlm_repl.py final   # prints the stored answer

Then report that final answer to the user, in the exact output format the query asked for.

Show full SKILL.md (500 more words)Show less

REPL interface (what's available inside exec)

Injected automatically every exec (you never import or define these):

namewhat it does
context / contentthe full context, as a str (two names for the same value)
llm_query(prompt, model=None, timeout=300, system=...)one sub-LM leaf call → text
llm_query_map(prompts, max_workers=8, ...)many leaf calls in parallel → list of texts, in order
rlm_query(context_text, query, ...)recursive RLM over a sub-context (depth>1); falls back to llm_query at max depth
FINAL(answer) / FINAL_VAR(name)set the final answer (literal / by variable name)
peek(start, end)a slice of the raw context
grep(pattern, max_matches, window)regex search → matches with surrounding snippets
chunked(seq, size)yield size-length slices of a list (lines, etc.)
chunk_indices(size, overlap) / write_chunks(dir, ...)character chunk spans / write chunks to files
add_buffer(text) / buffersappend to / read the persistent list of intermediate results

Your own variables persist between exec calls (anything pickleable). stdout is truncated (~8000 chars) before you see it — print summaries and samples, not bulk.

Standalone audit replay

Each audited exec writes:

  • steps/step_XXXX.py - a normal Python script containing the original REPL code plus a small prelude that recreates the RLM globals.
  • steps/step_XXXX.json - metadata such as hashes, output paths, and final status.
  • steps/step_XXXX.stdout.txt / .stderr.txt - the original captured output.
  • runtime/ - a copy of the runtime needed by the generated scripts.
  • replay_all.py - runs all saved steps from a clean replay_state.pkl.

Replay with:

bash
python .claude/rlm_runs/<run_id>/replay_all.py

Replay calls llm_query live, so sub-LM text can differ from the original run. The replay checkpoint is separate from the live REPL state and does not mutate .claude/rlm_state/state.pkl.

Guardrails — these are where RLMs win or lose

  • Never read the whole context into the conversation. No Read tool on the context file, no print(content). Work through the REPL and sub-LM calls. If you catch yourself wanting the full text in chat, chunk it and llm_query it instead.
  • Split semantics from arithmetic. LLM = meaning (classify/extract/summarise); Python = counting/aggregation/formatting. Counting with the LLM, or classifying with if "keyword" in line, both score badly.
  • Batch sub-calls; don't make one call per line. Put many items in each llm_query (e.g. 50–100 short lines per call) and parallelise with llm_query_map. Thousands of one-item calls are slow and costly for no accuracy gain. But keep batches small enough that the sub-LM doesn't drop or miscount items — verify classified == total and re-run any short/garbled batch.
  • Process the entire context before answering for aggregation tasks — the point is that you can't shortcut it. Check your coverage counts.
  • Return the answer from the REPL via FINAL/FINAL_VAR, then echo it to the user in the requested format. Don't stop at intermediate buffers.
  • Recursion (rlm_query) is for sub-tasks too big for one leaf, e.g. "analyse these 500 documents that each need their own chunking". It is slower and costlier; most tasks (including pure aggregation) only need llm_query. Default max_depth is 1.

Notes

  • The sub-LM is a nested headless Claude Code (claude -p) and reuses your existing login — no API key or SDK. llm_query runs it with tools off (a plain LLM); rlm_query runs it with bash + this skill on (its own REPL).
  • Keep all scratch/state under .claude/rlm_state/.

© brainqub3, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 19 other files (scripts) in .claude/skills/rlm of brainqub3/RLM.

  • SKILL.md
  • eval/.gitattributes
  • eval/.gitignore
  • eval/README.md
  • eval/_cache/build_eval.py
  • eval/_cache/pyarrow_fetch.py
  • eval/_cache/verify_eval.py
  • eval/_upstream_ref/constants.py
  • eval/_upstream_ref/eval_helpers.py
  • eval/_upstream_ref/task_constructors.py
  • eval/data/contexts/trec_coarse_cw6.txt
  • eval/data/contexts/trec_coarse_cw8.txt
  • eval/data/contexts_with_labels/GOLD_LABEL_STATS.md
  • eval/data/contexts_with_labels/trec_coarse_cw6.txt
  • eval/data/contexts_with_labels/trec_coarse_cw8.txt
  • … and 5 more

Open the folder on GitHubat commit 0039c00

Compare with similar skills

Rlm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rlm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rlm this skillbrainqub3/RLM393—~3kAutomated safety check: NotesMIT
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
NotebookLM Research AssistantPleasePrompto/notebooklm-skill7.8k14 repos~2.4kAutomated safety check: NotesMIT
Manim Video Productionbrowser-use/video-use28k6 repos~3kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • NotebookLM Research Assistant

    PleasePrompto/notebooklm-skill

    Lets Claude Code ask questions of your Google NotebookLM notebooks through browser automation and return answers grounded in your uploaded sources.

    7.8k GitHub starsUsed in 14 repos~2.4k tokens
    Knowledge ManagementAuto-check: notes
  • Manim Video Production

    browser-use/video-use

    Produces math and technical explainer videos with Manim Community Edition: concept animations, equation derivations, algorithm walkthroughs and data stories.

    28k GitHub starsUsed in 6 repos~3k tokens
    Media & CreativeAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • PPT Master

    hugohe3/ppt-master

    Generates editable PowerPoint decks, rebuilds slides from images, fills .pptx templates and polishes existing presentations through routed workflows.

    58k GitHub starsUsed in 1 repo~2.5k tokens
    Documents & OfficeAuto-check passed

Works with

Questions about Rlm

What does Rlm do?

Recursive Language Model (RLM) loop for processing a context that is too large to read into the conversation directly. Rlm is an agent skill from brainqub3/RLM. Recursive Language Model (RLM) loop for processing a context that is too large to read into the conversation directly.

When should I use Rlm?

Rlm fits situations like: classifying every item; multi-hop lookup; summarising the whole thing -- ESPECIALLY when the answer depends on almost every line and a single retrieval/grep wont do; it even if the user doesnt say RLM: phrases like this file is huge.

How do I install Rlm in Claude Code?

Run `npx skills add brainqub3/RLM --skill rlm -a claude-code`. Or copy the skill folder (.claude/skills/rlm in brainqub3/RLM) into .claude/skills/rlm in your project. Claude Code loads it when a task matches its description.

How do I install Rlm in Codex?

Run `npx skills add brainqub3/RLM --skill rlm -a codex`. Or copy the skill folder (.claude/skills/rlm in brainqub3/RLM) into .agents/skills/rlm in your project. Codex loads it when a task matches its description.

Can I use Rlm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brainqub3/RLM --skill rlm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rlm, .gemini/skills/rlm, .github/skills/rlm and .opencode/skills/rlm in your project.

What does Rlm need to run?

Going by SKILL.md and its folder, Rlm needs Python for the scripts in its folder and the command-line tools its instructions call (python and claude). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit, Grep, Glob.

Does Rlm access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rlm safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Rlm use?

Rlm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rlm use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Rlm?

Skills that share tags, products or a category with Rlm: MCP Server Builder (anthropics/skills, 180k stars), PDF Processing (anthropics/skills, 180k stars), NotebookLM Research Assistant (PleasePrompto/notebooklm-skill, 7.8k stars) and Manim Video Production (browser-use/video-use, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rlm?

brainqub3 (a GitHub organization) maintains it in brainqub3/RLM, which has 393 GitHub stars. The repository was last updated on June 22, 2026.

Source: brainqub3/RLM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.