Agent skill

RWKV Architecture Guide

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting.

MITAuto-check passedAI & LLM Engineering

Install RWKV Architecture Guide

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill rwkv-architecture -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs rwkv-architecture --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/01-model-architecture/rwkv .claude/skills/rwkv-architecture && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rwkv-architecture
GitHub stars
13k
Used in
2 other repos
Token cost
~1.8k tokens
SKILL.md length
322 words
Files
4 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting.

  • Streaming text generation with constant memory per token
  • SKILL.md covers Quick start, Common workflows, When to use vs alternatives and Common issues, plus 3 more sections
  • Calls pip; reaches download.pytorch.org
  • Processing very long documents without a growing KV cache

What it does

RWKV is described as combining Transformer-style parallel training with RNN-style sequential inference, giving linear-time inference, no KV cache and a fixed memory cost per token. The skill shows basic usage in both GPT mode and RNN mode, token-by-token streaming generation through the rwkv package's pipeline, and long-document processing, where state is carried forward instead of a growing cache.

Further sections cover fine-tuning with PyTorch Lightning, a memory and speed comparison against Transformers for million-token sequences, and guidance on when to choose RWKV versus Transformers, Mamba, RetNet or Hyena. Troubleshooting notes suggest gradient checkpointing with DeepSpeed ZeRO-3 and bf16 precision for out-of-memory errors during training, and enabling the CUDA kernel for slow inference. The description mentions RWKV-7 and models up to 14B parameters.

When your agent uses it

  • Streaming text generation with constant memory per token
  • Processing very long documents without a growing KV cache
  • Fine-tuning an RWKV model with PyTorch Lightning
  • Deciding between RWKV, Mamba and a Transformer

Example prompts

  • “Run RWKV in RNN mode and stream generated text token by token.”
  • “Process this very long document with RWKV and keep memory use flat.”
  • “Training RWKV runs out of GPU memory, so set up gradient checkpointing and DeepSpeed ZeRO-3.”

Requirements

  • Python with PyTorch and the rwkv package
  • Optional: a CUDA GPU for the faster kernel

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • download.pytorch.org

    Also links to:

    • arxiv.org
    • github.com
    • wiki.rwkv.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RWKV Architecture Guide loads about 1.8k tokens when it runs, and up to ~9.1k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 322 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 322 words, ~1,772 tokens.

Download SKILL.mdSave it as .claude/skills/rwkv-architecture/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
rwkv-architecture
description
RNN+Transformer hybrid with O(n) inference. Linear time, infinite context, no KV cache. Train like GPT (parallel), infer like RNN (sequential). Linux Foundation AI project. Production at Windows, Office, NeMo. RWKV-7 (March 2025). Models up to 14B parameters.
version
1.0.0
author
Orchestra Research
license
MIT
tags
RWKV, Model Architecture, RNN, Transformer Hybrid, Linear Complexity, Infinite Context, Efficient Inference, Linux Foundation, Alternative Architecture
dependencies
rwkv, torch, transformers

RWKV - Receptance Weighted Key Value

Quick start

RWKV (RwaKuv) combines Transformer parallelization (training) with RNN efficiency (inference).

Installation:

bash
# Install PyTorch
pip install torch --upgrade --extra-index-url https://download.pytorch.org/whl/cu121

# Install dependencies
pip install pytorch-lightning==1.9.5 deepspeed wandb ninja --upgrade

# Install RWKV
pip install rwkv

Basic usage (GPT mode + RNN mode):

python
import os
from rwkv.model import RWKV

os.environ["RWKV_JIT_ON"] = '1'
os.environ["RWKV_CUDA_ON"] = '1'  # Use CUDA kernel for speed

# Load model
model = RWKV(
    model='/path/to/RWKV-4-Pile-1B5-20220903-8040',
    strategy='cuda fp16'
)

# GPT mode (parallel processing)
out, state = model.forward([187, 510, 1563, 310, 247], None)
print(out.detach().cpu().numpy())  # Logits

# RNN mode (sequential processing, same result)
out, state = model.forward([187, 510], None)  # First 2 tokens
out, state = model.forward([1563], state)      # Next token
out, state = model.forward([310, 247], state)  # Last tokens
print(out.detach().cpu().numpy())  # Same logits as above!

Common workflows

Workflow 1: Text generation (streaming)

Efficient token-by-token generation:

python
from rwkv.model import RWKV
from rwkv.utils import PIPELINE

model = RWKV(model='RWKV-4-Pile-14B-20230313-ctx8192-test1050', strategy='cuda fp16')
pipeline = PIPELINE(model, "20B_tokenizer.json")

# Initial prompt
prompt = "The future of AI is"
state = None

# Generate token by token
for token in prompt:
    out, state = pipeline.model.forward(pipeline.encode(token), state)

# Continue generation
for _ in range(100):
    out, state = pipeline.model.forward(None, state)
    token = pipeline.sample_logits(out)
    print(pipeline.decode(token), end='', flush=True)

Key advantage: Constant memory per token (no growing KV cache)

Workflow 2: Long context processing (infinite context)

Process million-token sequences:

python
model = RWKV(model='RWKV-4-Pile-14B', strategy='cuda fp16')

# Process very long document
state = None
long_document = load_document()  # e.g., 1M tokens

# Stream through entire document
for chunk in chunks(long_document, chunk_size=1024):
    out, state = model.forward(chunk, state)

# State now contains information from entire 1M token document
# Memory usage: O(1) (constant, not O(n)!)
Workflow 3: Fine-tuning RWKV

Standard fine-tuning workflow:

python
# Training script
import pytorch_lightning as pl
from rwkv.model import RWKV
from rwkv.trainer import RWKVTrainer

# Configure model
config = {
    'n_layer': 24,
    'n_embd': 1024,
    'vocab_size': 50277,
    'ctx_len': 1024
}

# Setup trainer
trainer = pl.Trainer(
    accelerator='gpu',
    devices=8,
    precision='bf16',
    strategy='deepspeed_stage_2',
    max_epochs=1
)

# Train
model = RWKV(config)
trainer.fit(model, train_dataloader)
Workflow 4: RWKV vs Transformer comparison

Memory comparison (1M token sequence):

python
# Transformer (GPT)
# Memory: O(n²) for attention
# KV cache: 1M × hidden_dim × n_layers × 2 (keys + values)
# Example: 1M × 4096 × 24 × 2 = ~400GB (impractical!)

# RWKV
# Memory: O(1) per token
# State: hidden_dim × n_layers = 4096 × 24 = ~400KB
# 1,000,000× more efficient!

Speed comparison (inference):

python
# Transformer: O(n) per token (quadratic overall)
# First token: 1 computation
# Second token: 2 computations
# ...
# 1000th token: 1000 computations

# RWKV: O(1) per token (linear overall)
# Every token: 1 computation
# 1000th token: 1 computation (same as first!)

When to use vs alternatives

Use RWKV when:

  • Need very long context (100K+ tokens)
  • Want constant memory usage
  • Building streaming applications
  • Need RNN efficiency with Transformer performance
  • Memory-constrained deployment

Key advantages:

  • Linear time: O(n) vs O(n²) for Transformers
  • No KV cache: Constant memory per token
  • Infinite context: No fixed window limit
  • Parallelizable training: Like GPT
  • Sequential inference: Like RNN

Use alternatives instead:

  • Transformers: Need absolute best performance, have compute
  • Mamba: Want state-space models
  • RetNet: Need retention mechanism
  • Hyena: Want convolution-based approach

Common issues

Issue: Out of memory during training

Use gradient checkpointing and DeepSpeed:

python
trainer = pl.Trainer(
    strategy='deepspeed_stage_3',  # Full ZeRO-3
    precision='bf16'
)

Issue: Slow inference

Enable CUDA kernel:

python
os.environ["RWKV_CUDA_ON"] = '1'

Issue: Model not loading

Check model path and strategy:

python
model = RWKV(
    model='/absolute/path/to/model.pth',
    strategy='cuda fp16'  # Or 'cpu fp32' for CPU
)

Issue: State management in RNN mode

Always pass state between forward calls:

python
# WRONG: State lost
out1, _ = model.forward(tokens1, None)
out2, _ = model.forward(tokens2, None)  # No context from tokens1!

# CORRECT: State preserved
out1, state = model.forward(tokens1, None)
out2, state = model.forward(tokens2, state)  # Has context from tokens1

Advanced topics

Time-mixing and channel-mixing: See references/architecture-details.md for WKV operation, time-decay mechanism, and receptance gates.

State management: See references/state-management.md for att_x_prev, att_kv, ffn_x_prev states, and numerical stability considerations.

RWKV-7 improvements: See references/rwkv7.md for latest architectural improvements (March 2025) and multimodal capabilities.

Hardware requirements

  • GPU: NVIDIA (CUDA 11.6+) or CPU
  • VRAM (FP16):
    • 169M model: 1GB
    • 430M model: 2GB
    • 1.5B model: 4GB
    • 3B model: 8GB
    • 7B model: 16GB
    • 14B model: 32GB
  • Inference: O(1) memory per token
  • Training: Parallelizable like GPT

Performance (vs Transformers):

  • Speed: Similar training, faster inference
  • Memory: 1000× less for long sequences
  • Scaling: Linear vs quadratic

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in 01-model-architecture/rwkv of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/architecture-details.md
  • references/rwkv7.md
  • references/state-management.md

Open the folder on GitHubat commit 773a529

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

RWKV Architecture Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RWKV Architecture Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RWKV Architecture Guide this skillOrchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT
Quark Env Preflightamd/Quark182—~1.4kAutomated safety check: PassMIT
Magpie Kernel Evaluatoramd/skills408—~2.3kAutomated safety check: PassMIT
Quark Torch Debugamd/Quark182—~1.9kAutomated safety check: NotesMIT
Quark Installamd/Quark182—~3.5kAutomated safety check: NotesMIT
Alphagenome Predictionsgenomicsxai/alphagenome-pytorch162—~868Automated safety check: PassApache-2.0

Similar skills

  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    182 GitHub stars~1.4k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    408 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

    182 GitHub stars~1.9k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check: notes
  • Quark Install

    amd/Quark

    Installs or verifies AMD Quark and ensures the selected Python environment has an accelerator-matched PyTorch.

    182 GitHub stars~3.5k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check: notes
  • Alphagenome Predictions

    genomicsxai/alphagenome-pytorch

    Run AlphaGenome-PyTorch to get genomic track predictions — via the agt predict CLI (single locus, BED regions, whole chromosomes, raw FASTA sequences, or per-gene count tables/AnnData), variant…

    162 GitHub stars~868 tokensUpdated 25 days ago
    AI & LLM EngineeringAuto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    182 GitHub stars~3k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about RWKV Architecture Guide

What does RWKV Architecture Guide do?

Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting. RWKV is described as combining Transformer-style parallel training with RNN-style sequential inference, giving linear-time inference, no KV cache and a fixed memory cost per token. The skill shows basic usage in both GPT mode and RNN mode, token-by-token streaming generation through the rwkv package's pipeline, and long-document processing, where state is carried forward instead of a growing cache.

When should I use RWKV Architecture Guide?

RWKV Architecture Guide fits situations like: streaming text generation with constant memory per token; processing very long documents without a growing KV cache; fine-tuning an RWKV model with PyTorch Lightning; deciding between RWKV, Mamba and a Transformer.

How do I install RWKV Architecture Guide in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill rwkv-architecture -a claude-code`. Or copy the skill folder (01-model-architecture/rwkv in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/rwkv-architecture in your project. Claude Code loads it when a task matches its description.

How do I install RWKV Architecture Guide in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill rwkv-architecture -a codex`. Or copy the skill folder (01-model-architecture/rwkv in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/rwkv-architecture in your project. Codex loads it when a task matches its description.

Can I use RWKV Architecture Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill rwkv-architecture -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rwkv-architecture, .gemini/skills/rwkv-architecture, .github/skills/rwkv-architecture and .opencode/skills/rwkv-architecture in your project.

What does RWKV Architecture Guide need to run?

Going by SKILL.md and its folder, RWKV Architecture Guide needs the command-line tools its instructions call (pip). Our summary lists: Python with PyTorch and the rwkv package; Optional: a CUDA GPU for the faster kernel.

Does RWKV Architecture Guide access the network?

SKILL.md names 5 domains. In commands or code: download.pytorch.org; the agent is likely to contact it when it follows the instructions. As links in the text: arxiv.org, github.com, wiki.rwkv.com and huggingface.co. This is read from the text; nothing was executed.

Is RWKV Architecture Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RWKV Architecture Guide use?

RWKV Architecture Guide is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RWKV Architecture Guide use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.3k tokens, read only when the agent opens those files.

What are the alternatives to RWKV Architecture Guide?

Skills that share tags, products or a category with RWKV Architecture Guide: Quark Env Preflight (amd/Quark, 182 stars), Magpie Kernel Evaluator (amd/skills, 408 stars), Quark Torch Debug (amd/Quark, 182 stars) and Quark Install (amd/Quark, 182 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RWKV Architecture Guide?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.