Guides training and analyzing sparse autoencoders with SAELens to break neural network activations into interpretable features, including superposition and monosemanticity studies.
Install the "sparse-autoencoder-training" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/04-mechanistic-interpretability/saelens into .claude/skills/sparse-autoencoder-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sparse-autoencoder-training", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sparse-autoencoder-training -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "sparse-autoencoder-training" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/04-mechanistic-interpretability/saelens into .agents/skills/sparse-autoencoder-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sparse-autoencoder-training", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sparse-autoencoder-training -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "sparse-autoencoder-training" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/04-mechanistic-interpretability/saelens into .cursor/skills/sparse-autoencoder-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sparse-autoencoder-training", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sparse-autoencoder-training -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "sparse-autoencoder-training" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/04-mechanistic-interpretability/saelens into .gemini/skills/sparse-autoencoder-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sparse-autoencoder-training", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sparse-autoencoder-training -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "sparse-autoencoder-training" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/04-mechanistic-interpretability/saelens into .github/skills/sparse-autoencoder-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sparse-autoencoder-training", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sparse-autoencoder-training -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "sparse-autoencoder-training" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/04-mechanistic-interpretability/saelens into .opencode/skills/sparse-autoencoder-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sparse-autoencoder-training", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
sparse-autoencoder-training
GitHub stars
13k
Used in
6 other repos
Token cost
~3.2k tokens
SKILL.md length
581 words
Files
4 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT
At a glance
Guides training and analyzing sparse autoencoders with SAELens to break neural network activations into interpretable features, including superposition and monosemanticity studies.
Discovering interpretable features in a language model's activations
SKILL.md covers The Problem: Polysemanticity &…, When to Use SAELens, Installation and Core Concepts, plus 8 more sections
Calls pip
Training a sparse autoencoder on a chosen layer
What it does
The skill introduces SAELens, a library for training and analyzing sparse autoencoders that decompose polysemantic activations into sparse, more interpretable features. It explains the problem of polysemantic neurons and superposition, shows the encoder, sparse feature and decoder structure, and states the loss as reconstruction error plus an L1-weighted sparsity term. Installation is `pip install sae-lens`, with Python 3.10+ and `transformer-lens` 2.0.0 or newer.
Listed uses are discovering interpretable features, understanding what a model has learned, studying superposition and feature geometry, feature-based steering or ablation, and analyzing safety-relevant features. It points to TransformerLens or pyvene for basic activation analysis or causal interventions. The first workflow loads and analyzes pretrained SAEs, and reference files cover the API and tutorials. The skill cites Anthropic's monosemanticity work, in which human evaluators rated 70% of features as interpretable. The excerpt is truncated.
When your agent uses it
Discovering interpretable features in a language model's activations
Training a sparse autoencoder on a chosen layer
Studying superposition or feature geometry
Steering or ablating a model using SAE features
Example prompts
“Load a pretrained SAE for this model and show me the top features for these prompts.”
“Train a sparse autoencoder on one layer's activations with an L1 penalty and report the reconstruction loss.”
“Find features related to deceptive text and test whether ablating them changes the outputs.”
Requirements
Python 3.10 or newer
`pip install sae-lens` with `transformer-lens` 2.0.0 or newer
What it can do on your machine
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Shell commands in SKILL.md call:
pip
From the folder's file list and the shell code blocks in SKILL.md.
Network
Links to these hosts (documentation or services it may open):
github.com
neuronpedia.org
transformer-circuits.pub
lesswrong.com
arxiv.org
jbloomaus.github.io
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Sparse Autoencoder Training with SAELens loads about 3.2k tokens when it runs, and up to ~7.8k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 581 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~78
When it runs· the whole SKILL.md, loaded when a task matches
~3.2k
With references· SKILL.md plus every file in references/, read only if the agent opens them
~7.8k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/sparse-autoencoder-training/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
sparse-autoencoder-training
description
Provides guidance for training and analyzing Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features. Use when discovering interpretable features, analyzing superposition, or studying monosemantic representations in language models.
SAELens: Sparse Autoencoders for Mechanistic Interpretability
SAELens is the primary library for training and analyzing Sparse Autoencoders (SAEs) - a technique for decomposing polysemantic neural network activations into sparse, interpretable features. Based on Anthropic's groundbreaking research on monosemanticity.
Individual neurons in neural networks are polysemantic - they activate in multiple, semantically distinct contexts. This happens because models use superposition to represent more features than they have neurons, making interpretability difficult.
SAEs solve this by decomposing dense activations into sparse, monosemantic features - typically only a small number of features activate for any given input, and each feature corresponds to an interpretable concept.
When to Use SAELens
Use SAELens when you need to:
Discover interpretable features in model activations
Understand what concepts a model has learned
Study superposition and feature geometry
Perform feature-based steering or ablation
Analyze safety-relevant features (deception, bias, harmful content)
Consider alternatives when:
You need basic activation analysis → Use TransformerLens directly
You want causal intervention experiments → Use pyvene or TransformerLens
You need production steering → Consider direct activation engineering
Loss Function: MSE(original, reconstructed) + L1_coefficient × L1(features)
Key Validation (Anthropic Research)
In "Towards Monosemanticity", human evaluators found 70% of SAE features genuinely interpretable. Features discovered include:
DNA sequences, legal language, HTTP requests
Hebrew text, nutrition statements, code syntax
Sentiment, named entities, grammatical structures
Workflow 1: Loading and Analyzing Pre-trained SAEs
Step-by-Step
python
from transformer_lens import HookedTransformer
from sae_lens import SAE
# 1. Load model and pre-trained SAE
model = HookedTransformer.from_pretrained("gpt2-small", device="cuda")
sae, cfg_dict, sparsity = SAE.from_pretrained(
release="gpt2-small-res-jb",
sae_id="blocks.8.hook_resid_pre",
device="cuda"
)
# 2. Get model activations
tokens = model.to_tokens("The capital of France is Paris")
_, cache = model.run_with_cache(tokens)
activations = cache["resid_pre", 8] # [batch, pos, d_model]
# 3. Encode to SAE features
sae_features = sae.encode(activations) # [batch, pos, d_sae]
print(f"Active features: {(sae_features > 0).sum()}")
# 4. Find top features for each position
for pos in range(tokens.shape[1]):
top_features = sae_features[0, pos].topk(5)
token = model.to_str_tokens(tokens[0, pos:pos+1])[0]
print(f"Token '{token}': features {top_features.indices.tolist()}")
# 5. Reconstruct activations
reconstructed = sae.decode(sae_features)
reconstruction_error = (activations - reconstructed).norm()
Available Pre-trained SAEs
Release
Model
Layers
gpt2-small-res-jb
GPT-2 Small
Multiple residual streams
gemma-2b-res
Gemma 2B
Residual streams
Various on HuggingFace
Search tag saelens
Various
Checklist
Load model with TransformerLens
Load matching SAE for target layer
Encode activations to sparse features
Identify top-activating features per token
Validate reconstruction quality
Workflow 2: Training a Custom SAE
Step-by-Step
python
from sae_lens import SAE, LanguageModelSAERunnerConfig, SAETrainingRunner
# 1. Configure training
cfg = LanguageModelSAERunnerConfig(
# Model
model_name="gpt2-small",
hook_name="blocks.8.hook_resid_pre",
hook_layer=8,
d_in=768, # Model dimension
# SAE architecture
architecture="standard", # or "gated", "topk"
d_sae=768 * 8, # Expansion factor of 8
activation_fn="relu",
# Training
lr=4e-4,
l1_coefficient=8e-5, # Sparsity penalty
l1_warm_up_steps=1000,
train_batch_size_tokens=4096,
training_tokens=100_000_000,
# Data
dataset_path="monology/pile-uncopyrighted",
context_size=128,
# Logging
log_to_wandb=True,
wandb_project="sae-training",
# Checkpointing
checkpoint_path="checkpoints",
n_checkpoints=5,
)
# 2. Train
trainer = SAETrainingRunner(cfg)
sae = trainer.run()
# 3. Evaluate
print(f"L0 (avg active features): {trainer.metrics['l0']}")
print(f"CE Loss Recovered: {trainer.metrics['ce_loss_score']}")
Key Hyperparameters
Parameter
Typical Value
Effect
d_sae
4-16× d_model
More features, higher capacity
l1_coefficient
5e-5 to 1e-4
Higher = sparser, less accurate
lr
1e-4 to 1e-3
Standard optimizer LR
l1_warm_up_steps
500-2000
Prevents early feature death
Show full SKILL.md (246 more words)Show less
Evaluation Metrics
Metric
Target
Meaning
L0
50-200
Average active features per token
CE Loss Score
80-95%
Cross-entropy recovered vs original
Dead Features
<5%
Features that never activate
Explained Variance
>90%
Reconstruction quality
Checklist
Choose target layer and hook point
Set expansion factor (d_sae = 4-16× d_model)
Tune L1 coefficient for desired sparsity
Enable L1 warm-up to prevent dead features
Monitor metrics during training (W&B)
Validate L0 and CE loss recovery
Check dead feature ratio
Workflow 3: Feature Analysis and Steering
Analyzing Individual Features
python
from transformer_lens import HookedTransformer
from sae_lens import SAE
import torch
model = HookedTransformer.from_pretrained("gpt2-small", device="cuda")
sae, _, _ = SAE.from_pretrained(
release="gpt2-small-res-jb",
sae_id="blocks.8.hook_resid_pre",
device="cuda"
)
# Find what activates a specific feature
feature_idx = 1234
test_texts = [
"The scientist conducted an experiment",
"I love chocolate cake",
"The code compiles successfully",
"Paris is beautiful in spring",
]
for text in test_texts:
tokens = model.to_tokens(text)
_, cache = model.run_with_cache(tokens)
features = sae.encode(cache["resid_pre", 8])
activation = features[0, :, feature_idx].max().item()
print(f"{activation:.3f}: {text}")
Feature Steering
python
def steer_with_feature(model, sae, prompt, feature_idx, strength=5.0):
"""Add SAE feature direction to residual stream."""
tokens = model.to_tokens(prompt)
# Get feature direction from decoder
feature_direction = sae.W_dec[feature_idx] # [d_model]
def steering_hook(activation, hook):
# Add scaled feature direction at all positions
activation += strength * feature_direction
return activation
# Generate with steering
output = model.generate(
tokens,
max_new_tokens=50,
fwd_hooks=[("blocks.8.hook_resid_pre", steering_hook)]
)
return model.to_string(output[0])
Feature Attribution
python
# Which features most affect a specific output?
tokens = model.to_tokens("The capital of France is")
_, cache = model.run_with_cache(tokens)
# Get features at final position
features = sae.encode(cache["resid_pre", 8])[0, -1] # [d_sae]
# Get logit attribution per feature
# Feature contribution = feature_activation × decoder_weight × unembedding
W_dec = sae.W_dec # [d_sae, d_model]
W_U = model.W_U # [d_model, vocab]
# Contribution to "Paris" logit
paris_token = model.to_single_token(" Paris")
feature_contributions = features * (W_dec @ W_U[:, paris_token])
top_features = feature_contributions.topk(10)
print("Top features for 'Paris' prediction:")
for idx, val in zip(top_features.indices, top_features.values):
print(f" Feature {idx.item()}: {val.item():.3f}")
Common Issues & Solutions
Issue: High dead feature ratio
python
# WRONG: No warm-up, features die early
cfg = LanguageModelSAERunnerConfig(
l1_coefficient=1e-4,
l1_warm_up_steps=0, # Bad!
)
# RIGHT: Warm-up L1 penalty
cfg = LanguageModelSAERunnerConfig(
l1_coefficient=8e-5,
l1_warm_up_steps=1000, # Gradually increase
use_ghost_grads=True, # Revive dead features
)
We found 12 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 6 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
Sparse Autoencoder Training with SAELens next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Sparse Autoencoder Training with SAELens compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Sparse Autoencoder Training with SAELens this skillOrchestra-Research/AI-Research-SKILLs
Run AlphaGenome-PyTorch to get genomic track predictions — via the agt predict CLI (single locus, BED regions, whole chromosomes, raw FASTA sequences, or per-gene count tables/AnnData), variant…
Predict RBP binding from RNA sequence using deep learning models (RBPNet sequence-to-signal, RNAProt RNN, GraphProt2 GCN with structure, DeepCLIP, DeepRiPe multi-modal CNN) for variant-effect…
A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Questions about Sparse Autoencoder Training with SAELens
What does Sparse Autoencoder Training with SAELens do?
Guides training and analyzing sparse autoencoders with SAELens to break neural network activations into interpretable features, including superposition and monosemanticity studies. The skill introduces SAELens, a library for training and analyzing sparse autoencoders that decompose polysemantic activations into sparse, more interpretable features. It explains the problem of polysemantic neurons and superposition, shows the encoder, sparse feature and decoder structure, and states the loss as reconstruction error plus an L1-weighted sparsity term.
When should I use Sparse Autoencoder Training with SAELens?
Sparse Autoencoder Training with SAELens fits situations like: discovering interpretable features in a language model's activations; training a sparse autoencoder on a chosen layer; studying superposition or feature geometry; steering or ablating a model using SAE features.
How do I install Sparse Autoencoder Training with SAELens in Claude Code?
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill sparse-autoencoder-training -a claude-code`. Or copy the skill folder (04-mechanistic-interpretability/saelens in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/sparse-autoencoder-training in your project. Claude Code loads it when a task matches its description.
How do I install Sparse Autoencoder Training with SAELens in Codex?
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill sparse-autoencoder-training -a codex`. Or copy the skill folder (04-mechanistic-interpretability/saelens in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/sparse-autoencoder-training in your project. Codex loads it when a task matches its description.
Can I use Sparse Autoencoder Training with SAELens in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill sparse-autoencoder-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sparse-autoencoder-training, .gemini/skills/sparse-autoencoder-training, .github/skills/sparse-autoencoder-training and .opencode/skills/sparse-autoencoder-training in your project.
What does Sparse Autoencoder Training with SAELens need to run?
Going by SKILL.md and its folder, Sparse Autoencoder Training with SAELens needs the command-line tools its instructions call (pip). Our summary lists: Python 3.10 or newer; `pip install sae-lens` with `transformer-lens` 2.0.0 or newer.
Does Sparse Autoencoder Training with SAELens access the network?
SKILL.md names 6 domains. As links in the text: github.com, neuronpedia.org, transformer-circuits.pub, lesswrong.com, arxiv.org and jbloomaus.github.io. This is read from the text; nothing was executed.
Is Sparse Autoencoder Training with SAELens safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Sparse Autoencoder Training with SAELens use?
Sparse Autoencoder Training with SAELens is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Sparse Autoencoder Training with SAELens use?
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.6k tokens, read only when the agent opens those files.
What are the alternatives to Sparse Autoencoder Training with SAELens?
Skills that share tags, products or a category with Sparse Autoencoder Training with SAELens: Alphagenome Predictions (genomicsxai/alphagenome-pytorch, 162 stars), Bio Atac Seq Deep Learning Atac (GPTomics/bioSkills, 1.2k stars), Bio Clip Seq Clip Deep Learning (GPTomics/bioSkills, 1.2k stars) and Paddle Design Distributed (PaddlePaddle/Paddle, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Sparse Autoencoder Training with SAELens?
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,338 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.