The skill introduces TransformerLens, a library for mechanistic interpretability on GPT-style language models that exposes a HookPoint on every activation. It lists when to use it: reverse-engineering learned algorithms, activation patching and causal tracing, studying attention patterns and information flow, analyzing circuits such as induction heads and the IOI circuit, caching intermediate activations and direct logit attribution.
Setup is pip install transformer-lens. The main class is HookedTransformer, which wraps a model, and run_with_cache returns logits plus a cache of activations whose keys, such as resid_pre, resid_mid, resid_post and attn_out per layer, are tabulated with their shapes. A table lists supported model families, covering more than 50 models including GPT-2, LLaMA, Pythia, Mistral, Phi, Qwen, OPT and Gemma.
It also points elsewhere when TransformerLens fits poorly: nnsight or pyvene for non-transformer architectures or higher-level causal interventions, SAELens for sparse autoencoders, and nnsight with NDIF for remote runs on very large models. Reference files include api.md and tutorials.md.