Topic · AI & LLM Engineering
Best AI interpretability skills for Claude Code, Codex and other agents.
- skills
- 23
- official
- 1
AI interpretability skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | 1.Esmfold2 Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. | JimLiu/ | 227 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 2 | Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails… | RedWoodOG/ | 177 | 6 repos | ~3.8k | Automated safety check: Pass | MIT | 4 mo ago |
| 3 | Adds a new anomaly-detection model to anomalib under src/anomalib/models/. | open-edge-platform/ | 6.2k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Look up what a Biohub ESM-C sparse-autoencoder (SAE) feature means — its label, description, top-activating proteins, decoder neighbours, and activation statistics — by querying the Biohub… | softnanolab/ | 148 | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 5 | Guides training and analyzing sparse autoencoders with SAELens to break neural network activations into interpretable features, including superposition and monosemanticity studies. | Orchestra-Research/ | 13k | 6 repos | ~3.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Guides mechanistic interpretability work with TransformerLens: loading models, caching activations, using HookPoints, activation patching and attention-pattern analysis. | Orchestra-Research/ | 13k | 4 repos | ~3k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | Provides guidance for interpreting and manipulating neural network internals using nnsight with optional NDIF remote execution. | Orchestra-Research/ | 13k | 3 repos | ~3.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 8 | Guides causal experiments on PyTorch models with pyvene, such as causal tracing, activation patching and interchange intervention training, to test how a model works. | Orchestra-Research/ | 13k | 3 repos | ~3.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 9 | Explains machine learning predictions with SHAP: picking the right explainer, computing Shapley values and drawing waterfall, beeswarm, bar and force plots. | davila7/ | 32k | 12 repos | ~4.6k | Automated safety check: Pass | MIT | today |
| 10 | Write allocation-efficient buffer code in Corvus.JsonSchema using the codebase's established three-tier pooling pattern: stackalloc → ArrayPool → ThreadStatic caches. | corvus-dotnet/ | 199 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Applies the reasoning style of Geoffrey Hinton, deep learning pioneer and 2018 Turing Award winner. | K-Dense-AI/ | 282 | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 12 | Designing review workflows to surface and mitigate bias in AI outputs. | Owl-Listener/ | 180 | — | ~650 | Automated safety check: Pass | MIT | 3 mo ago |
| 13 | Reach for this skill whenever you are discussing reinforcement learning, agentic AI systems, AI alignment, continual learning, or the philosophical limits of large language models. | K-Dense-AI/ | 282 | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 14 | 14.Saelens Train sparse autoencoders to interpret model features. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 125 | 1 repo | ~3.7k | Automated safety check: Pass | MIT | today |
| 15 | NV-Tesseract Forecasting — transformer-based multivariate time series forecasting with DARR (context-enhanced kNN retrieval), interpretability, and fine-tuning. | NVIDIA/ | 3.5k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 | today |
| 16 | Designing for informed user consent, opt-out, and human override. | Owl-Listener/ | 180 | — | ~598 | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Operational guide for implementing a new Envilder runtime SDK. | macalbert/ | 138 | — | ~1.9k | Automated safety check: Pass | MIT | 2 days ago |
| 18 | Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values. | jaechang-hits/ | 370 | 1 repo | ~2.9k | Automated safety check: Pass | CC-BY-4.0 | 8 days ago |
| 19 | Designs primary, secondary, and exploratory endpoints for biomedical and clinical research protocols. | aipoch/ | 2k | — | ~4.3k | Automated safety check: Pass | MIT | 20 days ago |
| 20 | Designs retrospective or prospective clinical cohort study protocols for biomedical and clinical research. | aipoch/ | 2k | — | ~5.7k | Automated safety check: Pass | MIT | 20 days ago |
| 21 | Your AI research and engineering brain trust. An agent skill from majiayu000/claude-skill-registry. | majiayu000/ | 666 | 1 repo | ~3.7k | Automated safety check: Pass | MIT | today |
| 22 | A skill your agent uses when deciding whether a project is a strong AAAI submission across its broad AI scope, should be reframed or routed to a dedicated track such as AI for Social Impact or AI… | brycewang-stanford/ | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | 10 days ago |
| 23 | The creation of effective visualizations is a fundamental component of data analysis. | bioMate-AI/ | 804 | — | ~1.7k | Automated safety check: Pass | Unknown | 3 mo ago |
Questions, answered from the data.
What is the best AI interpretability skill?
Esmfold2 from JimLiu/science-skills ranks first of the 23 AI interpretability skills listed here, with the highest score: its repository has 227 GitHub stars, 4 other GitHub owners carry a copy, its SKILL.md loads about 2.5k tokens and it passes the automated safety check with no findings. Next come Obliteratus and Anomalib Adding A Model.
Which AI interpretability skills are official?
1 of the 23 AI interpretability skills are official, published by the vendor's own GitHub organization: Tao Finetune Nv Tesseract Forecasting.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67