Agent skill

Sae Feature Annotations

by softnanolab in softnanolab/bagel

Look up what a Biohub ESM-C sparse-autoencoder (SAE) feature means — its label, description, top-activating proteins, decoder neighbours, and activation statistics — by querying the Biohub…

MITAuto-check passedData & Analytics

Install Sae Feature Annotations

skills CLI
$ npx skills add softnanolab/bagel --skill sae-feature-annotations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install softnanolab/bagel sae-feature-annotations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/softnanolab/bagel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/sae-feature-annotations .claude/skills/sae-feature-annotations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sae-feature-annotations
GitHub stars
148
Token cost
~1.5k tokens
SKILL.md length
665 words
Files
4 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Look up what a Biohub ESM-C sparse-autoencoder (SAE) feature means — its label, description, top-activating proteins, decoder neighbours, and activation statistics — by querying the Biohub…

  • Works in 2 steps: The whole map, in one call — GET… → One feature, in full — GET…
  • Has SAE feature indices (e.g
  • SKILL.md covers The one thing to say first,…, One source: the live Biohub API, Two endpoints = two levels of… and The tool, plus 2 more sections
  • Runs Python scripts from its folder; reaches biohub.ai; needs ESM_API_KEY and FORGE_TOKEN

What it does

Sae Feature Annotations is an agent skill from softnanolab/bagel. Look up what a Biohub ESM-C sparse-autoencoder (SAE) feature means — its label, description, top-activating proteins, decoder neighbours, and activation statistics — by querying the Biohub feature-annotation API. Use this whenever the user has SAE feature indices (e.g. from boileroom's SAE model / pooledfeatures) and wants to interpret them: "what is SAE feature 12345", "which features fired on my protein and what do they correspond to", "annotate these feature indices", "what proteins most activate this…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `evals/evals.json`, `references/api.md` and `scripts/sae_features.py`).

It sits in Data & Analytics, covering AI interpretability and Statistics. The repository describes itself as: Protein Engineering via Exploration of an Energy Landscape. The licence is MIT.

When your agent uses it

  • Has SAE feature indices (e.g
  • The user mentions Biohub feature annotations
  • The feature viewer
  • The ESMC-6B layer-60 SAE codebook

Example prompts

  • “s SAE model / pooledfeatures) and wants to interpret them:”
  • “which features fired on my protein and what do they correspond to”
  • “annotate these feature indices”
  • “/sae-feature-annotations”

Requirements

  • Python 3
  • A credential in ESM_API_KEY
  • A credential in FORGE_TOKEN

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The whole map, in one call — GET /esm/protein/api/v1alpha1/features returns every
  2. One feature, in full — GET /esm/protein/api/v1alpha1/features/{feature_index}

What it can do on your machine

Read from SKILL.md and the folder at commit fa8c941. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • biohub.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ESM_API_KEY
    • FORGE_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sae Feature Annotations loads about 1.5k tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 241 tokens; SKILL.md has 665 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~241
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from softnanolab/bagel at commit fa8c941, republished under its MIT licence (© softnanolab). 665 words, ~1,533 tokens.

Download SKILL.mdSave it as .claude/skills/sae-feature-annotations/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
sae-feature-annotations
description
Look up what a Biohub ESM-C sparse-autoencoder (SAE) feature *means* — its label, description, top-activating proteins, decoder neighbours, and activation statistics — by querying the Biohub feature-annotation API. Use this whenever the user has SAE feature indices (e.g. from `boileroom`'s `SAE` model / `pooled_features`) and wants to interpret them: "what is SAE feature 12345", "which features fired on my protein and what do they correspond to", "annotate these feature indices", "what proteins most activate this feature", "find the zinc-binding SAE feature", "decoder nearest neighbours of feature N", or any request to turn ESM-C SAE feature numbers into biology. Also trigger when the user mentions Biohub feature annotations, the feature viewer, or the ESMC-6B layer-60 SAE codebook. These annotations are only valid for the `ESMC-6B-sae-layer60-k64-codebook16384` SAE — the skill says so up front and refuses to pretend otherwise.

Annotating ESM-C SAE features from Biohub

This skill turns SAE feature indices into biology. A sparse autoencoder trained on ESM-C representations decomposes each residue into a sparse set of interpretable features (directions in a codebook). Biohub publishes an annotation for each feature — a human label, top-activating proteins, decoder neighbours, activation stats — and this skill fetches and presents them by querying the Biohub API live.

The one thing to say first, every time

These annotations are only valid for the SAE ESMC-6B-sae-layer60-k64-codebook16384 — ESM-C 6B, transformer layer 60, TopK k=64, codebook of 2**14 = 16384 features. This is the SAE from Language Modeling Materializes a World Model of Protein Biology (Biohub, 2026), and the default Forge SAE that boileroom's SAE model uses.

A feature_index (0…16383) means something completely different under any other layer, k, or codebook. So before interpreting anything, confirm the user produced their features with this exact SAE (in boileroom that is the default feature_source="forge" path, i.e. forge_sae_model = "esmc-6b-2024-12-sae-layer60-k64-codebook16384"). If they used a local 300M/600M SAE or a different layer, tell them these labels do not apply and stop — do not hand them annotations that describe a different feature basis.

Lead with this. Do not bury it under the results.

One source: the live Biohub API

There is no bundled offline table. The two Biohub annotation endpoints are public reads, so the skill always queries them live. An API key is optional: if the user has one (ESM_API_KEY / FORGE_TOKEN, the same credential boileroom's Forge backend uses) it is sent; if not, the request is made with an anonymous placeholder token, which is enough for the annotation endpoints. Base URL defaults to https://biohub.ai.

If a deployment ever enforces auth and the anonymous call is rejected, the fix is to set ESM_API_KEY. That is the only case where a key matters for reading annotations — producing the features themselves is a different tool (see the last note).

Two endpoints = two levels of detail

  1. The whole map, in one call — GET /esm/protein/api/v1alpha1/features returns every feature's feature_index, label, and short description at once (all 16384 in a {"data": [...]} envelope). This is the feature -> annotation map, behind list and search.

  2. One feature, in full — GET /esm/protein/api/v1alpha1/features/{feature_index} returns the longform description, activation pattern, category, top-activating UniRef90 and SwissProt proteins, decoder nearest-neighbour feature indices, and activation statistics. This is behind get.

See references/api.md for the exact request/response shapes, base URL, and auth details.

Show full SKILL.md (269 more words)Show less

The tool

Everything goes through one stdlib-only script (no dependencies to install):

scripts/sae_features.py
  list   [--limit N]      # list features (map endpoint)
  search <text>           # find features whose label/description matches <text>
  get    <feature_index>  # full detail for one feature (detail endpoint)
  # shared flags: --base-url, --token, --json

Auth resolves from --token, then ESM_API_KEY, then FORGE_TOKEN, else the anonymous placeholder. Base URL defaults to https://biohub.ai (--base-url / BIOHUB_BASE_URL to override).

Workflow

  1. State the validity constraint (the section above) and confirm the user's features came from ESMC-6B-sae-layer60-k64-codebook16384. If not, stop and explain why the labels don't apply.

  2. Figure out what they need:

    • Just labels / short descriptions for some indices, or a keyword search → list / search.
    • Top-activating proteins, decoder neighbours, or activation stats for a feature → get.
  3. Query. Run the relevant subcommand. It works with or without an API key; prefer the user's own key via ESM_API_KEY (or --token) when they have one, but never ask them to paste a secret into the chat if an env var will do, and never store it.

  4. Present results plainly. Give the label and description for each index; for get, summarize the top proteins, decoder nearest-neighbour feature indices, and activation statistics. Keep the feature index next to every annotation so the mapping is unambiguous.

Notes

  • Don't invent labels. If the live call fails, report the error rather than guessing — there is no offline substitute.
  • If Biohub rejects the request or the endpoint shape differs from references/api.md, the auth/parse logic lives in one place (_http_get_json / _live_feature_map in the script) — adjust there. The defaults (Authorization: Bearer <token>, base https://biohub.ai, {"data": [...]} envelope) are verified against the live API.
  • This skill only reads annotations. Producing the SAE features themselves (running ESM-C + the SAE to get feature_index values for a sequence) is boileroom's SAE model, on the features/sae branch — a different tool.

© softnanolab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in .claude/skills/sae-feature-annotations of softnanolab/bagel.

  • SKILL.md
  • evals/evals.json
  • references/api.md
  • scripts/sae_features.py

Open the folder on GitHubat commit fa8c941

Compare with similar skills

Sae Feature Annotations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sae Feature Annotations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sae Feature Annotations this skillsoftnanolab/bagel148—~1.5kAutomated safety check: PassMIT
Nan Safe Correlationjaechang-hits/SciAgent-Skills3701 repos~2.9kAutomated safety check: PassCC-BY-4.0
AI Daily DigestvigorX777/ai-daily-digest1.6k—~1.3kAutomated safety check: PassNone
Anomalib Adding A Modelopen-edge-platform/anomalib6.2k—~1.9kAutomated safety check: PassApache-2.0
Optimize Op VerifyCVCUDA/CV-CUDA2.7k—~424Automated safety check: PassCustom licence
SHAP Model Explainabilitydavila7/claude-code-templates32k12 repos~4.6kAutomated safety check: PassMIT

Similar skills

  • Nan Safe Correlation

    jaechang-hits/SciAgent-Skills

    Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values.

    370 GitHub starsUsed in 1 repo~2.9k tokens
    Data & AnalyticsAuto-check passed
  • AI Daily Digest

    vigorX777/ai-daily-digest

    Fetches RSS feeds from 90 top Hacker News blogs (curated by Karpathy), uses AI to score and filter articles, and generates a daily digest in Markdown with Chinese-translated titles, category…

    1.6k GitHub stars~1.3k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Anomalib Adding A Model

    open-edge-platform/anomalib

    Adds a new anomaly-detection model to anomalib under src/anomalib/models/.

    6.2k GitHub stars~1.9k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Optimize Op Verify

    CVCUDA/CV-CUDA

    Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md.

    2.7k GitHub stars~424 tokensUpdated 21 days ago
    Data & AnalyticsAuto-check passed
  • SHAP Model Explainability

    davila7/claude-code-templates

    Explains machine learning predictions with SHAP: picking the right explainer, computing Shapley values and drawing waterfall, beeswarm, bar and force plots.

    32k GitHub starsUsed in 12 repos~4.6k tokens
    Data & AnalyticsAuto-check passed
  • Review a CV-CUDA operator's BENCHMARK coverage — drivers, layout axis, baselines, the basic-tier floor, row counts, and coverage statistics.

    2.7k GitHub stars~274 tokensUpdated 21 days ago
    Data & AnalyticsAuto-check passed

Questions about Sae Feature Annotations

What does Sae Feature Annotations do?

Look up what a Biohub ESM-C sparse-autoencoder (SAE) feature means — its label, description, top-activating proteins, decoder neighbours, and activation statistics — by querying the Biohub…. Sae Feature Annotations is an agent skill from softnanolab/bagel. Look up what a Biohub ESM-C sparse-autoencoder (SAE) feature means — its label, description, top-activating proteins, decoder neighbours, and activation statistics — by querying the Biohub feature-annotation API.

When should I use Sae Feature Annotations?

Sae Feature Annotations fits situations like: has SAE feature indices (e.g; the user mentions Biohub feature annotations; the feature viewer; the ESMC-6B layer-60 SAE codebook.

How do I install Sae Feature Annotations in Claude Code?

Run `npx skills add softnanolab/bagel --skill sae-feature-annotations -a claude-code`. Or copy the skill folder (.claude/skills/sae-feature-annotations in softnanolab/bagel) into .claude/skills/sae-feature-annotations in your project. Claude Code loads it when a task matches its description.

How do I install Sae Feature Annotations in Codex?

Run `npx skills add softnanolab/bagel --skill sae-feature-annotations -a codex`. Or copy the skill folder (.claude/skills/sae-feature-annotations in softnanolab/bagel) into .agents/skills/sae-feature-annotations in your project. Codex loads it when a task matches its description.

Can I use Sae Feature Annotations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softnanolab/bagel --skill sae-feature-annotations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sae-feature-annotations, .gemini/skills/sae-feature-annotations, .github/skills/sae-feature-annotations and .opencode/skills/sae-feature-annotations in your project.

What does Sae Feature Annotations need to run?

Going by SKILL.md and its folder, Sae Feature Annotations needs Python for the scripts in its folder and credentials named ESM_API_KEY and FORGE_TOKEN. Our summary lists: Python 3; A credential in ESM_API_KEY; A credential in FORGE_TOKEN.

Does Sae Feature Annotations access the network?

SKILL.md names 1 domain. In commands or code: biohub.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Sae Feature Annotations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Sae Feature Annotations use?

Sae Feature Annotations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sae Feature Annotations use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Sae Feature Annotations?

Skills that share tags, products or a category with Sae Feature Annotations: Nan Safe Correlation (jaechang-hits/SciAgent-Skills, 370 stars), AI Daily Digest (vigorX777/ai-daily-digest, 1.6k stars), Anomalib Adding A Model (open-edge-platform/anomalib, 6.2k stars) and Optimize Op Verify (CVCUDA/CV-CUDA, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sae Feature Annotations?

softnanolab (a GitHub organization) maintains it in softnanolab/bagel, which has 148 GitHub stars. The repository was last updated on October 6, 2026.

Source: softnanolab/bagel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.