ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design.

MITAuto-check passedAI & LLM Engineering

Install Esm

skills CLI
$ npx skills add adaptyvbio/protein-design-skills --skill esm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install adaptyvbio/protein-design-skills esm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/adaptyvbio/protein-design-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/esm .claude/skills/esm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
esm
GitHub stars
163
Token cost
~2k tokens
SKILL.md length
638 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
MIT

At a glance

ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design.

  • Computing pseudo-log-likelihood (PLL)
  • SKILL.md covers Which model to use, Prerequisites, ESM C: embeddings and scoring and ESMFold2: complex structure…, plus 4 more sections
  • Calls uv and pip; reaches github.com and biohub.ai
  • Mutation-effect scores

What it does

Esm is an agent skill from adaptyvbio/protein-design-skills. ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design. Use this skill when: (1) Computing pseudo-log-likelihood (PLL) or mutation-effect scores, (2) Getting protein embeddings for clustering or filtering, (3) Predicting complex structures with ESMFold2, (4) Designing binders by inverting ESMFold2, (5) Filtering designs by sequence plausibility. For diffusion-based structure prediction, use boltz or chai. For QC thresholds, use protein-qc. For gradient-based…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Protein structure and design and Embeddings. It works with Python. The repository describes itself as: Claude Code skills for protein design. The licence is MIT.

When your agent uses it

  • Computing pseudo-log-likelihood (PLL)
  • Mutation-effect scores
  • Getting protein embeddings for clustering
  • Predicting complex structures with ESMFold2

Example prompts

  • “/esm”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 59dd633. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • biohub.ai

    Also links to:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Esm loads about 2k tokens when it runs. Until then it costs about 137 tokens; SKILL.md has 638 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from adaptyvbio/protein-design-skills at commit 59dd633, republished under its MIT licence (© adaptyvbio). 638 words, ~2,024 tokens.

Download SKILL.mdSave it as .claude/skills/esm/SKILL.md (or your agent's skills folder).
name
esm
description
ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design. Use this skill when: (1) Computing pseudo-log-likelihood (PLL) or mutation-effect scores, (2) Getting protein embeddings for clustering or filtering, (3) Predicting complex structures with ESMFold2, (4) Designing binders by inverting ESMFold2, (5) Filtering designs by sequence plausibility. For diffusion-based structure prediction, use boltz or chai. For QC thresholds, use protein-qc. For gradient-based multi-objective design, use mosaic.
license
MIT
category
design-tools
tags
sequence-design, embeddings, scoring, structure-prediction, binder
proteinbase_slug
esm2-optimization
proteinbase_url
https://proteinbase.com/design-methods/esm2-optimization
biomodals_script
modal_esm2_predict_masked.py

ESM Protein Language Models

The ESM line is maintained at github.com/Biohub/esm (Chan Zuckerberg Biohub, MIT license; the older evolutionaryscale/esm URL redirects here). The current generation ships three artifacts: ESM C (language model), ESMFold2 (structure prediction), and ESM Atlas (a map of predicted structures). Weights are on huggingface.co/biohub; the hosted API is at biohub.ai.

This skill covers ESM C, ESMFold2, and legacy ESM2. ESM3 is not covered because its open weights are non-commercial.

Which model to use

TaskModel
Embeddings, PLL, mutation scoringESM C (ESMC-6B), or ESM2 for a lighter run
Complex structure predictionESMFold2
High-throughput single-sequence foldingESMFold2 fast mode
Binder designESMFold2 inversion (see below), or the mosaic / bindcraft skills
Variant effect / zero-shot scoringESM C or ESM2

Prerequisites

RequirementMinimumRecommended
Python3.10+3.11
PyTorch2.0+Latest
CUDA12.0+12.1+
GPU VRAM24GB (ESM2 / small ESMC)80GB (ESMC-6B, ESMFold2)

ESM C: embeddings and scoring

ESM C is the successor to ESM2. It improves long-range structural understanding as model scale grows and is the default choice for embeddings, pseudo-log-likelihood, and mutation-effect scoring.

Python (Hugging Face)
python
from transformers import AutoModelForMaskedLM, AutoTokenizer
import torch

model_id = "biohub/ESMC-6B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForMaskedLM.from_pretrained(
    model_id, output_hidden_states=True, torch_dtype=torch.bfloat16
).eval().cuda()

batch = tok(["MKTAYIAKQRQISFVK..."], return_tensors="pt").to("cuda")
with torch.no_grad():
    out = model(**batch)

logits = out.logits                      # for PLL / mutation scoring
embeddings = out.hidden_states[-1]       # per-residue representations

Install the package with pip install esm@git+https://github.com/Biohub/esm.git@main.

Hosted API
python
from esm.sdk import esmc_client
from esm.sdk.api import ESMProtein, LogitsConfig

model = esmc_client(model="esmc-600m-2024-12", url="https://biohub.ai", token="<API token>")
tensor = model.encode(ESMProtein(sequence="MKTAYIAKQRQISFVK..."))
out = model.logits(tensor, LogitsConfig(sequence=True, return_embeddings=True))

ESMC-6B has open weights; esmc-600m is the smaller API model. For mutation scoring and fine-tuning, see the esmc_mutation_scoring and esmc_finetune notebooks under cookbook/tutorials.

ESMFold2: complex structure prediction

ESMFold2 is built on ESMC-6B with a diffusion structure head. Unlike the original ESMFold, it predicts complexes (protein, DNA, ligand, and modified residues), takes an optional MSA, and has a single-sequence fast mode for high-throughput screening. It is validated for protein-protein interaction design and leads DockQ pass-rate on Foldbench protein-protein and antibody-antigen complexes.

Modal (biomodals)
bash
printf '>protein|A\nMKTAYIAKQRQISFVK...\n' > target.faa
uv run --with modal modal run modal_esmfold2.py --input-faa target.faa

The FASTA header tags protein|, dna|, rna|, and ligand| (SMILES) let you fold complexes. GPU defaults to A100-40GB (set with MODAL_GPU).

Python (local weights)
python
from transformers.models.esmfold2.modeling_esmfold2 import ESMFold2Model
from esm.models.esmfold2 import ProteinInput, StructurePredictionInput, ESMFold2InputBuilder

model = ESMFold2Model.from_pretrained("biohub/ESMFold2").cuda().eval()
spi = StructurePredictionInput(sequences=[ProteinInput(id="A", sequence="BINDER_SEQ")])
result = ESMFold2InputBuilder().fold(model, spi, num_loops=20, num_sampling_steps=100)
# result.plddt, result.ptm, result.iptm, result.complex.to_mmcif()

For single-sequence high-throughput folding, the fast variant is the SDK model string esmfold2-fast-2026-05 (HF repo biohub/ESMFold2-Fast). ESMFold2 is one option for complex validation alongside boltz and chai; ranking a shortlist across more than one predictor is more reliable than trusting a single model.

Show full SKILL.md (303 more words)Show less

Binder design by inverting ESMFold2

The binder_design cookbook runs gradient optimization through ESMFold2 (a BindCraft-style loop) with an ESMC language-model term for sequence plausibility. The published protocol is validated in the lab to nanomolar affinity across five targets and supports both minibinders and antibody-derived scFvs with framework scaffolds.

biomodals wraps this as modal_esmfold2_binder_design.py:

bash
uv run --with modal modal run modal_esmfold2_binder_design.py \
  --target-name pd-l1 --binder-name minibinder
  • Targets: presets cd45, ctla4, egfr, pd-l1, pdgfr, or pass --target-sequence.
  • Binders: presets minibinder and antibody frameworks (for example trastuzumab_framework_vhvl), or pass --binder-sequence with # for designable positions. Use --is-antibody for scFv designs.
  • Rank candidates by ipTM, filter minibinders to pI below 6, then validate the top shortlist with boltz or chai and rank with ipsae.

Adaptyv's own tests of these models showed ESMFold2-inversion binder design costing about $0.85 per accepted design, averaged across 7 targets.

For a framework that composes ESMFold2 with other predictors in one objective, use the mosaic skill.

ESM2 (legacy)

ESM2 still works well for quick embeddings and PLL when ESMC-6B is too large for the available GPU.

python
import torch, esm
model, alphabet = esm.pretrained.esm2_t33_650M_UR50D()
bc = alphabet.get_batch_converter()
model = model.eval().cuda()
_, _, toks = bc([("seq1", "MKTAYIAKQRQISFVK...")])
with torch.no_grad():
    rep = model(toks.cuda(), repr_layers=[33])["representations"][33]
ModelParametersUse
esm2_t12_35M35MFast screening
esm2_t33_650M650MStandard embeddings/PLL
esm2_t36_3B3BHighest-quality ESM2

PLL interpretation

PLL (pseudo-log-likelihood) scores how natural a sequence looks to the model. Higher is more natural. Designed sequences often score lower than natural ones, so treat PLL as a soft filter, not a hard cutoff.

Normalized PLLInterpretation
> 0.2Very natural
0.0 to 0.2Natural-like
-0.5 to 0.0Acceptable
< -0.5May be unnatural

Troubleshooting

IssueCauseFix
CUDA out of memoryESMC-6B / ESMFold2 too largeUse ESMC-600m API, ESM2, or an 80GB GPU
Wrong layer for embeddingsLayer index mismatchUse the last hidden state (layer 33 for ESM2-650M)
Invalid amino acidNon-standard residueCheck for non-canonical characters
Slow ESMFold2 on many designsFull MSA modeUse esmfold2-fast-2026-05 single-sequence mode

Next: Validate structures with boltz or chai, rank with ipsae, then filter with protein-qc.

© adaptyvbio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/esm of adaptyvbio/protein-design-skills.

Open the folder on GitHubat commit 59dd633

Compare with similar skills

Esm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Esm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Esm this skilladaptyvbio/protein-design-skills163—~2kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs13k3 repos~1.6kAutomated safety check: PassMIT
RAG Company Knowledge AssistantHermes-brasil/hermes-brasil152—~1.1kAutomated safety check: PassMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 3 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Company Knowledge Assistant

    Hermes-brasil/hermes-brasil

    Portuguese guide to building a retrieval-augmented generation assistant over a company's documents, with embeddings, section-based chunking, retrieval and a client workflow.

    152 GitHub stars~1.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Esm

    davila7/claude-code-templates

    Comprehensive toolkit for protein language models including ESM3 (generative multimodal protein design across sequence, structure, and function) and ESM C (efficient protein embeddings and…

    32k GitHub starsUsed in 10 repos~2.6k tokens
    AI & LLM EngineeringAuto-check: warnings

More from adaptyvbio/protein-design-skills

All 24 skills in this repo
  • Alphafold

    adaptyvbio/protein-design-skills

    Validate protein designs using AlphaFold2 structure prediction.

    163 GitHub starsUsed in 4 repos~1.2k tokens
    Auto-check passed
  • Bindcraft

    adaptyvbio/protein-design-skills

    End-to-end binder design using BindCraft hallucination. An agent skill from adaptyvbio/protein-design-skills.

    163 GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check passed
  • Boltzgen

    adaptyvbio/protein-design-skills

    All-atom protein design using BoltzGen diffusion model. An agent skill from adaptyvbio/protein-design-skills.

    163 GitHub starsUsed in 4 repos~2k tokens
    Auto-check passed
  • Chai

    adaptyvbio/protein-design-skills

    Structure prediction using Chai-1, a foundation model for molecular structure.

    163 GitHub starsUsed in 4 repos~1.5k tokens
    Auto-check passed
  • Protein Design Workflow

    adaptyvbio/protein-design-skills

    End-to-end guidance for protein design pipelines. An agent skill from adaptyvbio/protein-design-skills.

    163 GitHub starsUsed in 4 repos~1.2k tokens
    Auto-check passed
  • Protein Qc

    adaptyvbio/protein-design-skills

    Quality control metrics and filtering thresholds for protein design.

    163 GitHub starsUsed in 4 repos~3.2k tokens
    Auto-check passed

Works with

Questions about Esm

What does Esm do?

ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design. Esm is an agent skill from adaptyvbio/protein-design-skills. ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design.

When should I use Esm?

Esm fits situations like: computing pseudo-log-likelihood (PLL); mutation-effect scores; getting protein embeddings for clustering; predicting complex structures with ESMFold2.

How do I install Esm in Claude Code?

Run `npx skills add adaptyvbio/protein-design-skills --skill esm -a claude-code`. Or copy the skill folder (skills/esm in adaptyvbio/protein-design-skills) into .claude/skills/esm in your project. Claude Code loads it when a task matches its description.

How do I install Esm in Codex?

Run `npx skills add adaptyvbio/protein-design-skills --skill esm -a codex`. Or copy the skill folder (skills/esm in adaptyvbio/protein-design-skills) into .agents/skills/esm in your project. Codex loads it when a task matches its description.

Can I use Esm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adaptyvbio/protein-design-skills --skill esm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/esm, .gemini/skills/esm, .github/skills/esm and .opencode/skills/esm in your project.

What does Esm need to run?

Going by SKILL.md and its folder, Esm needs the command-line tools its instructions call (uv and pip). Our summary lists: Python 3.

Does Esm access the network?

SKILL.md names 3 domains. In commands or code: github.com and biohub.ai; the agent is likely to contact these when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is Esm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Esm use?

Esm is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Esm use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Esm?

Skills that share tags, products or a category with Esm: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Esmfold2 (JimLiu/science-skills, 227 stars) and Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Esm?

adaptyvbio (a GitHub organization) maintains it in adaptyvbio/protein-design-skills, which has 163 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on June 11, 2026.

Source: adaptyvbio/protein-design-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.