ESM2 protein language model for embeddings and sequence scoring.

MITAuto-check passedAI & LLM Engineering

Install Esm

skills CLI
$ npx skills add NeverSight/learn-skills.dev --skill esm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NeverSight/learn-skills.dev esm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NeverSight/learn-skills.dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data/skills-md/adaptyvbio/protein-design-skills/esm .claude/skills/esm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
esm
GitHub stars
216
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
233 words
Files
13
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

ESM2 protein language model for embeddings and sequence scoring.

  • Computing pseudo-log-likelihood (PLL) scores
  • SKILL.md covers Prerequisites, How to run, Key parameters and Output format, plus 6 more sections
  • Calls modal
  • Getting protein embeddings for clustering

What it does

Esm is an agent skill from NeverSight/learn-skills.dev. ESM2 protein language model for embeddings and sequence scoring. Use this skill when: (1) Computing pseudo-log-likelihood (PLL) scores, (2) Getting protein embeddings for clustering, (3) Filtering designs by sequence plausibility, (4) Zero-shot variant effect prediction, (5) Analyzing sequence-function relationships. For structure prediction, use chai or boltz. For QC thresholds, use protein-qc.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files (for example `stats.json`).

It sits in AI & LLM Engineering, covering Protein structure and design and Embeddings. The repository describes itself as: Curated high-quality AI Agent Skills. Search, install, copy and share. Works with Claude Code, Cursor, OpenClaw, and other AI coding tools. The licence is MIT.

When your agent uses it

  • Computing pseudo-log-likelihood (PLL) scores
  • Getting protein embeddings for clustering
  • Filtering designs by sequence plausibility
  • Zero-shot variant effect prediction

Example prompts

  • “/esm”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 08f9d22. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • modal

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Esm loads about 1.1k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 233 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NeverSight/learn-skills.dev at commit 08f9d22, republished under its MIT licence (© NeverSight). 233 words, ~1,147 tokens.

Download SKILL.mdSave it as .claude/skills/esm/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
esm
description
ESM2 protein language model for embeddings and sequence scoring. Use this skill when: (1) Computing pseudo-log-likelihood (PLL) scores, (2) Getting protein embeddings for clustering, (3) Filtering designs by sequence plausibility, (4) Zero-shot variant effect prediction, (5) Analyzing sequence-function relationships. For structure prediction, use chai or boltz. For QC thresholds, use protein-qc.
license
MIT
category
design-tools
tags
sequence-design, embeddings, scoring
proteinbase_slug
esm2-optimization
proteinbase_url
https://proteinbase.com/design-methods/esm2-optimization
biomodals_script
modal_esm2_predict_masked.py

ESM2 Protein Language Model

Prerequisites

RequirementMinimumRecommended
Python3.8+3.10
PyTorch1.10+2.0+
CUDA11.0+11.7+
GPU VRAM8GB24GB (A10G)
RAM16GB32GB

How to run

First time? See Installation Guide to set up Modal and biomodals.

Option 1: Modal
bash
cd biomodals
modal run modal_esm2_predict_masked.py \
  --input-faa sequences.fasta \
  --out-dir embeddings/

GPU: A10G (24GB) | Timeout: 300s default

python
import torch
import esm

# Load model
model, alphabet = esm.pretrained.esm2_t33_650M_UR50D()
batch_converter = alphabet.get_batch_converter()
model = model.eval().cuda()

# Process sequences
data = [("seq1", "MKTAYIAKQRQISFVK...")]
batch_labels, batch_strs, batch_tokens = batch_converter(data)

with torch.no_grad():
    results = model(batch_tokens.cuda(), repr_layers=[33])

# Get embeddings
embeddings = results["representations"][33]

Key parameters

ESM2 Models
ModelParametersSpeedQuality
esm2_t6_8M8MFastestFast screening
esm2_t12_35M35MFastGood
esm2_t33_650M650MMediumBetter
esm2_t36_3B3BSlowBest

Output format

embeddings/
├── embeddings.npy       # (N, 1280) array
├── pll_scores.csv       # PLL for each sequence
└── metadata.json        # Sequence info

Sample output

Successful run
$ modal run modal_esm2_predict_masked.py --input-faa designs.fasta
[INFO] Loading ESM2-650M model...
[INFO] Processing 100 sequences...
[INFO] Computing pseudo-log-likelihood...

embeddings/pll_scores.csv:
sequence_id,pll,pll_normalized,length
design_0,-0.82,0.15,78
design_1,-0.95,0.08,85
design_2,-1.23,-0.12,72
...

Summary:
  Mean PLL: -0.91
  Sequences with PLL > 0: 42/100 (42%)

What good output looks like:

  • PLL_normalized: > 0.0 (more natural-like)
  • Embeddings shape: (N, 1280) for 650M model
  • Higher PLL = more natural sequence

Decision tree

Should I use ESM2?
│
├─ What do you need?
│  ├─ Sequence plausibility score → ESM2 PLL ✓
│  ├─ Embeddings for clustering → ESM2 ✓
│  ├─ Variant effect prediction → ESM2 ✓
│  └─ Structure prediction → Use ESMFold
│
├─ What model size?
│  ├─ Fast screening → esm2_t12_35M
│  ├─ Standard use → esm2_t33_650M ✓
│  └─ Best quality → esm2_t36_3B
│
└─ Use case?
   ├─ QC filtering → PLL > 0.0 threshold
   ├─ Diversity analysis → Mean-pooled embeddings
   └─ Mutation scanning → Per-position log-odds

PLL interpretation

Normalized PLLInterpretation
> 0.2Very natural sequence
0.0 - 0.2Good, natural-like
-0.5 - 0.0Acceptable
< -0.5May be unnatural

Typical performance

Campaign SizeTime (A10G)Cost (Modal)Notes
100 sequences5-10 min~$1Quick screen
1000 sequences30-60 min~$5Standard
5000 sequences2-3h~$20Large batch

Throughput: ~100-200 sequences/minute with 650M model.


Verify

bash
wc -l embeddings/pll_scores.csv  # Should match input + 1 (header)

Troubleshooting

OOM errors: Use smaller model or batch sequences Slow processing: Use esm2_t12_35M for speed Low PLL scores: May indicate unusual/designed sequences

Error interpretation
ErrorCauseFix
RuntimeError: CUDA out of memorySequence too long or large batchReduce batch size
KeyError: representationWrong layer requestedUse layer 33 for 650M model
ValueError: sequenceInvalid amino acidCheck for non-standard AAs

Next: Structure prediction with chai or boltz → protein-qc for filtering.

© NeverSight, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files in data/skills-md/adaptyvbio/protein-design-skills/esm of NeverSight/learn-skills.dev.

  • SKILL.md
  • description_ar.txt
  • description_cn.txt
  • description_de.txt
  • description_en.txt
  • description_es.txt
  • description_fr.txt
  • description_it.txt
  • description_ja.txt
  • description_ko.txt
  • description_ru.txt
  • description_tw.txt
  • stats.json

Open the folder on GitHubat commit 08f9d22

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NeverSight/learn-skills.dev, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Esm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Esm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Esm this skillNeverSight/learn-skills.dev2161 repos~1.1kAutomated safety check: PassMIT
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Esmdavila7/claude-code-templates32k10 repos~2.6kAutomated safety check: WarnMIT
Esmmajiayu000/claude-skill-registry6661 repos~1.7kAutomated safety check: PassMIT
Esmadaptyvbio/protein-design-skills163—~2kAutomated safety check: PassMIT
Unimoljinzhezenggroup/computational-chemistry-agent-skills1481 repos~1.5kAutomated safety check: PassLGPL-3.0-or-later

Similar skills

  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Esm

    davila7/claude-code-templates

    Comprehensive toolkit for protein language models including ESM3 (generative multimodal protein design across sequence, structure, and function) and ESM C (efficient protein embeddings and…

    32k GitHub starsUsed in 10 repos~2.6k tokens
    AI & LLM EngineeringAuto-check: warnings
  • Esm

    majiayu000/claude-skill-registry

    Toolkit for protein language models (ESM3 for multimodal generative protein design; ESM C for efficient embeddings).

    666 GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Esm

    adaptyvbio/protein-design-skills

    ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design.

    163 GitHub stars~2k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Unimol

    jinzhezenggroup/computational-chemistry-agent-skills

    A standardized CLI wrapper for Uni-Mol molecular ML workflows that handles representation extraction (embeddings), model training (regression/classification), and property prediction with built-in…

    148 GitHub starsUsed in 1 repo~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Explore

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to survey a journal or field, fetch papers from OpenAlex, cluster topics, build exploration embeddings, or search named explore libraries under…

    576 GitHub stars~755 tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed

More from NeverSight/learn-skills.dev

All 43 skills in this repo
  • AI Marketing Videos

    NeverSight/learn-skills.dev

    Create AI marketing videos for ads, promos, product launches, and brand content.

    216 GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Google Calendar

    NeverSight/learn-skills.dev

    Interact with Google Calendar via the Google Calendar API – list upcoming events, create new events, update or delete them.

    216 GitHub starsUsed in 2 repos~826 tokens
    Auto-check passed
  • Agent Orchestrator

    NeverSight/learn-skills.dev

    Meta-agent skill for orchestrating complex tasks through autonomous sub-agents.

    216 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • AI Automation Workflows

    NeverSight/learn-skills.dev

    Build automated AI workflows combining multiple models and services.

    216 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • AI Content Pipeline

    NeverSight/learn-skills.dev

    Build multi-step AI content creation pipelines combining image, video, audio, and text.

    216 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • AI Podcast Creation

    NeverSight/learn-skills.dev

    Create AI-powered podcasts with text-to-speech, music, and audio editing.

    216 GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed

Questions about Esm

What does Esm do?

ESM2 protein language model for embeddings and sequence scoring. dev. ESM2 protein language model for embeddings and sequence scoring.

When should I use Esm?

Esm fits situations like: computing pseudo-log-likelihood (PLL) scores; getting protein embeddings for clustering; filtering designs by sequence plausibility; zero-shot variant effect prediction.

How do I install Esm in Claude Code?

Run `npx skills add NeverSight/learn-skills.dev --skill esm -a claude-code`. Or copy the skill folder (data/skills-md/adaptyvbio/protein-design-skills/esm in NeverSight/learn-skills.dev) into .claude/skills/esm in your project. Claude Code loads it when a task matches its description.

How do I install Esm in Codex?

Run `npx skills add NeverSight/learn-skills.dev --skill esm -a codex`. Or copy the skill folder (data/skills-md/adaptyvbio/protein-design-skills/esm in NeverSight/learn-skills.dev) into .agents/skills/esm in your project. Codex loads it when a task matches its description.

Can I use Esm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NeverSight/learn-skills.dev --skill esm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/esm, .gemini/skills/esm, .github/skills/esm and .opencode/skills/esm in your project.

What does Esm need to run?

Going by SKILL.md and its folder, Esm needs the command-line tools its instructions call (modal). Our summary lists: Python 3.

Does Esm access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Esm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Esm use?

Esm is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Esm use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Esm?

Skills that share tags, products or a category with Esm: Esmfold2 (JimLiu/science-skills, 227 stars), Esm (davila7/claude-code-templates, 32k stars), Esm (majiayu000/claude-skill-registry, 666 stars) and Esm (adaptyvbio/protein-design-skills, 163 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Esm?

NeverSight (a GitHub organization) maintains it in NeverSight/learn-skills.dev, which has 216 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on October 6, 2026.

Source: NeverSight/learn-skills.dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.