Agent skill

Esmfold2

by JimLiu in JimLiu/science-skills

Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Esmfold2

skills CLI
$ npx skills add JimLiu/science-skills --skill esmfold2 -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JimLiu/science-skills esmfold2 --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/esmfold2 .claude/skills/esmfold2 && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
esmfold2
GitHub stars
227
Used in
4 other repos
Token cost
~2.5k tokens
SKILL.md length
603 words
Files
3 (incl. references)
Skills in repo
27
Repo updated
First seen
Licence
Apache-2.0

At a glance

Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

  • Predicting complex structures with single-sequence input
  • SKILL.md covers Install, Usage — local model, Model variants on HF biohub/ and Throughput:…, plus 6 more sections
  • Calls uv and pip; reaches github.com
  • Validating designed binders with ESMFold2-Fast

What it does

Esmfold2 is an agent skill from JimLiu/science-skills. Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. 2026, github.com/Biohub/esm). Single-sequence and MSA modes; protein, DNA, RNA, ligand (CCD/SMILES), modified residues. FoldBench Ab-Ag 50-55%, PPI 70-77% DockQ-pass. Also covers the ESMC-{300M,600M,6B} protein language models from the same release: masked-LM logits, hidden states, mutation scoring, contact prediction, and the SAE interpretability head. MIT-licensed weights on HuggingFace org biohub. Use this skill when: (1) Predicting complex…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/design-hook.md` and `references/esmc.md`).

It sits in AI & LLM Engineering, covering Protein structure and design, Embeddings and AI interpretability. It works with GitHub, Hugging Face and CUDA. The licence is Apache-2.0.

When your agent uses it

  • Predicting complex structures with single-sequence input
  • Validating designed binders with ESMFold2-Fast
  • Running ESMFold2 with MSA input
  • Getting ESMC embeddings

Example prompts

  • “/esmfold2”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit fb309c3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Esmfold2 loads about 2.5k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 200 tokens; SKILL.md has 603 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~200
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JimLiu/science-skills at commit fb309c3, republished under its Apache-2.0 licence (© JimLiu). 603 words, ~2,541 tokens.

Download SKILL.mdSave it as .claude/skills/esmfold2/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
esmfold2
description
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. 2026, github.com/Biohub/esm). Single-sequence and MSA modes; protein, DNA, RNA, ligand (CCD/SMILES), modified residues. FoldBench Ab-Ag 50-55%, PPI 70-77% DockQ-pass. Also covers the ESMC-{300M,600M,6B} protein language models from the same release: masked-LM logits, hidden states, mutation scoring, contact prediction, and the SAE interpretability head. MIT-licensed weights on HuggingFace org `biohub`. Use this skill when: (1) Predicting complex structures with single-sequence input, (2) Validating designed binders with ESMFold2-Fast, (3) Running ESMFold2 with MSA input, (4) Getting ESMC embeddings or per-residue mutation scores, (5) Choosing kernel backend and sampling-step settings for paper-faithful throughput.
license
Apache-2.0
category
biomodels
requirements
gpu
metadata.display-name
ESMFold2

ESMFold2 (Biohub)

All-atom diffusion co-folding from the Biohub ESM release (2026). ESMFold2 = 48 pair layers with MSA support; ESMFold2-Fast = 24 layers, single-sequence only, ~1.7x faster.

License: MIT (code github.com/Biohub/esm + weights HF biohub/*). Paper: "Language Modeling Materializes a World Model of Protein Biology" (2026).

Install

CUDA 12.x GPU (H100/A100-class); Python 3.12 only. Fresh venv; needs egress to HF Hub, GitHub, PyPI:

bash
pip install --no-cache-dir uv
uv venv --python 3.12 /work/venv && source /work/venv/bin/activate
uv pip install \
  "torch>=2.5,<2.8" einops "biotite>=1.0" rdkit msgpack-numpy biopython \
  scikit-learn brotli attrs pandas cloudpathlib httpx tenacity zstd pydssp \
  pygtrie accelerate huggingface_hub safetensors "numpy<3" networkx \
  sentencepiece tokenizers regex packaging filelock pyyaml typing_extensions \
  "transformers @ git+https://github.com/Biohub/transformers.git@3a8956fb4d4ea16b0ec8e71deef2c2909b6a5cbf"
uv pip install --no-deps "esm @ git+https://github.com/Biohub/esm.git@f652b471"
# OPTIONAL — only affects ESMC attention; trunk speedup comes from set_kernel_backend("fused")
uv pip install ninja packaging wheel setuptools
MAX_JOBS=8 uv pip install --no-deps --no-build-isolation "flash-attn<3"
# Do NOT install transformer-engine — RuntimeError (not ImportError) on import
# slips ESMC's guard and kills ESMFold2Model import.

The bundled esmfold2_gpu Modal env (remote-compute-modal skill) is the canonical, version-pinned recipe.

Gotchas:

  • Default kernel backend is None (reference PyTorch, ~12x slower than paper). Call model.set_kernel_backend('fused') after from_pretrained(). See section below.
  • Match torch CUDA build to your driver; the pin <2.8 targets CUDA 12.2.
  • Weights via Xet bridge ~300 MB/s: ESMFold2 1.36 GB, ESMFold2-Fast 0.76 GB. Set HF_HOME=/work/hf_cache.

Usage — local model

python
from esm.models.esmfold2 import (
    ESMFold2InputBuilder, StructurePredictionInput,
    ProteinInput, DNAInput, RNAInput, LigandInput, Modification,
)
from transformers.models.esmfold2.modeling_esmfold2 import ESMFold2Model

model = ESMFold2Model.from_pretrained("biohub/ESMFold2").cuda().eval()
# or "biohub/ESMFold2-Fast" (24 layers, no MSA, ~1.7x faster)
# or "biohub/ESMFold2-Experimental{,-Fast}{,-Cutoff2025}" (4 design-critic models)

spi = StructurePredictionInput(sequences=[
    ProteinInput(id="A", sequence=target_seq),
    ProteinInput(id="B", sequence=binder_seq),
    # DNAInput(id="C", sequence="ACGT", modifications=[Modification(position=5, ccd="C36")]),
    # RNAInput(id="D", sequence="ACGU"),
    # LigandInput(id="L", ccd=["SAH"]),  # or smiles="..."
])
# Homodimer: ProteinInput(id=["A","B"], sequence=seq)

results = ESMFold2InputBuilder().fold(
    model, spi,
    num_loops=10,             # paper FoldBench eval: 10; 20-loop variant: 20
    num_sampling_steps=68,    # paper eval: 68 (truncated EDM)
    num_diffusion_samples=5,  # paper eval: 5/seed
    seed=0,
)
# fold() returns list[Prediction], one per diffusion sample. Each carries
# .plddt [L], .ptm, .iptm, .pae [L,L], .pair_chains_iptm, .complex.to_mmcif().
# Rank by ipTM for complexes / mean pLDDT for monomers:
best = max(results, key=lambda r: float(r.iptm if r.iptm is not None
                                        else r.plddt.mean()))
open("pred.cif", "w").write(best.complex.to_mmcif())

Paper-faithful FoldBench settings: 10 loops, 68 sampling steps, 25 seeds x 5 diffusion samples; rank by ipTM (complexes) or pLDDT (monomers); MSA mode adds msa_depth=1024 with 10% column masking and ESMC dropout 0.3.

Model variants on HF biohub/

reposizepair layersMSAuse
ESMFold20.94 GB + ccd.pkl 0.42 GB48yesfull eval
ESMFold2-Fast0.76 GB24nofast single-seq
ESMFold2-Experimental{,-Fast}0.90 / 0.72 GB48 / 24—design search (Alg 11)
ESMFold2-Experimental{,-Fast}-Cutoff20250.90 / 0.72 GB——design search + critic
ESMFold2-Experimental-Fast-base{300M,600M,6B}-step{250k..1500k}———15 critic ensemble

Throughput: set_kernel_backend("fused") is REQUIRED

Default is the slow path. ESMFold2Model.from_pretrained(...) loads with _kernel_backend=None (reference PyTorch) and chunk_size=64. You MUST call:

python
model = ESMFold2Model.from_pretrained("biohub/ESMFold2").cuda().eval()
model.set_kernel_backend("fused")   # vendored Triton TriMul/LN+SwiGLU/pair-bias kernels
model.set_chunk_size(None)           # optimal & OOM-safe L<=1024; use 256 above

"fused" gives ~1.5–6× trunk speedup over the reference backend, growing with L; end-to-end fold() is diffusion-bound at short L so fused breaks even around L≈300–400. Fused vs reference outputs are numerically consistent (pLDDT within noise). "fused" (Triton, bundled with the GPU torch wheel) is inference-only — auto-disables under backprop. Above ~L=1400 (chunk_size=128) it hits illegal memory access — fall back to set_kernel_backend(None) + set_chunk_size(64); validated through L=1024.

Do NOT use set_kernel_backend("cuequivariance"): the cuequivariance-torch==0.10.0 wheel lacks the compiled ops and silently falls back to the reference path. apply_torch_compile() is an alternative (NOT additive — call set_kernel_backend(None) first).

ESMFold2-Experimental* — design hook

Experimental variants expose res_type_soft for gradient-guided design — see references/design-hook.md. Do NOT use the fused backend with them (fp32/bf16 dtype crash; the reference path is correct).

Show full SKILL.md (260 more words)Show less

Gotcha: cusolver SVD poison + structseq constructor

The Kabsch alignment in modeling_esmfold2_common.py calls torch.linalg.svd(H32, driver="gesvd") on batched 3x3 matrices. NaN/Inf inputs (degenerate diffusion samples) corrupt the cusolver workspace — all subsequent CUDA calls fail with "illegal memory access". Monkeypatch: redirect small batched SVDs to CPU:

python
_orig_svd = torch.linalg.svd
def _safe_svd(A, full_matrices=True, driver=None):
    if A.is_cuda and A.shape[-1] <= 4 and A.shape[-2] <= 4:
        Acpu = A.detach().float().cpu()
        if not torch.isfinite(Acpu).all():
            Acpu = torch.nan_to_num(Acpu, nan=0.0, posinf=1e6, neginf=-1e6)
        out = _orig_svd(Acpu, full_matrices=full_matrices)
        # torch.return_types.linalg_svd is a C structseq -> ctor takes ONE tuple.
        return type(out)(tuple(t.to(A.device, A.dtype) for t in out))
    return _orig_svd(A, full_matrices=full_matrices, driver=driver)
torch.linalg.svd = _safe_svd

Note type(out)(tuple(...)), not type(out)(*(...)) — torch.return_types.* are C structseqs whose constructor takes a single tuple argument.

With-MSA mode

ESMFold2 supports per-chain MSA input via ProteinInput(id, sequence, msa=MSA). The MSA object lives at esm.utils.msa.msa.MSA:

python
from esm.utils.msa.msa import MSA
# ProteinInput, StructurePredictionInput as imported above

msa_A = MSA.from_a3m("/path/chain_A.a3m", max_sequences=2048)
msa_B = MSA.from_a3m("/path/chain_B.a3m", max_sequences=2048)
inp = StructurePredictionInput(sequences=[
    ProteinInput(id="A", sequence=seq_A, msa=msa_A),
    ProteinInput(id="B", sequence=seq_B, msa=msa_B),
])

Gotchas:

  • MSA.from_a3m(remove_insertions=True) asserts equal row lengths after insertion removal. ColabFold a3m files often carry trailing null bytes and off-by-one rows vs the query — tr -d '\000' and force row 0 to the exact query sequence (or MSA.from_sequences on manually cleaned, query-length rows).
  • ESMFold2-Fast does NOT support MSA (single-seq only).
  • With-MSA mode improves AbAg interface pass-rate per the paper's evaluation.

Paper-matched inference configuration

The paper's FoldBench protocol (section A.2.11):

ParameterPaper defaultPaper "20lp"Notes
num_loops (folding-trunk recycles)1020+2pp on AbAg
num_sampling_steps (diffusion)6868EDM-tuned; do NOT use 200
seeds x diffusion samples25 x 525 x 5Fig S6/S7 oracle = best-of-125

Training data cutoff

ESMFold2 and ESMFold2-Fast both use a Sept 2021 PDB training cutoff (HF biohub/ESMFold2 README).

ESMC language model

ESMC is the Biohub successor to ESM-2; three sizes: 300M (30L), 600M (36L), 6B (80L, d=2560). HF path: AutoModelForMaskedLM.from_pretrained("biohub/ESMC-6B").

Mask token is <mask> (id 32) — use tok.mask_token. The native-SDK _ convention does NOT apply to the HF tokenizer: _ is not in the vocab and encodes to <unk>, silently corrupting mutation scores.

Full API, mutation scoring, SAE features, contact prediction: see references/esmc.md.

© JimLiu, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/esmfold2 of JimLiu/science-skills.

  • SKILL.md
  • references/design-hook.md
  • references/esmc.md

Open the folder on GitHubat commit fb309c3

Used in 4 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in JimLiu/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Esmfold2 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Esmfold2 compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Esmfold2 this skillJimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Kermt EmbedNVIDIA/skills3.5k1 repos~1.9kAutomated safety check: PassApache-2.0
Kermt Continue PretrainNVIDIA/skills3.5k1 repos~4.1kAutomated safety check: PassApache-2.0
Mteb LeaderboardlazyFrogLOL/Harness_Engineering128—~2.1kAutomated safety check: PassNone
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0

Similar skills

  • Kermt Embed

    NVIDIA/skills

    Official

    Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint.

    3.5k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Continue KERMT pretraining on a custom SMILES corpus with a groverbase, cmim, or hybrid checkpoint.

    3.5k GitHub starsUsed in 1 repo~4.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Mteb Leaderboard

    lazyFrogLOL/Harness_Engineering

    Guidance for querying ML model leaderboards and benchmarks (MTEB, HuggingFace, embedding benchmarks).

    128 GitHub stars~2.1k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • ML Dataset Discovery

    OpenLAIR/dr-claw

    Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.

    1.2k GitHub stars~741 tokensUpdated 19 days ago
    AI & LLM EngineeringAuto-check passed

More from JimLiu/science-skills

All 27 skills in this repo
  • Compute Env Setup

    JimLiu/science-skills

    Set up a compute environment on a remote provider so Claude Science jobs can run there.

    227 GitHub starsUsed in 2 repos~4.4k tokens
    Auto-check passed
  • Borzoi

    JimLiu/science-skills

    Predict genome-wide functional tracks (RNA-seq, CAGE, DNase, ChIP) from DNA sequence with Borzoi.

    227 GitHub starsUsed in 4 repos~973 tokens
    Auto-check passed
  • Evo2

    JimLiu/science-skills

    Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check passed
  • Fair Esm2

    JimLiu/science-skills

    Embed proteins with Meta AI's ESM-2 (fair-esm package). An agent skill from JimLiu/science-skills.

    227 GitHub starsUsed in 4 repos~1.2k tokens
    Auto-check passed
  • Openfold3

    JimLiu/science-skills

    Structure prediction using OpenFold3, an open-weights PyTorch reproduction of AlphaFold3 from the AlQuraishi Lab.

    227 GitHub starsUsed in 4 repos~1.8k tokens
    Auto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check passed

Questions about Esmfold2

What does Esmfold2 do?

Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. Esmfold2 is an agent skill from JimLiu/science-skills. Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

When should I use Esmfold2?

Esmfold2 fits situations like: predicting complex structures with single-sequence input; validating designed binders with ESMFold2-Fast; running ESMFold2 with MSA input; getting ESMC embeddings.

How do I install Esmfold2 in Claude Code?

Run `npx skills add JimLiu/science-skills --skill esmfold2 -a claude-code`. Or copy the skill folder (skills/esmfold2 in JimLiu/science-skills) into .claude/skills/esmfold2 in your project. Claude Code loads it when a task matches its description.

How do I install Esmfold2 in Codex?

Run `npx skills add JimLiu/science-skills --skill esmfold2 -a codex`. Or copy the skill folder (skills/esmfold2 in JimLiu/science-skills) into .agents/skills/esmfold2 in your project. Codex loads it when a task matches its description.

Can I use Esmfold2 in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JimLiu/science-skills --skill esmfold2 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/esmfold2, .gemini/skills/esmfold2, .github/skills/esmfold2 and .opencode/skills/esmfold2 in your project.

What does Esmfold2 need to run?

Going by SKILL.md and its folder, Esmfold2 needs the command-line tools its instructions call (uv and pip). Our summary lists: Python 3.

Does Esmfold2 access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Esmfold2 safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Esmfold2 use?

Esmfold2 is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Esmfold2 use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to Esmfold2?

Skills that share tags, products or a category with Esmfold2: Kermt Embed (NVIDIA/skills, 3.5k stars), Kermt Continue Pretrain (NVIDIA/skills, 3.5k stars), Mteb Leaderboard (lazyFrogLOL/Harness_Engineering, 128 stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Esmfold2?

JimLiu (a GitHub user) maintains it in JimLiu/science-skills, which has 227 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on July 1, 2026.

Source: JimLiu/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.