Chroma Vector Database
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design.
$ npx skills add adaptyvbio/protein-design-skills --skill esm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install adaptyvbio/protein-design-skills esm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/adaptyvbio/protein-design-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/esm .claude/skills/esm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "esm" agent skill from https://github.com/adaptyvbio/protein-design-skills/tree/main/skills/esm into .claude/skills/esm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "esm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/adaptyvbio/protein-design-skills/tree/main/skills/esmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add adaptyvbio/protein-design-skills --skill esm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install adaptyvbio/protein-design-skills esm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adaptyvbio/protein-design-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/esm .agents/skills/esm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "esm" agent skill from https://github.com/adaptyvbio/protein-design-skills/tree/main/skills/esm into .agents/skills/esm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "esm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add adaptyvbio/protein-design-skills --skill esm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install adaptyvbio/protein-design-skills esm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adaptyvbio/protein-design-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/esm .cursor/skills/esm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "esm" agent skill from https://github.com/adaptyvbio/protein-design-skills/tree/main/skills/esm into .cursor/skills/esm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "esm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/adaptyvbio/protein-design-skills.git --path skills/esm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add adaptyvbio/protein-design-skills --skill esm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install adaptyvbio/protein-design-skills esm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adaptyvbio/protein-design-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/esm .gemini/skills/esm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "esm" agent skill from https://github.com/adaptyvbio/protein-design-skills/tree/main/skills/esm into .gemini/skills/esm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "esm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install adaptyvbio/protein-design-skills esmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add adaptyvbio/protein-design-skills --skill esm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/adaptyvbio/protein-design-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/esm .github/skills/esm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "esm" agent skill from https://github.com/adaptyvbio/protein-design-skills/tree/main/skills/esm into .github/skills/esm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "esm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add adaptyvbio/protein-design-skills --skill esm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install adaptyvbio/protein-design-skills esm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adaptyvbio/protein-design-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/esm .opencode/skills/esm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "esm" agent skill from https://github.com/adaptyvbio/protein-design-skills/tree/main/skills/esm into .opencode/skills/esm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "esm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
esmESM protein language models for embeddings, sequence scoring, structure prediction, and binder design.
Esm is an agent skill from adaptyvbio/protein-design-skills. ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design. Use this skill when: (1) Computing pseudo-log-likelihood (PLL) or mutation-effect scores, (2) Getting protein embeddings for clustering or filtering, (3) Predicting complex structures with ESMFold2, (4) Designing binders by inverting ESMFold2, (5) Filtering designs by sequence plausibility. For diffusion-based structure prediction, use boltz or chai. For QC thresholds, use protein-qc. For gradient-based…
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Protein structure and design and Embeddings. It works with Python. The repository describes itself as: Claude Code skills for protein design. The licence is MIT.
Read from SKILL.md and the folder at commit 59dd633. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvpipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.combiohub.aiAlso links to:
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Esm loads about 2k tokens when it runs. Until then it costs about 137 tokens; SKILL.md has 638 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from adaptyvbio/protein-design-skills at commit 59dd633, republished under its MIT licence (© adaptyvbio). 638 words, ~2,024 tokens.
.claude/skills/esm/SKILL.md (or your agent's skills folder).The ESM line is maintained at github.com/Biohub/esm
(Chan Zuckerberg Biohub, MIT license; the older evolutionaryscale/esm URL
redirects here). The current generation ships three artifacts: ESM C (language
model), ESMFold2 (structure prediction), and ESM Atlas (a map of predicted
structures). Weights are on huggingface.co/biohub;
the hosted API is at biohub.ai.
This skill covers ESM C, ESMFold2, and legacy ESM2. ESM3 is not covered because its open weights are non-commercial.
| Task | Model |
|---|---|
| Embeddings, PLL, mutation scoring | ESM C (ESMC-6B), or ESM2 for a lighter run |
| Complex structure prediction | ESMFold2 |
| High-throughput single-sequence folding | ESMFold2 fast mode |
| Binder design | ESMFold2 inversion (see below), or the mosaic / bindcraft skills |
| Variant effect / zero-shot scoring | ESM C or ESM2 |
| Requirement | Minimum | Recommended |
|---|---|---|
| Python | 3.10+ | 3.11 |
| PyTorch | 2.0+ | Latest |
| CUDA | 12.0+ | 12.1+ |
| GPU VRAM | 24GB (ESM2 / small ESMC) | 80GB (ESMC-6B, ESMFold2) |
ESM C is the successor to ESM2. It improves long-range structural understanding as model scale grows and is the default choice for embeddings, pseudo-log-likelihood, and mutation-effect scoring.
from transformers import AutoModelForMaskedLM, AutoTokenizer
import torch
model_id = "biohub/ESMC-6B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForMaskedLM.from_pretrained(
model_id, output_hidden_states=True, torch_dtype=torch.bfloat16
).eval().cuda()
batch = tok(["MKTAYIAKQRQISFVK..."], return_tensors="pt").to("cuda")
with torch.no_grad():
out = model(**batch)
logits = out.logits # for PLL / mutation scoring
embeddings = out.hidden_states[-1] # per-residue representationsInstall the package with pip install esm@git+https://github.com/Biohub/esm.git@main.
from esm.sdk import esmc_client
from esm.sdk.api import ESMProtein, LogitsConfig
model = esmc_client(model="esmc-600m-2024-12", url="https://biohub.ai", token="<API token>")
tensor = model.encode(ESMProtein(sequence="MKTAYIAKQRQISFVK..."))
out = model.logits(tensor, LogitsConfig(sequence=True, return_embeddings=True))ESMC-6B has open weights; esmc-600m is the smaller API model. For mutation
scoring and fine-tuning, see the esmc_mutation_scoring and esmc_finetune
notebooks under cookbook/tutorials.
ESMFold2 is built on ESMC-6B with a diffusion structure head. Unlike the original ESMFold, it predicts complexes (protein, DNA, ligand, and modified residues), takes an optional MSA, and has a single-sequence fast mode for high-throughput screening. It is validated for protein-protein interaction design and leads DockQ pass-rate on Foldbench protein-protein and antibody-antigen complexes.
printf '>protein|A\nMKTAYIAKQRQISFVK...\n' > target.faa
uv run --with modal modal run modal_esmfold2.py --input-faa target.faaThe FASTA header tags protein|, dna|, rna|, and ligand| (SMILES) let you fold
complexes. GPU defaults to A100-40GB (set with MODAL_GPU).
from transformers.models.esmfold2.modeling_esmfold2 import ESMFold2Model
from esm.models.esmfold2 import ProteinInput, StructurePredictionInput, ESMFold2InputBuilder
model = ESMFold2Model.from_pretrained("biohub/ESMFold2").cuda().eval()
spi = StructurePredictionInput(sequences=[ProteinInput(id="A", sequence="BINDER_SEQ")])
result = ESMFold2InputBuilder().fold(model, spi, num_loops=20, num_sampling_steps=100)
# result.plddt, result.ptm, result.iptm, result.complex.to_mmcif()For single-sequence high-throughput folding, the fast variant is the SDK model string
esmfold2-fast-2026-05 (HF repo biohub/ESMFold2-Fast). ESMFold2 is one option for
complex validation alongside boltz and chai; ranking a shortlist across more than
one predictor is more reliable than trusting a single model.
The binder_design cookbook runs gradient optimization through ESMFold2 (a BindCraft-style loop) with an ESMC language-model term for sequence plausibility. The published protocol is validated in the lab to nanomolar affinity across five targets and supports both minibinders and antibody-derived scFvs with framework scaffolds.
biomodals wraps this as modal_esmfold2_binder_design.py:
uv run --with modal modal run modal_esmfold2_binder_design.py \
--target-name pd-l1 --binder-name minibindercd45, ctla4, egfr, pd-l1, pdgfr, or pass --target-sequence.minibinder and antibody frameworks (for example
trastuzumab_framework_vhvl), or pass --binder-sequence with # for designable
positions. Use --is-antibody for scFv designs.boltz or chai and rank with ipsae.Adaptyv's own tests of these models showed ESMFold2-inversion binder design costing about $0.85 per accepted design, averaged across 7 targets.
For a framework that composes ESMFold2 with other predictors in one objective, use the
mosaic skill.
ESM2 still works well for quick embeddings and PLL when ESMC-6B is too large for the available GPU.
import torch, esm
model, alphabet = esm.pretrained.esm2_t33_650M_UR50D()
bc = alphabet.get_batch_converter()
model = model.eval().cuda()
_, _, toks = bc([("seq1", "MKTAYIAKQRQISFVK...")])
with torch.no_grad():
rep = model(toks.cuda(), repr_layers=[33])["representations"][33]| Model | Parameters | Use |
|---|---|---|
| esm2_t12_35M | 35M | Fast screening |
| esm2_t33_650M | 650M | Standard embeddings/PLL |
| esm2_t36_3B | 3B | Highest-quality ESM2 |
PLL (pseudo-log-likelihood) scores how natural a sequence looks to the model. Higher is more natural. Designed sequences often score lower than natural ones, so treat PLL as a soft filter, not a hard cutoff.
| Normalized PLL | Interpretation |
|---|---|
| > 0.2 | Very natural |
| 0.0 to 0.2 | Natural-like |
| -0.5 to 0.0 | Acceptable |
| < -0.5 | May be unnatural |
| Issue | Cause | Fix |
|---|---|---|
| CUDA out of memory | ESMC-6B / ESMFold2 too large | Use ESMC-600m API, ESM2, or an 80GB GPU |
| Wrong layer for embeddings | Layer index mismatch | Use the last hidden state (layer 33 for ESM2-650M) |
| Invalid amino acid | Non-standard residue | Check for non-canonical characters |
| Slow ESMFold2 on many designs | Full MSA mode | Use esmfold2-fast-2026-05 single-sequence mode |
Next: Validate structures with boltz or chai, rank with ipsae, then filter
with protein-qc.
© adaptyvbio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/esm of adaptyvbio/protein-design-skills.
Open the folder on GitHubat commit 59dd633
Esm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Esm this skilladaptyvbio/protein-design-skills | 163 | — | ~2k | Automated safety check: Pass | MIT | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~2.3k | Automated safety check: Pass | MIT | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Esmfold2JimLiu/science-skills | 227 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~1.6k | Automated safety check: Pass | MIT | |
| RAG Company Knowledge AssistantHermes-brasil/hermes-brasil | 152 | — | ~1.1k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
Orchestra-Research/AI-Research-SKILLs
Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.
Hermes-brasil/hermes-brasil
Portuguese guide to building a retrieval-augmented generation assistant over a company's documents, with embeddings, section-based chunking, retrieval and a client workflow.
davila7/claude-code-templates
Comprehensive toolkit for protein language models including ESM3 (generative multimodal protein design across sequence, structure, and function) and ESM C (efficient protein embeddings and…
adaptyvbio/protein-design-skills
Validate protein designs using AlphaFold2 structure prediction.
adaptyvbio/protein-design-skills
End-to-end binder design using BindCraft hallucination. An agent skill from adaptyvbio/protein-design-skills.
adaptyvbio/protein-design-skills
All-atom protein design using BoltzGen diffusion model. An agent skill from adaptyvbio/protein-design-skills.
adaptyvbio/protein-design-skills
Structure prediction using Chai-1, a foundation model for molecular structure.
adaptyvbio/protein-design-skills
End-to-end guidance for protein design pipelines. An agent skill from adaptyvbio/protein-design-skills.
adaptyvbio/protein-design-skills
Quality control metrics and filtering thresholds for protein design.
Works with
Categories
ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design. Esm is an agent skill from adaptyvbio/protein-design-skills. ESM protein language models for embeddings, sequence scoring, structure prediction, and binder design.
Esm fits situations like: computing pseudo-log-likelihood (PLL); mutation-effect scores; getting protein embeddings for clustering; predicting complex structures with ESMFold2.
Run `npx skills add adaptyvbio/protein-design-skills --skill esm -a claude-code`. Or copy the skill folder (skills/esm in adaptyvbio/protein-design-skills) into .claude/skills/esm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add adaptyvbio/protein-design-skills --skill esm -a codex`. Or copy the skill folder (skills/esm in adaptyvbio/protein-design-skills) into .agents/skills/esm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adaptyvbio/protein-design-skills --skill esm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/esm, .gemini/skills/esm, .github/skills/esm and .opencode/skills/esm in your project.
Going by SKILL.md and its folder, Esm needs the command-line tools its instructions call (uv and pip). Our summary lists: Python 3.
SKILL.md names 3 domains. In commands or code: github.com and biohub.ai; the agent is likely to contact these when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Esm is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Esm: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Esmfold2 (JimLiu/science-skills, 227 stars) and Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
adaptyvbio (a GitHub organization) maintains it in adaptyvbio/protein-design-skills, which has 163 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on June 11, 2026.
Source: adaptyvbio/protein-design-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.