Agent skill

Proteinmpnn

by adaptyvbio in adaptyvbio/protein-design-skills

Design protein sequences using ProteinMPNN inverse folding. An agent skill from adaptyvbio/protein-design-skills.

MITAuto-check passedResearch & Science

Install Proteinmpnn

skills CLI
$ npx skills add adaptyvbio/protein-design-skills --skill proteinmpnn -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install adaptyvbio/protein-design-skills proteinmpnn --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/adaptyvbio/protein-design-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteinmpnn .claude/skills/proteinmpnn && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
proteinmpnn
GitHub stars
164
Used in
3 other repos
Token cost
~1.8k tokens
SKILL.md length
346 words
Files
2 (incl. references)
Skills in repo
24
Repo updated
First seen
Licence
MIT

At a glance

Design protein sequences using ProteinMPNN inverse folding. An agent skill from adaptyvbio/protein-design-skills.

  • Designing sequences for RFdiffusion backbones
  • SKILL.md covers Prerequisites, How to run, Config Schema and Common mistakes, plus 8 more sections
  • Calls python, git and modal; reaches github.com
  • Redesigning existing protein sequences

What it does

Proteinmpnn is an agent skill from adaptyvbio/protein-design-skills. Design protein sequences using ProteinMPNN inverse folding. Use this skill when: (1) Designing sequences for RFdiffusion backbones, (2) Redesigning existing protein sequences, (3) Fixing specific residues while designing others, (4) Optimizing sequences for expression or stability, (5) Multi-state or negative design. For backbone generation, use rfdiffusion or bindcraft. For ligand-aware design, use ligandmpnn. For solubility optimization, use solublempnn.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/temperature-guide.md`).

It sits in Research & Science, covering Protein structure and design. The repository describes itself as: Claude Code skills for protein design. The licence is MIT.

When your agent uses it

  • Designing sequences for RFdiffusion backbones
  • Redesigning existing protein sequences
  • Fixing specific residues while designing others
  • Optimizing sequences for expression

Example prompts

  • “/proteinmpnn”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 59dd633. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git
    • modal

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Proteinmpnn loads about 1.8k tokens when it runs, and up to ~2.6k if it reads all its reference files. Until then it costs about 118 tokens; SKILL.md has 346 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from adaptyvbio/protein-design-skills at commit 59dd633, republished under its MIT licence (© adaptyvbio). 346 words, ~1,832 tokens.

Download SKILL.mdSave it as .claude/skills/proteinmpnn/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
proteinmpnn
description
Design protein sequences using ProteinMPNN inverse folding. Use this skill when: (1) Designing sequences for RFdiffusion backbones, (2) Redesigning existing protein sequences, (3) Fixing specific residues while designing others, (4) Optimizing sequences for expression or stability, (5) Multi-state or negative design. For backbone generation, use rfdiffusion or bindcraft. For ligand-aware design, use ligandmpnn. For solubility optimization, use solublempnn.
license
MIT
category
design-tools
tags
sequence-design, inverse-folding
biomodals_script
modal_ligandmpnn.py

ProteinMPNN Sequence Design

Prerequisites

RequirementMinimumRecommended
Python3.8+3.10
CUDA11.0+11.7+
GPU VRAM8GB16GB (T4)
RAM8GB16GB

How to run

First time? See Getting started to set up Modal and biomodals.

bash
git clone https://github.com/dauparas/ProteinMPNN.git
cd ProteinMPNN

python protein_mpnn_run.py \
  --pdb_path backbone.pdb \
  --out_folder output/ \
  --num_seq_per_target 16 \
  --sampling_temp "0.1"

GPU: T4 (16GB) sufficient | Time: ~50-100 sequences/minute

Option 2: Modal (via LigandMPNN wrapper)
bash
cd biomodals
# modal_ligandmpnn.py takes --input-pdb and forwards run.py args via --params-str
modal run modal_ligandmpnn.py \
  --input-pdb backbone.pdb \
  --params-str "--model_type protein_mpnn --number_of_batches 16 --temperature 0.1"

GPU (Modal): A10G default | Timeout: 900s default

Note: LigandMPNN includes ProteinMPNN functionality (select with --model_type protein_mpnn).

Config Schema

Core Parameters
ParameterDefaultRangeDescription
--pdb_pathrequiredpathSingle PDB input
--pdb_path_chainsallA,BChains to design (comma-sep)
--out_folderrequiredpathOutput directory
--num_seq_per_target11-1000Sequences per structure
--sampling_temp"0.1""0.0001-1.0"Temperature (string!)
--seed0intRandom seed
--batch_size11-32Batch size
Temperature Guide
0.1  -> Low diversity, high recovery (production)
0.2  -> Moderate diversity (default)
0.3  -> Higher diversity (exploration)
0.5+ -> Very diverse, lower quality

IMPORTANT: Temperature must be passed as a string, not float.

Common mistakes

Temperature Parameter

✅ Correct:

bash
--sampling_temp "0.1"    # String with quotes

❌ Wrong:

bash
--sampling_temp 0.1      # Float without quotes - may cause errors
--sampling_temp 0.1,0.2  # Multiple temps need proper format
Fixed Positions JSONL

✅ Correct:

json
{"A": [1, 2, 3, 10, 11], "B": [5, 6]}

❌ Wrong:

json
{"A": "1,2,3,10,11"}     # String instead of list
{A: [1, 2, 3]}           # Missing quotes on key
{"A": [1,2,3,]}          # Trailing comma
Chain Selection

✅ Correct:

bash
--pdb_path_chains A,B    # No spaces

❌ Wrong:

bash
--pdb_path_chains A, B   # Space after comma
--pdb_path_chains "A,B"  # Quotes may cause issues
Amino Acid Biases
bash
# Bias toward certain AAs (positive = favor)
--bias_AA_jsonl '{"A": {"A": 1.5, "W": -2.0}}'

# Omit specific AAs globally
--omit_AAs "CM"  # No cysteine or methionine

# Per-position omission
--omit_AA_jsonl '{"A": {"1": "C", "2": "CM"}}'
Multi-Chain Design
bash
# Design chains A and B together
--pdb_path_chains A,B

# Tie chains (same sequence)
--tied_positions_jsonl tied.jsonl

Variants Comparison

VariantUse CaseKey Difference
ProteinMPNNGeneralOriginal model
SolubleMPNNExpressionTrained on soluble proteins
LigandMPNNSmall moleculesLigand-aware context

Output format

output/
├── seqs/
│   └── backbone.fa          # FASTA sequences
└── backbone_pdb/
    └── backbone_0001.pdb    # PDBs with designed sequence
FASTA Header Format
>backbone_0001, score=1.234, global_score=1.234, seq_recovery=0.85
MKTAYIAKQRQISFVKSHFSRQLE...

Common workflows

Binder Sequence Design
bash
python protein_mpnn_run.py \
  --pdb_path binder_backbone.pdb \
  --out_folder output/ \
  --num_seq_per_target 16 \
  --sampling_temp "0.1" \
  --pdb_path_chains B  # Design binder chain only
Interface Redesign
bash
# Fix core, design interface
python protein_mpnn_run.py \
  --pdb_path complex.pdb \
  --fixed_positions_jsonl core_positions.jsonl \
  --num_seq_per_target 32
Multi-State Design
bash
# Design for multiple conformations
python protein_mpnn_run.py \
  --pdb_path_multi state1.pdb,state2.pdb \
  --num_seq_per_target 16

Sample output

Successful run
$ python protein_mpnn_run.py --pdb_path backbone.pdb --out_folder output/ --num_seq_per_target 8
Loading model weights...
Designing sequences for backbone.pdb
Generated 8 sequences in 2.3 seconds

output/seqs/backbone.fa:
>backbone_0001, score=1.234, global_score=1.189, seq_recovery=0.82
MKTAYIAKQRQISFVKSHFSRQLEERGLTKE...
>backbone_0002, score=1.198, global_score=1.156, seq_recovery=0.79
MKTAYIAKQRQISFVKSQFSRQLDERGLTKE...

What good output looks like:

  • Score: 1.0-2.0 (lower = more confident)
  • Seq recovery: 0.3-0.6 for de novo, 0.7-0.9 for redesign
  • Diverse sequences (not all identical) when temp > 0.1

Decision tree

Should I use ProteinMPNN?
│
├─ Have a backbone structure?
│  ├─ Yes → Continue below
│  └─ No → Use RFdiffusion first
│
├─ What's in the binding site?
│  ├─ Nothing / protein only → ProteinMPNN ✓
│  ├─ Small molecule / ligand → Use LigandMPNN
│  └─ Metal / cofactor → Use LigandMPNN
│
├─ Priority?
│  ├─ Solubility/expression → Consider SolubleMPNN
│  ├─ Speed → ProteinMPNN ✓
│  └─ AF2 optimization → Consider ColabDesign
│
└─ Need fixed positions?
   ├─ Yes → Use --fixed_positions_jsonl
   └─ No → ProteinMPNN ✓ (design all)

Typical performance

Campaign SizeTime (T4)Cost (Modal)Notes
100 backbones × 8 seq15-20 min~$2Standard
500 backbones × 8 seq1-1.5h~$8Large campaign
1000 backbones × 16 seq3-4h~$18Comprehensive

Throughput: ~50-100 sequences/minute on T4 GPU.


Verify

bash
grep -c "^>" output/seqs/*.fa  # Should match backbone_count × num_seq_per_target

Troubleshooting

Low sequence diversity: Increase sampling_temp to 0.2-0.3 Poor recovery: Decrease sampling_temp to 0.1 OOM errors: Reduce batch_size Unwanted cysteines: Use --omit_AAs "C"

Error interpretation
ErrorCauseFix
RuntimeError: CUDA out of memoryLong protein or large batchReduce batch_size or use larger GPU
KeyError: 'A'Chain not in PDBCheck chain IDs in your PDB file
JSONDecodeErrorInvalid JSONL formatValidate JSON syntax (see Common Mistakes)
IndexError: list indexEmpty chain or residue listCheck PDB has atoms, not just HEADER

Next: Structure prediction for validation → protein-qc for filtering.

© adaptyvbio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/proteinmpnn of adaptyvbio/protein-design-skills.

  • SKILL.md
  • references/temperature-guide.md

Open the folder on GitHubat commit 59dd633

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in adaptyvbio/protein-design-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Proteinmpnn next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Proteinmpnn compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Proteinmpnn this skilladaptyvbio/protein-design-skills1643 repos~1.8kAutomated safety check: PassMIT
Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills3.2k2 repos~1.2kAutomated safety check: PassApache-2.0
Pymol VisualizationChatMol/ChatMol373—~1.2kAutomated safety check: PassMIT
Complexa Binder DesignNVIDIA-BioNeMo/bionemo-agent-toolkit478—~3.1kAutomated safety check: NotesApache-2.0
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT

Similar skills

  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Pymol Visualization

    ChatMol/ChatMol

    Generate publication-quality molecular visualization images using PyMOL.

    373 GitHub stars~1.2k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Complexa Binder Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided…

    478 GitHub stars~3.1k tokensUpdated today
    Research & ScienceAuto-check: notes
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Protein Binder Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Orchestrate an end-to-end de novo protein binder design campaign against a protein target by composing BioNeMo NIM skills.

    478 GitHub stars~1.4k tokensUpdated today
    Research & ScienceAuto-check: notes

More from adaptyvbio/protein-design-skills

All 24 skills in this repo
  • Alphafold

    adaptyvbio/protein-design-skills

    Validate protein designs using AlphaFold2 structure prediction.

    164 GitHub starsUsed in 3 repos~1.2k tokens
    Auto-check passed
  • Bindcraft

    adaptyvbio/protein-design-skills

    End-to-end binder design using BindCraft hallucination. An agent skill from adaptyvbio/protein-design-skills.

    164 GitHub starsUsed in 3 repos~1.3k tokens
    Auto-check passed
  • Binder Design

    adaptyvbio/protein-design-skills

    Guidance for choosing the right protein binder design tool. An agent skill from adaptyvbio/protein-design-skills.

    164 GitHub starsUsed in 3 repos~1.8k tokens
    Auto-check passed
  • Boltzgen

    adaptyvbio/protein-design-skills

    All-atom protein design using BoltzGen diffusion model. An agent skill from adaptyvbio/protein-design-skills.

    164 GitHub starsUsed in 3 repos~2k tokens
    Auto-check passed
  • Chai

    adaptyvbio/protein-design-skills

    Structure prediction using Chai-1, a foundation model for molecular structure.

    164 GitHub starsUsed in 3 repos~1.5k tokens
    Auto-check passed
  • Protein Design Workflow

    adaptyvbio/protein-design-skills

    End-to-end guidance for protein design pipelines. An agent skill from adaptyvbio/protein-design-skills.

    164 GitHub starsUsed in 3 repos~1.2k tokens
    Auto-check passed

Questions about Proteinmpnn

What does Proteinmpnn do?

Design protein sequences using ProteinMPNN inverse folding. An agent skill from adaptyvbio/protein-design-skills. Proteinmpnn is an agent skill from adaptyvbio/protein-design-skills. Design protein sequences using ProteinMPNN inverse folding.

When should I use Proteinmpnn?

Proteinmpnn fits situations like: designing sequences for RFdiffusion backbones; redesigning existing protein sequences; fixing specific residues while designing others; optimizing sequences for expression.

How do I install Proteinmpnn in Claude Code?

Run `npx skills add adaptyvbio/protein-design-skills --skill proteinmpnn -a claude-code`. Or copy the skill folder (skills/proteinmpnn in adaptyvbio/protein-design-skills) into .claude/skills/proteinmpnn in your project. Claude Code loads it when a task matches its description.

How do I install Proteinmpnn in Codex?

Run `npx skills add adaptyvbio/protein-design-skills --skill proteinmpnn -a codex`. Or copy the skill folder (skills/proteinmpnn in adaptyvbio/protein-design-skills) into .agents/skills/proteinmpnn in your project. Codex loads it when a task matches its description.

Can I use Proteinmpnn in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adaptyvbio/protein-design-skills --skill proteinmpnn -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/proteinmpnn, .gemini/skills/proteinmpnn, .github/skills/proteinmpnn and .opencode/skills/proteinmpnn in your project.

What does Proteinmpnn need to run?

Going by SKILL.md and its folder, Proteinmpnn needs the command-line tools its instructions call (python, git and modal). Our summary lists: Python 3.

Does Proteinmpnn access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Proteinmpnn safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Proteinmpnn use?

Proteinmpnn is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Proteinmpnn use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 752 tokens, read only when the agent opens those files.

What are the alternatives to Proteinmpnn?

Skills that share tags, products or a category with Proteinmpnn: Alphafold Database Fetch And Analyze (google-deepmind/science-skills, 3.2k stars), Pymol Visualization (ChatMol/ChatMol, 373 stars), Complexa Binder Design (NVIDIA-BioNeMo/bionemo-agent-toolkit, 478 stars) and DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Proteinmpnn?

adaptyvbio (a GitHub organization) maintains it in adaptyvbio/protein-design-skills, which has 164 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on June 11, 2026.

Source: adaptyvbio/protein-design-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.