Agent skill

Protein Design MCP

by PKU-YuanGroup in PKU-YuanGroup/OpenAI4S

Compose auditable protein-design operations through the configured OpenAI4S protein-design MCP connector: target-conditioned RFdiffusion backbone generation, constrained ProteinMPNN sequence design…

MITAuto-check passedResearch & Science

Install Protein Design MCP

skills CLI
$ npx skills add PKU-YuanGroup/OpenAI4S --skill protein-design-mcp -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PKU-YuanGroup/OpenAI4S protein-design-mcp --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/protein-design-mcp .claude/skills/protein-design-mcp && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
protein-design-mcp
GitHub stars
622
Token cost
~2.3k tokens
SKILL.md length
1,136 words
Files
3
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Compose auditable protein-design operations through the configured OpenAI4S protein-design MCP connector: target-conditioned RFdiffusion backbone generation, constrained ProteinMPNN sequence design…

  • Redesigning proteins
  • SKILL.md covers Discover the connector, Select compute before…, Acquire checkpoints only when… and Choose only the operations the…, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Creating target-binding proteins

What it does

Protein Design MCP is an agent skill from PKU-YuanGroup/OpenAI4S. Compose auditable protein-design operations through the configured OpenAI4S protein-design MCP connector: target-conditioned RFdiffusion backbone generation, constrained ProteinMPNN sequence design, monomer or complex structure prediction, Rosetta scoring and relaxation, ESM-2 sequence naturalness scoring, and OpenMM minimization. Use when designing or redesigning proteins, creating target-binding proteins, preserving sequence motifs, validating candidate structures or complexes, refining structures, or ranking…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `README.md` and `README_zh.md`).

It sits in Research & Science, covering Protein structure and design. It works with Model Context Protocol. The repository describes itself as: Open-source AI agent for scientific research. Analyze data in Python/R with Claude, GPT, Gemini, and more. The licence is MIT.

When your agent uses it

  • Redesigning proteins
  • Creating target-binding proteins
  • Preserving sequence motifs
  • Validating candidate structures

Example prompts

  • “/protein-design-mcp”

Requirements

  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 4a72e87. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Protein Design MCP loads about 2.3k tokens when it runs. Until then it costs about 149 tokens; SKILL.md has 1,136 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~149
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PKU-YuanGroup/OpenAI4S at commit 4a72e87, republished under its MIT licence (© PKU-YuanGroup). 1,136 words, ~2,291 tokens.

Download SKILL.mdSave it as .claude/skills/protein-design-mcp/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
protein-design-mcp
description
Compose auditable protein-design operations through the configured OpenAI4S protein-design MCP connector: target-conditioned RFdiffusion backbone generation, constrained ProteinMPNN sequence design, monomer or complex structure prediction, Rosetta scoring and relaxation, ESM-2 sequence naturalness scoring, and OpenMM minimization. Use when designing or redesigning proteins, creating target-binding proteins, preserving sequence motifs, validating candidate structures or complexes, refining structures, or ranking protein-design candidates with reproducible model evidence.
origin
openai4s
category
biomodels

Compose protein-design operations over MCP

Select atomic tools according to the scientific objective. Do not force every task through one pipeline, and do not treat a model score as experimental proof.

Discover the connector

Find the enabled connector with host.mcp.list(), then inspect it with host.mcp.tools(server). The connector can expose:

  • generate_backbone
  • design_sequence
  • predict_structure
  • predict_complex
  • rosetta_score
  • rosetta_relax
  • rosetta_interface_score
  • score_stability
  • energy_minimize

Discovery shows that the server started. Before running a model, also verify its execution route, external environment, pinned revision, checkpoint, compute resources and required network-isolation mechanism.

Select compute before provisioning

Call host.accelerator_status() before any GPU-only operation. It probes the daemon's local GPUs first and then lists configured SSH GPU routes. These are different from a BYOC provider catalogue and from model-backend readiness.

When both a local route and one or more SSH routes are candidates, ask the user to choose local or ssh:<alias> before downloading, installing or launching anything. Do not silently prefer either route. When only one route exists, state the selected execution_target and record it in the bring-up evidence. An empty SSH registry is not evidence that the local machine has no GPU, and a missing Docker executable does not make a natively usable local GPU disappear.

Acquire checkpoints only when needed

Do not ask for, locate or download checkpoints during generic connector discovery. Only enter this flow after selecting an operation whose tool schema requires checkpoint_path. Before that call, inspect whether the user already supplied a path. If not, stop provisioning and ask whether they have an existing local checkpoint; request its path plus any known digest. Do not search for or download weights while that question is unanswered, and do not make them download a second copy merely because it is outside a conventional directory.

If the user says there is no local checkpoint, use the normal approved network and tool bring-up controls to download it. The framework does not maintain a closed list of allowed scientific sources: resolve the source selected for this run to an immutable version, prefer an upstream-published checksum, compute the downloaded file's SHA-256 independently, and retain source URL, size and digest. An observed digest with no independently trusted reference proves transfer identity, not that the file is the intended model; report that distinction.

After acquiring code, environment or weights, run a small real inference canary on the selected execution target. The canary must exercise the same adapter, backend revision and checkpoint that the formal call will use, produce the output types the adapter promises, and pass the adapter's parser and digest checks. Set run_mode="canary" on this attempt. Record failed bring-up attempts. Only after its result contains a verified bringup_admission may a new attempt use run_mode="formal". A formal attempt made too early ends in a durable failure, so retry it with a new attempt_id after admission rather than reusing the failed ID. Only after a canary reaches a verified terminal success may the backend be used in the user's formal work. Then retry the original scientific operation instead of ending the task with “backend not configured.” If permission, licensing, source integrity, disk capacity or the canary genuinely blocks bring-up, report that specific blocker and do not fabricate results.

Choose only the operations the task needs

Typical compositions include:

  • target-conditioned binder design: generate_backbone → design_sequence → predict_structure and predict_complex → interface scoring;
  • backbone sequence redesign: design_sequence → predict_structure → optional physical scoring;
  • fixed-motif sequence design: design_sequence with explicit per-chain fixed positions → structure validation;
  • structure refinement: rosetta_relax or energy_minimize → score the input and refined structures with the same method;
  • candidate ranking: combine sequence, monomer, complex and physical evidence while retaining the individual scores and provenance.

These are examples, not mandatory pipelines. Start from the user's design objective and constraints, then choose the smallest informative set of calls.

Record reproducible attempts

Give each model execution a stable attempt_id and explicit seed. Pin the backend revision and checkpoint SHA-256, use a dedicated output directory, and retain the resolved configuration, command, residue maps, raw outputs and terminal record. Reusing the same attempt and configuration is idempotent; changing the configuration under an existing attempt ID is a conflict.

One call to generate_backbone produces one design. Run multiple attempts with distinct IDs and seeds when sampling a population. A failed attempt remains part of the provenance rather than being silently discarded.

Show full SKILL.md (431 more words)Show less

Generate target-conditioned binder backbones

The current generate_backbone contract requires a target PDB, explicit target chain or chains, validated target hotspot residues and a binder length. It verifies the local RFdiffusion checkpoint and returns both PDB and .trb mapping outputs.

Do not describe this operation as epitope-free or purely function-guided de novo design: the hotspot list supplies structural contact-region information. If the task does not provide a contact region, epitope selection is a separate scientific step and its assumptions must be reported.

The current schema also does not express unconditional monomer generation, motif-scaffolding contigs, symmetric oligomer generation or membrane-specific constraints. Use another suitable atomic connector or extend this schema before claiming those backbone-generation capabilities.

Design sequences without losing constraints

Call design_sequence with every input chain represented in fixed_positions. Values are "all" or chain-local, 1-based sequence positions. Include only mutable chains in design_chains; mark fixed target or context chains as "all".

The connector uses ProteinMPNN's --pdb_path_chains and --fixed_positions_jsonl inputs and independently rejects output when a fixed chain or motif changes, a chain length changes, or the residue map does not close. Inspect this validation before using a sequence downstream.

Predict structures and complexes without self-conditioning

Use predict_structure for monomer evidence and predict_complex for blind sequence-only complex evidence. Formal prediction calls require:

  • msa_mode="single_sequence";
  • a local checkpoint bundle with verified digests;
  • fixed model type, recycles, model count and seed;
  • templates and initial guesses disabled;
  • an operator-configured OS-level network-isolation prefix.

Preserve raw confidence values and PAE. Treat pLDDT, pTM, ipTM and interface PAE as model confidence, not as proof of folding, binding, affinity or function. Do not feed a generated complex back as a template or initial guess for its own validation.

Add physical and sequence evidence carefully

Use rosetta_interface_score for dG_separated, dSASA, packstat, interface residue count and interface_delta_unsat_hbonds. The final field is a change in unsatisfied hydrogen bonds, not a count of formed interface hydrogen bonds.

Treat score_stability as ESM-2 masked pseudo-log-likelihood or sequence naturalness, not thermodynamic stability. Treat energy_minimize as local force-field refinement, not evidence that a candidate folds or binds. rosetta_relax is optional; when using it, compare consistently scored input and relaxed structures and retain both.

Rank and report without collapsing evidence

Apply hard task constraints before ranking. Keep evidence types separate, report failed attempts, and preserve structural diversity instead of selecting only near-duplicates with the best value from one model.

When these tools are used inside a benchmark, keep any withheld references or labels inaccessible during candidate generation and ranking. This is an optional benchmark-integrity rule, not a restriction on ordinary protein design use and not a responsibility assigned to this connector.

© PKU-YuanGroup, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/protein-design-mcp of PKU-YuanGroup/OpenAI4S.

  • SKILL.md
  • README.md
  • README_zh.md

Open the folder on GitHubat commit 4a72e87

Compare with similar skills

Protein Design MCP next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Protein Design MCP compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Protein Design MCP this skillPKU-YuanGroup/OpenAI4S622—~2.3kAutomated safety check: PassMIT
Tooluniverseynulihao/AgentSkillOS6182 repos~2.5kAutomated safety check: PassNone
TamarindK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: PassMIT
Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills3.2k2 repos~1.2kAutomated safety check: PassApache-2.0
Deep Researchjordan-gibbs/hyperresearch3.8k—~1.2kAutomated safety check: PassMIT
Annotate Paper54yyyu/zotero-mcp5.3k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Tooluniverse

    ynulihao/AgentSkillOS

    A skill your agent uses when working with scientific research tools and workflows across bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery.

    618 GitHub starsUsed in 2 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Tamarind

    K-Dense-AI/scientific-agent-skills

    Provides access to a collection of open-source molecular design and structural biology tools on the Tamarind Bio platform, via its REST API or MCP server — no local GPUs required.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Research & ScienceAuto-check passed
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Deep Research

    jordan-gibbs/hyperresearch

    Deep research with hyperresearch, for Claude Code and OpenAI Codex.

    3.8k GitHub stars~1.2k tokensUpdated today
    Research & ScienceAuto-check passed
  • Annotate Paper

    54yyyu/zotero-mcp

    Read the open paper and write study annotations into its PDF with zotero-cli - a context box on the title, a four-part summary on the abstract, role-coded abstract highlights, one box per figure…

    5.3k GitHub stars~1.5k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Paper Search

    openags/paper-search-mcp

    Search, download, and read academic papers from 20+ sources (arXiv, PubMed, Semantic Scholar, CrossRef, etc).

    2.8k GitHub stars~1.2k tokensUpdated 9 days ago
    Research & ScienceAuto-check: notes

More from PKU-YuanGroup/OpenAI4S

All 17 skills in this repo
  • Single Cell Rna Analysis

    PKU-YuanGroup/OpenAI4S

    Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…

    622 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Bioprobench

    PKU-YuanGroup/OpenAI4S

    Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

    622 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Reaction Atom Mapping

    PKU-YuanGroup/OpenAI4S

    Map atoms and changed bonds for a complete reaction with RXNMapper.

    622 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Reaction Forward Prediction

    PKU-YuanGroup/OpenAI4S

    Predict ranked products from reactants and reagents with ReactionT5v2-forward; use for outcome prediction or round-trip recovery.

    622 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Reaction Yield Estimation

    PKU-YuanGroup/OpenAI4S

    Estimate yield for a fully specified reactant/reagent/product record with ReactionT5v2-yield.

    622 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Rfdiffusion

    PKU-YuanGroup/OpenAI4S

    Generate de novo protein backbones with RFdiffusion for protein-target binders, hotspot-conditioned interfaces, motif scaffolding, partial diffusion, or symmetric assemblies.

    622 GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed

Questions about Protein Design MCP

What does Protein Design MCP do?

Compose auditable protein-design operations through the configured OpenAI4S protein-design MCP connector: target-conditioned RFdiffusion backbone generation, constrained ProteinMPNN sequence design…. Protein Design MCP is an agent skill from PKU-YuanGroup/OpenAI4S. Compose auditable protein-design operations through the configured OpenAI4S protein-design MCP connector: target-conditioned RFdiffusion backbone generation, constrained ProteinMPNN sequence design, monomer or complex structure prediction, Rosetta scoring and relaxation, ESM-2 sequence naturalness scoring, and OpenMM minimization.

When should I use Protein Design MCP?

Protein Design MCP fits situations like: redesigning proteins; creating target-binding proteins; preserving sequence motifs; validating candidate structures.

How do I install Protein Design MCP in Claude Code?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill protein-design-mcp -a claude-code`. Or copy the skill folder (skills/protein-design-mcp in PKU-YuanGroup/OpenAI4S) into .claude/skills/protein-design-mcp in your project. Claude Code loads it when a task matches its description.

How do I install Protein Design MCP in Codex?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill protein-design-mcp -a codex`. Or copy the skill folder (skills/protein-design-mcp in PKU-YuanGroup/OpenAI4S) into .agents/skills/protein-design-mcp in your project. Codex loads it when a task matches its description.

Can I use Protein Design MCP in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PKU-YuanGroup/OpenAI4S --skill protein-design-mcp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/protein-design-mcp, .gemini/skills/protein-design-mcp, .github/skills/protein-design-mcp and .opencode/skills/protein-design-mcp in your project.

What does Protein Design MCP need to run?

SKILL.md names no scripts, command-line tools or credentials: Protein Design MCP is instructions for the agent only. Our summary lists: Docker.

Does Protein Design MCP access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Protein Design MCP safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Protein Design MCP use?

Protein Design MCP is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Protein Design MCP use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Protein Design MCP?

Skills that share tags, products or a category with Protein Design MCP: Tooluniverse (ynulihao/AgentSkillOS, 618 stars), Tamarind (K-Dense-AI/scientific-agent-skills, 48k stars), Alphafold Database Fetch And Analyze (google-deepmind/science-skills, 3.2k stars) and Deep Research (jordan-gibbs/hyperresearch, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Protein Design MCP?

PKU-YuanGroup (a GitHub organization) maintains it in PKU-YuanGroup/OpenAI4S, which has 622 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: PKU-YuanGroup/OpenAI4S on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.