Agent skill

Protein Design Qc

by BioTender-max in BioTender-max/awesome-bio-agent-skills

Protein design quality control, filtering thresholds, and ranking guidance.

MITAuto-check passedResearch & Science

Install Protein Design Qc

skills CLI
$ npx skills add BioTender-max/awesome-bio-agent-skills --skill protein-design-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BioTender-max/awesome-bio-agent-skills protein-design-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BioTender-max/awesome-bio-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bioclaw_hub/protein-design-qc .claude/skills/protein-design-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
protein-design-qc
GitHub stars
200
Token cost
~2.6k tokens
SKILL.md length
530 words
Files
6 (incl. references)
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Protein design quality control, filtering thresholds, and ranking guidance.

  • Evaluating design quality for binding
  • SKILL.md covers Critical Limitation, QC Organization, Quick Reference: All Thresholds and Sequential Filtering Pipeline, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Setting filtering thresholds for pLDDT

What it does

Protein Design Qc is an agent skill from BioTender-max/awesome-bio-agent-skills. Protein design quality control, filtering thresholds, and ranking guidance. Use this skill when: (1) Evaluating design quality for binding, expression, or structure, (2) Setting filtering thresholds for pLDDT, ipTM, PAE, (3) Checking sequence liabilities (cysteines, deamidation, polybasic clusters), (4) Creating multi-stage filtering pipelines, (5) Computing PyRosetta interface metrics (dG, SC, dSASA), (6) Checking biophysical properties (instability, GRAVY, pI), (7) Ranking designs with composite scoring. This…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `references/binding-qc.md` and `references/composite-scoring.md`).

It sits in Research & Science, covering Protein structure and design and Design review and critique. The repository describes itself as: A curated collection of AI agent skills for biomedical research, covering genomics, proteomics, single-cell analysis, clinical AI, and protein design. The licence is MIT.

When your agent uses it

  • Evaluating design quality for binding
  • Setting filtering thresholds for pLDDT
  • Checking sequence liabilities (cysteines
  • Polybasic clusters)

Example prompts

  • “/protein-design-qc”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 8cbdd18. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Protein Design Qc loads about 2.6k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 159 tokens; SKILL.md has 530 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from BioTender-max/awesome-bio-agent-skills at commit 8cbdd18, republished under its MIT licence (© BioTender-max). 530 words, ~2,561 tokens.

Download SKILL.mdSave it as .claude/skills/protein-design-qc/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
protein-design-qc
description
Protein design quality control, filtering thresholds, and ranking guidance. Use this skill when: (1) Evaluating design quality for binding, expression, or structure, (2) Setting filtering thresholds for pLDDT, ipTM, PAE, (3) Checking sequence liabilities (cysteines, deamidation, polybasic clusters), (4) Creating multi-stage filtering pipelines, (5) Computing PyRosetta interface metrics (dG, SC, dSASA), (6) Checking biophysical properties (instability, GRAVY, pI), (7) Ranking designs with composite scoring. This skill provides research-backed thresholds from binder design competitions and published benchmarks.
license
MIT
category
evaluation
tags
qc, filtering, metrics, thresholds

Protein Design Quality Control And Ranking

Plain-language role: Use this skill to decide which designs pass QC and which ones should move forward.

Critical Limitation

Individual metrics have weak predictive power for binding. Research shows:

  • Individual metric ROC AUC: 0.64-0.66 (slightly better than random)
  • Metrics are pre-screening filters, not affinity predictors
  • Composite scoring is essential for meaningful ranking

These thresholds filter out poor designs but do NOT predict binding affinity.

QC Organization

QC is organized by purpose and level:

PurposeWhat it assessesKey metrics
BindingInterface quality, binding geometryipTM, PAE, SC, dG, dSASA
ExpressionManufacturability, solubilityInstability, GRAVY, pI, cysteines
StructuralFold confidence, consistencypLDDT, pTM, scRMSD

Each category has two levels:

  • Metric-level: Calculated values with thresholds (pLDDT > 0.85)
  • Design-level: Pattern/motif detection (odd cysteines, NG sites)

Quick Reference: All Thresholds

CategoryMetricStandardStringentSource
StructuralpLDDT> 0.85> 0.90AF2/Chai/Boltz
pTM> 0.70> 0.80AF2/Chai/Boltz
scRMSD< 2.0 Å< 1.5 ÅDesign vs pred
BindingipTM> 0.50> 0.60AF2/Chai/Boltz
PAE_interaction< 12 Å< 10 ÅAF2/Chai/Boltz
Shape Comp (SC)> 0.50> 0.60PyRosetta
interface_dG< -10< -15PyRosetta
ExpressionInstability< 40< 30BioPython
GRAVY< 0.4< 0.2BioPython
ESM2 PLL> 0.0> 0.2ESM2
Design-Level Checks (Expression)
PatternRiskAction
Odd cysteine countUnpaired disulfidesRedesign
NG/NS/NT motifsDeamidationFlag/avoid
K/R >= 3 consecutiveProteolysisFlag
>= 6 hydrophobic runAggregationRedesign

See: references/binding-qc.md, references/expression-qc.md, references/structural-qc.md


Sequential Filtering Pipeline

python
import pandas as pd

designs = pd.read_csv('designs.csv')

# Stage 1: Structural confidence
designs = designs[designs['pLDDT'] > 0.85]

# Stage 2: Self-consistency
designs = designs[designs['scRMSD'] < 2.0]

# Stage 3: Binding quality
designs = designs[(designs['ipTM'] > 0.5) & (designs['PAE_interaction'] < 10)]

# Stage 4: Sequence plausibility
designs = designs[designs['esm2_pll_normalized'] > 0.0]

# Stage 5: Expression checks (design-level)
designs = designs[designs['cysteine_count'] % 2 == 0]  # Even cysteines
designs = designs[designs['instability_index'] < 40]

Composite Scoring (Required for Ranking)

Individual metrics alone are too weak. Use composite scoring:

python
def composite_score(row):
    return (
        0.30 * row['pLDDT'] +
        0.20 * row['ipTM'] +
        0.20 * (1 - row['PAE_interaction'] / 20) +
        0.15 * row['shape_complementarity'] +
        0.15 * row['esm2_pll_normalized']
    )

designs['score'] = designs.apply(composite_score, axis=1)
top_designs = designs.nlargest(100, 'score')

For advanced composite scoring, see references/composite-scoring.md.


Tool-Specific Filtering

BindCraft Filter Levels
LevelUse CaseStringency
DefaultStandard designMost stringent
RelaxedNeed more designsHigher failure rate
PeptideDesigns < 30 AA~5-10x lower success
BoltzGen Filtering
bash
boltzgen run ... \
  --budget 60 \
  --alpha 0.01 \
  --filter_biased true \
  --refolding_rmsd_threshold 2.0 \
  --additional_filters 'ALA_fraction<0.3'
  • alpha=0.0: Quality-only ranking
  • alpha=0.01: Default (slight diversity)
  • alpha=1.0: Diversity-only

Design-Level Severity Scoring

For pattern-based checks, use severity scoring:

Severity LevelScoreAction
LOW0-15Proceed
MODERATE16-35Review flagged issues
HIGH36-60Redesign recommended
CRITICAL61+Redesign required

Show full SKILL.md (220 more words)Show less

Experimental Correlation

MetricAUCUse
ipTM~0.64Pre-screening
PAE~0.65Pre-screening
ESM2 PLL~0.72Best single metric
Composite~0.75+Always use

Key insight: Metrics work as filters (eliminating failures) not predictors (ranking successes).


Campaign Health Assessment

Quick assessment of your design campaign:

Pass RateStatusInterpretation
> 15%ExcellentAbove average, proceed
10-15%GoodNormal, proceed
5-10%MarginalBelow average, review issues
< 5%PoorSignificant problems, diagnose

Failure Recovery Trees

Too Few Pass pLDDT Filter (< 5% with pLDDT > 0.85)
Low pLDDT across campaign
├── Check scRMSD distribution
│   ├── High scRMSD (>2.5Å): Backbone issue
│   │   └── Fix: Regenerate backbones with lower noise_scale (0.5-0.8)
│   └── Low scRMSD but low pLDDT: Disordered regions
│       └── Fix: Check design length, simplify topology
├── Try more sequences per backbone
│   └── modal run modal_proteinmpnn.py --num-seq-per-target 32 --sampling-temp 0.1
├── Use SolubleMPNN instead of ProteinMPNN
│   └── Better for expression-optimized sequences
└── Consider different design tool
    └── BindCraft (integrated design) may work better
Too Few Pass ipTM Filter (< 5% with ipTM > 0.5)
Low ipTM across campaign
├── Review hotspot selection
│   ├── Are hotspots surface-exposed? (SASA > 20Ų)
│   ├── Are hotspots conserved? (check MSA)
│   └── Try 3-6 different hotspot combinations
├── Increase binder length (more contact area)
│   └── Try 80-100 AA instead of 60-80 AA
├── Check interface geometry
│   ├── Is target flat? → Try helical binders
│   └── Is target concave? → Try smaller binders
└── Try all-atom design tool
    └── BoltzGen (all-atom, better packing)
High scRMSD (> 50% with scRMSD > 2.0Å)
Sequences don't specify intended structure
├── ProteinMPNN issue
│   ├── Lower temperature: --sampling-temp 0.1
│   ├── Increase sequences: --num-seq-per-target 32
│   └── Check fixed_positions aren't over-constraining
├── Backbone geometry issue
│   ├── Backbones may be unusual/strained
│   ├── Regenerate with lower noise_scale (0.5-0.8)
│   └── Reduce diffuser.T to 30-40
└── Try different sequence design
    └── ColabDesign (AF2 gradient-based) may work better
Everything Passes But No Experimental Hits
In silico metrics don't predict affinity
├── Generate MORE designs (10x current)
│   └── Computational metrics have high false positive rate
├── Increase diversity
│   ├── Higher ProteinMPNN temperature (0.2-0.3)
│   ├── Different backbone topologies
│   └── Different hotspot combinations
├── Try different design approach
│   ├── BindCraft (different algorithm)
│   ├── ColabDesign (AF2 hallucination)
│   └── BoltzGen (all-atom diffusion)
└── Check if target is druggable
    └── Some targets are inherently difficult
Too Many Designs Pass (> 50%)
Suspiciously high pass rate
├── Check if thresholds are too lenient
│   └── Use stringent thresholds: pLDDT > 0.90, ipTM > 0.60
├── Verify prediction quality
│   ├── Are predictions actually running? Check output files
│   └── Are complexes being predicted, not just monomers?
├── Check for data issues
│   ├── Same sequence being predicted multiple times?
│   └── Wrong FASTA format (missing chain separator)?
└── Apply diversity filter
    └── Cluster at 70% identity, take top per cluster

Diagnostic Commands

Quick Campaign Assessment
python
import pandas as pd

df = pd.read_csv('designs.csv')

# Pass rates at each stage
print(f"Total designs: {len(df)}")
print(f"pLDDT > 0.85: {(df['pLDDT'] > 0.85).mean():.1%}")
print(f"ipTM > 0.50: {(df['ipTM'] > 0.50).mean():.1%}")
print(f"scRMSD < 2.0: {(df['scRMSD'] < 2.0).mean():.1%}")
print(f"All filters: {((df['pLDDT'] > 0.85) & (df['ipTM'] > 0.5) & (df['scRMSD'] < 2.0)).mean():.1%}")

# Identify top issue
if (df['pLDDT'] > 0.85).mean() < 0.1:
    print("ISSUE: Low pLDDT - check backbone or sequence quality")
elif (df['ipTM'] > 0.50).mean() < 0.1:
    print("ISSUE: Low ipTM - check hotspots or interface geometry")
elif (df['scRMSD'] < 2.0).mean() < 0.5:
    print("ISSUE: High scRMSD - sequences don't specify backbone")

Templates and Demo

  • Executable scorer: scripts/protein_qc_score.py
  • Sample metrics table: templates/protein-design-qc/design_metrics.csv
  • Demo input table: examples/minimal-binder-campaign/inputs/design_metrics.csv

Inputs

  • A CSV table of design metrics such as pLDDT, ipTM, PAE, scRMSD, and optional ESM or expression fields.
  • Optional threshold overrides for stricter or more permissive filtering.
  • A design campaign context so summary pass rates can be interpreted correctly.

Outputs

  • Per-design pass or fail annotations across structural, binding, and expression checks.
  • A composite score and ranked output table suitable for experimental triage.
  • A campaign summary highlighting which filter stage is removing most candidates.

Next Step

Advance the top-ranked designs into ipsae or experimental testing, and use the summary to decide whether the generation stage needs adjustment.

© BioTender-max, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/bioclaw_hub/protein-design-qc of BioTender-max/awesome-bio-agent-skills.

  • SKILL.md
  • README.md
  • references/binding-qc.md
  • references/composite-scoring.md
  • references/expression-qc.md
  • references/structural-qc.md

Open the folder on GitHubat commit 8cbdd18

Compare with similar skills

Protein Design Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Protein Design Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Protein Design Qc this skillBioTender-max/awesome-bio-agent-skills200—~2.6kAutomated safety check: PassMIT
Protein Qcadaptyvbio/protein-design-skills1643 repos~3.2kAutomated safety check: PassMIT
Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills3.2k2 repos~1.2kAutomated safety check: PassApache-2.0
Alphafoldadaptyvbio/protein-design-skills1643 repos~1.2kAutomated safety check: PassMIT
Pymol VisualizationChatMol/ChatMol373—~1.2kAutomated safety check: PassMIT
Complexa Binder DesignNVIDIA-BioNeMo/bionemo-agent-toolkit479—~3.1kAutomated safety check: NotesApache-2.0

Similar skills

  • Protein Qc

    adaptyvbio/protein-design-skills

    Quality control metrics and filtering thresholds for protein design.

    164 GitHub starsUsed in 3 repos~3.2k tokens
    Research & ScienceAuto-check passed
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Alphafold

    adaptyvbio/protein-design-skills

    Validate protein designs using AlphaFold2 structure prediction.

    164 GitHub starsUsed in 3 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Pymol Visualization

    ChatMol/ChatMol

    Generate publication-quality molecular visualization images using PyMOL.

    373 GitHub stars~1.2k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Complexa Binder Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided…

    479 GitHub stars~3.1k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Bindcraft

    adaptyvbio/protein-design-skills

    End-to-end binder design using BindCraft hallucination. An agent skill from adaptyvbio/protein-design-skills.

    164 GitHub starsUsed in 3 repos~1.3k tokens
    Research & ScienceAuto-check passed

More from BioTender-max/awesome-bio-agent-skills

All 23 skills in this repo
  • AI Scientist Evaluator

    BioTender-max/awesome-bio-agent-skills

    Critically review, score, compare, and rank one or more AI scientist outputs for biology, bioinformatics, computational life science, or adjacent research tasks.

    200 GitHub stars~2.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Exa Search

    BioTender-max/awesome-bio-agent-skills

    Web toolkit powered by Exa, tuned for scientific and technical content.

    200 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check: notes
  • Jgi Lakehouse

    BioTender-max/awesome-bio-agent-skills

    Queries JGI Lakehouse (Dremio) for genomics metadata from GOLD, IMG, Mycocosm, Phytozome.

    200 GitHub stars~3.7k tokensUpdated 3 mo ago
    Auto-check passed
  • Pacsomatic

    BioTender-max/awesome-bio-agent-skills

    Operator toolkit for nf-core/pacsomatic matched tumor-normal workflows from BAM inputs.

    200 GitHub stars~1.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Scientific Impact Assessment

    BioTender-max/awesome-bio-agent-skills

    Assess paper and journal impact using OpenAlex citation counts, optional Altmetric data, and curated journal impact-factor references.

    200 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed
  • Arxiv Search

    BioTender-max/awesome-bio-agent-skills

    Search arXiv preprints through the official arXiv API and turn arXiv IDs into local Markdown summaries.

    200 GitHub stars~2.9k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Protein Design Qc

What does Protein Design Qc do?

Protein design quality control, filtering thresholds, and ranking guidance. Protein Design Qc is an agent skill from BioTender-max/awesome-bio-agent-skills. Protein design quality control, filtering thresholds, and ranking guidance.

When should I use Protein Design Qc?

Protein Design Qc fits situations like: evaluating design quality for binding; setting filtering thresholds for pLDDT; checking sequence liabilities (cysteines; polybasic clusters).

How do I install Protein Design Qc in Claude Code?

Run `npx skills add BioTender-max/awesome-bio-agent-skills --skill protein-design-qc -a claude-code`. Or copy the skill folder (skills/bioclaw_hub/protein-design-qc in BioTender-max/awesome-bio-agent-skills) into .claude/skills/protein-design-qc in your project. Claude Code loads it when a task matches its description.

How do I install Protein Design Qc in Codex?

Run `npx skills add BioTender-max/awesome-bio-agent-skills --skill protein-design-qc -a codex`. Or copy the skill folder (skills/bioclaw_hub/protein-design-qc in BioTender-max/awesome-bio-agent-skills) into .agents/skills/protein-design-qc in your project. Codex loads it when a task matches its description.

Can I use Protein Design Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BioTender-max/awesome-bio-agent-skills --skill protein-design-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/protein-design-qc, .gemini/skills/protein-design-qc, .github/skills/protein-design-qc and .opencode/skills/protein-design-qc in your project.

What does Protein Design Qc need to run?

SKILL.md names no scripts, command-line tools or credentials: Protein Design Qc is instructions for the agent only. Our summary lists: Python 3.

Does Protein Design Qc access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Protein Design Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Protein Design Qc use?

Protein Design Qc is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Protein Design Qc use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.1k tokens, read only when the agent opens those files.

What are the alternatives to Protein Design Qc?

Skills that share tags, products or a category with Protein Design Qc: Protein Qc (adaptyvbio/protein-design-skills, 164 stars), Alphafold Database Fetch And Analyze (google-deepmind/science-skills, 3.2k stars), Alphafold (adaptyvbio/protein-design-skills, 164 stars) and Pymol Visualization (ChatMol/ChatMol, 373 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Protein Design Qc?

BioTender-max (a GitHub user) maintains it in BioTender-max/awesome-bio-agent-skills, which has 200 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on July 1, 2026.

Source: BioTender-max/awesome-bio-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.