Agent skill

Protein Structure Retrieval

by lamm-mit in lamm-mit/scienceclaw

ToolUniverse workflow — Protein Structure Retrieval. An agent skill from lamm-mit/scienceclaw.

Apache-2.0Auto-check passedResearch & Science

Install Protein Structure Retrieval

skills CLI
$ npx skills add lamm-mit/scienceclaw --skill protein-structure-retrieval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lamm-mit/scienceclaw protein-structure-retrieval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lamm-mit/scienceclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/protein-structure-retrieval .claude/skills/protein-structure-retrieval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
protein-structure-retrieval
GitHub stars
244
Token cost
~2.8k tokens
SKILL.md length
621 words
Files
3 (incl. scripts)
Skills in repo
86
Repo updated
First seen
Licence
Apache-2.0

At a glance

ToolUniverse workflow — Protein Structure Retrieval. An agent skill from lamm-mit/scienceclaw.

  • Works in 4 steps: Clarification (When Needed) → Protein Disambiguation → Data Retrieval (Internal) → …
  • Tasks that involve Protein structure and design
  • SKILL.md covers Workflow Overview, Phase 0: Clarification (When…, Phase 1: Protein Disambiguation and Phase 2: Data Retrieval…, plus 6 more sections
  • Runs Python scripts from its folder; reaches rcsb.org and ebi.ac.uk

What it does

Protein Structure Retrieval is an agent skill from lamm-mit/scienceclaw. ToolUniverse workflow — Protein Structure Retrieval

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/run.py`).

It sits in Research & Science, covering Protein structure and design. It works with AlphaFold and UniProt. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Protein structure and design

Example prompts

  • “/protein-structure-retrieval”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Clarification (When Needed)
  2. Protein Disambiguation
  3. Data Retrieval (Internal)
  4. Report Structure Profile

What it can do on your machine

Read from SKILL.md and the folder at commit ab9aba1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • rcsb.org
    • ebi.ac.uk
    • alphafold.ebi.ac.uk

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Protein Structure Retrieval loads about 2.8k tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 621 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from lamm-mit/scienceclaw at commit ab9aba1, republished under its Apache-2.0 licence (© lamm-mit). 621 words, ~2,834 tokens.

Download SKILL.mdSave it as .claude/skills/protein-structure-retrieval/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
protein-structure-retrieval
description
ToolUniverse workflow — Protein Structure Retrieval
source
https://github.com/mims-harvard/ToolUniverse/tree/main/skills/tooluniverse-protein-structure-retrieval

name: tooluniverse-protein-structure-retrieval description: Retrieves protein structure data from RCSB PDB, PDBe, and AlphaFold with protein disambiguation, quality assessment, and comprehensive structural profiles. Creates detailed structure reports with experimental metadata, ligand information, and download links. Use when users need protein structures, 3D models, crystallography data, or mention PDB IDs (4-character codes like 1ABC) or UniProt accessions.

Protein Structure Data Retrieval

Retrieve protein structures with proper disambiguation, quality assessment, and comprehensive metadata.

IMPORTANT: Always use English terms in tool calls (protein names, organism names), even if the user writes in another language. Only try original-language terms as a fallback if English returns no results. Respond in the user's language.

Workflow Overview

Phase 0: Clarify (if needed)
    ↓
Phase 1: Disambiguate Protein Identity
    ↓
Phase 2: Retrieve Structures (Internal)
    ↓
Phase 3: Report Structure Profile

Phase 0: Clarification (When Needed)

Ask the user ONLY if:

  • Protein name matches multiple genes/families (e.g., "kinase" → which kinase?)
  • Organism not specified for conserved proteins
  • Intent unclear: need experimental structure vs AlphaFold prediction?

Skip clarification for:

  • Specific PDB IDs (4-character codes)
  • UniProt accessions
  • Unambiguous protein names with organism

Phase 1: Protein Disambiguation

1.1 Resolve Protein Identity
python
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()

# Strategy depends on input type
if user_provided_pdb_id:
    # Direct structure retrieval
    pdb_id = user_provided_pdb_id.upper()
    
elif user_provided_uniprot:
    # Get UniProt info, then search structures
    uniprot_id = user_provided_uniprot
    # Can also get AlphaFold structure
    af_structure = tu.tools.alphafold_get_structure_by_uniprot(
        uniprot_id=uniprot_id
    )
    
elif user_provided_protein_name:
    # Search by name
    result = tu.tools.search_structures_by_protein_name(
        protein_name=protein_name
    )
1.2 Identity Resolution Checklist
  • Protein name/gene identified
  • Organism confirmed
  • UniProt accession (if available)
  • Isoform/variant specified (if relevant)
1.3 Handle Naming Collisions

Common ambiguous terms:

TermAmbiguityResolution
"kinase"Hundreds of kinasesAsk which kinase (EGFR, CDK2, etc.)
"receptor"Many receptor typesSpecify receptor family
"protease"Multiple familiesAsk serine/cysteine/metallo/etc.
"hemoglobin"ClearProceed (α/β chain specified if needed)
"insulin"ClearProceed

Phase 2: Data Retrieval (Internal)

Retrieve all data silently. Do NOT narrate the search process.

2.1 Search Structures
python
# Search by protein name
result = tu.tools.search_structures_by_protein_name(
    protein_name=protein_name
)

# Filter results by quality
high_res = [
    entry for entry in result["data"]
    if entry.get("resolution") and entry["resolution"] < 2.5
]
2.2 Get Structure Details

For each relevant structure:

python
pdb_id = "4INS"

# Basic metadata
metadata = tu.tools.get_protein_metadata_by_pdb_id(pdb_id=pdb_id)

# Experimental details
exp_details = tu.tools.get_protein_experimental_details_by_pdb_id(
    pdb_id=pdb_id
)

# Resolution (if X-ray)
resolution = tu.tools.get_protein_resolution_by_pdb_id(pdb_id=pdb_id)

# Bound ligands
ligands = tu.tools.get_protein_ligands_by_pdb_id(pdb_id=pdb_id)

# Similar structures
similar = tu.tools.get_similar_structures_by_pdb_id(
    pdb_id=pdb_id,
    cutoff=2.0
)
2.3 PDBe Additional Data
python
# Entry summary
summary = tu.tools.pdbe_get_entry_summary(pdb_id=pdb_id)

# Molecular entities
molecules = tu.tools.pdbe_get_molecules(pdb_id=pdb_id)

# Binding sites
binding_sites = tu.tools.pdbe_get_binding_sites(pdb_id=pdb_id)
2.4 AlphaFold Predictions
python
# When no experimental structure exists, or for comparison
if uniprot_id:
    af_structure = tu.tools.alphafold_get_structure_by_uniprot(
        uniprot_id=uniprot_id
    )
Fallback Chains
PrimaryFallbackNotes
RCSB searchPDBe searchRegional availability
get_protein_metadatapdbe_get_entry_summaryAlternative source
Experimental structureAlphaFold predictionNo experimental structure
get_protein_ligandspdbe_get_binding_sitesLigand info unavailable

Phase 3: Report Structure Profile

Output Structure

Present as a Structure Profile Report. Hide search process.

markdown
# Protein Structure Profile: [Protein Name]

**Search Summary**
- Query: [protein name/PDB ID]
- Organism: [species]
- Structures Found: [N] experimental, [M] AlphaFold

---

## Best Available Structure

### [PDB ID]: [Title]

| Attribute | Value |
|-----------|-------|
| **PDB ID** | [pdb_id] |
| **UniProt** | [uniprot_id] |
| **Organism** | [species] |
| **Method** | X-ray / Cryo-EM / NMR |
| **Resolution** | [X.XX] Å |
| **Release Date** | [date] |

**Quality Assessment**: ●●● High / ●●○ Medium / ●○○ Low

### Experimental Details
| Parameter | Value |
|-----------|-------|
| **Method** | [X-ray crystallography] |
| **Resolution** | [1.9 Å] |
| **R-factor** | [0.18] |
| **R-free** | [0.21] |
| **Space Group** | [P 21 21 21] |

### Structure Composition
| Component | Count | Details |
|-----------|-------|---------|
| **Chains** | [N] | [A (enzyme), B (inhibitor)] |
| **Residues** | [N] | [coverage %] |
| **Ligands** | [N] | [list ligand names] |
| **Waters** | [N] | |
| **Metals** | [N] | [Zn, Mg, etc.] |

### Bound Ligands
| Ligand ID | Name | Type | Binding Site |
|-----------|------|------|--------------|
| [ATP] | Adenosine triphosphate | Substrate | Active site |
| [MG] | Magnesium ion | Cofactor | Catalytic |

### Binding Site Details
For drug discovery applications:

**Site 1: Active Site**
- Location: Chain A, residues 45-89
- Key residues: Asp45, Glu67, His89
- Pocket volume: [X] ų
- Druggability: High/Medium/Low

---

## Alternative Structures

Ranked by quality and relevance:

| Rank | PDB ID | Resolution | Method | Ligands | Notes |
|------|--------|------------|--------|---------|-------|
| 1 | [4INS] | 1.9 Å | X-ray | Zn | Best resolution |
| 2 | [3I40] | 2.1 Å | X-ray | Zn, phenol | With inhibitor |
| 3 | [1TRZ] | 2.3 Å | X-ray | None | Porcine |

---

## AlphaFold Prediction

### AF-[UniProt]-F1

| Attribute | Value |
|-----------|-------|
| **UniProt** | [uniprot_id] |
| **Model Version** | [v4] |
| **Confidence (pLDDT)** | [average score] |

**Confidence Distribution**:
- Very High (>90): [X]% of residues
- High (70-90): [X]% of residues
- Low (50-70): [X]% of residues
- Very Low (<50): [X]% of residues

**Use Cases**:
- ✓ Overall fold reliable
- ✓ Core domain structure
- ⚠ Loop regions uncertain
- ✗ Not suitable for binding site analysis

---

## Structure Comparison

| Property | [PDB_1] | [PDB_2] | AlphaFold |
|----------|---------|---------|-----------|
| Resolution | 1.9 Å | 2.5 Å | N/A (predicted) |
| Completeness | 98% | 85% | 100% |
| Ligands | Yes | No | No |
| Confidence | Experimental | Experimental | High (85 avg) |

---

## Download Links

### Coordinate Files
| Format | PDB ID | Link |
|--------|--------|------|
| PDB | [4INS] | [link] |
| mmCIF | [4INS] | [link] |
| AlphaFold | [UniProt] | [link] |

### Database Links
- RCSB PDB: https://www.rcsb.org/structure/[pdb_id]
- PDBe: https://www.ebi.ac.uk/pdbe/entry/pdb/[pdb_id]
- AlphaFold: https://alphafold.ebi.ac.uk/entry/[uniprot_id]

Retrieved: [date]

Quality Assessment Tiers

Experimental Structures
TierSymbolCriteria
Excellent●●●●X-ray <1.5Å, complete, R-free <0.22
High●●●○X-ray <2.0Å OR Cryo-EM <3.0Å
Good●●○○X-ray 2.0-3.0Å OR Cryo-EM 3.0-4.0Å
Moderate●○○○X-ray >3.0Å OR NMR ensemble
Low○○○○>4.0Å, incomplete, or problematic
Resolution Guide
ResolutionUse Case
<1.5 ÅAtomic detail, H-bond analysis
1.5-2.0 ÅDrug design, mechanism studies
2.0-2.5 ÅStructure-based design
2.5-3.5 ÅOverall architecture, fold
>3.5 ÅDomain arrangement only
Show full SKILL.md (252 more words)Show less
AlphaFold Confidence
pLDDT ScoreInterpretation
>90Very high confidence, experimental-like
70-90Good backbone confidence
50-70Uncertain, flexible regions
<50Low confidence, likely disordered

Completeness Checklist

Every structure report MUST include:

For Specific PDB ID (Required)
  • PDB ID and title
  • Experimental method
  • Resolution (or N/A for NMR)
  • Organism
  • Quality assessment
  • Download links
For Protein Name Search (Required)
  • Search summary with result count
  • Top structures with quality ranking
  • Best structure recommendation
  • AlphaFold alternative (if no experimental structure)
Always Include
  • Ligand information (or "No ligands bound")
  • Data sources with links
  • Retrieval date

Common Use Cases

Drug Discovery Target

User: "Get structure for EGFR kinase with inhibitor" → Filter for ligand-bound structures, emphasize binding site

Model Building

User: "Find best template for homology modeling of protein X" → High-resolution structures, note sequence coverage

Structure Comparison

User: "Compare available SARS-CoV-2 main protease structures" → All structures with systematic comparison table

AlphaFold When No Experimental

User: "Structure of protein with UniProt P12345" → Check PDB first, then AlphaFold, note confidence


Error Handling

ErrorResponse
"PDB ID not found"Verify 4-character format, check if obsoleted
"No structures for protein"Offer AlphaFold prediction, suggest similar proteins
"Download failed"Retry once, provide alternative link
"Resolution unavailable"Likely NMR/model, note in assessment

Tool Reference

RCSB PDB (Experimental Structures)

ToolPurpose
search_structures_by_protein_nameName-based search
get_protein_metadata_by_pdb_idBasic info
get_protein_experimental_details_by_pdb_idMethod details
get_protein_resolution_by_pdb_idQuality metric
get_protein_ligands_by_pdb_idBound molecules
download_pdb_structure_fileCoordinate files
get_similar_structures_by_pdb_idHomologs

PDBe (European PDB)

ToolPurpose
pdbe_get_entry_summaryOverview
pdbe_get_moleculesMolecular entities
pdbe_get_experiment_infoExperimental data
pdbe_get_binding_sitesLigand pockets

AlphaFold (Predictions)

ToolPurpose
alphafold_get_structure_by_uniprotGet prediction
alphafold_search_structuresSearch predictions

© lamm-mit, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/protein-structure-retrieval of lamm-mit/scienceclaw.

  • SKILL.md
  • scripts/__pycache__/run.cpython-313.pyc
  • scripts/run.py

Open the folder on GitHubat commit ab9aba1

Compare with similar skills

Protein Structure Retrieval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Protein Structure Retrieval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Protein Structure Retrieval this skilllamm-mit/scienceclaw244—~2.8kAutomated safety check: PassApache-2.0
Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills3.2k2 repos~1.2kAutomated safety check: PassApache-2.0
Bio DB ToolsDrugClaw/DrugClaw125—~1.4kAutomated safety check: PassApache-2.0
Ggetdavila7/claude-code-templates32k10 repos~6.3kAutomated safety check: PassMIT
Alphafold Databasedavila7/claude-code-templates32k10 repos~4kAutomated safety check: PassMIT
Foldseek Structural Searchgoogle-deepmind/science-skills3.2k1 repos~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Bio DB Tools

    DrugClaw/DrugClaw

    Query public biology databases and APIs including UniProt, RCSB PDB, AlphaFold DB, ClinVar, dbSNP, gnomAD, Ensembl, GEO, InterPro, KEGG, OpenTargets, Reactome, and STRING.

    125 GitHub stars~1.4k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Alphafold Database

    davila7/claude-code-templates

    Access AlphaFold's 200M+ AI-predicted protein structures. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~4k tokens
    Research & ScienceAuto-check passed
  • Foldseek Structural Search

    google-deepmind/science-skills

    Performs 3D structural searches of proteins against various databases (PDB, AlphaFold, CATH, MGnify, etc.) using the Foldseek API.

    3.2k GitHub starsUsed in 1 repo~1.3k tokens
    Research & ScienceAuto-check passed
  • Retrieves protein structure data from RCSB PDB, PDBe, and AlphaFold with protein disambiguation, quality assessment, and comprehensive structural profiles.

    1.1k GitHub starsUsed in 2 repos~2.8k tokens
    Research & ScienceAuto-check passed

More from lamm-mit/scienceclaw

All 86 skills in this repo
  • Fred Economic Data

    lamm-mit/scienceclaw

    Query FRED (Federal Reserve Economic Data) API for 800,000+ economic time series from 100+ sources.

    244 GitHub starsUsed in 4 repos~3k tokens
    Auto-check passed
  • Drug Research

    lamm-mit/scienceclaw

    Generates comprehensive drug research reports with compound disambiguation, evidence grading, and mandatory completeness sections.

    244 GitHub starsUsed in 3 repos~1.7k tokens
    Auto-check passed
  • Imaging Data Commons

    lamm-mit/scienceclaw

    Query and download public cancer imaging data from NCI Imaging Data Commons using idc-index.

    244 GitHub starsUsed in 5 repos~11k tokens
    Auto-check passed
  • Rowan

    lamm-mit/scienceclaw

    Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.

    244 GitHub starsUsed in 4 repos~3.1k tokens
    Auto-check: warnings
  • Infographics

    lamm-mit/scienceclaw

    Create professional infographics using Nano Banana Pro AI with smart iterative refinement.

    244 GitHub starsUsed in 6 repos~4.4k tokens
    Auto-check: notes
  • Disease Research

    lamm-mit/scienceclaw

    Generate comprehensive disease research reports using 100+ ToolUniverse tools.

    244 GitHub stars~946 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Protein Structure Retrieval

What does Protein Structure Retrieval do?

ToolUniverse workflow — Protein Structure Retrieval. An agent skill from lamm-mit/scienceclaw. Protein Structure Retrieval is an agent skill from lamm-mit/scienceclaw.

When should I use Protein Structure Retrieval?

Protein Structure Retrieval fits situations like: tasks that involve Protein structure and design.

How do I install Protein Structure Retrieval in Claude Code?

Run `npx skills add lamm-mit/scienceclaw --skill protein-structure-retrieval -a claude-code`. Or copy the skill folder (skills/protein-structure-retrieval in lamm-mit/scienceclaw) into .claude/skills/protein-structure-retrieval in your project. Claude Code loads it when a task matches its description.

How do I install Protein Structure Retrieval in Codex?

Run `npx skills add lamm-mit/scienceclaw --skill protein-structure-retrieval -a codex`. Or copy the skill folder (skills/protein-structure-retrieval in lamm-mit/scienceclaw) into .agents/skills/protein-structure-retrieval in your project. Codex loads it when a task matches its description.

Can I use Protein Structure Retrieval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lamm-mit/scienceclaw --skill protein-structure-retrieval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/protein-structure-retrieval, .gemini/skills/protein-structure-retrieval, .github/skills/protein-structure-retrieval and .opencode/skills/protein-structure-retrieval in your project.

What does Protein Structure Retrieval need to run?

Going by SKILL.md and its folder, Protein Structure Retrieval needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Protein Structure Retrieval access the network?

SKILL.md names 3 domains. In commands or code: rcsb.org, ebi.ac.uk and alphafold.ebi.ac.uk; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Protein Structure Retrieval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Protein Structure Retrieval use?

Protein Structure Retrieval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Protein Structure Retrieval use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Protein Structure Retrieval?

Skills that share tags, products or a category with Protein Structure Retrieval: Alphafold Database Fetch And Analyze (google-deepmind/science-skills, 3.2k stars), Bio DB Tools (DrugClaw/DrugClaw, 125 stars), Gget (davila7/claude-code-templates, 32k stars) and Alphafold Database (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Protein Structure Retrieval?

lamm-mit (a GitHub user) maintains it in lamm-mit/scienceclaw, which has 244 GitHub stars. The repository holds 86 skills in this directory. The repository was last updated on August 21, 2026.

Source: lamm-mit/scienceclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.