Agent skill

Pdbe API

by pemsley in pemsley/coot

Query the PDBe (Protein Data Bank in Europe) REST API and Solr search API from within Coot to access structure metadata, validation data, revision history, search capabilities, and download…

GPL-3.0Auto-check passedBackend & APIs

Install Pdbe API

skills CLI
$ npx skills add pemsley/coot --skill pdbe-api -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pemsley/coot pdbe-api --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pemsley/coot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/mcp/docs/skills/pdbe-api .claude/skills/pdbe-api && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdbe-api
GitHub stars
168
Token cost
~7.5k tokens
SKILL.md length
1,329 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
GPL-3.0

At a glance

Query the PDBe (Protein Data Bank in Europe) REST API and Solr search API from within Coot to access structure metadata, validation data, revision history, search capabilities, and download…

  • Works in 8 steps: Structure Summary and Metadata → Molecule/Entity Information → Compound/Ligand Information → …
  • Tasks that involve REST APIs
  • SKILL.md covers Overview, Core Function, Main API Endpoints and Common Query Patterns, plus 4 more sections
  • Reaches ebi.ac.uk and files.rcsb.org

What it does

Pdbe API is an agent skill from pemsley/coot. Query the PDBe (Protein Data Bank in Europe) REST API and Solr search API from within Coot to access structure metadata, validation data, revision history, search capabilities, and download coordinate files

Its SKILL.md is about 7.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Backend & APIs, covering REST APIs. The repository describes itself as: Software for macromolecular model-building. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve REST APIs

Example prompts

  • “/pdbe-api”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Structure Summary and Metadata
  2. Molecule/Entity Information
  3. Compound/Ligand Information
  4. Downloading Coordinate Files
  5. Validation Reports
  6. Structure Status and Revision History
  7. Ligand Binding Sites
  8. Assembly Information

What it can do on your machine

Read from SKILL.md and the folder at commit 6e3c026. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ebi.ac.uk
    • files.rcsb.org

    Also links to:

    • pdbe.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pdbe API loads about 7.5k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 1,329 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pemsley/coot at commit 6e3c026, republished under its GPL-3.0 licence (© pemsley). 1,329 words, ~7,458 tokens.

Download SKILL.mdSave it as .claude/skills/pdbe-api/SKILL.md (or your agent's skills folder).
name
pdbe-api
description
Query the PDBe (Protein Data Bank in Europe) REST API and Solr search API from within Coot to access structure metadata, validation data, revision history, search capabilities, and download coordinate files

PDBe API Access from Coot

Overview

The PDBe (Protein Data Bank in Europe) provides comprehensive REST and Solr-based APIs for programmatic access to structure data, validation reports, compound information, revision history, and search capabilities. Coot can access these APIs directly using the coot_get_url_as_string_py() function, which now supports both text and binary data.

Core Function

coot.coot_get_url_as_string_py(url) - Fetch URL content

Returns:

  • Python str for text/JSON content (valid UTF-8)
  • Python bytes for binary content (gzipped files, images, etc.)
python
import json

# Example 1: Get structure summary (returns string)
result = coot.coot_get_url_as_string_py("https://www.ebi.ac.uk/pdbe/api/pdb/entry/summary/4wa9")
data = json.loads(result)

# Example 2: Download gzipped coordinates (returns bytes)
import gzip
compressed = coot.coot_get_url_as_string_py("https://files.rcsb.org/download/4wa9.cif.gz")
decompressed = gzip.decompress(compressed)
imol = coot.read_coordinates_as_string(decompressed.decode('utf-8'), "4wa9")

Main API Endpoints

Entry-based API

Base URL: https://www.ebi.ac.uk/pdbe/api/

Documentation: https://www.ebi.ac.uk/pdbe/api/doc/

Aggregated API

Base URL: https://www.ebi.ac.uk/pdbe/graph-api/

Documentation: https://pdbe.org/graph-api

Search API (Solr)

Base URL: https://www.ebi.ac.uk/pdbe/search/pdb/select?

Documentation: https://www.ebi.ac.uk/pdbe/api/doc/search.html

Common Query Patterns

1. Structure Summary and Metadata

Get basic information about a structure including deposition date, revision date, authors, and experimental method:

python
import json

pdb_id = "4wa9"
url = f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/summary/{pdb_id}"
result = coot.coot_get_url_as_string_py(url)

# Handle both string and bytes responses
if isinstance(result, bytes):
   result = result.decode('utf-8')

data = json.loads(result)

# Extract key information
entry = data[pdb_id][0]
print(f"Title: {entry['title']}")
print(f"Release date: {entry['release_date']}")
print(f"Revision date: {entry['revision_date']}")
print(f"Method: {entry['experimental_method']}")
print(f"Authors: {entry['entry_authors']}")

Key fields in response:

  • title - Structure title
  • release_date - Original deposition date (YYYYMMDD)
  • revision_date - Most recent revision date (YYYYMMDD)
  • experimental_method - List of experimental methods
  • entry_authors - List of authors
  • number_of_entities - Count of different entity types (protein, ligand, water, etc.)
2. Molecule/Entity Information

Get organism and molecule details:

python
import json

pdb_id = "4wa9"
url = f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/molecules/{pdb_id}"
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

data = json.loads(result)

if pdb_id in data:
   for entity in data[pdb_id]:
      molecule_name = entity.get('molecule_name', ['N/A'])[0] if isinstance(entity.get('molecule_name', []), list) else entity.get('molecule_name', 'N/A')
      source = entity.get('source', [{}])[0] if isinstance(entity.get('source', []), list) else entity.get('source', {})
      organism = source.get('organism_scientific_name', 'N/A')
      expression_host = source.get('expression_host_scientific_name', 'N/A')
      
      print(f"Molecule: {molecule_name}")
      print(f"  Source organism: {organism}")
      if expression_host != 'N/A' and expression_host != organism:
         print(f"  Expression host: {expression_host}")
3. Compound/Ligand Information

Get detailed information about a specific compound including formula, SMILES, InChI, and revision history:

python
import json

comp_id = "AXI"  # 3-letter code
url = f"https://www.ebi.ac.uk/pdbe/api/pdb/compound/summary/{comp_id}"
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

data = json.loads(result)

compound = data[comp_id][0]
print(f"Name: {compound['name']}")
print(f"Formula: {compound['formula']}")
print(f"Weight: {compound['weight']}")
print(f"Creation date: {compound['creation_date']}")
print(f"Revision date: {compound['revision_date']}")
print(f"InChI: {compound['inchi']}")
print(f"SMILES: {compound['smiles'][0]['name']}")

Use case: Check if a ligand definition was recently revised, which might explain geometry changes.

4. Downloading Coordinate Files

Download and load PDB/mmCIF files directly:

python
import gzip

# Download current version from RCSB
pdb_id = "4wa9"
url = f"https://files.rcsb.org/download/{pdb_id}.cif.gz"

print(f"Downloading {pdb_id}...")
compressed = coot.coot_get_url_as_string_py(url)
print(f"Downloaded {len(compressed)} bytes (compressed)")

# Decompress
decompressed = gzip.decompress(compressed)
print(f"Decompressed to {len(decompressed)} bytes")

# Load into Coot
imol = coot.read_coordinates_as_string(decompressed.decode('utf-8'), f"{pdb_id}")
print(f"Loaded as molecule {imol}")

Note: The wwPDB versioned archive exists but is not currently accessible via HTTPS through this API. Use the current version from RCSB or PDBe.

5. Validation Reports

Get residue-wise outliers including clashes, geometry outliers, and density fit issues:

python
import json

pdb_id = "4wa9"
url = f"https://www.ebi.ac.uk/pdbe/api/validation/residuewise_outlier_summary/entry/{pdb_id}"
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

# Note: Response may have multiple JSON objects, parse carefully
# Access specific chain/residue validation data

Available validation endpoints:

  • /validation/residuewise_outlier_summary/entry/{pdb_id} - Residue-level outliers
  • /validation/rama_sidechain_listing/entry/{pdb_id} - Ramachandran and rotamer outliers
  • /validation/global_percentiles/entry/{pdb_id} - Overall quality metrics
6. Structure Status and Revision History

Check if a structure has been superseded or revised:

python
import json

pdb_id = "4wa9"
url = f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/status/{pdb_id}"
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

data = json.loads(result)

status = data[pdb_id][0]
print(f"Status: {status['status_code']}")  # REL = released, OBS = obsolete
print(f"Since: {status['since']}")
print(f"Superseded by: {status['superceded_by']}")
print(f"Obsoletes: {status['obsoletes']}")
7. Ligand Binding Sites

Get information about ligand binding sites and interactions:

python
import json

pdb_id = "4wa9"
url = f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/ligand_monomers/{pdb_id}"
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

data = json.loads(result)

# Access ligand information per chain
for entity in data[pdb_id]:
    print(f"Chain: {entity['chain_id']}")
    for ligand in entity.get('ligands', []):
        print(f"  Ligand: {ligand['chem_comp_id']}")
        print(f"  Residue: {ligand['author_residue_number']}")
8. Assembly Information

Get biological assembly information:

python
import json

pdb_id = "2hyy"
url = f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/assembly/{pdb_id}"
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

data = json.loads(result)

for assembly in data[pdb_id]:
    print(f"Assembly {assembly['assembly_id']}: {assembly['name']}")
    print(f"  Form: {assembly['form']}")
    print(f"  Preferred: {assembly['preferred']}")

Solr Search API

The Solr search API allows complex queries across the entire PDB. However, it has important limitations.

What Solr Search CAN Do (Well-Indexed Fields)

✅ Metadata searches:

  • By release/deposition date: release_year:2025
  • By experimental method: experimental_method:"X-ray diffraction"
  • By resolution: resolution:[* TO 2.0] or resolution:[1.5 TO 2.5]
  • By organism: organism_scientific_name:"Homo sapiens"

✅ Presence/absence queries:

  • Has protein: number_of_protein_chains:[1 TO *]
  • Has carbohydrate: has_carb_polymer:Y
  • Has bound molecules: has_bound_molecule:Y
  • Has modified residues: has_modified_residues:Y

✅ Component searches:

  • Specific ligand: chem_comp_id:ATP
  • Ligand name: ligand_name:imatinib
  • Molecule name: molecule_name:*kinase*

✅ Author/citation:

  • By author: entry_authors:"Smith J"
  • By UniProt: uniprot_accession:P12345

✅ Combined queries:

python
# Example: Human kinases with resolution < 2Å from 2024
query = 'release_year:2024 AND organism_scientific_name:"Homo sapiens" AND molecule_name:*kinase* AND resolution:[* TO 2.0]'
What Solr Search CANNOT Do

❌ Detailed connectivity: Cannot search for "THR covalently bonded to NAG" or other specific atom-level connections

❌ Geometry queries: Cannot search for "bonds longer than X" or "angles outside range Y"

❌ Spatial relationships: Cannot search for "atoms within 5Å of ligand"

❌ Sequence motifs: Cannot search for "structures with GXGXXG motif"

❌ Complex structural features: Cannot search for "beta-barrel with 8 strands"

❌ Validation specifics: Cannot search for "residues with Ramachandran outliers at position X"

The Pattern: Solr indexes metadata and simple categorical data, not structural details or relationships.

For analyses requiring connectivity or geometry (like finding O-glycosylated threonines), you must:

  1. Use Solr to find candidates (e.g., structures with NAG + resolution < 2.5Å)
  2. Download those structures
  3. Parse mmCIF connectivity tables locally
  4. Extract geometric parameters
Basic Search Syntax
python
import json

# Simple search for high-resolution X-ray structures from 2024
query = "release_year:2024 AND experimental_method:\"X-ray diffraction\" AND resolution:[* TO 1.5]"
url = f"https://www.ebi.ac.uk/pdbe/search/pdb/select?q={query}&wt=json&rows=10&fl=pdb_id,title,resolution"

result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

data = json.loads(result)

print(f"Found {data['response']['numFound']} structures")
for doc in data['response']['docs']:
    print(f"  {doc['pdb_id']}: {doc.get('resolution', 'N/A')}Å")
    print(f"    {doc.get('title', 'N/A')[:70]}")
Common Solr Search Fields

Identifiers & Metadata:

  • pdb_id - PDB entry ID
  • molecule_name - Molecule name
  • molecule_type - Entity type: use "Protein" (capital P) — NOT polypeptide(L) or protein
  • molecule_sequence - One-letter sequence string (stored but not full-text indexed — wildcard search like *C*C* returns 0 results; fetch and filter in Python instead)
  • polymer_length - Length of the polymer entity in residues (supports range queries: [5 TO 30])
  • number_of_polymer_residues - Total residues across all chains in the entry
  • number_of_protein_chains - Number of protein chains
  • organism_scientific_name - Source organism
  • experimental_method - Experimental method (e.g., "X-ray diffraction")
  • resolution - Structure resolution
  • ligand_name - Ligand/compound name
  • citation_title - Publication title
  • deposition_date - Deposition date
  • revision_date - Revision date

Experimental Details:

  • experimental_method - Method (e.g., "X-ray diffraction", "Electron Microscopy", "Solution NMR")
  • resolution - Structure resolution (numeric, use ranges like [1.0 TO 2.0])
  • em_resolution - EM-specific resolution
  • data_quality - Overall quality metric

Molecular Content:

  • molecule_name - Molecule name (supports wildcards: *kinase*)
  • molecule_type - Type (Protein, DNA, RNA, etc.)
  • organism_scientific_name - Source organism
  • organism_synonyms - Alternative organism names
  • genus - Organism genus
  • expression_host_scientific_name - Expression system

Ligands & Modifications:

  • chem_comp_id - Chemical component 3-letter code
  • ligand_name - Ligand name
  • has_bound_molecule - Y/N
  • has_carb_polymer - Y/N (has carbohydrate)
  • has_modified_residues - Y/N
  • number_of_bound_molecules - Count

Authors & Citations:

  • entry_authors - Entry authors
  • citation_authors - Publication authors
  • citation_title - Paper title
  • citation_year - Publication year
  • pubmed_id - PubMed ID

Protein Details:

  • uniprot_accession - UniProt accession
  • uniprot_id - UniProt ID
  • gene_name - Gene name
  • go_id - Gene Ontology ID

Structure Properties:

  • number_of_protein_chains - Count
  • number_of_polymer_entities - Count
  • assembly_composition - Assembly type
  • symmetry_group - Symmetry
Practical Search Examples

Example 1: High-resolution X-ray structures from 2024

python
import json

url = "https://www.ebi.ac.uk/pdbe/search/pdb/select?q=release_year:2024%20AND%20experimental_method:\"X-ray%20diffraction\"%20AND%20resolution:[*%20TO%201.5]&wt=json&rows=5&fl=pdb_id,title,resolution"
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

Sorting by `sort=molecular_weight+asc` gives a 400 error. Use `polymer_length` instead
as a proxy for size, or use `number_of_polymer_residues` for entry-level size.

**5. Use `fq` (filter query) for range constraints**

Range filtering on numeric fields works well as a filter query:
```python
# Filter to entities with 5-30 residues:
url = "...&q=molecule_type:Protein&fq=polymer_length:[5+TO+30]&sort=polymer_length+asc..."

Example 2: Human kinase structures

python
query = "organism_scientific_name:\"Homo sapiens\" AND molecule_name:*kinase*"
url = f"https://www.ebi.ac.uk/pdbe/search/pdb/select?q={query}&wt=json&rows=10&fl=pdb_id,title,resolution"

Example 3: Cryo-EM structures better than 3Å from 2025

python
query = "release_year:2025 AND experimental_method:\"Electron Microscopy\" AND resolution:[* TO 3.0]"
url = f"https://www.ebi.ac.uk/pdbe/search/pdb/select?q={query}&wt=json&rows=10&fl=pdb_id,title,em_resolution"

Example 4: Structures with carbohydrates

python
query = "has_carb_polymer:Y"
url = f"https://www.ebi.ac.uk/pdbe/search/pdb/select?q={query}&wt=json&rows=10&fl=pdb_id,title"

Example 5: Structures of a specific protein from different species

python
# Find ABL1 structures from different mammals
query = "molecule_name:*ABL1* OR molecule_name:*ABL*kinase*"
url = f"https://www.ebi.ac.uk/pdbe/search/pdb/select?q={query}&wt=json&rows=100&fl=pdb_id,organism_scientific_name,title"

6. Discover available fields by fetching a sample document with fl=*

When you don't know what fields are in the index:

python
url = "https://www.ebi.ac.uk/pdbe/search/pdb/select?q=*:*&wt=json&rows=1&fl=*"
data = json.loads(coot.coot_get_url_as_string_py(url))
for k in sorted(data['response']['docs'][0].keys()):
    print(k)

Fields prefixed q_ and t_ are query/text variants of the base fields — ignore them when exploring the schema.

7. There is no disulfide or bond_types field in the Solr index

To find structures with disulfide bonds, you must:

  • Use Solr to find small proteins with ≥2 Cys in their sequence (fetch + filter in Python), then
  • Use the PDBe REST API (/pdb/entry/molecules/{pdb_id}) to confirm the sequence and structure.
Show full SKILL.md (471 more words)Show less
Advanced Search Examples

Find structures with specific ligand:

python
query = "ligand_name:axitinib"

Find high-resolution kinase structures:

python
query = "molecule_name:kinase AND resolution:[0 TO 2.0]"

Find structures revised in 2024:

python
query = "revision_date:[20240101 TO 20241231]"

Practical Workflows

Detecting Structure Revisions

Check if a structure has been significantly revised since release:

python
import json
from datetime import datetime

def check_structure_revision(pdb_id):
    """Check if structure was revised and when"""
    url = f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/summary/{pdb_id}"
    result = coot.coot_get_url_as_string_py(url)
    
    if isinstance(result, bytes):
        result = result.decode('utf-8')
    
    data = json.loads(result)
    
    entry = data[pdb_id][0]
    release = entry['release_date']
    revision = entry['revision_date']
    
    # Convert to datetime for comparison
    release_dt = datetime.strptime(release, "%Y%m%d")
    revision_dt = datetime.strptime(revision, "%Y%m%d")
    
    days_diff = (revision_dt - release_dt).days
    years_diff = days_diff / 365.25
    
    print(f"PDB {pdb_id}:")
    print(f"  Released: {release}")
    print(f"  Revised: {revision}")
    print(f"  Time since release: {years_diff:.1f} years")
    
    if days_diff > 30:
        print(f"  WARNING: Structure revised {days_diff} days after release")
        return True
    return False

# Example usage
check_structure_revision("4wa9")
Checking Ligand Revisions

Determine if a ligand definition was updated, which might explain geometry changes:

python
import json

def check_ligand_revision(comp_id):
    """Check when a ligand was last revised"""
    url = f"https://www.ebi.ac.uk/pdbe/api/pdb/compound/summary/{comp_id}"
    result = coot.coot_get_url_as_string_py(url)
    
    if isinstance(result, bytes):
        result = result.decode('utf-8')
    
    data = json.loads(result)
    
    compound = data[comp_id][0]
    print(f"Compound {comp_id} ({compound['name']}):")
    print(f"  Created: {compound['creation_date']}")
    print(f"  Revised: {compound['revision_date']}")
    
    if compound['creation_date'] != compound['revision_date']:
        print(f"  WARNING: Ligand definition was revised")
        return True
    return False

# Example usage
check_ligand_revision("AXI")
Downloading and Comparing Structures

Download structures from different species and compare them:

python
import gzip
import json

def download_and_load_structure(pdb_id):
    """Download and load a structure from RCSB"""
    url = f"https://files.rcsb.org/download/{pdb_id}.cif.gz"
    
    print(f"Downloading {pdb_id}...")
    compressed = coot.coot_get_url_as_string_py(url)
    
    # Check if it's an error response (HTML)
    if isinstance(compressed, str) and compressed.startswith("<!DOCTYPE"):
        print(f"ERROR: Could not download {pdb_id}")
        return None
    
    decompressed = gzip.decompress(compressed)
    imol = coot.read_coordinates_as_string(decompressed.decode('utf-8'), pdb_id)
    print(f"Loaded as molecule {imol}")
    
    return imol

def compare_species_structures(pdb_id1, pdb_id2):
    """Download two structures and superpose them"""
    # Download both structures
    imol1 = download_and_load_structure(pdb_id1)
    imol2 = download_and_load_structure(pdb_id2)
    
    if imol1 is None or imol2 is None:
        print("Failed to download one or both structures")
        return
    
    # Superpose (using CA atoms from chain A, residues 240-400)
    print(f"\nSuperposing {pdb_id2} onto {pdb_id1}...")
    sel1 = "//A/240-400/CA"
    sel2 = "//A/240-400/CA"
    result = coot.superpose_with_atom_selection(imol1, imol2, sel1, sel2, 0)
    
    if result >= 0:
        print(f"Success! Structures superposed.")
    else:
        print("Superposition failed!")
    
    return imol1, imol2

# Example: Compare human and mouse ABL1
# First find structures using Solr
query = "molecule_name:*ABL1*"
url = f"https://www.ebi.ac.uk/pdbe/search/pdb/select?q={query}&wt=json&rows=100&fl=pdb_id,organism_scientific_name"
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
   result = result.decode('utf-8')

data = json.loads(result)

# Find human and mouse structures
human_pdbs = []
mouse_pdbs = []
for doc in data['response']['docs']:
   org = doc.get('organism_scientific_name', ['Unknown'])[0]
   if 'Homo sapiens' in org:
      human_pdbs.append(doc['pdb_id'])
   elif 'Mus musculus' in org:
      mouse_pdbs.append(doc['pdb_id'])

print(f"Human ABL1 structures: {len(human_pdbs)}")
print(f"Mouse ABL1 structures: {len(mouse_pdbs)}")

# Compare first human and mouse structures
if human_pdbs and mouse_pdbs:
   compare_species_structures(human_pdbs[0], mouse_pdbs[0])

Search for structures with the same ligand and protein:

python
import json

def find_related_structures(protein_name, ligand_name=None):
    """Find structures containing specific protein-ligand combination"""
    if ligand_name:
        query = f'molecule_name:*{protein_name}* AND chem_comp_id:{ligand_name}'
    else:
        query = f'molecule_name:*{protein_name}*'
    
    url = f"https://www.ebi.ac.uk/pdbe/search/pdb/select?q={query}&wt=json&rows=50&fl=pdb_id,title,resolution,organism_scientific_name"
    
    result = coot.coot_get_url_as_string_py(url)
    
    if isinstance(result, bytes):
        result = result.decode('utf-8')
    
    data = json.loads(result)
    
    print(f"Found {data['response']['numFound']} structures")
    for doc in data['response']['docs']:
        org = doc.get('organism_scientific_name', ['N/A'])
        if isinstance(org, list):
            org = org[0] if org else 'N/A'
        
        print(f"  {doc['pdb_id']}: {doc.get('title', 'N/A')[:60]}")
        print(f"    Resolution: {doc.get('resolution', 'N/A')} Å")
        print(f"    Organism: {org}")

# Example usage
find_related_structures("ABL1", "STI")  # ABL1 with imatinib

Error Handling

Always wrap API calls in try/except blocks and handle both string and bytes responses:

python
import json

def safe_pdbe_query(url):
    """Safely query PDBe API with error handling"""
    try:
        result = coot.coot_get_url_as_string_py(url)
        if not result or result == "":
            print(f"Empty response from {url}")
            return None
        
        # Handle bytes response
        if isinstance(result, bytes):
            result = result.decode('utf-8')
        
        # Check for HTML error pages
        if result.startswith("<!DOCTYPE") or result.startswith("<html"):
            print(f"Received HTML error page instead of JSON")
            print(result[:200])
            return None
        
        data = json.loads(result)
        return data
    
    except json.JSONDecodeError as e:
        print(f"JSON parsing error: {e}")
        print(f"Response was: {result[:200]}...")
        return None
    
    except Exception as e:
        print(f"Error querying PDBe API: {e}")
        return None

# Example usage
data = safe_pdbe_query("https://www.ebi.ac.uk/pdbe/api/pdb/entry/summary/4wa9")
if data:
    print("Success!")

Common Issues and Solutions

Issue: Binary vs Text Data

The function returns bytes for binary data (gzipped files) and str for text (JSON). Always check the type:

python
result = coot.coot_get_url_as_string_py(url)

if isinstance(result, bytes):
    # Binary data - might be gzipped
    if result.startswith(b'\x1f\x8b'):  # gzip magic bytes
        import gzip
        decompressed = gzip.decompress(result)
        content = decompressed.decode('utf-8')
    else:
        content = result.decode('utf-8')
else:
    # Already a string
    content = result
Issue: JSON parsing errors with validation endpoints

Some validation endpoints return multiple JSON objects or malformed responses. Handle carefully:

python
# Instead of json.loads(), parse line by line or handle errors
try:
    data = json.loads(result)
except json.JSONDecodeError:
    # Try alternative parsing or just display raw result
    print("Could not parse JSON, raw response:")
    print(result[:1000])
Issue: Unicode decode errors

If you get UnicodeDecodeError, the response might contain non-UTF-8 bytes. This should be handled automatically by the function now, but if you encounter issues:

python
try:
    result = coot.coot_get_url_as_string_py(url)
except Exception as e:
    print(f"Error fetching URL: {e}")
Issue: URL encoding for complex queries

Always encode special characters in Solr queries:

python
import urllib.parse

query = "molecule_name:\"Protein kinase\" AND resolution:[0 TO 2.0]"
encoded = urllib.parse.quote(query)
url = f"https://www.ebi.ac.uk/pdbe/search/pdb/select?q={encoded}&wt=json"
Issue: Rate limiting

The PDBe API may rate limit excessive requests. Add delays between batch queries:

python
import time

pdb_ids = ["4wa9", "2hyy", "1iep"]
for pdb_id in pdb_ids:
    data = safe_pdbe_query(f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/summary/{pdb_id}")
    # Process data...
    time.sleep(0.5)  # Wait 500ms between requests

Quick Reference

Most Useful Endpoints
PurposeEndpoint
Structure summary/pdb/entry/summary/{pdb_id}
Molecule/organism info/pdb/entry/molecules/{pdb_id}
Compound info/pdb/compound/summary/{comp_id}
Validation outliers/validation/residuewise_outlier_summary/entry/{pdb_id}
Structure status/pdb/entry/status/{pdb_id}
Ligand binding sites/pdb/entry/ligand_monomers/{pdb_id}
Search structures/search/pdb/select?q={query}
Download coordinateshttps://files.rcsb.org/download/{pdb_id}.cif.gz
Common Solr Query Patterns
QueryPurpose
pdb_id:4wa9Specific PDB entry
molecule_name:*kinase*By protein name (wildcards)
chem_comp_id:ATPStructures with specific ligand
resolution:[0 TO 2.0]High resolution structures
release_year:2025Structures from 2025
revision_year:2024Recently revised structures
experimental_method:"X-ray diffraction"By experimental method
organism_scientific_name:"Homo sapiens"By organism
has_carb_polymer:YHas carbohydrate
has_bound_molecule:YHas ligands
Combining Queries with AND/OR
python
# Human kinases with resolution < 2Å from 2024
query = 'release_year:2024 AND organism_scientific_name:"Homo sapiens" AND molecule_name:*kinase* AND resolution:[* TO 2.0]'

# ABL1 from human OR mouse
query = 'molecule_name:*ABL1* AND (organism_scientific_name:"Homo sapiens" OR organism_scientific_name:"Mus musculus")'

Integration with Coot Workflows

Example: Automated Structure Quality Check
python
import json

def structure_quality_report(imol):
    """Generate quality report using PDBe API data"""
    
    # Get PDB ID from molecule
    pdb_file = coot.molecule_name(imol)
    # Extract PDB ID from filename (assumes format like "pdb4wa9.ent" or "4wa9")
    import re
    match = re.search(r'(\d\w{3})', pdb_file.lower())
    if not match:
        print("Could not extract PDB ID from filename")
        return
    
    pdb_id = match.group(1)
    
    # Get structure info
    data = safe_pdbe_query(f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/summary/{pdb_id}")
    if not data:
        return
    
    entry = data[pdb_id][0]
    
    print("=" * 60)
    print(f"STRUCTURE QUALITY REPORT: {pdb_id.upper()}")
    print("=" * 60)
    print(f"Title: {entry['title']}")
    print(f"Method: {entry['experimental_method']}")
    print(f"Released: {entry['release_date']}")
    print(f"Revised: {entry['revision_date']}")
    
    # Check for significant revisions
    if entry['revision_date'] != entry['release_date']:
        from datetime import datetime
        release = datetime.strptime(entry['release_date'], "%Y%m%d")
        revision = datetime.strptime(entry['revision_date'], "%Y%m%d")
        days = (revision - release).days
        print(f"\n⚠️  STRUCTURE REVISED {days} days after release")
        print("   Check PDBe for revision details")
    
    print("=" * 60)

# Usage: structure_quality_report(0)

Resources

Summary

The PDBe API provides rich programmatic access to structure metadata, validation data, and search capabilities. Using coot.coot_get_url_as_string_py(), you can:

  1. Download coordinate files - Get structures in mmCIF/PDB format (gzipped)
  2. Check revision history - Detect structures and ligands that have been revised
  3. Access validation reports - Get quality metrics and outlier information
  4. Search across the PDB - Find related structures, compare organisms, filter by properties
  5. Get compound information - Access chemical details, SMILES, InChI
  6. Verify structure status - Check for supersession or obsolescence
  7. Integrate external data - Bring PDB metadata into Coot workflows

Key Capabilities:

  • Binary data support (download gzipped files)
  • Comprehensive metadata access
  • Powerful search with well-understood limitations
  • Cross-species structure comparison
  • Revision tracking and provenance checking

Key Limitations:

  • Solr search cannot query detailed connectivity or geometry
  • Versioned coordinates not accessible via HTTPS (use current versions)
  • For analyses requiring atom-level connectivity, download and parse structures locally

This enables powerful automated quality checks, cross-species structure comparison, data-driven validation, and integration of PDB metadata into Coot-based structural biology workflows.

© pemsley, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in mcp/docs/skills/pdbe-api of pemsley/coot.

Open the folder on GitHubat commit 6e3c026

Compare with similar skills

Pdbe API next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pdbe API compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pdbe API this skillpemsley/coot168—~7.5kAutomated safety check: PassGPL-3.0
Uniprot Databaseaipoch/medical-research-skills2k—~1.3kAutomated safety check: PassMIT
String Database Ppijaechang-hits/SciAgent-Skills3701 repos~4.2kAutomated safety check: PassCC-BY-4.0
Bio Clinical Databases Clinvar LookupFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~1.4kAutomated safety check: PassNone
Ena Databaseaipoch/medical-research-skills2k—~1.9kAutomated safety check: PassMIT
Pride Databasemajiayu000/claude-skill-registry6662 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • Uniprot Database

    aipoch/medical-research-skills

    Direct REST API access to UniProt for protein search, entry retrieval, and identifier mapping; use when you need programmatic UniProtKB queries or cross-database ID conversion.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Backend & APIsAuto-check passed
  • String Database Ppi

    jaechang-hits/SciAgent-Skills

    Query STRING REST API for PPIs (59M proteins, 20B interactions, 5000+ species).

    370 GitHub starsUsed in 1 repo~4.2k tokens
    Backend & APIsAuto-check passed
  • Bio Clinical Databases Clinvar Lookup

    FreedomIntelligence/OpenClaw-Medical-Skills

    Query ClinVar for variant pathogenicity classifications, review status, and disease associations via REST API or local VCF.

    3.1k GitHub stars~1.4k tokensUpdated 2 mo ago
    Backend & APIsAuto-check passed
  • Ena Database

    aipoch/medical-research-skills

    Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need…

    2k GitHub stars~1.9k tokensUpdated 21 days ago
    Backend & APIsAuto-check passed
  • Pride Database

    majiayu000/claude-skill-registry

    Search the PRIDE Archive v3 REST API for proteomics datasets: discover projects by keyword + faceted filters (organism, instrument, disease, software), fetch project metadata, list and download…

    666 GitHub starsUsed in 2 repos~8.2k tokens
    Backend & APIsAuto-check passed
  • Claude To Medrixflow

    Citrus-bit/Anaxa

    Interact with MedrixFlow AI agent platform via its HTTP API.

    120 GitHub stars~1.7k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from pemsley/coot

All 11 skills in this repo
  • Coot Inline Graphs

    pemsley/coot

    Create interactive inline Chart.js graphs directly in the chat from live Coot data.

    168 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Coot Rdkit

    pemsley/coot

    RDKit molecular manipulation and visualization within Coot's Python environment.

    168 GitHub stars~981 tokensUpdated today
    Auto-check passed
  • Coot Refinement

    pemsley/coot

    Best practices for protein structure refinement and validation in Coot.

    168 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Coot Essential API

    pemsley/coot

    API documentation to be loaded at startup - when starting a Coot session, immediately call getfunctiondescriptions() with the functions listed in this skill.

    168 GitHub stars~4.5k tokensUpdated today
    Auto-check passed
  • Coot Figure Making

    pemsley/coot

    Best practices for creating publication-quality molecular graphics figures in Coot using user-defined colors, ribbons, and molecular representations

    168 GitHub stars~7.3k tokensUpdated today
    Auto-check passed
  • Best Practices for Model-Building Tools and Refinement. An agent skill from pemsley/coot.

    168 GitHub stars~13k tokensUpdated today
    Auto-check passed

Questions about Pdbe API

What does Pdbe API do?

Query the PDBe (Protein Data Bank in Europe) REST API and Solr search API from within Coot to access structure metadata, validation data, revision history, search capabilities, and download…. Pdbe API is an agent skill from pemsley/coot.

When should I use Pdbe API?

Pdbe API fits situations like: tasks that involve REST APIs.

How do I install Pdbe API in Claude Code?

Run `npx skills add pemsley/coot --skill pdbe-api -a claude-code`. Or copy the skill folder (mcp/docs/skills/pdbe-api in pemsley/coot) into .claude/skills/pdbe-api in your project. Claude Code loads it when a task matches its description.

How do I install Pdbe API in Codex?

Run `npx skills add pemsley/coot --skill pdbe-api -a codex`. Or copy the skill folder (mcp/docs/skills/pdbe-api in pemsley/coot) into .agents/skills/pdbe-api in your project. Codex loads it when a task matches its description.

Can I use Pdbe API in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pemsley/coot --skill pdbe-api -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdbe-api, .gemini/skills/pdbe-api, .github/skills/pdbe-api and .opencode/skills/pdbe-api in your project.

What does Pdbe API need to run?

SKILL.md names no scripts, command-line tools or credentials: Pdbe API is instructions for the agent only. Our summary lists: Python 3.

Does Pdbe API access the network?

SKILL.md names 4 domains. In commands or code: ebi.ac.uk and files.rcsb.org; the agent is likely to contact these when it follows the instructions. As links in the text: pdbe.org and github.com. This is read from the text; nothing was executed.

Is Pdbe API safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pdbe API use?

Pdbe API is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pdbe API use?

About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pdbe API?

Skills that share tags, products or a category with Pdbe API: Uniprot Database (aipoch/medical-research-skills, 2k stars), String Database Ppi (jaechang-hits/SciAgent-Skills, 370 stars), Bio Clinical Databases Clinvar Lookup (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars) and Ena Database (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pdbe API?

pemsley (a GitHub user) maintains it in pemsley/coot, which has 168 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 7, 2026.

Source: pemsley/coot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.