Agent skill

Pdb Database

by jaechang-hits in jaechang-hits/SciAgent-Skills

Query RCSB PDB (200K+ structures) via the public REST + GraphQL APIs with plain requests (no SDK).

BSD-3-ClauseAuto-check passedResearch & Science

Install Pdb Database

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill pdb-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills pdb-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pdb-database .claude/skills/pdb-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdb-database
GitHub stars
370
Used in
1 other repo
Token cost
~7.7k tokens
SKILL.md length
1,353 words
Files
1
Skills in repo
163
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Query RCSB PDB (200K+ structures) via the public REST + GraphQL APIs with plain requests (no SDK).

  • Works in 7 steps: Search → fetch: Use the Search API to… → Use full_text vs text deliberately:… → mmCIF over PDB format: PDB format is… → …
  • Tasks that involve Protein structure and design
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 6 more sections
  • Calls pip; reaches search.rcsb.org and data.rcsb.org

What it does

Pdb Database is an agent skill from jaechang-hits/SciAgent-Skills. Query RCSB PDB (200K+ structures) via the public REST + GraphQL APIs with plain requests (no SDK). Search by text, attribute, sequence, or 3D structure similarity (Search API); retrieve metadata via GraphQL (Data API); download PDB/mmCIF from files.rcsb.org. For AlphaFold predictions use alphafold-database-access; for protein sequences only use uniprot-protein-database.

Its SKILL.md is about 7.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Protein structure and design, GraphQL and Vector databases. It works with AlphaFold, GraphQL and UniProt. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.

When your agent uses it

  • Tasks that involve Protein structure and design
  • Tasks that involve GraphQL
  • Tasks that involve Vector databases

Example prompts

  • “/pdb-database”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Search → fetch: Use the Search API to get a list of IDs, then GraphQL entries(entry_ids: [...]) for batch metadata. Avoid one GraphQL…
  2. Use full_text vs text deliberately: free-text keyword search needs "service": "full_text". Structured attribute filters need "service"…
  3. mmCIF over PDB format: PDB format is being phased out and has a 99,999 atom limit. Always download .cif for new code.
  4. Set realistic paginate.rows: rows: 100 is a good default for batch work; the API may slow down beyond ~10000. Loop with paginate.start for…
  5. Rate limit with time.sleep(0.2) in batch loops: No published hard cap, but the public infrastructure is shared. On HTTP 429, back off…
  6. Inspect the payload before posting: print(json.dumps(payload, indent=2)) is the cheapest way to debug HTTP 400 errors.
  7. entries(entry_ids: [...]) does not validate every ID: if one ID is wrong, the whole array returns null entries. Validate IDs separately if…

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • search.rcsb.org
    • data.rcsb.org
    • files.rcsb.org
    • rcsb.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pdb Database loads about 7.7k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 1,353 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~7.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,353 words, ~7,746 tokens.

Download SKILL.mdSave it as .claude/skills/pdb-database/SKILL.md (or your agent's skills folder).
name
pdb-database
description
Query RCSB PDB (200K+ structures) via the public REST + GraphQL APIs with plain `requests` (no SDK). Search by text, attribute, sequence, or 3D structure similarity (Search API); retrieve metadata via GraphQL (Data API); download PDB/mmCIF from files.rcsb.org. For AlphaFold predictions use alphafold-database-access; for protein sequences only use uniprot-protein-database.
license
BSD-3-Clause

PDB Database

Why no SDK? The rcsb-api Python SDK is convenient sugar over three public, no-auth REST endpoints (search.rcsb.org, data.rcsb.org, files.rcsb.org). When the SDK is unavailable, every operation can be reproduced with plain requests and a small JSON payload. This SKILL.md uses the REST path throughout so the code runs in any environment with requests installed.

Overview

RCSB PDB is the worldwide repository for 3D structural data of biological macromolecules with 200,000+ experimentally determined structures. Programmatic access is via three free, no-auth endpoints:

APIBase URLMethodPurpose
Searchhttps://search.rcsb.org/rcsbsearch/v2/queryPOST JSONFind PDB IDs by text, attribute filters, sequence, or 3D similarity
Datahttps://data.rcsb.org/graphqlPOST GraphQLRetrieve structured metadata (entries, polymer entities, assemblies, ligands)
Fileshttps://files.rcsb.org/download/{id}.{format}GETDownload coordinate files (mmCIF, PDB, FASTA)

Use this skill for programmatic structural biology queries, drug target analysis, and protein family comparisons.

When to Use

  • Searching for protein or nucleic acid crystal/cryo-EM/NMR structures by keyword or property
  • Finding structures similar to a query sequence (MMseqs2) or 3D geometry (BioZernike)
  • Retrieving experimental metadata (resolution, method, organism, deposition date) for structure sets
  • Downloading coordinate files (PDB, mmCIF) for molecular dynamics, docking, or visualization
  • Building structure-based datasets for machine learning or drug discovery pipelines
  • Comparing protein-ligand complexes across a target family
  • For AlphaFold predicted structures, use alphafold-database-access instead
  • For protein sequence/annotation queries without structures, use uniprot-protein-database instead

Prerequisites

  • Python packages: requests (only requirement). Optional: biopython for parsing downloaded coordinate files.
  • No API key required: RCSB PDB is freely accessible.
  • Rate limits: No published hard limit. Polite delays of time.sleep(0.2-0.5) between requests are sufficient; implement exponential backoff on HTTP 429.
bash
pip install requests
# Optional, for coordinate parsing:
pip install biopython

Quick Start

Typical search-then-fetch pattern: hit the Search API, get a list of PDB IDs, then resolve metadata via the GraphQL Data API.

python
import requests

SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"
DATA   = "https://data.rcsb.org/graphql"

# 1. Search: human X-ray structures of "kinase" at resolution < 2.0 Å
payload = {
    "query": {
        "type": "group", "logical_operator": "and",
        "nodes": [
            {"type": "terminal", "service": "full_text",
             "parameters": {"value": "kinase"}},
            {"type": "terminal", "service": "text",
             "parameters": {"attribute": "rcsb_entity_source_organism.scientific_name",
                            "operator": "exact_match", "value": "Homo sapiens"}},
            {"type": "terminal", "service": "text",
             "parameters": {"attribute": "rcsb_entry_info.resolution_combined",
                            "operator": "less", "value": 2.0}},
        ],
    },
    "return_type": "entry",
    "request_options": {"paginate": {"rows": 10}},
}
r = requests.post(SEARCH, json=payload, timeout=30)
r.raise_for_status()
result = r.json()
pdb_ids = [hit["identifier"] for hit in result["result_set"]]
print(f"Total matches: {result['total_count']}, first batch: {pdb_ids}")

# 2. Fetch metadata for the first hit via GraphQL
gql = """{ entry(entry_id: "%s") {
  struct { title }
  exptl { method }
  rcsb_entry_info { resolution_combined deposited_atom_count polymer_entity_count }
} }""" % pdb_ids[0]
r2 = requests.post(DATA, json={"query": gql}, timeout=30)
entry = r2.json()["data"]["entry"]
print(entry["struct"]["title"])
print(f"Method: {entry['exptl'][0]['method']}, Resolution: {entry['rcsb_entry_info']['resolution_combined']} Å")

Core API

Free-text search uses service: "full_text" and searches across all indexed fields.

python
import requests
SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"

def text_search(keyword, rows=25):
    payload = {
        "query": {"type": "terminal", "service": "full_text",
                  "parameters": {"value": keyword}},
        "return_type": "entry",
        "request_options": {"paginate": {"rows": rows}},
    }
    r = requests.post(SEARCH, json=payload, timeout=30)
    r.raise_for_status()
    data = r.json()
    return [hit["identifier"] for hit in data["result_set"]], data["total_count"]

ids, total = text_search("hemoglobin")
print(f"Found {total} structures; first batch: {ids[:5]}")

Attribute search uses service: "text" with structured attribute/operator/value parameters.

python
import requests
SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"

def attribute_search(attribute, operator, value, return_type="entry", rows=25):
    payload = {
        "query": {"type": "terminal", "service": "text",
                  "parameters": {"attribute": attribute,
                                 "operator": operator,
                                 "value": value}},
        "return_type": return_type,
        "request_options": {"paginate": {"rows": rows}},
    }
    r = requests.post(SEARCH, json=payload, timeout=30)
    r.raise_for_status()
    return r.json()

# Human proteins
human = attribute_search("rcsb_entity_source_organism.scientific_name",
                          "exact_match", "Homo sapiens", rows=5)
print(f"Human structures: {human['total_count']}")

# X-ray only
xray = attribute_search("exptl.method", "exact_match", "X-RAY DIFFRACTION", rows=5)
print(f"X-ray structures: {xray['total_count']}")

# Resolution range: 1.5–2.5 Å
res = attribute_search(
    "rcsb_entry_info.resolution_combined", "range",
    {"from": 1.5, "to": 2.5, "include_lower": True, "include_upper": True},
    rows=5
)
print(f"1.5–2.5 Å: {res['total_count']}")

Find structures with similar sequences using MMseqs2. Service is "sequence"; target selects protein vs. nucleic acid.

python
import requests
SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"

kras_seq = ("MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGQ"
            "EEYSAMRDQYMRTGEGFLCVFAINNTKSFEDIHHYREQIKRVKDSEDVPMVLVGNKCDLPS"
            "RTVDTKQAQDLARSYGIPFIETSAKTRQGVDDAFYTLVREIRKHKEKMSK")

payload = {
    "query": {
        "type": "terminal", "service": "sequence",
        "parameters": {
            "target": "pdb_protein_sequence",  # or "pdb_dna_sequence", "pdb_rna_sequence"
            "value": kras_seq,
            "evalue_cutoff": 0.1,
            "identity_cutoff": 0.9,
        },
    },
    "return_type": "polymer_entity",
    "request_options": {"paginate": {"rows": 10}},
}
r = requests.post(SEARCH, json=payload, timeout=30)
r.raise_for_status()
data = r.json()
print(f"KRAS-like hits: {data['total_count']}")
for hit in data["result_set"][:5]:
    print(f"  {hit['identifier']}  score={hit.get('score', 'n/a')}")

Find structures with similar 3D geometry using BioZernike descriptors. Service is "structure"; pass the reference entry + assembly ID.

python
import requests
SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"

payload = {
    "query": {
        "type": "terminal", "service": "structure",
        "parameters": {
            "value": {"entry_id": "4HHB", "assembly_id": "1"},
            "operator": "strict_shape_match",  # or "relaxed_shape_match"
        },
    },
    "return_type": "polymer_entity",
    "request_options": {"paginate": {"rows": 10}},
}
r = requests.post(SEARCH, json=payload, timeout=30)
r.raise_for_status()
data = r.json()
print(f"Structurally similar to 4HHB: {data['total_count']}")
for hit in data["result_set"][:5]:
    print(f"  {hit['identifier']}  score={hit.get('score', 'n/a')}")
Module 4: Data Retrieval (GraphQL)

The GraphQL endpoint at data.rcsb.org/graphql is the canonical way to retrieve structured metadata for known PDB IDs. One request can pull fields across the full data hierarchy (entry → polymer_entity → assembly → chem_comp).

python
import requests
DATA = "https://data.rcsb.org/graphql"

# Entry-level metadata
gql = """{ entry(entry_id: "4HHB") {
  struct { title }
  exptl { method }
  rcsb_entry_info { resolution_combined deposited_atom_count polymer_entity_count nonpolymer_entity_count }
  rcsb_accession_info { deposit_date initial_release_date }
} }"""
r = requests.post(DATA, json={"query": gql}, timeout=30)
entry = r.json()["data"]["entry"]
print(f"Title       : {entry['struct']['title']}")
print(f"Method      : {entry['exptl'][0]['method']}")
print(f"Resolution  : {entry['rcsb_entry_info']['resolution_combined']} Å")
print(f"Atoms       : {entry['rcsb_entry_info']['deposited_atom_count']}")
python
# Polymer entity (sequence, organism, MW)
gql = """{ polymer_entity(entry_id: "4HHB", entity_id: "1") {
  entity_poly { pdbx_seq_one_letter_code }
  rcsb_polymer_entity { formula_weight }
  rcsb_entity_source_organism { scientific_name ncbi_taxonomy_id }
} }"""
r = requests.post(DATA, json={"query": gql}, timeout=30)
pe = r.json()["data"]["polymer_entity"]
print(f"Sequence (first 50): {pe['entity_poly']['pdbx_seq_one_letter_code'][:50]}")
print(f"Organism : {pe['rcsb_entity_source_organism'][0]['scientific_name']}")
print(f"MW       : {pe['rcsb_polymer_entity']['formula_weight']}")
python
# Batch: pull metadata for many entries in one request
gql = """{ entries(entry_ids: ["4HHB", "1A3N", "1HHB"]) {
  rcsb_id
  struct { title }
  exptl { method }
  rcsb_entry_info { resolution_combined }
} }"""
r = requests.post(DATA, json={"query": gql}, timeout=30)
for e in r.json()["data"]["entries"]:
    res = e["rcsb_entry_info"]["resolution_combined"]
    print(f"  {e['rcsb_id']}: {e['exptl'][0]['method']:<25} {res} Å  — {e['struct']['title'][:40]}")
Module 5: File Download

Coordinate files (mmCIF, PDB, FASTA, assembly variants) are served directly from files.rcsb.org.

python
import requests

def download_structure(pdb_id, fmt="cif", output_dir="."):
    """Download mmCIF / PDB / FASTA. URLs: .pdb, .cif, /fasta/entry/{ID}, .pdb1 (assembly)."""
    url = f"https://files.rcsb.org/download/{pdb_id}.{fmt}"
    r = requests.get(url, timeout=60)
    if r.status_code == 200:
        path = f"{output_dir}/{pdb_id}.{fmt}"
        # mmCIF / PDB are text; assemblies and biological units are also text
        with open(path, "w") as f:
            f.write(r.text)
        print(f"Downloaded {path}  ({len(r.text)/1024:.1f} KB)")
        return path
    print(f"HTTP {r.status_code} for {pdb_id}.{fmt}")
    return None

download_structure("4HHB", fmt="cif")
download_structure("4HHB", fmt="pdb")
python
# FASTA sequence for an entry
r = requests.get("https://www.rcsb.org/fasta/entry/4HHB", timeout=30)
r.raise_for_status()
print(r.text[:400])
Module 6: Query Composition (group + logical_operator)

Combine terminal queries with type: "group" and a logical_operator of "and" / "or". Nested groups give arbitrary boolean expressions; negation is via "node_id" references with "operator": "negate" on the group (rare — usually expressed as the inverse attribute filter).

python
import requests, datetime
SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"

# AND: high-resolution human structures
q_and = {
    "type": "group", "logical_operator": "and",
    "nodes": [
        {"type": "terminal", "service": "text",
         "parameters": {"attribute": "rcsb_entity_source_organism.scientific_name",
                        "operator": "exact_match", "value": "Homo sapiens"}},
        {"type": "terminal", "service": "text",
         "parameters": {"attribute": "rcsb_entry_info.resolution_combined",
                        "operator": "less", "value": 2.0}},
    ],
}

# OR: human or mouse
q_or = {
    "type": "group", "logical_operator": "or",
    "nodes": [
        {"type": "terminal", "service": "text",
         "parameters": {"attribute": "rcsb_entity_source_organism.scientific_name",
                        "operator": "exact_match", "value": "Homo sapiens"}},
        {"type": "terminal", "service": "text",
         "parameters": {"attribute": "rcsb_entity_source_organism.scientific_name",
                        "operator": "exact_match", "value": "Mus musculus"}},
    ],
}

# Combined: recent (last 30 days) + high-quality
one_month_ago = (datetime.date.today() - datetime.timedelta(days=30)).isoformat()
today = datetime.date.today().isoformat()
q_recent_hq = {
    "type": "group", "logical_operator": "and",
    "nodes": [
        {"type": "terminal", "service": "text",
         "parameters": {"attribute": "rcsb_entry_info.resolution_combined",
                        "operator": "less", "value": 2.0}},
        {"type": "terminal", "service": "text",
         "parameters": {"attribute": "refine.ls_R_factor_R_free",
                        "operator": "less", "value": 0.25}},
        {"type": "terminal", "service": "text",
         "parameters": {"attribute": "rcsb_accession_info.initial_release_date",
                        "operator": "range",
                        "value": {"from": one_month_ago, "to": today,
                                  "include_lower": True, "include_upper": True}}},
    ],
}

payload = {"query": q_recent_hq, "return_type": "entry",
           "request_options": {"paginate": {"rows": 5}}}
r = requests.post(SEARCH, json=payload, timeout=30)
print(f"Recent high-quality: {r.json()['total_count']} structures")
Module 7: Pagination + Batch with Rate Limiting

Search responses include total_count. Paginate with request_options.paginate.start and rows (max ~10000 per page in practice; 100–500 is a good batch size).

python
import requests, time
SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"

def search_all(query_node, return_type="entry", page=100, max_results=None, delay=0.3):
    """Paginate through every result; rate-limit between pages."""
    out, start = [], 0
    while True:
        payload = {"query": query_node, "return_type": return_type,
                   "request_options": {"paginate": {"start": start, "rows": page}}}
        r = requests.post(SEARCH, json=payload, timeout=30)
        r.raise_for_status()
        d = r.json()
        batch = [h["identifier"] for h in d.get("result_set", [])]
        if not batch:
            break
        out.extend(batch)
        if max_results and len(out) >= max_results:
            return out[:max_results]
        if len(batch) < page or len(out) >= d.get("total_count", 0):
            break
        start += page
        time.sleep(delay)
    return out

# Example: every insulin entry
ids = search_all(
    {"type": "terminal", "service": "full_text", "parameters": {"value": "insulin"}},
    max_results=300,
)
print(f"Insulin entries collected: {len(ids)}")
python
# Batch metadata fetch via GraphQL `entries(...)` to avoid one round-trip per ID
DATA = "https://data.rcsb.org/graphql"

def batch_metadata(pdb_ids, chunk=50):
    """Fetch (title, method, resolution) for many entries with one POST per chunk."""
    all_rows = []
    for i in range(0, len(pdb_ids), chunk):
        ids_arr = pdb_ids[i:i+chunk]
        ids_str = ", ".join(f'"{p}"' for p in ids_arr)
        gql = f"""{{ entries(entry_ids: [{ids_str}]) {{
            rcsb_id
            struct {{ title }}
            exptl {{ method }}
            rcsb_entry_info {{ resolution_combined }}
        }} }}"""
        r = requests.post(DATA, json={"query": gql}, timeout=60)
        r.raise_for_status()
        for e in r.json()["data"]["entries"]:
            res = e["rcsb_entry_info"]["resolution_combined"]
            all_rows.append({
                "pdb_id": e["rcsb_id"],
                "method": e["exptl"][0]["method"] if e["exptl"] else None,
                "resolution": res[0] if isinstance(res, list) and res else res,
                "title": e["struct"]["title"],
            })
    return all_rows

rows = batch_metadata(ids[:20])
for r in rows[:5]:
    print(f"  {r['pdb_id']}: {r['method']:<25} {r['resolution']} Å  — {r['title'][:50]}")

Key Concepts

Search Service Cheat Sheet
ServiceUse caseRequired parameters
full_textFree-text keyword across all indexed fieldsvalue (string)
textStructured attribute filterattribute, operator, value
sequenceMMseqs2 sequence similaritytarget ∈ {pdb_protein_sequence, pdb_dna_sequence, pdb_rna_sequence}, value (sequence), evalue_cutoff, identity_cutoff
seqmotifPattern / regex / PROSITE motifvalue (pattern), pattern_type ∈ {simple, prosite, regex}
structure3D shape similarity (BioZernike)value ({entry_id, assembly_id}), operator ∈ {strict_shape_match, relaxed_shape_match}
strucmotif3D residue-arrangement motifvalue (residue list), rmsd_cutoff
chemicalLigand similarity by SMILES/InChIvalue, match_type ∈ {graph-exact, graph-relaxed, fingerprint-similarity, sub-structure-stereo-relaxed}
AttributeQuery Operators (service: "text")
OperatorValue shapeExample
exact_matchstring"Homo sapiens"
contains_words / contains_phrasestring"tyrosine kinase"
equals / greater / less / greater_or_equal / less_or_equalnumber2.0
range{from, to, include_lower, include_upper}{"from": 1.5, "to": 2.5, "include_lower": True, "include_upper": True}
exists(none)—
inarray["X-RAY DIFFRACTION", "ELECTRON MICROSCOPY"]
Return Types

return_type controls the granularity of identifiers in result_set:

return_typeIdentifier shapeExample
entry4HHBOne per PDB ID
polymer_entity4HHB_1One per polymer chain entity
non_polymer_entity4HHB_2Ligands, cofactors
assembly4HHB-1Biological unit
polymer_instance4HHB.AIndividual chain coordinates
mol_definitionHEMChemical component (PDB ligand code)
Common Data API GraphQL Roots
RootIdentifier shapeReturns
entry(entry_id: ...)"4HHB"Entry-level metadata
entries(entry_ids: [...])arrayBatch entry lookup
polymer_entity(entry_id: ..., entity_id: ...)"4HHB", "1"Sequence + organism
polymer_entity_instance(entry_id: ..., asym_id: ...)"4HHB", "A"Chain-level coords/metadata
assembly(entry_id: ..., assembly_id: ...)"4HHB", "1"Biological assembly
chem_comp(comp_id: ...)"HEM"Small molecule reference
File Formats
FormatURL patternNotes
mmCIFhttps://files.rcsb.org/download/{id}.cifRecommended; no atom-count limit
PDBhttps://files.rcsb.org/download/{id}.pdbLegacy; 99,999 atom limit
Assembly (mmCIF)https://files.rcsb.org/download/{id}-assembly{N}.cifBiological unit
FASTAhttps://www.rcsb.org/fasta/entry/{id}Sequence only

Common Workflows

Workflow 1: Drug Target Structure Set

Goal: Find high-resolution human EGFR structures with bound ligands.

python
import requests, time

SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"
DATA   = "https://data.rcsb.org/graphql"

payload = {
    "query": {
        "type": "group", "logical_operator": "and",
        "nodes": [
            {"type": "terminal", "service": "full_text",
             "parameters": {"value": "EGFR epidermal growth factor receptor"}},
            {"type": "terminal", "service": "text",
             "parameters": {"attribute": "rcsb_entity_source_organism.scientific_name",
                            "operator": "exact_match", "value": "Homo sapiens"}},
            {"type": "terminal", "service": "text",
             "parameters": {"attribute": "rcsb_entry_info.resolution_combined",
                            "operator": "less", "value": 2.5}},
        ],
    },
    "return_type": "entry",
    "request_options": {"paginate": {"rows": 50}},
}
r = requests.post(SEARCH, json=payload, timeout=30)
r.raise_for_status()
pdb_ids = [h["identifier"] for h in r.json()["result_set"]]
print(f"EGFR ≤2.5 Å human structures: {len(pdb_ids)}")

# Filter to entries with bound ligands via batch GraphQL
ids_str = ", ".join(f'"{p}"' for p in pdb_ids[:20])
gql = f"""{{ entries(entry_ids: [{ids_str}]) {{
    rcsb_id
    struct {{ title }}
    rcsb_entry_info {{ resolution_combined nonpolymer_entity_count }}
}} }}"""
r2 = requests.post(DATA, json={"query": gql}, timeout=60)
for e in r2.json()["data"]["entries"]:
    n_lig = e["rcsb_entry_info"]["nonpolymer_entity_count"] or 0
    if n_lig > 0:
        res = e["rcsb_entry_info"]["resolution_combined"]
        res_v = res[0] if isinstance(res, list) else res
        print(f"  {e['rcsb_id']}: {res_v} Å, ligands={n_lig}  — {e['struct']['title'][:60]}")
    time.sleep(0.05)
Workflow 2: Protein Family — Sequence-Similar Structures

Goal: Find all PDB structures with sequence similar to a query (KRAS), then summarize their resolution + experimental method.

python
import requests, time

SEARCH = "https://search.rcsb.org/rcsbsearch/v2/query"
DATA   = "https://data.rcsb.org/graphql"

kras_seq = ("MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGQ"
            "EEYSAMRDQYMRTGEGFLCVFAINNTKSFEDIHHYREQIKRVKDSEDVPMVLVGNKCDLPS"
            "RTVDTKQAQDLARSYGIPFIETSAKTRQGVDDAFYTLVREIRKHKEKMSK")

payload = {
    "query": {
        "type": "terminal", "service": "sequence",
        "parameters": {"target": "pdb_protein_sequence", "value": kras_seq,
                       "evalue_cutoff": 1e-5, "identity_cutoff": 0.5},
    },
    "return_type": "polymer_entity",
    "request_options": {"paginate": {"rows": 20}},
}
r = requests.post(SEARCH, json=payload, timeout=30)
hits = r.json()["result_set"]
print(f"KRAS family hits: {len(hits)}")

# Unique PDB IDs from polymer_entity identifiers (e.g., "4OBE_1" -> "4OBE")
entry_ids = sorted({h["identifier"].split("_")[0] for h in hits})

# Batch metadata
ids_str = ", ".join(f'"{p}"' for p in entry_ids)
gql = f"""{{ entries(entry_ids: [{ids_str}]) {{
    rcsb_id
    struct {{ title }}
    exptl {{ method }}
    rcsb_entry_info {{ resolution_combined }}
}} }}"""
r2 = requests.post(DATA, json={"query": gql}, timeout=60)
for e in sorted(r2.json()["data"]["entries"],
                key=lambda x: (x["rcsb_entry_info"]["resolution_combined"] or [99])[0] if isinstance(x["rcsb_entry_info"]["resolution_combined"], list) else (x["rcsb_entry_info"]["resolution_combined"] or 99)):
    res = e["rcsb_entry_info"]["resolution_combined"]
    res_v = res[0] if isinstance(res, list) else res
    print(f"  {e['rcsb_id']}: {res_v} Å  {e['exptl'][0]['method']:<25} {e['struct']['title'][:50]}")
Workflow 3: Download + Parse with BioPython

Goal: Download mmCIF, then enumerate chains with BioPython.

python
import requests
from Bio.PDB import MMCIFParser

pdb_id = "4HHB"
r = requests.get(f"https://files.rcsb.org/download/{pdb_id}.cif", timeout=60)
r.raise_for_status()
with open(f"{pdb_id}.cif", "w") as f:
    f.write(r.text)

parser = MMCIFParser(QUIET=True)
structure = parser.get_structure(pdb_id, f"{pdb_id}.cif")
for model in structure:
    for chain in model:
        std_res = [r for r in chain if r.id[0] == " "]
        atoms = sum(len(list(r.get_atoms())) for r in std_res)
        print(f"Chain {chain.id}: {len(std_res)} residues, {atoms} atoms")
Show full SKILL.md (577 more words)Show less

Key Parameters

ParameterEndpointDefaultRange / OptionsEffect
valuesearch sequencerequiredprotein/DNA/RNA sequence stringQuery sequence for MMseqs2
evalue_cutoffsearch sequence0.11e-10–10E-value threshold
identity_cutoffsearch sequence0.90.0–1.0Minimum identity fraction
targetsearch sequence"pdb_protein_sequence"pdb_protein_sequence, pdb_dna_sequence, pdb_rna_sequenceSequence type
operatorsearch text (attribute)requiredsee Attribute OperatorsComparison kind
operatorsearch structurestrict_shape_matchstrict_shape_match, relaxed_shape_match3D match stringency
return_typeall searchentryentry, polymer_entity, assembly, polymer_instance, mol_definition, …Identifier granularity
paginate.start / paginate.rowsrequest_options0 / 25up to ~10000 rows/page in practicePagination window
GraphQL field entries(entry_ids: [...])data.rcsb.org/graphql—array of PDB IDsBatch entry metadata

Best Practices

  1. Search → fetch: Use the Search API to get a list of IDs, then GraphQL entries(entry_ids: [...]) for batch metadata. Avoid one GraphQL request per ID.

  2. Use full_text vs text deliberately: free-text keyword search needs "service": "full_text". Structured attribute filters need "service": "text". They are not interchangeable.

  3. mmCIF over PDB format: PDB format is being phased out and has a 99,999 atom limit. Always download .cif for new code.

  4. Set realistic paginate.rows: rows: 100 is a good default for batch work; the API may slow down beyond ~10000. Loop with paginate.start for full traversal.

  5. Rate limit with time.sleep(0.2) in batch loops: No published hard cap, but the public infrastructure is shared. On HTTP 429, back off exponentially.

  6. Inspect the payload before posting: print(json.dumps(payload, indent=2)) is the cheapest way to debug HTTP 400 errors.

  7. entries(entry_ids: [...]) does not validate every ID: if one ID is wrong, the whole array returns null entries. Validate IDs separately if you can't trust the source.

Common Recipes

Recipe: Get FASTA for a PDB ID
python
import requests
r = requests.get("https://www.rcsb.org/fasta/entry/4HHB", timeout=30)
print(r.text)
Recipe: All chains in an entry
python
import requests
DATA = "https://data.rcsb.org/graphql"
gql = """{ entry(entry_id: "4HHB") {
  polymer_entities {
    rcsb_id
    rcsb_polymer_entity_container_identifiers { auth_asym_ids }
    entity_poly { rcsb_entity_polymer_type pdbx_seq_one_letter_code_can }
  }
} }"""
r = requests.post(DATA, json={"query": gql}, timeout=30)
for pe in r.json()["data"]["entry"]["polymer_entities"]:
    chains = pe["rcsb_polymer_entity_container_identifiers"]["auth_asym_ids"]
    seq    = pe["entity_poly"]["pdbx_seq_one_letter_code_can"][:50]
    print(f"  {pe['rcsb_id']} chains={chains}  type={pe['entity_poly']['rcsb_entity_polymer_type']}  seq={seq}…")
Recipe: List all ligands in an entry
python
import requests
DATA = "https://data.rcsb.org/graphql"
gql = """{ entry(entry_id: "1IEP") {
  nonpolymer_entities {
    rcsb_id
    nonpolymer_comp { chem_comp { id name formula } }
  }
} }"""
r = requests.post(DATA, json={"query": gql}, timeout=30)
for npe in r.json()["data"]["entry"]["nonpolymer_entities"]:
    cc = npe["nonpolymer_comp"]["chem_comp"]
    print(f"  {npe['rcsb_id']}: {cc['id']} ({cc['name']})  {cc['formula']}")
Recipe: Inspect available search attributes

The Search API exposes a JSON schema at https://search.rcsb.org/rcsbsearch/v2/metadata/schema. Use it to look up valid attribute paths.

python
import requests
r = requests.get("https://search.rcsb.org/rcsbsearch/v2/metadata/schema", timeout=30)
schema = r.json()
# Schema lists hundreds of attribute paths; sample a few
sample_paths = [k for k in schema if "resolution" in k.lower()][:5]
print(sample_paths)

Troubleshooting

ProblemCauseSolution
HTTP 400 — Invalid request to the [ text ] service on a free-text queryWrong service nameUse "service": "full_text" for keyword search; "service": "text" is for structured attribute filters
HTTP 400 with cryptic schema messageBad operator/value shapeCheck the AttributeQuery Operators table; range needs the {from,to,include_lower,include_upper} dict
Empty result_setFilters too strictRelax filters one at a time; verify attribute names via the schema endpoint
HTTP 404 on entries(entry_ids: ["XYZW"])The entry doesn't existRCSB returns null rather than 404 inside the GraphQL response — check each data.entries[i] for null
HTTP 429 Too Many RequestsBurst paceAdd time.sleep(0.3) between requests; exponential backoff on 429
HTTP 500 from searchServer-side glitchRetry after 5–10 s; check status.rcsb.org
Downloaded .pdb file truncated>99,999 atoms (legacy format limit)Download .cif instead
GraphQL response has errors arrayField name typo or wrong rootRead the error message; the API is strict about field names — check the schema browser at https://data.rcsb.org/index.html#graphql-api
  • alphafold-database-access — AI-predicted structures; use when no experimental structure exists
  • uniprot-protein-database — protein annotations, sequences, ID mapping (UniProt accession needed for AlphaFold)
  • biopython-molecular-biology — parse downloaded PDB/mmCIF files, extract coordinates, compute distances
  • autodock-vina-docking — downstream molecular docking using PDB structures as receptors
  • rdkit-cheminformatics — analyze ligands extracted from PDB complexes

References

© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/structural-biology-drug-discovery/pdb-database of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pdb Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pdb Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pdb Database this skilljaechang-hits/SciAgent-Skills3701 repos~7.7kAutomated safety check: PassBSD-3-Clause
Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills3.2k2 repos~1.2kAutomated safety check: PassApache-2.0
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT
Tooluniverseynulihao/AgentSkillOS6173 repos~2.5kAutomated safety check: PassNone
Bio DB ToolsDrugClaw/DrugClaw125—~1.4kAutomated safety check: PassApache-2.0
Ggetdavila7/claude-code-templates32k11 repos~6.3kAutomated safety check: PassMIT

Similar skills

  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 8 days ago
    Research & ScienceAuto-check passed
  • Tooluniverse

    ynulihao/AgentSkillOS

    A skill your agent uses when working with scientific research tools and workflows across bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery.

    617 GitHub starsUsed in 3 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Bio DB Tools

    DrugClaw/DrugClaw

    Query public biology databases and APIs including UniProt, RCSB PDB, AlphaFold DB, ClinVar, dbSNP, gnomAD, Ensembl, GEO, InterPro, KEGG, OpenTargets, Reactome, and STRING.

    125 GitHub stars~1.4k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Alphafold Database

    davila7/claude-code-templates

    Access AlphaFold's 200M+ AI-predicted protein structures. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~4k tokens
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 163 skills in this repo
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    370 GitHub stars~3.2k tokensUpdated 9 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    370 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    370 GitHub stars~6.9k tokensUpdated 9 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub stars~2.3k tokensUpdated 9 days ago
    Auto-check passed
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 2 repos~5.8k tokens
    Auto-check passed

Questions about Pdb Database

What does Pdb Database do?

Query RCSB PDB (200K+ structures) via the public REST + GraphQL APIs with plain requests (no SDK). Pdb Database is an agent skill from jaechang-hits/SciAgent-Skills. Query RCSB PDB (200K+ structures) via the public REST + GraphQL APIs with plain requests (no SDK).

When should I use Pdb Database?

Pdb Database fits situations like: tasks that involve Protein structure and design; tasks that involve GraphQL; tasks that involve Vector databases.

How do I install Pdb Database in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill pdb-database -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/pdb-database in jaechang-hits/SciAgent-Skills) into .claude/skills/pdb-database in your project. Claude Code loads it when a task matches its description.

How do I install Pdb Database in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill pdb-database -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/pdb-database in jaechang-hits/SciAgent-Skills) into .agents/skills/pdb-database in your project. Codex loads it when a task matches its description.

Can I use Pdb Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill pdb-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdb-database, .gemini/skills/pdb-database, .github/skills/pdb-database and .opencode/skills/pdb-database in your project.

What does Pdb Database need to run?

Going by SKILL.md and its folder, Pdb Database needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Pdb Database access the network?

SKILL.md names 4 domains. In commands or code: search.rcsb.org, data.rcsb.org, files.rcsb.org and rcsb.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Pdb Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pdb Database use?

Pdb Database is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pdb Database use?

About 7.7k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pdb Database?

Skills that share tags, products or a category with Pdb Database: Alphafold Database Fetch And Analyze (google-deepmind/science-skills, 3.2k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars), Tooluniverse (ynulihao/AgentSkillOS, 617 stars) and Bio DB Tools (DrugClaw/DrugClaw, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pdb Database?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 163 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.