Agent skill

Interpro Database

by jaechang-hits in jaechang-hits/SciAgent-Skills

Query InterPro REST API for protein domain architecture, family classification, and member-DB integration.

CC-BY-4.0Auto-check passedResearch & Science

Install Interpro Database

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill interpro-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills interpro-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/interpro-database .claude/skills/interpro-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
interpro-database
GitHub stars
370
Used in
1 other repo
Token cost
~7.6k tokens
SKILL.md length
1,304 words
Files
1
Skills in repo
163
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Query InterPro REST API for protein domain architecture, family classification, and member-DB integration.

  • Works in 5 steps: Use reviewed proteins for curated domain… → Chunk large taxonomy or protein lists:… → Add time.sleep(1.0) between paginated… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 9 more sections
  • Calls pip; reaches ebi.ac.uk and rest.uniprot.org

What it does

Interpro Database is an agent skill from jaechang-hits/SciAgent-Skills. Query InterPro REST API for protein domain architecture, family classification, and member-DB integration. Search entries, retrieve a protein's domains, list family members, get taxonomic distribution, link to PDB. Unifies Pfam, PANTHER, PIRSF, PRINTS, PROSITE, SMART, CDD, NCBIfam. Use uniprot-protein-database for sequences; pdb-database for 3D structures.

Its SKILL.md is about 7.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics and REST APIs. It works with UniProt. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve REST APIs

Example prompts

  • “/interpro-database”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Use reviewed proteins for curated domain lists: The unreviewed TrEMBL set is 5–10× larger and contains automated predictions. For…
  2. Chunk large taxonomy or protein lists: Retrieving all 10,000+ proteins for a broad family like the protein kinase superfamily can take…
  3. Add time.sleep(1.0) between paginated calls: The InterPro API is shared EBI infrastructure with no published rate limit. A 1-second pause…
  4. Prefer InterPro accessions over member DB accessions for cross-database queries: A Pfam PF00069 and PANTHER PTHR24340 both model kinase…
  5. Check type before interpreting entry_protein_locations: Only domain, repeat, and site entries carry meaningful position information…

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ebi.ac.uk
    • rest.uniprot.org

    Also links to:

    • doi.org
    • interpro-documentation.readthedocs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Interpro Database loads about 7.6k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 1,304 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~7.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 1,304 words, ~7,570 tokens.

Download SKILL.mdSave it as .claude/skills/interpro-database/SKILL.md (or your agent's skills folder).
name
interpro-database
description
Query InterPro REST API for protein domain architecture, family classification, and member-DB integration. Search entries, retrieve a protein's domains, list family members, get taxonomic distribution, link to PDB. Unifies Pfam, PANTHER, PIRSF, PRINTS, PROSITE, SMART, CDD, NCBIfam. Use uniprot-protein-database for sequences; pdb-database for 3D structures.
license
CC-BY-4.0

InterPro Database

Overview

InterPro is the EBI's integrated protein family, domain, and functional site database. It consolidates signatures from 13 member databases (Pfam, PANTHER, PIRSF, PRINTS, PROSITE, SMART, CDD, NCBIfam, and others) into unified InterPro entries, each describing a homologous superfamily, domain, family, repeat, or conserved site. The REST API at https://www.ebi.ac.uk/interpro/api/ is free and requires no authentication.

When to Use

  • Identifying all domains and families present in a protein by UniProt accession (domain architecture)
  • Searching for proteins that contain a specific domain or belong to a specific family
  • Finding the taxonomic distribution of organisms that encode a given domain or family
  • Cross-linking a domain to experimental 3D structures in the PDB
  • Checking which source databases (Pfam, PANTHER, SMART, etc.) cover an InterPro entry
  • Discovering InterPro entries by keyword (e.g., "kinase domain") when you do not yet know the accession
  • For protein sequence retrieval, functional annotations (GO, pathways, active sites), and ID mapping use uniprot-protein-database
  • For downloading domain-aligned sequences or building HMM profiles use Pfam directly; InterPro is the meta-layer

Prerequisites

  • Python packages: requests, pandas, matplotlib
  • Data requirements: UniProt accessions (e.g., P04637) or InterPro accessions (e.g., IPR011009)
  • Environment: internet connection; no API key required
  • Rate limits: no published hard limit; use time.sleep(1.0) between requests for batch queries; paginate with ?cursor= or ?page_size=
bash
pip install requests pandas matplotlib

Quick Start

python
import requests

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def interpro_get(path: str, params: dict = None) -> dict:
    """Send a GET request to the InterPro API and return parsed JSON."""
    r = requests.get(
        f"{INTERPRO_BASE}/{path}",
        params=params,
        headers={"Accept": "application/json"},
        timeout=30
    )
    r.raise_for_status()
    return r.json()

# Get domain architecture for TP53 (P04637)
# Note: `protein/uniprot/{acc}/` returns only {metadata}; the entries-per-protein
# data lives at `entry/interpro/protein/uniprot/{acc}/` and is keyed `results`.
data = interpro_get("entry/interpro/protein/uniprot/P04637/")
entries = data.get("results", [])
print(f"InterPro entries for TP53: {data.get('count')}  (this page: {len(entries)})")
for e in entries[:4]:
    m = e["metadata"]
    print(f"  {m['accession']}  {m['type']:<25}  {m['name']}")
# InterPro entries for TP53: 9
#   IPR002117  family                     p53 tumour suppressor family
#   IPR036674  homologous_superfamily     p53-like tetramerisation domain superfamily

Core API

Query 1: Entry Search

Search for InterPro entries by name keyword or fetch a specific entry by accession.

python
import requests

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def search_entries(query: str, entry_type: str = None,
                   page_size: int = 20) -> list:
    """Search InterPro entries by keyword; optionally filter by type."""
    params = {"search": query, "page_size": page_size}
    if entry_type:
        params["type"] = entry_type   # family, domain, homologous_superfamily, repeat, site
    r = requests.get(
        f"{INTERPRO_BASE}/entry/interpro/",
        params=params,
        headers={"Accept": "application/json"},
        timeout=30
    )
    r.raise_for_status()
    return r.json().get("results", [])

hits = search_entries("serine kinase", entry_type="domain")
print(f"InterPro domain entries matching 'serine kinase': {len(hits)}")
for h in hits[:5]:
    m = h["metadata"]
    print(f"  {m['accession']}  {m['type']:<10}  {m['name']}")
# InterPro domain entries matching 'serine kinase': 8
#   IPR000719  domain    Protein kinase domain
#   IPR008271  domain    Serine/threonine/tyrosine kinase, active site
python
# Fetch a specific InterPro entry by accession
r = requests.get(
    f"{INTERPRO_BASE}/entry/interpro/IPR000719/",
    headers={"Accept": "application/json"},
    timeout=30
)
r.raise_for_status()
meta = r.json()["metadata"]
print(f"Accession    : {meta['accession']}")
print(f"Name         : {meta['name']}")
print(f"Type         : {meta['type']}")
print(f"Member DBs   : {list(meta.get('member_databases', {}).keys())}")
go_terms = meta.get("go_terms", [])
print(f"GO terms     : {[g['identifier'] for g in go_terms[:3]]}")
# Accession    : IPR000719
# Name         : Protein kinase domain
# Type         : domain
# Member DBs   : ['pfam', 'smart', 'cdd', 'ncbifam', 'panther']
# GO terms     : ['GO:0004672', 'GO:0005524', 'GO:0006468']
Query 2: Protein Domain Architecture

Retrieve all InterPro entries (domains, families, sites) matched in a protein by UniProt accession.

python
import requests

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def get_protein_domain_architecture(uniprot_acc: str) -> dict:
    """Return all InterPro entry matches for a protein. Uses the
    `entry/interpro/protein/uniprot/{acc}/` endpoint, which returns
    {count, next, previous, results}. Each result has metadata + a
    nested `proteins[0].entry_protein_locations` for the per-protein match."""
    r = requests.get(
        f"{INTERPRO_BASE}/entry/interpro/protein/uniprot/{uniprot_acc}/",
        headers={"Accept": "application/json"},
        timeout=60
    )
    r.raise_for_status()
    return r.json()

data = get_protein_domain_architecture("P04637")   # TP53
results = data.get("results", [])
# Pull length/source from the first match's nested protein record
prot0 = results[0]["proteins"][0] if results and results[0].get("proteins") else {}
print(f"Protein length : {prot0.get('protein_length')}")
print(f"Source DB      : {prot0.get('source_database')}")
print(f"InterPro entries: {data.get('count')}")
for entry in results[:6]:
    m = entry["metadata"]
    # Locations are nested under proteins[0].entry_protein_locations
    locs = entry["proteins"][0].get("entry_protein_locations", []) if entry.get("proteins") else []
    loc_str = ", ".join(
        f"{frag['start']}-{frag['end']}"
        for loc in locs for frag in loc.get("fragments", [])
    )
    print(f"  {m['accession']}  {m['type']:<25}  {m['name'][:35]:<35}  [{loc_str}]")
python
# Compare domain architectures of two proteins side-by-side
import pandas as pd

def domain_set(uniprot_acc: str) -> set:
    data = get_protein_domain_architecture(uniprot_acc)
    return {e["metadata"]["accession"] for e in data.get("results", [])}

brca1_domains = domain_set("P38398")   # BRCA1
tp53_domains   = domain_set("P04637")  # TP53

shared = brca1_domains & tp53_domains
unique_brca1 = brca1_domains - tp53_domains
unique_tp53  = tp53_domains - brca1_domains
print(f"Shared InterPro entries: {len(shared)}")
print(f"BRCA1-unique           : {len(unique_brca1)}")
print(f"TP53-unique            : {len(unique_tp53)}")
Query 3: Entry Proteins

List proteins that contain a specific InterPro entry (family or domain).

python
import requests, time

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def get_entry_proteins(interpro_acc: str, reviewed_only: bool = True,
                       page_size: int = 50) -> list:
    """Return proteins (UniProt) containing a given InterPro entry.
    Path order is `/protein/{db}/entry/interpro/{IPR}/` — the inverse
    `entry/interpro/{IPR}/protein/{db}/` times out (408) on large families."""
    db = "reviewed" if reviewed_only else "uniprot"
    r = requests.get(
        f"{INTERPRO_BASE}/protein/{db}/entry/interpro/{interpro_acc}/",
        params={"page_size": page_size},
        headers={"Accept": "application/json"},
        timeout=60
    )
    r.raise_for_status()
    return r.json().get("results", [])

proteins = get_entry_proteins("IPR011009")   # Protein kinase-like domain SF
print(f"Reviewed proteins with IPR011009 (page 1): {len(proteins)}")
for p in proteins[:4]:
    m = p["metadata"]
    # metadata fields: accession, gene, length, name, source_database, source_organism
    print(f"  {m['accession']}  {(m.get('gene') or ''):<8}  "
          f"len={m.get('length', '?')}  "
          f"org={(m.get('source_organism') or {}).get('scientificName', '')[:30]}")
python
# Paginate all proteins for a family using cursor
def get_all_entry_proteins(interpro_acc: str,
                            reviewed_only: bool = True) -> list:
    INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"
    db = "reviewed" if reviewed_only else "uniprot"
    # Path-inverted: /protein/{db}/entry/interpro/{IPR}/ is the working order
    url = f"{INTERPRO_BASE}/protein/{db}/entry/interpro/{interpro_acc}/"
    all_proteins = []
    params = {"page_size": 200}
    while url:
        r = requests.get(url, params=params,
                         headers={"Accept": "application/json"}, timeout=60)
        r.raise_for_status()
        data = r.json()
        all_proteins.extend(data.get("results", []))
        url = data.get("next")
        params = None   # next URL already has params encoded
        if url:
            time.sleep(1.0)
    return all_proteins

proteins = get_all_entry_proteins("IPR000719")   # Protein kinase domain
print(f"Total reviewed proteins with protein kinase domain: {len(proteins)}")
Query 4: Entry Taxonomy

Get the taxonomic distribution of proteins annotated with a given InterPro entry.

python
import requests

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def get_entry_taxonomy(interpro_acc: str,
                        page_size: int = 50) -> list:
    """Return taxonomic summary for proteins in a given InterPro entry.
    Path-inverted: `/taxonomy/uniprot/entry/interpro/{IPR}/`. Each result
    has `metadata` (taxon: accession=taxId, name, parent, children, rank)
    and `entries[]` (representative protein-match locations for that taxon)."""
    r = requests.get(
        f"{INTERPRO_BASE}/taxonomy/uniprot/entry/interpro/{interpro_acc}/",
        params={"page_size": page_size},
        headers={"Accept": "application/json"},
        timeout=90
    )
    r.raise_for_status()
    return r.json().get("results", [])

# Use a smaller entry (p53 DBD); IPR000719 (kinase) has ~270k taxa and times out.
taxa = get_entry_taxonomy("IPR011615")
print(f"Top taxa for IPR011615 (p53 DNA-binding domain):")
for t in taxa[:8]:
    m = t["metadata"]
    print(f"  taxId={m['accession']:>10}  {m.get('name', ''):<30}  "
          f"rank={m.get('rank') or 'n/a'}")
Query 5: Structure Integration

Retrieve PDB structures associated with an InterPro entry.

python
import requests

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def get_entry_structures(interpro_acc: str, page_size: int = 25) -> list:
    """Return PDB structures that include a match to a given InterPro entry.
    Path-inverted: `/structure/pdb/entry/interpro/{IPR}/`. The flat form with
    `?entry_interpro=...` is silently slow / 408s on this resource."""
    r = requests.get(
        f"{INTERPRO_BASE}/structure/pdb/entry/interpro/{interpro_acc}/",
        params={"page_size": page_size},
        headers={"Accept": "application/json"},
        timeout=60
    )
    r.raise_for_status()
    return r.json().get("results", [])

structures = get_entry_structures("IPR011009")   # Protein kinase-like SF
print(f"PDB structures linked to IPR011009 (page 1): {len(structures)}")
for s in structures[:5]:
    m = s["metadata"]
    print(f"  {m['accession'].upper()}  resolution={m.get('resolution', 'N/A')} Å  "
          f"experiment={m.get('experiment_type', 'N/A')}")
# PDB structures linked to IPR011009: ~8,000+
#   1A06  resolution=2.5 Å  experiment=x-ray
#   ...
Query 6: Domain Sequence Retrieval

Download the FASTA sequences of proteins in an InterPro family for alignment or phylogenetics.

python
import requests, time

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def get_family_fasta(interpro_acc: str,
                      reviewed_only: bool = True,
                      max_sequences: int = 100) -> str:
    """Retrieve FASTA sequences for proteins in an InterPro entry."""
    db = "reviewed" if reviewed_only else "uniprot"
    proteins = []
    # Path-inverted: protein-list-for-entry is /protein/{db}/entry/interpro/{IPR}/
    url = f"{INTERPRO_BASE}/protein/{db}/entry/interpro/{interpro_acc}/"
    params = {"page_size": min(max_sequences, 200)}
    while url and len(proteins) < max_sequences:
        r = requests.get(url, params=params,
                         headers={"Accept": "application/json"}, timeout=60)
        r.raise_for_status()
        data = r.json()
        proteins.extend(data.get("results", []))
        url = data.get("next") if len(proteins) < max_sequences else None
        params = None
        if url:
            time.sleep(1.0)

    # Fetch FASTA from UniProt for each accession
    accessions = [p["metadata"]["accession"] for p in proteins[:max_sequences]]
    fasta_url = "https://rest.uniprot.org/uniprotkb/stream"
    query = " OR ".join(f"accession:{acc}" for acc in accessions)
    r = requests.get(fasta_url,
                     params={"query": query, "format": "fasta"},
                     timeout=120)
    r.raise_for_status()
    return r.text

fasta = get_family_fasta("IPR000719", reviewed_only=True, max_sequences=20)
seq_count = fasta.count(">")
print(f"FASTA sequences retrieved: {seq_count}")
print(fasta[:300])   # preview first sequence header + start

Key Concepts

InterPro Entry Types

InterPro classifies entries into five types. The type determines what biological relationship the match implies:

TypeDescriptionExample
familyHomologous group of proteins sharing common ancestry and functionIPR000719 (Protein kinase)
domainDiscrete structural and functional unit that can occur in multiple protein contextsIPR011009 (Protein kinase-like SF)
homologous_superfamilyStructurally similar domains that may have diverged in sequenceIPR011993 (Pleckstrin-like)
repeatShort, repeated sequence unit that occurs multiple times within a proteinIPR001440 (TPR repeat)
siteShort conserved motif: active site, binding site, or post-translational modification siteIPR008271 (Ser/Thr kinase active site)
Member Database Hierarchy

Each InterPro entry integrates signatures from one or more member databases. The InterPro accession (IPR...) is the unified meta-entry; member database accessions point to the underlying models:

Member DBAccession prefixModeling approach
PfamPFHidden Markov Models (profile HMMs)
PANTHERPTHRPhylogenetic trees + HMMs
PIRSFPIRSFFull-length HMMs
PRINTSPRFingerprint motif groups
PROSITEPSPatterns and profiles
SMARTSMHMMs with database integration
CDDcdPosition-specific scoring matrices (PSSMs)
NCBIfamNFNCBI-curated HMMs
Pagination

The InterPro API paginates results at the collection level. Each response includes a next URL (or null when exhausted) and a count field. For large families (e.g., kinases: 10,000+ proteins) always iterate using the next cursor.

python
import requests, time

def iterate_interpro(url: str, page_size: int = 200) -> list:
    """Generic paginator for any InterPro list endpoint."""
    results = []
    params = {"page_size": page_size}
    while url:
        r = requests.get(url, params=params,
                         headers={"Accept": "application/json"}, timeout=60)
        r.raise_for_status()
        data = r.json()
        results.extend(data.get("results", []))
        url = data.get("next")
        params = None
        if url:
            time.sleep(1.0)
    return results

Common Workflows

Workflow 1: Domain Architecture Report for a Protein Set

Goal: Retrieve all InterPro domains for a list of proteins and produce a summary table showing which domains each protein carries.

python
import requests, time, pandas as pd

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def get_domains(uniprot_acc: str) -> list:
    """List InterPro entries for a protein. Uses the
    `entry/interpro/protein/uniprot/{acc}/` endpoint (keyed `results`)."""
    r = requests.get(
        f"{INTERPRO_BASE}/entry/interpro/protein/uniprot/{uniprot_acc}/",
        headers={"Accept": "application/json"}, timeout=60
    )
    if r.status_code == 404:
        return []
    r.raise_for_status()
    data = r.json()
    return [
        {
            "protein": uniprot_acc,
            "accession": e["metadata"]["accession"],
            "name": e["metadata"]["name"],
            "type": e["metadata"]["type"],
            "source_db": list(e["metadata"].get("member_databases", {}).keys()),
        }
        for e in data.get("results", [])
    ]

proteins = ["P04637", "P38398", "Q00987", "P10415"]  # TP53, BRCA1, MDM2, BCL2
rows = []
for acc in proteins:
    rows.extend(get_domains(acc))
    time.sleep(1.0)

df = pd.DataFrame(rows)
print(f"Total domain matches: {len(df)}")
print(df.groupby(["protein", "type"])["accession"].count().unstack(fill_value=0))

# Pivot: proteins × domain accessions
pivot = df[df["type"] == "domain"].pivot_table(
    index="protein", columns="accession", aggfunc="size", fill_value=0
)
pivot.to_csv("domain_architecture_matrix.csv")
print(f"\nDomain × protein matrix: {pivot.shape}")
Workflow 2: Find Kinase Family Members with PDB Structures

Goal: Retrieve proteins in a kinase domain family that have experimental structures in the PDB, ranked by resolution.

python
import requests, time, pandas as pd

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

# Step 1: Get PDB structures linked to the protein kinase-like SF entry.
# Use the path-inverted form; the flat `?entry_interpro=` filter 408s.
r = requests.get(
    f"{INTERPRO_BASE}/structure/pdb/entry/interpro/IPR011009/",
    params={"page_size": 200},
    headers={"Accept": "application/json"}, timeout=60
)
r.raise_for_status()
structures = r.json().get("results", [])
print(f"PDB structures with IPR011009 (kinase-like SF, page 1): {len(structures)}")

rows = []
for s in structures:
    m = s["metadata"]
    rows.append({
        "pdb_id": m["accession"],
        "resolution": m.get("resolution"),
        "experiment": m.get("experiment_type", ""),
        "name": m.get("name", ""),
    })

df = pd.DataFrame(rows)
df = df.dropna(subset=["resolution"]).sort_values("resolution")
print(f"\nTop 10 highest-resolution kinase structures:")
print(df[["pdb_id", "resolution", "experiment", "name"]].head(10).to_string(index=False))
df.to_csv("kinase_structures.csv", index=False)
print(f"\nSaved kinase_structures.csv ({len(df)} X-ray / cryo-EM structures)")
Workflow 3: Taxonomic Coverage Bar Chart for a Domain

Goal: Visualize how many reviewed proteins in each major kingdom carry a given InterPro domain.

python
import requests, time
import pandas as pd
import matplotlib.pyplot as plt

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def get_taxonomy_counts(interpro_acc: str, page_size: int = 100,
                        max_pages: int = 5) -> pd.DataFrame:
    """Walk the taxonomy results for an InterPro entry. The API does not
    expose a per-taxon protein count at this endpoint — instead each
    paged record is one (taxon × representative protein-match) row.
    Aggregate client-side by taxon name to approximate frequency."""
    rows, url = [], f"{INTERPRO_BASE}/taxonomy/uniprot/entry/interpro/{interpro_acc}/"
    params = {"page_size": page_size}
    for _ in range(max_pages):
        if not url:
            break
        r = requests.get(url, params=params,
                         headers={"Accept": "application/json"}, timeout=90)
        r.raise_for_status()
        data = r.json()
        for t in data.get("results", []):
            m = t["metadata"]
            rows.append({
                "taxon_id": m["accession"],
                "name": m.get("name", ""),
                "rank": m.get("rank") or "",
            })
        url = data.get("next")
        params = None
        if url:
            time.sleep(1.0)
    return pd.DataFrame(rows)

IPR_ACC = "IPR011615"   # p53 DNA-binding domain (smaller; kinase 408s)
df = get_taxonomy_counts(IPR_ACC, max_pages=3)
print(f"Tax entries pulled for {IPR_ACC}: {len(df)}")

# Aggregate by name and take top 15 (each row = one rep. protein-match)
top = (df.groupby("name").size().sort_values(ascending=False).head(15)
       .reset_index(name="rep_matches"))
fig, ax = plt.subplots(figsize=(10, 5))
bars = ax.barh(top["name"], top["rep_matches"], color="#2171B5")
ax.bar_label(bars, fmt="%d", padding=3, fontsize=8)
ax.set_xlabel("Representative protein-matches")
ax.set_title(f"Taxonomic distribution of {IPR_ACC} (p53 DNA-binding domain)")
ax.invert_yaxis()
plt.tight_layout()
plt.savefig(f"{IPR_ACC}_taxonomy.png", dpi=150, bbox_inches="tight")
print(f"Saved {IPR_ACC}_taxonomy.png")

Key Parameters

ParameterEndpointDefaultRange / OptionsEffect
searchentry/interpro/—free-text stringKeyword filter on entry name and short name
typeentry/interpro/all typesfamily, domain, homologous_superfamily, repeat, siteFilter entries by InterPro type
page_sizeall list endpoints201–200Results returned per page
entry_interprostructure/pdb/—IPR######Filter structures by linked InterPro entry
source_databaseprotein/—reviewed, uniprot, tremblFilter proteins by UniProt curation level
reviewed (URL path)entry/{ipr}/{acc}/protein/uniprotreviewed, uniprotSwiss-Prot reviewed only vs all UniProtKB
relationsentry/interpro/{acc}/—contains, contained_by, child_of, parent_ofNavigate the InterPro hierarchy
nextall list endpoints—URL from responseCursor-based pagination; use the full URL from the next field
Show full SKILL.md (590 more words)Show less

Best Practices

  1. Use reviewed proteins for curated domain lists: The unreviewed TrEMBL set is 5–10× larger and contains automated predictions. For benchmarking, family analysis, or training sets, restrict to reviewed (Swiss-Prot) entries to avoid noise from unreviewed predictions.

  2. Chunk large taxonomy or protein lists: Retrieving all 10,000+ proteins for a broad family like the protein kinase superfamily can take minutes and produce large payloads. Limit queries with page_size=200 and the next cursor; store intermediate results to disk.

  3. Add time.sleep(1.0) between paginated calls: The InterPro API is shared EBI infrastructure with no published rate limit. A 1-second pause per page is a safe minimum for batch scripts.

  4. Prefer InterPro accessions over member DB accessions for cross-database queries: A Pfam PF00069 and PANTHER PTHR24340 both model kinase domains but with different protein coverage. Using the parent InterPro IPR000719 gives the union of all member DB matches in one query.

  5. Check type before interpreting entry_protein_locations: Only domain, repeat, and site entries carry meaningful position information. family and homologous_superfamily entries typically span the full protein and their coordinates are less informative.

Common Recipes

Recipe: Quick Domain Check for a Protein

When to use: Given a UniProt accession, rapidly list which InterPro domains it contains.

python
import requests

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def list_protein_domains(uniprot_acc: str) -> list:
    """Return list of (accession, type, name) tuples for a protein."""
    r = requests.get(
        f"{INTERPRO_BASE}/entry/interpro/protein/uniprot/{uniprot_acc}/",
        headers={"Accept": "application/json"}, timeout=60
    )
    r.raise_for_status()
    return [
        (e["metadata"]["accession"], e["metadata"]["type"], e["metadata"]["name"])
        for e in r.json().get("results", [])
    ]

domains = list_protein_domains("P00533")   # EGFR
print(f"InterPro entries in EGFR (P00533): {len(domains)}")
for acc, etype, name in domains:
    print(f"  {acc}  {etype:<25}  {name}")
# InterPro entries in EGFR (P00533): 10
#   IPR009030  homologous_superfamily   Growth factor receptor, cysteine-rich
#   IPR000719  domain                   Protein kinase domain
Recipe: Find All Proteins in a Family with Source DB Coverage

When to use: Map how many proteins in a domain family are covered by each member database (Pfam vs PANTHER vs SMART, etc.).

python
import requests, time
import pandas as pd

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

interpro_acc = "IPR000719"   # Protein kinase domain
r = requests.get(
    f"{INTERPRO_BASE}/entry/interpro/{interpro_acc}/",
    headers={"Accept": "application/json"}, timeout=30
)
r.raise_for_status()
member_dbs = r.json()["metadata"].get("member_databases", {})
print(f"Member databases for {interpro_acc}:")
for db, details in member_dbs.items():
    print(f"  {db}: {details}")

# Visualize member database source breakdown
labels = list(member_dbs.keys())
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(7, 4))
ax.bar(labels, [1] * len(labels), color="#4472C4")   # presence/absence per DB
ax.set_ylabel("Integrated (1=yes)")
ax.set_title(f"Member databases in {interpro_acc}")
plt.tight_layout()
plt.savefig(f"{interpro_acc}_member_dbs.png", dpi=150, bbox_inches="tight")
Recipe: Get GO Terms for an InterPro Entry

When to use: Bridge from structural domain to functional GO annotation.

python
import requests

INTERPRO_BASE = "https://www.ebi.ac.uk/interpro/api"

def get_go_terms_for_entry(interpro_acc: str) -> list:
    """Return GO terms associated with an InterPro entry."""
    r = requests.get(
        f"{INTERPRO_BASE}/entry/interpro/{interpro_acc}/",
        headers={"Accept": "application/json"}, timeout=30
    )
    r.raise_for_status()
    go_terms = r.json()["metadata"].get("go_terms", [])
    return [
        {"id": g["identifier"], "name": g["name"],
         "category": g.get("category", {}).get("name", "")}
        for g in go_terms
    ]

go_terms = get_go_terms_for_entry("IPR000719")
print(f"GO terms for IPR000719 (protein kinase domain): {len(go_terms)}")
for g in go_terms:
    print(f"  {g['id']}  [{g['category'][:2].upper()}]  {g['name']}")
# GO terms for IPR000719 (protein kinase domain): 3
#   GO:0004672  [MO]  protein kinase activity
#   GO:0005524  [MO]  ATP binding
#   GO:0006468  [BI]  protein phosphorylation

Troubleshooting

ProblemCauseSolution
HTTP 404 on protein lookupAccession not found in InterProVerify the UniProt accession exists; isoform accessions (P12345-2) may not be indexed separately
Empty entries list for a proteinProtein has no InterPro matches (e.g., intrinsically disordered)Check UniProt directly; not all proteins have classified domains
protein/uniprot/{acc}/ returns only metadata (no entries)That endpoint is protein-only; entry matches live elsewhereUse entry/interpro/protein/uniprot/{acc}/ and read the results[] key
entry/interpro/{IPR}/protein/{db}/ returns 408 / hangsThe path with entry/... first does a slow joinInvert the path: protein/{db}/entry/interpro/{IPR}/
structure/pdb/?entry_interpro={IPR} times out (408)Same join order issueUse structure/pdb/entry/interpro/{IPR}/
entry/interpro/{IPR}/taxonomy/uniprot/ 408s for large familiesSameUse taxonomy/uniprot/entry/interpro/{IPR}/; for very large entries (e.g. IPR000719 kinase) the inverted form may still 408 — fall back to a more specific sub-family entry
HTTP 400 on entry searchInvalid query parameters or unsupported type valueUse one of: family, domain, homologous_superfamily, repeat, site
Pagination stops earlynext is null before expected countThis is correct; all results have been returned
Very slow response for large familiesProtein set has thousands of membersIncrease page_size to 200; persist results after each page
ConnectionError or TimeoutTransient network or server issueRetry with exponential backoff; EBI services occasionally have brief downtimes
Member DB accessions missingEntry is new and member DB integration is pendingUse the InterPro accession for queries; member DB-level details update with each release
  • uniprot-protein-database — UniProt REST API for protein sequences, Swiss-Prot functional annotations (active sites, PTMs, disease associations), and ID mapping
  • esm-protein-language-model — Generate protein language model embeddings for sequences; useful after identifying a protein family with InterPro
  • pdb-database — Retrieve and download experimental 3D structures by PDB ID; cross-reference structure IDs discovered via InterPro structure queries

References

© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/proteomics-protein-engineering/interpro-database of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Interpro Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Interpro Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Interpro Database this skilljaechang-hits/SciAgent-Skills3701 repos~7.6kAutomated safety check: PassCC-BY-4.0
UniProt Database Accessdavila7/claude-code-templates32k15 repos~1.7kAutomated safety check: PassMIT
Pride Databasemajiayu000/claude-skill-registry6662 repos~8.2kAutomated safety check: PassApache-2.0
Database Lookupmajiayu000/claude-skill-registry6661 repos~7kAutomated safety check: NotesMIT
Bio Ensembl RESTGPTomics/bioSkills1.2k2 repos~3.6kAutomated safety check: PassMIT
Pride FetchClawBio/ClawBio1.2k—~4.2kAutomated safety check: PassMIT

Similar skills

  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    32k GitHub starsUsed in 15 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • Pride Database

    majiayu000/claude-skill-registry

    Search the PRIDE Archive v3 REST API for proteomics datasets: discover projects by keyword + faceted filters (organism, instrument, disease, software), fetch project metadata, list and download…

    666 GitHub starsUsed in 2 repos~8.2k tokens
    Backend & APIsAuto-check passed
  • Database Lookup

    majiayu000/claude-skill-registry

    Search 78 public scientific, biomedical, materials science, and economic databases via REST APIs.

    666 GitHub starsUsed in 1 repo~7k tokens
    Research & ScienceAuto-check: notes
  • Bio Ensembl REST

    GPTomics/bioSkills

    Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Pride Fetch

    ClawBio/ClawBio

    Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.

    1.2k GitHub stars~4.2k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Ensembl Database

    aipoch/medical-research-skills

    Access Ensembl REST API for vertebrate genomic data; use when you need gene/ID lookups, sequence retrieval, variant effect prediction (VEP), or homology/assembly coordinate mapping.

    2k GitHub stars~1.5k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 163 skills in this repo
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    370 GitHub stars~3.2k tokensUpdated 9 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    370 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    370 GitHub stars~6.9k tokensUpdated 9 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub stars~2.3k tokensUpdated 9 days ago
    Auto-check passed
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 2 repos~5.8k tokens
    Auto-check passed

Works with

Questions about Interpro Database

What does Interpro Database do?

Query InterPro REST API for protein domain architecture, family classification, and member-DB integration. Interpro Database is an agent skill from jaechang-hits/SciAgent-Skills. Query InterPro REST API for protein domain architecture, family classification, and member-DB integration.

When should I use Interpro Database?

Interpro Database fits situations like: tasks that involve Bioinformatics; tasks that involve REST APIs.

How do I install Interpro Database in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill interpro-database -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/interpro-database in jaechang-hits/SciAgent-Skills) into .claude/skills/interpro-database in your project. Claude Code loads it when a task matches its description.

How do I install Interpro Database in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill interpro-database -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/interpro-database in jaechang-hits/SciAgent-Skills) into .agents/skills/interpro-database in your project. Codex loads it when a task matches its description.

Can I use Interpro Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill interpro-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/interpro-database, .gemini/skills/interpro-database, .github/skills/interpro-database and .opencode/skills/interpro-database in your project.

What does Interpro Database need to run?

Going by SKILL.md and its folder, Interpro Database needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Interpro Database access the network?

SKILL.md names 4 domains. In commands or code: ebi.ac.uk and rest.uniprot.org; the agent is likely to contact these when it follows the instructions. As links in the text: doi.org and interpro-documentation.readthedocs.io. This is read from the text; nothing was executed.

Is Interpro Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Interpro Database use?

Interpro Database is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Interpro Database use?

About 7.6k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Interpro Database?

Skills that share tags, products or a category with Interpro Database: UniProt Database Access (davila7/claude-code-templates, 32k stars), Pride Database (majiayu000/claude-skill-registry, 666 stars), Database Lookup (majiayu000/claude-skill-registry, 666 stars) and Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Interpro Database?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 163 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.