Agent skill

Interpro Database

by LeonChaoX in LeonChaoX/qinyan-academic-skills

Query InterPro for protein family, domain, and functional site annotations.

CC0-1.0Auto-check passedResearch & Science

Install Interpro Database

skills CLI
$ npx skills add LeonChaoX/qinyan-academic-skills --skill interpro-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeonChaoX/qinyan-academic-skills interpro-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeonChaoX/qinyan-academic-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'skills/08-蛋白质工程与结构生物学/interpro-database' .claude/skills/interpro-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
interpro-database
GitHub stars
944
Used in
1 other repo
Token cost
~2.7k tokens
SKILL.md length
540 words
Files
2 (incl. references)
Skills in repo
31
Repo updated
First seen
Licence
CC0-1.0

At a glance

Query InterPro for protein family, domain, and functional site annotations.

  • Works in 8 steps: InterPro REST API → Look Up a Protein → Get Specific InterPro Entry → …
  • Protein function prediction
  • SKILL.md covers Overview, When to Use This Skill, Core Capabilities and Query Workflows, plus 4 more sections
  • Reaches ebi.ac.uk

What it does

Interpro Database is an agent skill from LeonChaoX/qinyan-academic-skills. Query InterPro for protein family, domain, and functional site annotations. Integrates Pfam, PANTHER, PRINTS, SMART, SUPERFAMILY, and 11 other member databases. Use for protein function prediction, domain architecture analysis, evolutionary classification, and GO term mapping.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/domain_analysis.md`).

It sits in Research & Science, covering Protein structure and design. The repository describes itself as: A curated, multilingual library of 182 installable AI agent skills for end-to-end academic research—spanning literature discovery, scientific writing, grant development… The licence is CC0-1.0.

When your agent uses it

  • Protein function prediction
  • Domain architecture analysis
  • Evolutionary classification
  • GO term mapping

Example prompts

  • “/interpro-database”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. InterPro REST API
  2. Look Up a Protein
  3. Get Specific InterPro Entry
  4. Search Proteins by InterPro Entry
  5. Domain Architecture
  6. GO Term Mapping
  7. Batch Protein Lookup
  8. Search by Text or Taxonomy

What it can do on your machine

Read from SKILL.md and the folder at commit df5a498. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ebi.ac.uk

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Interpro Database loads about 2.7k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 540 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LeonChaoX/qinyan-academic-skills at commit df5a498, republished under its CC0-1.0 licence (© LeonChaoX). 540 words, ~2,697 tokens.

Download SKILL.mdSave it as .claude/skills/interpro-database/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
interpro-database
description
Query InterPro for protein family, domain, and functional site annotations. Integrates Pfam, PANTHER, PRINTS, SMART, SUPERFAMILY, and 11 other member databases. Use for protein function prediction, domain architecture analysis, evolutionary classification, and GO term mapping.
license
CC0-1.0
metadata.skill-author
Kuan-lin Huang

InterPro Database

Overview

InterPro (https://www.ebi.ac.uk/interpro/) is a comprehensive resource for protein family and domain classification maintained by EMBL-EBI. It integrates signatures from 13 member databases including Pfam, PANTHER, PRINTS, ProSite, SMART, TIGRFAM, SUPERFAMILY, CDD, and others, providing a unified view of protein functional annotations for over 100 million protein sequences.

InterPro classifies proteins into:

  • Families: Groups of proteins sharing common ancestry and function
  • Domains: Independently folding structural/functional units
  • Homologous superfamilies: Structurally similar protein regions
  • Repeats: Short tandem sequences
  • Sites: Functional sites (active, binding, PTM)

Key resources:

When to Use This Skill

Use InterPro when:

  • Protein function prediction: What function(s) does an uncharacterized protein likely have?
  • Domain architecture: What domains make up a protein, and in what order?
  • Protein family classification: Which family/superfamily does a protein belong to?
  • GO term annotation: Map protein sequences to Gene Ontology terms via InterPro
  • Evolutionary analysis: Are two proteins in the same homologous superfamily?
  • Structure prediction context: What domains should a new protein structure be compared against?
  • Pipeline annotation: Batch-annotate proteomes or novel sequences

Core Capabilities

1. InterPro REST API

Base URL: https://www.ebi.ac.uk/interpro/api/

python
import requests

BASE_URL = "https://www.ebi.ac.uk/interpro/api"

def interpro_get(endpoint, params=None):
    url = f"{BASE_URL}/{endpoint}"
    headers = {"Accept": "application/json"}
    response = requests.get(url, params=params, headers=headers)
    response.raise_for_status()
    return response.json()
2. Look Up a Protein
python
def get_protein_entries(uniprot_id):
    """Get all InterPro entries that match a UniProt protein."""
    data = interpro_get(f"protein/UniProt/{uniprot_id}/entry/InterPro/")
    return data

# Example: Human p53 (TP53)
result = get_protein_entries("P04637")
entries = result.get("results", [])

for entry in entries:
    meta = entry["metadata"]
    print(f"  {meta['accession']} ({meta['type']}): {meta['name']}")
    # e.g., IPR011615 (domain): p53, tetramerisation domain
    #       IPR010991 (domain): p53, DNA-binding domain
    #       IPR013872 (family): p53 family
3. Get Specific InterPro Entry
python
def get_entry(interpro_id):
    """Fetch details for an InterPro entry."""
    return interpro_get(f"entry/InterPro/{interpro_id}/")

# Example: Get Pfam domain PF00397 (WW domain)
ww_entry = get_entry("IPR001202")
print(f"Name: {ww_entry['metadata']['name']}")
print(f"Type: {ww_entry['metadata']['type']}")

# Also supports member database IDs:
def get_pfam_entry(pfam_id):
    return interpro_get(f"entry/Pfam/{pfam_id}/")

pfam = get_pfam_entry("PF00397")
4. Search Proteins by InterPro Entry
python
def get_proteins_for_entry(interpro_id, database="UniProt", page_size=25):
    """Get all proteins annotated with an InterPro entry."""
    params = {"page_size": page_size}
    data = interpro_get(f"entry/InterPro/{interpro_id}/protein/{database}/", params)
    return data

# Example: Find all human kinase-domain proteins
kinase_proteins = get_proteins_for_entry("IPR000719")  # Protein kinase domain
print(f"Total proteins: {kinase_proteins['count']}")
5. Domain Architecture
python
def get_domain_architecture(uniprot_id):
    """Get the complete domain architecture of a protein."""
    data = interpro_get(f"protein/UniProt/{uniprot_id}/")
    return data

# Example: Get full domain architecture for EGFR
egfr = get_domain_architecture("P00533")

# The response includes locations of all matching entries on the sequence
for entry in egfr.get("entries", []):
    for fragment in entry.get("entry_protein_locations", []):
        for loc in fragment.get("fragments", []):
            print(f"  {entry['accession']}: {loc['start']}-{loc['end']}")
6. GO Term Mapping
python
def get_go_terms_for_protein(uniprot_id):
    """Get GO terms associated with a protein via InterPro."""
    data = interpro_get(f"protein/UniProt/{uniprot_id}/")

    # GO terms are embedded in the entry metadata
    go_terms = []
    for entry in data.get("entries", []):
        go = entry.get("metadata", {}).get("go_terms", [])
        go_terms.extend(go)

    # Deduplicate
    seen = set()
    unique_go = []
    for term in go_terms:
        if term["identifier"] not in seen:
            seen.add(term["identifier"])
            unique_go.append(term)

    return unique_go

# GO terms include:
# {"identifier": "GO:0004672", "name": "protein kinase activity", "category": {"code": "F", "name": "Molecular Function"}}
7. Batch Protein Lookup
python
def batch_lookup_proteins(uniprot_ids, database="UniProt"):
    """Look up multiple proteins and collect their InterPro entries."""
    import time
    results = {}
    for uid in uniprot_ids:
        try:
            data = interpro_get(f"protein/{database}/{uid}/entry/InterPro/")
            entries = data.get("results", [])
            results[uid] = [
                {
                    "accession": e["metadata"]["accession"],
                    "name": e["metadata"]["name"],
                    "type": e["metadata"]["type"]
                }
                for e in entries
            ]
        except Exception as e:
            results[uid] = {"error": str(e)}
        time.sleep(0.3)  # Rate limiting
    return results

# Example
proteins = ["P04637", "P00533", "P38398", "Q9Y6I9"]
domain_info = batch_lookup_proteins(proteins)
for uid, entries in domain_info.items():
    print(f"\n{uid}:")
    for e in entries[:3]:
        print(f"  - {e['accession']} ({e['type']}): {e['name']}")
8. Search by Text or Taxonomy
python
def search_entries(query, entry_type=None, taxonomy_id=None):
    """Search InterPro entries by text."""
    params = {"search": query, "page_size": 20}
    if entry_type:
        params["type"] = entry_type  # family, domain, homologous_superfamily, etc.

    endpoint = "entry/InterPro/"
    if taxonomy_id:
        endpoint = f"entry/InterPro/taxonomy/UniProt/{taxonomy_id}/"

    return interpro_get(endpoint, params)

# Search for kinase-related entries
kinase_entries = search_entries("kinase", entry_type="domain")

Query Workflows

Workflow 1: Characterize an Unknown Protein
  1. Run InterProScan locally or via the web (https://www.ebi.ac.uk/interpro/search/sequence/) to scan a protein sequence
  2. Parse results to identify domain architecture
  3. Look up each InterPro entry for biological context
  4. Get GO terms from associated InterPro entries for functional inference
python
# After running InterProScan and getting a UniProt ID:
def characterize_protein(uniprot_id):
    """Complete characterization workflow."""

    # 1. Get all annotations
    entries = get_protein_entries(uniprot_id)

    # 2. Group by type
    by_type = {}
    for e in entries.get("results", []):
        t = e["metadata"]["type"]
        by_type.setdefault(t, []).append({
            "accession": e["metadata"]["accession"],
            "name": e["metadata"]["name"]
        })

    # 3. Get GO terms
    go_terms = get_go_terms_for_protein(uniprot_id)

    return {
        "families": by_type.get("family", []),
        "domains": by_type.get("domain", []),
        "superfamilies": by_type.get("homologous_superfamily", []),
        "go_terms": go_terms
    }
Workflow 2: Find All Members of a Protein Family
  1. Identify the InterPro family entry ID (e.g., IPR000719 for protein kinases)
  2. Query all UniProt proteins annotated with that entry
  3. Filter by organism/taxonomy if needed
  4. Download FASTA sequences for phylogenetic analysis
Show full SKILL.md (224 more words)Show less
Workflow 3: Comparative Domain Analysis
  1. Collect proteins of interest (e.g., all paralogs)
  2. Get domain architecture for each protein
  3. Compare domain compositions and orders
  4. Identify domain gain/loss events

API Endpoint Summary

EndpointDescription
/protein/UniProt/{id}/Full annotation for a protein
/protein/UniProt/{id}/entry/InterPro/InterPro entries for a protein
/entry/InterPro/{id}/Details of an InterPro entry
/entry/Pfam/{id}/Pfam entry details
/entry/InterPro/{id}/protein/UniProt/Proteins with an entry
/entry/InterPro/Search/list InterPro entries
/taxonomy/UniProt/{tax_id}/Proteins from a taxon
/structure/PDB/{pdb_id}/Structures mapped to InterPro

Member Databases

DatabaseFocus
PfamProtein domains (HMM profiles)
PANTHERProtein families and subfamilies
PRINTSProtein fingerprints
ProSitePatternsAmino acid patterns
ProSiteProfilesProtein profile patterns
SMARTProtein domain analysis
TIGRFAMJCVI curated protein families
SUPERFAMILYStructural classification
CDDConserved Domain Database (NCBI)
HAMAPMicrobial protein families
NCBIfamNCBI curated TIGRFAMs
Gene3DCATH structural classification
PIRSRPIR site rules

Best Practices

  • Use UniProt accession numbers (not gene names) for the most reliable lookups
  • Distinguish types: family gives broad classification; domain gives specific structural/functional units
  • InterProScan is faster for novel sequences: For sequences not in UniProt, submit to the web service
  • Handle pagination: Large result sets require iterating through pages
  • Combine with UniProt data: InterPro entries often include links to UniProt, PDB, and GO

Additional Resources

© LeonChaoX, CC0-1.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/08-蛋白质工程与结构生物学/interpro-database of LeonChaoX/qinyan-academic-skills.

  • SKILL.md
  • references/domain_analysis.md

Open the folder on GitHubat commit df5a498

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in LeonChaoX/qinyan-academic-skills, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Interpro Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Interpro Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Interpro Database this skillLeonChaoX/qinyan-academic-skills9441 repos~2.7kAutomated safety check: PassCC0-1.0
Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills3.2k2 repos~1.2kAutomated safety check: PassApache-2.0
Alphafoldadaptyvbio/protein-design-skills1643 repos~1.2kAutomated safety check: PassMIT
Pymol VisualizationChatMol/ChatMol373—~1.2kAutomated safety check: PassMIT
Complexa Binder DesignNVIDIA-BioNeMo/bionemo-agent-toolkit479—~3.1kAutomated safety check: NotesApache-2.0
Bindcraftadaptyvbio/protein-design-skills1643 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Alphafold

    adaptyvbio/protein-design-skills

    Validate protein designs using AlphaFold2 structure prediction.

    164 GitHub starsUsed in 3 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Pymol Visualization

    ChatMol/ChatMol

    Generate publication-quality molecular visualization images using PyMOL.

    373 GitHub stars~1.2k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Complexa Binder Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided…

    479 GitHub stars~3.1k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Bindcraft

    adaptyvbio/protein-design-skills

    End-to-end binder design using BindCraft hallucination. An agent skill from adaptyvbio/protein-design-skills.

    164 GitHub starsUsed in 3 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes

More from LeonChaoX/qinyan-academic-skills

All 31 skills in this repo
  • Paper Slide Deck

    LeonChaoX/qinyan-academic-skills

    Generate professional slide deck images from academic papers and content.

    944 GitHub starsUsed in 2 repos~4.6k tokens
    Auto-check passed
  • Parallel Web

    LeonChaoX/qinyan-academic-skills

    Search the web, extract URL content, and run deep research using the Parallel Chat API and Extract API.

    944 GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check: notes
  • Research Proposal

    LeonChaoX/qinyan-academic-skills

    Generate academic research proposals for PhD applications. An agent skill from LeonChaoX/qinyan-academic-skills.

    944 GitHub starsUsed in 2 repos~4.8k tokens
    Auto-check passed
  • Medical Imaging Review

    LeonChaoX/qinyan-academic-skills

    Write comprehensive literature reviews for medical imaging AI research.

    944 GitHub starsUsed in 3 repos~1.1k tokens
    Auto-check: notes
  • Depmap

    LeonChaoX/qinyan-academic-skills

    Query the Cancer Dependency Map (DepMap) for cancer cell line gene dependency scores (CRISPR Chronos), drug sensitivity data, and gene effect profiles.

    944 GitHub starsUsed in 2 repos~2.8k tokens
    Auto-check passed
  • Phylogenetics

    LeonChaoX/qinyan-academic-skills

    Build and analyze phylogenetic trees using MAFFT (multiple alignment), IQ-TREE 2 (maximum likelihood), and FastTree (fast NJ/ML).

    944 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed

Questions about Interpro Database

What does Interpro Database do?

Query InterPro for protein family, domain, and functional site annotations. Interpro Database is an agent skill from LeonChaoX/qinyan-academic-skills. Query InterPro for protein family, domain, and functional site annotations.

When should I use Interpro Database?

Interpro Database fits situations like: protein function prediction; domain architecture analysis; evolutionary classification; GO term mapping.

How do I install Interpro Database in Claude Code?

Run `npx skills add LeonChaoX/qinyan-academic-skills --skill interpro-database -a claude-code`. Or copy the skill folder (skills/08-蛋白质工程与结构生物学/interpro-database in LeonChaoX/qinyan-academic-skills) into .claude/skills/interpro-database in your project. Claude Code loads it when a task matches its description.

How do I install Interpro Database in Codex?

Run `npx skills add LeonChaoX/qinyan-academic-skills --skill interpro-database -a codex`. Or copy the skill folder (skills/08-蛋白质工程与结构生物学/interpro-database in LeonChaoX/qinyan-academic-skills) into .agents/skills/interpro-database in your project. Codex loads it when a task matches its description.

Can I use Interpro Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeonChaoX/qinyan-academic-skills --skill interpro-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/interpro-database, .gemini/skills/interpro-database, .github/skills/interpro-database and .opencode/skills/interpro-database in your project.

What does Interpro Database need to run?

SKILL.md names no scripts, command-line tools or credentials: Interpro Database is instructions for the agent only. Our summary lists: Python 3.

Does Interpro Database access the network?

SKILL.md names 2 domains. In commands or code: ebi.ac.uk; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Interpro Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Interpro Database use?

Interpro Database is published under the CC0-1.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Interpro Database use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Interpro Database?

Skills that share tags, products or a category with Interpro Database: Alphafold Database Fetch And Analyze (google-deepmind/science-skills, 3.2k stars), Alphafold (adaptyvbio/protein-design-skills, 164 stars), Pymol Visualization (ChatMol/ChatMol, 373 stars) and Complexa Binder Design (NVIDIA-BioNeMo/bionemo-agent-toolkit, 479 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Interpro Database?

LeonChaoX (a GitHub user) maintains it in LeonChaoX/qinyan-academic-skills, which has 944 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeonChaoX/qinyan-academic-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.