Agent skill

Interpro Database

by majiayu000 in majiayu000/claude-skill-registry

Query InterPro for protein family, domain, and functional site annotations.

CC0-1.0Auto-check passedResearch & Science

Install Interpro Database

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill interpro-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry interpro-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/interpro-database .claude/skills/interpro-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
interpro-database
GitHub stars
666
Used in
2 other repos
Token cost
~2.7k tokens
SKILL.md length
540 words
Files
2
Skills in repo
1,273
Repo updated
First seen
Licence
CC0-1.0

At a glance

Query InterPro for protein family, domain, and functional site annotations.

  • Works in 8 steps: InterPro REST API → Look Up a Protein → Get Specific InterPro Entry → …
  • Protein function prediction
  • SKILL.md covers Overview, When to Use This Skill, Core Capabilities and Query Workflows, plus 4 more sections
  • Reaches ebi.ac.uk

What it does

Interpro Database is an agent skill from majiayu000/claude-skill-registry. Query InterPro for protein family, domain, and functional site annotations. Integrates Pfam, PANTHER, PRINTS, SMART, SUPERFAMILY, and 11 other member databases. Use for protein function prediction, domain architecture analysis, evolutionary classification, and GO term mapping.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in Research & Science, covering Protein structure and design. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is CC0-1.0.

When your agent uses it

  • Protein function prediction
  • Domain architecture analysis
  • Evolutionary classification
  • GO term mapping

Example prompts

  • “/interpro-database”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. InterPro REST API
  2. Look Up a Protein
  3. Get Specific InterPro Entry
  4. Search Proteins by InterPro Entry
  5. Domain Architecture
  6. GO Term Mapping
  7. Batch Protein Lookup
  8. Search by Text or Taxonomy

What it can do on your machine

Read from SKILL.md and the folder at commit 2d14a69. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ebi.ac.uk

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Interpro Database loads about 2.7k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 540 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 2d14a69, republished under its CC0-1.0 licence (© majiayu000). 540 words, ~2,697 tokens.

Download SKILL.mdSave it as .claude/skills/interpro-database/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
interpro-database
description
Query InterPro for protein family, domain, and functional site annotations. Integrates Pfam, PANTHER, PRINTS, SMART, SUPERFAMILY, and 11 other member databases. Use for protein function prediction, domain architecture analysis, evolutionary classification, and GO term mapping.
license
CC0-1.0
metadata.skill-author
Kuan-lin Huang

InterPro Database

Overview

InterPro (https://www.ebi.ac.uk/interpro/) is a comprehensive resource for protein family and domain classification maintained by EMBL-EBI. It integrates signatures from 13 member databases including Pfam, PANTHER, PRINTS, ProSite, SMART, TIGRFAM, SUPERFAMILY, CDD, and others, providing a unified view of protein functional annotations for over 100 million protein sequences.

InterPro classifies proteins into:

  • Families: Groups of proteins sharing common ancestry and function
  • Domains: Independently folding structural/functional units
  • Homologous superfamilies: Structurally similar protein regions
  • Repeats: Short tandem sequences
  • Sites: Functional sites (active, binding, PTM)

Key resources:

When to Use This Skill

Use InterPro when:

  • Protein function prediction: What function(s) does an uncharacterized protein likely have?
  • Domain architecture: What domains make up a protein, and in what order?
  • Protein family classification: Which family/superfamily does a protein belong to?
  • GO term annotation: Map protein sequences to Gene Ontology terms via InterPro
  • Evolutionary analysis: Are two proteins in the same homologous superfamily?
  • Structure prediction context: What domains should a new protein structure be compared against?
  • Pipeline annotation: Batch-annotate proteomes or novel sequences

Core Capabilities

1. InterPro REST API

Base URL: https://www.ebi.ac.uk/interpro/api/

python
import requests

BASE_URL = "https://www.ebi.ac.uk/interpro/api"

def interpro_get(endpoint, params=None):
    url = f"{BASE_URL}/{endpoint}"
    headers = {"Accept": "application/json"}
    response = requests.get(url, params=params, headers=headers)
    response.raise_for_status()
    return response.json()
2. Look Up a Protein
python
def get_protein_entries(uniprot_id):
    """Get all InterPro entries that match a UniProt protein."""
    data = interpro_get(f"protein/UniProt/{uniprot_id}/entry/InterPro/")
    return data

# Example: Human p53 (TP53)
result = get_protein_entries("P04637")
entries = result.get("results", [])

for entry in entries:
    meta = entry["metadata"]
    print(f"  {meta['accession']} ({meta['type']}): {meta['name']}")
    # e.g., IPR011615 (domain): p53, tetramerisation domain
    #       IPR010991 (domain): p53, DNA-binding domain
    #       IPR013872 (family): p53 family
3. Get Specific InterPro Entry
python
def get_entry(interpro_id):
    """Fetch details for an InterPro entry."""
    return interpro_get(f"entry/InterPro/{interpro_id}/")

# Example: Get Pfam domain PF00397 (WW domain)
ww_entry = get_entry("IPR001202")
print(f"Name: {ww_entry['metadata']['name']}")
print(f"Type: {ww_entry['metadata']['type']}")

# Also supports member database IDs:
def get_pfam_entry(pfam_id):
    return interpro_get(f"entry/Pfam/{pfam_id}/")

pfam = get_pfam_entry("PF00397")
4. Search Proteins by InterPro Entry
python
def get_proteins_for_entry(interpro_id, database="UniProt", page_size=25):
    """Get all proteins annotated with an InterPro entry."""
    params = {"page_size": page_size}
    data = interpro_get(f"entry/InterPro/{interpro_id}/protein/{database}/", params)
    return data

# Example: Find all human kinase-domain proteins
kinase_proteins = get_proteins_for_entry("IPR000719")  # Protein kinase domain
print(f"Total proteins: {kinase_proteins['count']}")
5. Domain Architecture
python
def get_domain_architecture(uniprot_id):
    """Get the complete domain architecture of a protein."""
    data = interpro_get(f"protein/UniProt/{uniprot_id}/")
    return data

# Example: Get full domain architecture for EGFR
egfr = get_domain_architecture("P00533")

# The response includes locations of all matching entries on the sequence
for entry in egfr.get("entries", []):
    for fragment in entry.get("entry_protein_locations", []):
        for loc in fragment.get("fragments", []):
            print(f"  {entry['accession']}: {loc['start']}-{loc['end']}")
6. GO Term Mapping
python
def get_go_terms_for_protein(uniprot_id):
    """Get GO terms associated with a protein via InterPro."""
    data = interpro_get(f"protein/UniProt/{uniprot_id}/")

    # GO terms are embedded in the entry metadata
    go_terms = []
    for entry in data.get("entries", []):
        go = entry.get("metadata", {}).get("go_terms", [])
        go_terms.extend(go)

    # Deduplicate
    seen = set()
    unique_go = []
    for term in go_terms:
        if term["identifier"] not in seen:
            seen.add(term["identifier"])
            unique_go.append(term)

    return unique_go

# GO terms include:
# {"identifier": "GO:0004672", "name": "protein kinase activity", "category": {"code": "F", "name": "Molecular Function"}}
7. Batch Protein Lookup
python
def batch_lookup_proteins(uniprot_ids, database="UniProt"):
    """Look up multiple proteins and collect their InterPro entries."""
    import time
    results = {}
    for uid in uniprot_ids:
        try:
            data = interpro_get(f"protein/{database}/{uid}/entry/InterPro/")
            entries = data.get("results", [])
            results[uid] = [
                {
                    "accession": e["metadata"]["accession"],
                    "name": e["metadata"]["name"],
                    "type": e["metadata"]["type"]
                }
                for e in entries
            ]
        except Exception as e:
            results[uid] = {"error": str(e)}
        time.sleep(0.3)  # Rate limiting
    return results

# Example
proteins = ["P04637", "P00533", "P38398", "Q9Y6I9"]
domain_info = batch_lookup_proteins(proteins)
for uid, entries in domain_info.items():
    print(f"\n{uid}:")
    for e in entries[:3]:
        print(f"  - {e['accession']} ({e['type']}): {e['name']}")
8. Search by Text or Taxonomy
python
def search_entries(query, entry_type=None, taxonomy_id=None):
    """Search InterPro entries by text."""
    params = {"search": query, "page_size": 20}
    if entry_type:
        params["type"] = entry_type  # family, domain, homologous_superfamily, etc.

    endpoint = "entry/InterPro/"
    if taxonomy_id:
        endpoint = f"entry/InterPro/taxonomy/UniProt/{taxonomy_id}/"

    return interpro_get(endpoint, params)

# Search for kinase-related entries
kinase_entries = search_entries("kinase", entry_type="domain")

Query Workflows

Workflow 1: Characterize an Unknown Protein
  1. Run InterProScan locally or via the web (https://www.ebi.ac.uk/interpro/search/sequence/) to scan a protein sequence
  2. Parse results to identify domain architecture
  3. Look up each InterPro entry for biological context
  4. Get GO terms from associated InterPro entries for functional inference
python
# After running InterProScan and getting a UniProt ID:
def characterize_protein(uniprot_id):
    """Complete characterization workflow."""

    # 1. Get all annotations
    entries = get_protein_entries(uniprot_id)

    # 2. Group by type
    by_type = {}
    for e in entries.get("results", []):
        t = e["metadata"]["type"]
        by_type.setdefault(t, []).append({
            "accession": e["metadata"]["accession"],
            "name": e["metadata"]["name"]
        })

    # 3. Get GO terms
    go_terms = get_go_terms_for_protein(uniprot_id)

    return {
        "families": by_type.get("family", []),
        "domains": by_type.get("domain", []),
        "superfamilies": by_type.get("homologous_superfamily", []),
        "go_terms": go_terms
    }
Workflow 2: Find All Members of a Protein Family
  1. Identify the InterPro family entry ID (e.g., IPR000719 for protein kinases)
  2. Query all UniProt proteins annotated with that entry
  3. Filter by organism/taxonomy if needed
  4. Download FASTA sequences for phylogenetic analysis
Show full SKILL.md (224 more words)Show less
Workflow 3: Comparative Domain Analysis
  1. Collect proteins of interest (e.g., all paralogs)
  2. Get domain architecture for each protein
  3. Compare domain compositions and orders
  4. Identify domain gain/loss events

API Endpoint Summary

EndpointDescription
/protein/UniProt/{id}/Full annotation for a protein
/protein/UniProt/{id}/entry/InterPro/InterPro entries for a protein
/entry/InterPro/{id}/Details of an InterPro entry
/entry/Pfam/{id}/Pfam entry details
/entry/InterPro/{id}/protein/UniProt/Proteins with an entry
/entry/InterPro/Search/list InterPro entries
/taxonomy/UniProt/{tax_id}/Proteins from a taxon
/structure/PDB/{pdb_id}/Structures mapped to InterPro

Member Databases

DatabaseFocus
PfamProtein domains (HMM profiles)
PANTHERProtein families and subfamilies
PRINTSProtein fingerprints
ProSitePatternsAmino acid patterns
ProSiteProfilesProtein profile patterns
SMARTProtein domain analysis
TIGRFAMJCVI curated protein families
SUPERFAMILYStructural classification
CDDConserved Domain Database (NCBI)
HAMAPMicrobial protein families
NCBIfamNCBI curated TIGRFAMs
Gene3DCATH structural classification
PIRSRPIR site rules

Best Practices

  • Use UniProt accession numbers (not gene names) for the most reliable lookups
  • Distinguish types: family gives broad classification; domain gives specific structural/functional units
  • InterProScan is faster for novel sequences: For sequences not in UniProt, submit to the web service
  • Handle pagination: Large result sets require iterating through pages
  • Combine with UniProt data: InterPro entries often include links to UniProt, PDB, and GO

Additional Resources

© majiayu000, CC0-1.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/analysis/interpro-database of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 2d14a69

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Interpro Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Interpro Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Interpro Database this skillmajiayu000/claude-skill-registry6662 repos~2.7kAutomated safety check: PassCC0-1.0
Msa Structure Prediction PipelineNVIDIA/skills3.5k1 repos~1.6kAutomated safety check: NotesApache-2.0
Fda Databasejaechang-hits/SciAgent-Skills3701 repos~4.5kAutomated safety check: PassCC0-1.0
Pdb Structure APIwentorai/research-plugins2981 repos~1.9kAutomated safety check: PassMIT
Dailymed Databasejaechang-hits/SciAgent-Skills3701 repos~6kAutomated safety check: PassCC0-1.0
Ddinter Databasejaechang-hits/SciAgent-Skills3701 repos~7.3kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Official

    NOTE: your protein sequence and the retrieved MSA alignment are transmitted to external NVIDIA-hosted APIs (health.api.nvidia.com) on every call.

    3.5k GitHub starsUsed in 1 repo~1.6k tokens
    Research & ScienceAuto-check: notes
  • Fda Database

    jaechang-hits/SciAgent-Skills

    Query openFDA REST API for adverse events (FAERS), labeling, product info, recalls, enforcement.

    370 GitHub starsUsed in 1 repo~4.5k tokens
    Research & ScienceAuto-check passed
  • Pdb Structure API

    wentorai/research-plugins

    Search and retrieve 3D protein structures from the RCSB Protein Data Bank

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Research & ScienceAuto-check passed
  • Dailymed Database

    jaechang-hits/SciAgent-Skills

    Query FDA drug labels (DailyMed) via REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~6k tokens
    Research & ScienceAuto-check passed
  • Ddinter Database

    jaechang-hits/SciAgent-Skills

    Query DDInter drug-drug interactions via REST API (1.7M+ interactions, 2,400+ drugs).

    370 GitHub starsUsed in 1 repo~7.3k tokens
    Research & ScienceAuto-check passed
  • Unichem Database

    jaechang-hits/SciAgent-Skills

    Cross-reference compound IDs across 20+ databases (ChEMBL, DrugBank, PubChem, ChEBI, PDB, SureChEMBL, HMDB, DrugCentral, BindingDB) via UniChem REST API.

    370 GitHub starsUsed in 1 repo~8.8k tokens
    Research & ScienceAuto-check passed

More from majiayu000/claude-skill-registry

All 1,273 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Pyzotero

    majiayu000/claude-skill-registry

    Interact with Zotero reference management libraries using the pyzotero Python client.

    666 GitHub starsUsed in 5 repos~1.6k tokens
    Auto-check: notes
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed

Questions about Interpro Database

What does Interpro Database do?

Query InterPro for protein family, domain, and functional site annotations. Interpro Database is an agent skill from majiayu000/claude-skill-registry. Query InterPro for protein family, domain, and functional site annotations.

When should I use Interpro Database?

Interpro Database fits situations like: protein function prediction; domain architecture analysis; evolutionary classification; GO term mapping.

How do I install Interpro Database in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill interpro-database -a claude-code`. Or copy the skill folder (skills/analysis/interpro-database in majiayu000/claude-skill-registry) into .claude/skills/interpro-database in your project. Claude Code loads it when a task matches its description.

How do I install Interpro Database in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill interpro-database -a codex`. Or copy the skill folder (skills/analysis/interpro-database in majiayu000/claude-skill-registry) into .agents/skills/interpro-database in your project. Codex loads it when a task matches its description.

Can I use Interpro Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill interpro-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/interpro-database, .gemini/skills/interpro-database, .github/skills/interpro-database and .opencode/skills/interpro-database in your project.

What does Interpro Database need to run?

SKILL.md names no scripts, command-line tools or credentials: Interpro Database is instructions for the agent only. Our summary lists: Python 3.

Does Interpro Database access the network?

SKILL.md names 2 domains. In commands or code: ebi.ac.uk; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Interpro Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Interpro Database use?

Interpro Database is published under the CC0-1.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Interpro Database use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Interpro Database?

Skills that share tags, products or a category with Interpro Database: Msa Structure Prediction Pipeline (NVIDIA/skills, 3.5k stars), Fda Database (jaechang-hits/SciAgent-Skills, 370 stars), Pdb Structure API (wentorai/research-plugins, 298 stars) and Dailymed Database (jaechang-hits/SciAgent-Skills, 370 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Interpro Database?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 1,273 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.