Agent skill

Interpro Database

by google-deepmind in google-deepmind/science-skills

Identify domains, families, and sites in proteins; find all proteins in a family or sharing a domain; explore species distribution for a domain; annotate genomes with protein families and GO terms.

Apache-2.0Auto-check passedResearch & Science

Install Interpro Database

skills CLI
$ npx skills add google-deepmind/science-skills --skill interpro-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-deepmind/science-skills interpro-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/interpro_database .claude/skills/interpro-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
interpro-database
GitHub stars
3.2k
Used in
1 other repo
Token cost
~4.2k tokens
SKILL.md length
1,728 words
Files
5 (incl. scripts, references)
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Identify domains, families, and sites in proteins; find all proteins in a family or sharing a domain; explore species distribution for a domain; annotate genomes with protein families and GO terms.

  • Works in 4 steps: Find matching architectures (ida_search) → Fetch proteins for those architectures… → Determining all protein domains → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Prerequisites, Overview, Core Rules and Valid Source Databases…, plus 6 more sections
  • Runs Python scripts from its folder; calls uv

What it does

Interpro Database is an agent skill from google-deepmind/science-skills. Identify domains, families, and sites in proteins; find all proteins in a family or sharing a domain; explore species distribution for a domain; annotate genomes with protein families and GO terms. InterPro combines 14 databases (e.g., Pfam, CDD) into one searchable resource. InterPro-N significantly expands annotation and sequence coverage with deep learning. Includes domain architecture (IDA) search.

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/api_reference.md` and `scripts/interpro_client.py`).

It sits in Research & Science, covering Bioinformatics and Deep learning. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve Deep learning

Example prompts

  • “/interpro-database”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Find matching architectures (ida_search)
  2. Fetch proteins for those architectures (ida)
  3. Determining all protein domains
  4. Fetching all PDB structures for an Entry

What it can do on your machine

Read from SKILL.md and the folder at commit 6883275. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ebi.ac.uk

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Interpro Database loads about 4.2k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 1,728 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google-deepmind/science-skills at commit 6883275, republished under its Apache-2.0 licence (© google-deepmind). 1,728 words, ~4,244 tokens.

Download SKILL.mdSave it as .claude/skills/interpro-database/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
interpro-database
description
Identify domains, families, and sites in proteins; find all proteins in a family or sharing a domain; explore species distribution for a domain; annotate genomes with protein families and GO terms. InterPro combines 14 databases (e.g., Pfam, CDD) into one searchable resource. InterPro-N significantly expands annotation and sequence coverage with deep learning. Includes domain architecture (IDA) search.

InterPro Database Access

Prerequisites

  1. uv: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH.
  2. User Notification: If .licenses/interpro_database_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://www.ebi.ac.uk/interpro/ and https://www.ebi.ac.uk/about/terms-of-use/, then (2) create the file recording the notification text and timestamp.

Overview

InterPro combines signatures from multiple, diverse databases into a single searchable resource, reducing redundancy and helping users interpret their sequence analysis results. By uniting these member databases (e.g., Pfam, CDD, SMART), InterPro capitalises on their individual strengths to produce a powerful diagnostic tool and integrated resource.

Use interpro-database to:

  • Identify what domains, families, and sites are found in a particular protein.
  • Identify all proteins that belong to a protein family or contain a particular domain, even when the names and activities of the proteins are highly variable.
  • Examine the species in which a particular protein family or domain is found.
  • Annotate genomes with protein family information and Gene Ontology (GO) terms.

This skill provides a robust utility, interpro_client.py, to interact with the InterPro API seamlessly. It natively handles rate limiting (HTTP 429), background query sleep tracking (HTTP 408), terminal errors (HTTP 404/410), and lazy pagination.

Core Rules

  • Use the Wrapper: ALWAYS execute the scripts/interpro_client.py helper script to query the database rather than accessing the database directly. The scripts automatically enforce fair use and implement retry logic.
  • For exploratory queries: ALWAYS use the CLI with a strict --limit. This allows you to rapidly understand the data schema without polluting your context window or fetching millions of results.
  • Output to file: Use the CLI with --output to output to a file rather than attempting to print it all to the console. Process the output using jq or code.
  • For more complex pipelines import the module natively into your Python scripts to consume the generator directly, preventing the need to deserialize CLI strings in large workflows.
  • Notification: If this skill is used, ensure this is mentioned in the output.

Examples:

bash
uv run ./scripts/interpro_client.py fetch protein --source_db reviewed --limit 2 --query_params tax_id=9606 --output exploratory_results.jsonl
python
import sys
sys.path.append('scripts')
from interpro_client import fetch_interpro_data
import itertools

# fetch_interpro_data lazily yields results page-by-page
results = fetch_interpro_data(
    endpoint="entry",
    source_db="pfam",
    query_params={"page_size": 10}
)
for match in itertools.islice(results, 10):
    print(match["metadata"]["accession"])
4 Ways to Construct Endpoints:

The arguments strictly map to the four common API path constructions. Do not format your own / separated strings:

  1. /{endpoint} (e.g. /entry) uv run ./scripts/interpro_client.py fetch entry --limit 10 --output entries.jsonl
  2. /{endpoint}/{sourceDB} (e.g. /entry/pfam) uv run ./scripts/interpro_client.py fetch entry --source_db pfam --limit 10 --output pfam_entries.jsonl
  3. /{endpoint}/{sourceDB}/{accession} (e.g. /entry/pfam/PF00001) uv run ./scripts/interpro_client.py fetch entry --source_db pfam --accession PF00001 --limit 10 --output pf00001_entry.jsonl
  4. /{endpoint}/{sourceDB}/{linked_endpoint}/{sourceDB}/{accession} (e.g. /entry/interpro/protein/uniprot/P04637) uv run ./scripts/interpro_client.py fetch entry \ --source_db interpro \ --linked_endpoint protein \ --linked_source_db uniprot \ --linked_accession P04637 \ --limit 10 --output p04637_entries.jsonl

Valid Source Databases (--source_db)

Each endpoint only accepts specific source_db values. Using an invalid value returns a 404 error.

  • /entry (16 values): interpro, pfam, cathgene3d, ssf, panther, cdd, profile, smart, ncbifam, prosite, prints, hamap, pirsf, sfld, antifam.
  • /protein (3 values): uniprot (all), reviewed (SwissProt), unreviewed (TrEMBL).
  • /structure (1 value): pdb.
  • /taxonomy (1 value): uniprot.
  • /proteome (1 value): uniprot.
  • /set (2 values): pfam, cdd.

Quick Reference / Core Endpoints & Parameters

For a complete, exhaustive list of all query parameters, see the Full API Reference.

The API is fully open and supports 6 core endpoints. You can combine them using the linked parameters described above. Below is a nested list of the specific query parameters available for each endpoint:

  • /entry (Domain, family, active site, repeat, or homologous superfamily entries)

    • integrated: Filter by integrated status (e.g., pfam).
    • type: Filter by type (e.g., family, domain, homologous_superfamily).
    • go_term / go_category: Filter by Gene Ontology.
    • ida_search / ida_ignore / exact / ordered: Filter by domain architecture (see IDA Search section).
    • extra_fields: Request additional data (e.g., counters for match coordinates).
    • group_by / sort_by: Aggregate or sort results (valid values depend on context, see Full API Reference).
    • Example: uv run ./scripts/interpro_client.py count entry --source_db pfam --query_params type=domain --output count.jsonl
  • /protein (Protein records matching entries or domains)

    • tax_id: Filter by taxonomy ID (does not search lineage).
    • match_presence: Filter by proteins having InterPro matches (true/false).
    • is_fragment: Filter complete vs. fragment sequences.
    • group_by: Aggregate results (e.g., taxonomy).
    • extra_fields: Request sequence or match details.
    • isoforms / residues / structureinfo: Include specific sub-features.
    • conservation / extra_features: Append residue conservation flags or Mobidb/coil features (only valid for /protein/{source_db}/{accession}).
    • Example: uv run ./scripts/interpro_client.py fetch protein --source_db uniprot --limit 20 --query_params tax_id=9606 --output human_proteins.jsonl
  • /structure (PDB structures linked to InterPro entries)

    • experiment_type: Filter by experimental method (e.g., X-RAY DIFFRACTION).
    • resolution: Filter by resolution limit.
    • extra_fields: Include additional structural metadata.
    • group_by: Aggregate results.
    • Example: ./scripts/interpro_client.py fetch structure --source_db pdb --accession 1ATP --limit 10 --output 1atp_structures.jsonl
  • /taxonomy (Taxonomy distribution nodes)

    • key_species: Filter to limit to key species.
    • with_names: Include scientific names.
    • filter_by_entry / filter_by_entry_db: Filter intersection with specific entries.
    • extra_fields: Additional taxonomic metadata.
    • Example: ./scripts/interpro_client.py fetch taxonomy --source_db uniprot --accession 9606 --limit 10 --output human_taxonomy.jsonl
  • /proteome (Complete proteomes linked to InterPro)

    • extra_fields: General query expansion.
    • Example: uv run ./scripts/interpro_client.py fetch proteome --source_db uniprot --accession UP000005640 --limit 10 --output proteome.jsonl
  • /set (Curated sets of related entries, e.g., Pfam clans)

    • extra_fields: Additional metadata (only valid for /set/{sourceDB}).
    • Example: uv run ./scripts/interpro_client.py fetch set --source_db pfam --accession CL0001 --limit 10 --output pfam_clan.jsonl

InterPro provides powerful tools for searching proteins by their domain architecture (the exact combination and order of domains). Because the API does not allow querying proteins directly by multiple domains at once (e.g., "give me proteins with PF00069 AND PF00017"), finding proteins with specific domain combinations requires a two-step process.

The ida_search parameter is used on the root /entry endpoint to find all Domain Architectures (IDAs) containing the domains you specify.

  • Constraints:
    • Valid ONLY on the root /entry endpoint.
    • Cannot be combined with non-IDA parameters.
  • Modifiers (Only valid with ida_search):
    • ida_ignore: Ignores the given domains in the search (query param).
    • ordered: Ensures domains appear in the exact specified order (flag).
    • exact: Ensures the architecture matches exactly (no additional domains) (flag). Requires ordered flag to be present.

Example: Find architectures containing both a kinase domain (PF00069) and an SH2 domain (PF00017), in that exact order:

bash
uv run scripts/interpro_client.py fetch entry
  --query_params ida_search=PF00069,PF00017
  --flags ordered exact
  --output architectures.jsonl

Note: This returns the architectures and their unique ida_ids, not all individual proteins.

Step 2: Fetch proteins for those architectures (ida)

Once you have the ida_ids (e.g., 619edbb...) from Step 1, you can fetch all the actual proteins that share that precise layout by filtering the /protein endpoint.

Constraints:

  • Valid on /protein and /entry/{sourceDB}/{accession} endpoints.

Example: Fetch proteins matching one of the architecture IDs from Step 1:

bash
uv run scripts/interpro_client.py fetch protein
  --source_db uniprot
  --query_params ida=619edbb2b445bfa3ad51bd894e3c115b025a5f25
  --output matching_proteins.jsonl

(When building pipelines or querying comprehensively, you would loop through all the ida_ids from Step 1 and run Step 2 for each one).

Show full SKILL.md (653 more words)Show less

InterPro Entry Types

Each InterPro entry is assigned a type indicating what you can infer when a protein matches the entry:

  • Domain: Distinct functional, structural or sequence units that may exist in a variety of biological contexts. Example: PH domain or classical C2H2 zinc finger.
  • Family: A group of proteins sharing a common evolutionary origin reflected by related functions, sequence similarities, or primary/secondary/tertiary structures.
  • Homologous Superfamily: Proteins sharing an evolutionary origin reflected by structural similarity but often displaying very low sequence similarity. Usually comprises signatures from the SUPERFAMILY and CATH-Gene3D databases.
  • Repeat: A short sequence that is typically repeated within a protein, often <50 amino acids long. Example: Leucine Rich Repeats or WD40 repeats.
  • Site: Includes Active site (sequence containing conserved residues for catalytic activity) and Binding site (sequence containing conserved residues forming a protein interaction site).

InterPro-N Predictions

InterPro-N is a deep-learning-based extension of the standard InterPro database. It utilizes an AI architecture inspired by computer vision to treat protein sequence annotation as a "panoptic segmentation" task, labeling residues and distinguishing between domains.

When to use InterPro-N

Standard InterPro signatures are the "gold standard" and should not be discarded in favor of InterPro-N predictions. Use InterPro-N primarily to fill in gaps or refine results.

In addition to standard InterPro:

  • Analyzing "Dark Matter" (Uncharacterised Proteins): Use when a protein returns no hits in standard InterPro. InterPro-N excels at identifying remote homologs.
  • Resolving Complex Repeats: Use for proteins with multiple tandem repeats (e.g., TPR or WD40) where standard HMMs might merge or miss them.
  • Predicting Discontinuous Domains: Use when a domain sequence is interrupted by a completely different inserted sequence.

Instead of standard InterPro (specific scenarios):

  • Precise Boundary Delineation: When you need more accurate start-and-stop coordinates for a domain than fuzzy standard hits provide.
  • Large-Scale Metagenomic Screening: For initial high-recovery screening of fragmented or highly divergent sequences.
Fallback Strategy: Checking InterPro-N

When you are asked to find annotations for a protein and standard InterPro queries return no results or no annotations, you MUST check InterPro-N as a fallback.

Example Scenario: If a user asks to "List the SSF annotations for protein X" and the standard query returns no hits, you should retry the query with the interpro_n flag.

This fallback is crucial because InterPro-N can identify remote homologs and domains in "dark matter" proteins that standard methods miss.

If found, ALWAYS report to the user that these annotations are deep learning predictions from InterPro-N.

How to Use

InterPro-N predictions are accessed by passing the interpro_n flag to the protein endpoint with uniprot as the source database.

Via CLI:

bash
uv run ./scripts/interpro_client.py fetch protein
    --source_db uniprot
    --accession A0A096LNN2
    --flags interpro_n
    --output A0A096LNN2_interpro_n.jsonl

Via Python Pipeline:

python
results = fetch_interpro_data(
    endpoint="protein",
    source_db="uniprot",
    accession="A0A096LNN2",
    flags=["interpro_n"])

Strict Lookup Rules

  1. Always Use UniProt Accessions, NEVER Gene Names: When looking up proteins in InterPro, you MUST use their UniProt Accessions (e.g. P04637). InterPro does not natively support or reliably map gene names (e.g. TP53). If the user provides a gene name, you must use a database like Ensembl or UniProt first to resolve it to an accession.

  2. NEVER Iterate to Count: When asked for an aggregate count (e.g., "How many domains are there?"), you MUST read the count field from the initial API JSON response using the get_interpro_count() helper. NEVER iterate over the fetch_interpro_data generator to tally elements. Iterating over an endpoint with 50,000+ entries just to count them silently hangs the agent and abuses the API. Every time. No exceptions.

    ✅ Correct:

    Via CLI:

    bash
    uv run ./scripts/interpro_client.py count entry
        --source_db interpro
        --query_params type=domain
        --output count.json

    Via Python Pipeline:

    python
    from interpro_client import get_interpro_count
    cnt = get_interpro_count(
        endpoint="entry",
        source_db="interpro",
        query_params={"type": "domain"},
    )

    ❌ Wrong (Iterating over fetch):

    bash
    # NEVER DO THIS:
    uv run ./scripts/interpro_client.py fetch entry
        --source_db interpro
        --query_params type=domain
        --output output.jsonl
        && wc -l output.jsonl

Quick examples

For detailed examples of the invocations and JSON output schemas returned by various endpoints, see the Example Responses Reference. This TSV contains command-line calls, Python equivalents, and the corresponding JSON payload structures.

1. Determining all protein domains
bash
# Fetches InterPro Entries within UniProt protein P04637
# URL equivalent: /entry/interpro/protein/uniprot/P04637
uv run ./scripts/interpro_client.py fetch entry
    --source_db interpro
    --linked_endpoint protein
    --linked_source_db uniprot
    --linked_accession P04637
    --output p04637_domains.jsonl
2. Fetching all PDB structures for an Entry
bash
# URL equivalent: /structure/pdb/entry/interpro/IPR011615
# Only fetch the first 5 structures
uv run ./scripts/interpro_client.py fetch structure
    --source_db pdb
    --linked_endpoint entry
    --linked_source_db interpro
    --linked_accession IPR011615
    --output ipr011615_structures.jsonl

© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/interpro_database of google-deepmind/science-skills.

  • SKILL.md
  • references/api_reference.md
  • references/citation.bib
  • references/example_responses.tsv
  • scripts/interpro_client.py

Open the folder on GitHubat commit 6883275

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in google-deepmind/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Interpro Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Interpro Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Interpro Database this skillgoogle-deepmind/science-skills3.2k1 repos~4.2kAutomated safety check: PassApache-2.0
tangermeme Genomic Model Analysisjmschrei/tangermeme311—~1.6kAutomated safety check: PassMIT
FlexynesisBIMSBbioinfo/flexynesis110—~2.3kAutomated safety check: PassCustom licence
Cellxgene Censusdavila7/claude-code-templates32k11 repos~3.8kAutomated safety check: PassMIT
Pixi Environment Builderxuzhougeng/wisp-science1k—~3.7kAutomated safety check: PassAGPL-3.0
Bio Chipseq Motif AnalysisGPTomics/bioSkills1.2k2 repos~4.2kAutomated safety check: PassMIT

Similar skills

  • Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.

    311 GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Flexynesis

    BIMSBbioinfo/flexynesis

    Run flexynesis, a deep-learning suite for multi-omics data integration and clinical outcome prediction (drug response, cancer subtyping, survival analysis).

    110 GitHub stars~2.3k tokensUpdated 28 days ago
    Research & ScienceAuto-check passed
  • Cellxgene Census

    davila7/claude-code-templates

    Query CZ CELLxGENE Census (61M+ cells). An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.8k tokens
    Research & ScienceAuto-check passed
  • Pixi Environment Builder

    xuzhougeng/wisp-science

    A skill your agent uses when creating, migrating, or debugging pixi environments, especially for scientific Python, bioinformatics, single-cell analysis, CUDA/PyTorch, Jupyter/VS Code kernels…

    1k GitHub stars~3.7k tokensUpdated today
    Research & ScienceAuto-check passed
  • Bio Chipseq Motif Analysis

    GPTomics/bioSkills

    Discovers de novo motifs and tests known motif enrichment in ChIP-seq, ATAC-seq, or other peak sequences using HOMER, MEME-ChIP (STREME, CentriMo, TOMTOM, FIMO), monaLisa, and AME.

    1.2k GitHub starsUsed in 2 repos~4.2k tokens
    Research & ScienceAuto-check passed
  • Bio Imaging Mass Cytometry Cell Segmentation

    FreedomIntelligence/OpenClaw-Medical-Skills

    Cell segmentation from multiplexed tissue images. An agent skill from FreedomIntelligence/OpenClaw-Medical-Skills.

    3.1k GitHub starsUsed in 1 repo~1.7k tokens
    Research & ScienceAuto-check passed

More from google-deepmind/science-skills

All 40 skills in this repo
  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 3 repos~3.4k tokens
    Auto-check: notes
  • Gtex Database

    google-deepmind/science-skills

    A skill your agent uses when you want to retrieve quantitative RNA expression data and variant eQTL information from the GTEx (Genotype-Tissue Expression) Project across 54 non-diseased tissue sites.

    3.2k GitHub starsUsed in 3 repos~1.5k tokens
    Auto-check passed
  • Human Protein Atlas Database

    google-deepmind/science-skills

    A skill your agent uses when you want to retrieve semi-quantitative protein expression and spatial localisation data from the Human Protein Atlas (HPA).

    3.2k GitHub starsUsed in 3 repos~1.6k tokens
    Auto-check passed
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Auto-check: notes

Questions about Interpro Database

What does Interpro Database do?

Identify domains, families, and sites in proteins; find all proteins in a family or sharing a domain; explore species distribution for a domain; annotate genomes with protein families and GO terms. Interpro Database is an agent skill from google-deepmind/science-skills. Identify domains, families, and sites in proteins; find all proteins in a family or sharing a domain; explore species distribution for a domain; annotate genomes with protein families and GO terms.

When should I use Interpro Database?

Interpro Database fits situations like: tasks that involve Bioinformatics; tasks that involve Deep learning.

How do I install Interpro Database in Claude Code?

Run `npx skills add google-deepmind/science-skills --skill interpro-database -a claude-code`. Or copy the skill folder (skills/interpro_database in google-deepmind/science-skills) into .claude/skills/interpro-database in your project. Claude Code loads it when a task matches its description.

How do I install Interpro Database in Codex?

Run `npx skills add google-deepmind/science-skills --skill interpro-database -a codex`. Or copy the skill folder (skills/interpro_database in google-deepmind/science-skills) into .agents/skills/interpro-database in your project. Codex loads it when a task matches its description.

Can I use Interpro Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill interpro-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/interpro-database, .gemini/skills/interpro-database, .github/skills/interpro-database and .opencode/skills/interpro-database in your project.

What does Interpro Database need to run?

Going by SKILL.md and its folder, Interpro Database needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Interpro Database access the network?

SKILL.md names 1 domain. As links in the text: ebi.ac.uk. This is read from the text; nothing was executed.

Is Interpro Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Interpro Database use?

Interpro Database is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Interpro Database use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.1k tokens, read only when the agent opens those files.

What are the alternatives to Interpro Database?

Skills that share tags, products or a category with Interpro Database: tangermeme Genomic Model Analysis (jmschrei/tangermeme, 311 stars), Flexynesis (BIMSBbioinfo/flexynesis, 110 stars), Cellxgene Census (davila7/claude-code-templates, 32k stars) and Pixi Environment Builder (xuzhougeng/wisp-science, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Interpro Database?

google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,220 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on September 15, 2026.

Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.