Agent skill

Uniprot Database

by google-deepmind in google-deepmind/science-skills

Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef.

Apache-2.0Auto-check passedResearch & Science

Install Uniprot Database

skills CLI
$ npx skills add google-deepmind/science-skills --skill uniprot-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-deepmind/science-skills uniprot-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/uniprot_database .claude/skills/uniprot-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
uniprot-database
GitHub stars
3.2k
Used in
1 other repo
Token cost
~3.1k tokens
SKILL.md length
1,298 words
Files
6 (incl. scripts, references)
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef.

  • Works in 2 steps: uv: Read the uv skill and follow its… → User Notification: If…
  • Searching for proteins
  • SKILL.md covers Prerequisites, Overview, Core Rules and Use Cases, plus 5 more sections
  • Runs Python scripts from its folder; calls uv; reaches purl.uniprot.org and w3.org

What it does

Uniprot Database is an agent skill from google-deepmind/science-skills. Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef. Use when searching for proteins, mapping identifiers, or retrieving functional annotations and publications. Don't use for sequence alignment, protein folding, or sequence similarity search (use specialized skills for those tasks).

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/id_mapping_databases.md`, `references/search_query_fields.md` and `references/sparql_examples.md`).

It sits in Research & Science, covering Protein structure and design and Vector databases. It works with UniProt. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.

When your agent uses it

  • Searching for proteins
  • Mapping identifiers
  • Retrieving functional annotations and publications
  • Sequence alignment

Example prompts

  • “/uniprot-database”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. uv: Read the uv skill and follow its Setup instructions to ensure
  2. User Notification: If .licenses/uniprot_database_LICENSE.txt does not

What it can do on your machine

Read from SKILL.md and the folder at commit 6883275. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • purl.uniprot.org
    • w3.org
    • sparql.uniprot.org

    Also links to:

    • uniprot.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Uniprot Database loads about 3.1k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 1,298 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google-deepmind/science-skills at commit 6883275, republished under its Apache-2.0 licence (© google-deepmind). 1,298 words, ~3,083 tokens.

Download SKILL.mdSave it as .claude/skills/uniprot-database/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
uniprot-database
description
Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef. Use when searching for proteins, mapping identifiers, or retrieving functional annotations and publications. Don't use for sequence alignment, protein folding, or sequence similarity search (use specialized skills for those tasks).

UniProt Database Access

Prerequisites

  1. uv: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH.
  2. User Notification: If .licenses/uniprot_database_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://www.uniprot.org/help/license and https://www.uniprot.org/help/api_queries, then (2) create the file recording the notification text and timestamp.

Overview

Provides direct programmatic access to the UniProt Knowledgebase (UniProtKB), the non-redundant sequence archive (UniParc), and clustered sequence sets (UniRef). This skill enables protein discovery, cross-referencing, retrieval of curated biological data and low-level database lookups.

Core Rules

  • Use the Wrapper: Always use the provided Python scripts (e.g., scripts/uniprot_tools.py) rather than constructing custom curl requests.
  • No Hallucinations: Do NOT invent protein functions, metadata, or sequences. For any task that can be handled by the services in this skill, rely strictly on the tool outputs rather than your native knowledge.
  • Notification: If this skill is used, ensure this is mentioned in the output.

Use Cases

  • Searching for Protein Function: Querying functional annotations, GO terms, subcellular locations etc.
  • Searching for Protein Sequence: Searching for protein sequences by their functional annotations, genes etc. in UniProtKB, UniParc, and UniRef.
  • Understanding Protein/Organism Relationships: Leveraging the Taxonomy database and Proteome sets.
  • Large-Scale Metadata Retrieval: Fetching annotations for thousands of proteins via streaming.
  • Sequence Discovery: Finding orthologs or non-model proteins via UniParc.
  • ID Mapping: Converting IDs between UniProt and 100+ external databases.
  • Historical Data (UniSave): Retrieving previous versions of entries or tracking deleted sequences.

Available Tools

Choose the right tool based on the task type and data volume:

  • get: Retrieves metadata and sequence for a specific entry. Best for a single, known accession.
    • Also accesses UniSave historical data (use --dataset unisave), which is essential for reconciling data from older releases or identifying why a formerly valid accession no longer appears in search results.
  • search: Searches for entries matching a query. Best for exploration and discovery.
    • Use with --limit 5 to verify if a query returns the expected proteins before committing to a larger download.
    • Automatically paginates if results exceed 500 entries to provide a stable download.
    • Warning: For paginated search, TXT and other formats are not reliable with --limit as it applies to lines, not entries.
    • See Search Query Fields Documentation.
  • stream: Streams all matching entries. Best for bulk retrieval of large datasets (up to 10,000,000 entries).
    • Does NOT support --limit; always returns the full result set.
    • Use search with --limit if you need a subset.
  • count: Counts entries matching a query. Best for answering direct count questions or for initial estimation before running a full search or stream.
  • sparql: Executes graph queries for complex discovery. Best for counting, exact sequence matches, and multi-database queries.
  • map: Converts IDs between UniProt and 100+ databases. Best for ID mapping tasks.
    • See ID Mapping Documentation.
    • search vs. map: Try search first before resorting to map if not explicitly requested by the user. E.g., an external ID might be searchable in UniParc but fail to map to UniProtKB.

Workflows

Typical Protein Research Workflow

Copy this checklist and track progress:

  • Step 1: Identify target protein(s) and organism(s).
  • Step 2: Search UniProtKB for reviewed entries (reviewed:true).
  • Step 3: If no reviewed entries, search unreviewed or use UniParc for sequence discovery.
  • Step 4: Map external IDs (e.g., Ensembl, PDB) to UniProt Accessions if necessary.
  • Step 5: Retrieve functional metadata or sequence in desired format (JSON, FASTA).
Handling Search Misses (e.g. Gene Search in Non-Model Organisms)

If a direct query (e.g., gene:SYMBOL) fails:

  1. Pivot to Protein Name: Search for the common protein name (e.g., protein_name:Alpha-crystallin A).
  2. Use UniParc: Search the UniParc dataset, which integrates sequences from across all of life, even if they aren't fully annotated in UniProtKB.
  3. Check Orthologs/Canonical: Resolve the Human/Mouse ortholog first to find the correct naming/mnemonic.
Bulk Retrieval Priorities

[!IMPORTANT] Always prefer stream or sparql for bulk data. search is suitable for exploration; if results exceed 500 entries, it automatically paginates to provide a stable download.

  • Priority 0: count: ALWAYS check the result count before running a search or stream.
  • Priority 1: stream: The primary method for bulk data retrieval (up to 10M entries). Does NOT support --limit; always returns all results.
  • Priority 2: sparql: Best for complex filtering and exact matching during retrieval.
Sequence-Based Search (Exact Match)

[!IMPORTANT] Use SPARQL when searching for a protein by its full amino acid sequence. The REST API /search endpoint does not support direct sequence-string lookups. For any non-exact match use specialized sequence similarity search skills. Use UniParc if you cannot find query in UniProt.

SPARQL Query Pattern (UniProt):

text
PREFIX up: <http://purl.uniprot.org/core/>
PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
SELECT ?protein ?name WHERE {
  ?protein a up:Protein ;
           up:sequence/rdf:value "SEQUENCE_HERE" .
  OPTIONAL {
    ?protein up:recommendedName/up:fullName ?name .
  }
}

SPARQL Query Pattern (UniParc):

text
PREFIX up: <http://purl.uniprot.org/core/>
PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>

SELECT ?uniparc ?val WHERE {
  GRAPH <http://sparql.uniprot.org/uniparc> {
    ?uniparc a up:Sequence ;
             rdf:value ?val .
    FILTER (?val = "SEQUENCE_HERE")
  }
}
Counting Entries Efficiently

[!IMPORTANT] Use count or SPARQL for counting entries (e.g., "How many proteins in Human?").

Counting Pattern (Proteins per Organism):

text
PREFIX up: <http://purl.uniprot.org/core/>
PREFIX taxon: <http://purl.uniprot.org/taxonomy/>
SELECT (COUNT(?protein) AS ?count) WHERE {
  ?protein a up:Protein ;
           up:reviewed true ;
           up:organism taxon:9606 .
}
Show full SKILL.md (510 more words)Show less
REST Search Syntax
  • No Commas in Lists: Commas are treated as literals. Use capitalized OR to separate items.
    • Grouped: accession:(P12345 OR P67890)
    • Repeated: accession:P12345 OR accession:P67890
  • Space = AND: E.g., gene:p53 human searches for both.

Example Commands

Below are example commands for each mode of uniprot_tools.py.

Count total number of entries for a given query.

bash
uv run scripts/uniprot_tools.py count "taxonomy_id:9606"

Search for entries.

bash
uv run scripts/uniprot_tools.py search "gene:p53 AND reviewed:true" --limit 5

Retrieve a single entry by accession.

bash
uv run scripts/uniprot_tools.py get P04637

Retrieve Historical/Deleted Entry (UniSave).

bash
uv run scripts/uniprot_tools.py get P04637 --dataset unisave

Stream large result sets for bulk retrieval (returns ALL matched entries, no --limit support).

bash
uv run scripts/uniprot_tools.py stream "taxonomy_id:9606 AND reviewed:true" --format tsv --fields accession,gene_names > human_reviewed.tsv

Map IDs from one database to another.

bash
uv run scripts/uniprot_tools.py map "P04637" --from_db UniProtKB_AC-ID --to_db Gene_Name

Execute graph queries with SPARQL.

bash
uv run scripts/uniprot_tools.py sparql 'PREFIX up: <http://purl.uniprot.org/core/> SELECT ?protein WHERE { ?protein a up:Protein ; up:reviewed true . } LIMIT 5'

Common Mistakes

  • Using name: instead of protein_name:: name: is not a supported query term, use protein_name: instead.
  • Ignoring UniParc: Non-model organisms might only exist in UniParc.
  • Confusing Accession with UPI: UniProtKB Accessions (e.g., P04637) are linked to functional metadata; UniParc IDs (UPI...) are for sequences only. You can find cross-references from UniParc IDs to UniProtKB Accessions using the ID Mapping tool.
  • Using UniProtKB-AC as Target in ID Mapping: Use UniProtKB instead.
  • Giving up on Complex Queries: If a complex search query fails, try to use SPARQL instead of giving up.
  • Using IDs Without Verifying Meaning: NEVER assume you know the meaning of an ID (e.g. keyword, GO term, Pfam ID etc.). ALWAYS look up the natural language description/meaning of an ID in UniProt before using it for search to ensure it matches your intended search term.
  • Ignoring Citation Noise in Broad Searches: Broad text searches (search "term") frequently return false positives (e.g., common maintenance proteins) because UniProt searches full metadata, including publication titles. ALWAYS prefer field-specific filters like cc_function: or protein_name: for functional discovery.
  • Forgetting to Quote Short Search Terms: Short, unquoted terms (e.g., lanM) can match substrings in organism names (e.g., Lancefieldella) or other fields. Use quotes and field prefixes (e.g., gene:lanM) to isolate true hits.
  • Manipulating Protein Sequences Directly: Always use code and tools for sequence-based operations. Do not attempt to edit, truncate, or modify protein sequences manually.
  • Over-using Search for Bulk Data: DO NOT use search for retrieving millions of entries if stream or sparql can do the job. Streaming is more efficient for very large datasets. Note that stream has a hard limit of 10,000,000 outputs and does NOT support --limit.
  • Forgetting to Check Data Volume: ALWAYS perform a count before running a search without --limit or before using stream. Unlimited queries can take a long time and consume significant resources if millions of entries are returned.
  • Using --limit with stream: The stream command does NOT support --limit. If you need a limited number of results, use search with --limit instead.
  • Forgetting the License Notice: Do not neglect to state that the UniProt Database was used and to advise the user to review the licensing terms when presenting results for the first time. Even if the task is concise, this attribution is required in the first response containing UniProt data.

Reference Materials

© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/uniprot_database of google-deepmind/science-skills.

  • SKILL.md
  • references/citation.bib
  • references/id_mapping_databases.md
  • references/search_query_fields.md
  • references/sparql_examples.md
  • scripts/uniprot_tools.py

Open the folder on GitHubat commit 6883275

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in google-deepmind/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Uniprot Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Uniprot Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Uniprot Database this skillgoogle-deepmind/science-skills3.2k1 repos~3.1kAutomated safety check: PassApache-2.0
Pdb Databasejaechang-hits/SciAgent-Skills3701 repos~7.7kAutomated safety check: PassBSD-3-Clause
Tooluniverseynulihao/AgentSkillOS6173 repos~2.5kAutomated safety check: PassNone
Bio DB ToolsDrugClaw/DrugClaw125—~1.4kAutomated safety check: PassApache-2.0
Ggetdavila7/claude-code-templates32k11 repos~6.3kAutomated safety check: PassMIT
Alphafold Databasedavila7/claude-code-templates32k10 repos~4kAutomated safety check: PassMIT

Similar skills

  • Pdb Database

    jaechang-hits/SciAgent-Skills

    Query RCSB PDB (200K+ structures) via the public REST + GraphQL APIs with plain requests (no SDK).

    370 GitHub starsUsed in 1 repo~7.7k tokens
    Research & ScienceAuto-check passed
  • Tooluniverse

    ynulihao/AgentSkillOS

    A skill your agent uses when working with scientific research tools and workflows across bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery.

    617 GitHub starsUsed in 3 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Bio DB Tools

    DrugClaw/DrugClaw

    Query public biology databases and APIs including UniProt, RCSB PDB, AlphaFold DB, ClinVar, dbSNP, gnomAD, Ensembl, GEO, InterPro, KEGG, OpenTargets, Reactome, and STRING.

    125 GitHub stars~1.4k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Alphafold Database

    davila7/claude-code-templates

    Access AlphaFold's 200M+ AI-predicted protein structures. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~4k tokens
    Research & ScienceAuto-check passed
  • Pdb

    adaptyvbio/protein-design-skills

    Fetch and analyze protein structures from RCSB PDB. An agent skill from adaptyvbio/protein-design-skills.

    163 GitHub starsUsed in 4 repos~1.4k tokens
    Research & ScienceAuto-check passed

More from google-deepmind/science-skills

All 40 skills in this repo
  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 3 repos~3.4k tokens
    Auto-check: notes
  • Gtex Database

    google-deepmind/science-skills

    A skill your agent uses when you want to retrieve quantitative RNA expression data and variant eQTL information from the GTEx (Genotype-Tissue Expression) Project across 54 non-diseased tissue sites.

    3.2k GitHub starsUsed in 3 repos~1.5k tokens
    Auto-check passed
  • Human Protein Atlas Database

    google-deepmind/science-skills

    A skill your agent uses when you want to retrieve semi-quantitative protein expression and spatial localisation data from the Human Protein Atlas (HPA).

    3.2k GitHub starsUsed in 3 repos~1.6k tokens
    Auto-check passed
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Auto-check: notes

Works with

Questions about Uniprot Database

What does Uniprot Database do?

Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef. Uniprot Database is an agent skill from google-deepmind/science-skills. Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef.

When should I use Uniprot Database?

Uniprot Database fits situations like: searching for proteins; mapping identifiers; retrieving functional annotations and publications; sequence alignment.

How do I install Uniprot Database in Claude Code?

Run `npx skills add google-deepmind/science-skills --skill uniprot-database -a claude-code`. Or copy the skill folder (skills/uniprot_database in google-deepmind/science-skills) into .claude/skills/uniprot-database in your project. Claude Code loads it when a task matches its description.

How do I install Uniprot Database in Codex?

Run `npx skills add google-deepmind/science-skills --skill uniprot-database -a codex`. Or copy the skill folder (skills/uniprot_database in google-deepmind/science-skills) into .agents/skills/uniprot-database in your project. Codex loads it when a task matches its description.

Can I use Uniprot Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill uniprot-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/uniprot-database, .gemini/skills/uniprot-database, .github/skills/uniprot-database and .opencode/skills/uniprot-database in your project.

What does Uniprot Database need to run?

Going by SKILL.md and its folder, Uniprot Database needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Uniprot Database access the network?

SKILL.md names 4 domains. In commands or code: purl.uniprot.org, w3.org and sparql.uniprot.org; the agent is likely to contact these when it follows the instructions. As links in the text: uniprot.org. This is read from the text; nothing was executed.

Is Uniprot Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Uniprot Database use?

Uniprot Database is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Uniprot Database use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.5k tokens, read only when the agent opens those files.

What are the alternatives to Uniprot Database?

Skills that share tags, products or a category with Uniprot Database: Pdb Database (jaechang-hits/SciAgent-Skills, 370 stars), Tooluniverse (ynulihao/AgentSkillOS, 617 stars), Bio DB Tools (DrugClaw/DrugClaw, 125 stars) and Gget (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Uniprot Database?

google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,220 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on September 15, 2026.

Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.