Agent skill

Protein Sequence Similarity Search

by google-deepmind in google-deepmind/science-skills

Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback).

Apache-2.0Auto-check: notesResearch & Science

Install Protein Sequence Similarity Search

skills CLI
$ npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-deepmind/science-skills protein-sequence-similarity-search --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/protein_sequence_similarity_search .claude/skills/protein-sequence-similarity-search && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
protein-sequence-similarity-search
GitHub stars
3.2k
Used in
1 other repo
Token cost
~2.7k tokens
SKILL.md length
1,226 words
Files
4 (incl. scripts, references)
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback).

  • Works in 4 steps: uv: Read the uv skill and follow its… → User Notification: If → .env file: Make sure the .env file… → …
  • This whenever the user provides a protein sequence
  • SKILL.md covers Prerequisites, Goal, Core Rules and Search Method Selection, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Protein Sequence Similarity Search is an agent skill from google-deepmind/science-skills. Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback). Trigger this whenever the user provides a protein sequence or FASTA file and asks to find homologues, sequence matches, or wants to infer protein function based on sequence similarity, but not when the user wants to infer protein function based on structural similarity.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `scripts/mmseqs2_search.py` and `scripts/uniprot_blast.py`).

It sits in Research & Science, covering Vector databases. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.

When your agent uses it

  • This whenever the user provides a protein sequence
  • FASTA file and asks to find homologues
  • Sequence matches
  • Wants to infer protein function based on sequence similarity

Example prompts

  • “Use the protein-sequence-similarity-search skill to search for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive…”
  • “/protein-sequence-similarity-search”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. uv: Read the uv skill and follow its Setup instructions to ensure
  2. User Notification: If
  3. .env file: Make sure the .env file exists in your home directory.
  4. USER_EMAIL (optional but recommended): Recommended by the EBI for

What it can do on your machine

Read from SKILL.md and the folder at commit 8ab7672. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ebi.ac.uk

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Protein Sequence Similarity Search loads about 2.7k tokens when it runs, and up to ~3.2k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 1,226 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:22
    3.  **`.env` file**: Make sure the `.env` file exists in your home directory.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google-deepmind/science-skills at commit 8ab7672, republished under its Apache-2.0 licence (© google-deepmind). 1,226 words, ~2,669 tokens.

Download SKILL.mdSave it as .claude/skills/protein-sequence-similarity-search/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
protein-sequence-similarity-search
description
Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback). Trigger this whenever the user provides a protein sequence or FASTA file and asks to find homologues, sequence matches, or wants to infer protein function based on sequence similarity, but not when the user wants to infer protein function based on structural similarity.

Prerequisites

  1. uv: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH.
  2. User Notification: If .licenses/protein_sequence_similarity_search_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://www.ebi.ac.uk/jdispatcher/sss/ncbiblast and https://colabfold.com, then (2) create the file recording the notification text and timestamp.
  3. .env file: Make sure the .env file exists in your home directory. Create one if it does not exist.
  4. USER_EMAIL (optional but recommended): Recommended by the EBI for BLAST job tracking, but the skill works without it. You MUST use the safe credentials protocol in the credentials skill to check for and request this credential if this skill looks relevant to the user's request.

Goal

Take a user-provided amino acid sequence (or a path to a .fasta file), search for sequence homologues using the fastest available method, generate a Markdown-formatted table of the top hits, interpret key alignment metrics, summarize the inferred protein functions, and save results locally for future programmatic analysis.

Core Rules

  • Strict Validation: For BLAST, only use database codes listed in the table below.
  • No Hallucinations: If a script throws an error or returns no hits, inform the user clearly. Do NOT invent sequence homologues.
  • Do Not Parse Output Files: Do not parse the JSON, a3m, or any other raw output files. Rely on the generated .md file for your summary. The JSON and other outputs are for subsequent tool use only.
  • Always State the Method: Every report must clearly state whether the search used the quick MMseqs2 (ColabFold API) or the slower EBI BLAST method.
  • Notification: If this skill is used, ensure this is mentioned in the output. Explicitly state that the corresponding program (MMSEQS2 or EBI BLAST) and Sequence Databases were used.

Search Method Selection

Choose the search method based on the user's request:

If the user says "quick search" or "fast search", no specific method requested / general homologue search, of if you are unsure: Run MMseqs2 (fast, default) using mmseqs2_search.py

If MMseqs2 fails (exit code 2: RATELIMIT or API error) or User explicitly requests "BLAST" or a specific BLAST database (e.g. uniprotkb_swissprot, pdb, uniprotkb_human): Run BLAST using uniprot_blast.py

Instructions

  1. Identify the query from the user. It can be a raw sequence string (e.g., "MKVLY...") or a path to a local file (e.g., "./data/sequence.fasta").

  2. Determine the search method using the list above.

Path A: MMseqs2 Search (Default)
  1. Generate File Names: Generate descriptive output file names based on the input (e.g., proteinA_mmseqs2.json and proteinA_mmseqs2.md).

  2. Execute the MMseqs2 script:

    • Default:
    uv run scripts/mmseqs2_search.py <SEQUENCE_OR_FILE> -o <generated-filename.md> -j <generated-filename.json>
    • With mgnify:
    uv run scripts/mmseqs2_search.py <SEQUENCE_OR_FILE> -o <generated-filename.md> -j <generated-filename.json> --include-mgnify
  3. The script will query the ColabFold MMseqs2 API and poll for completion. This is typically fast (under 2 minutes).

  4. If the script exits with code 2 (API failure, rate limit), automatically fall back to BLAST (Path B below). Inform the user: "MMseqs2 search failed, falling back to BLAST."

  5. Read the Results: Open and read the generated .md file.

Path B: BLAST Search (Explicit or Fallback)
  1. Database Selection & Validation: Determine the most appropriate database(s) based on the user's prompt.

    • Consult the Available BLAST Databases table below.
    • If the user specifies a taxonomic group (e.g., "Find homologues in microbes"), select the corresponding Database Code (e.g., uniprotkb_bacteria).
    • If the user explicitly requests curated hits, use uniprotkb_swissprot.
    • If no specific database is requested, do not specify --databases.
    • Validation: Ensure the database code exactly matches an entry in the table. If the user requests a database not on the list, do not proceed and provide the allowed list.
  2. Generate File Names: (e.g., proteinA_ebi_blast.json and proteinA_ebi_blast.md).

  3. This API requires the user email address to be set in the USER_EMAIL environment variable for inclusion in request header. You MUST use the safe credentials protocol in the credentials skill to check for and request this credential if this skill looks relevant to the user's request.

  4. Execute the BLAST script:

    • Default (uniprotkb):
    uv run scripts/uniprot_blast.py <SEQUENCE_OR_FILE> -o <generated-filename.md> -j <generated-filename.json>
    • Custom database:
    uv run scripts/uniprot_blast.py <SEQUENCE_OR_FILE> -o <generated-filename.md> -j <generated-filename.json> --databases <db1,db2>
  5. The script will query the EBI BLAST API and poll the server. Note: This can take up to 15 minutes; wait patiently.

  6. Read the Results: Open and read the generated .md file.

Show full SKILL.md (505 more words)Show less
Common Steps (Both Methods)
  1. Interpret the Metrics: Summarize the top 3 to 5 sequence homologues. Assess match quality using:
    • Q-Cov (Query Coverage): High percentages mean the match covers most of the query sequence.
    • E-value: Lower E-values (e.g., 1e-50) indicate extreme statistical significance.
    • Seq Identity: Provides evolutionary context (highly conserved vs. distant homologue).
  2. Perform Functional Analysis:
    • If the results table includes protein descriptions, analyze them directly: report specific protein names/functions of the top homologues and summarize the variety of functions, domains, or protein families found.
    • If the results contain only UniProt accession IDs without descriptions (common with MMseqs2), look up the protein names and functions for the top 3–5 hits using the uniprot-database skill or other appropriate methods before summarizing.
  3. Inform the user of both newly created files (.json and .md) and their locations.

Available BLAST Databases

  • uniprotkb – UniProt Knowledgebase (The UniProt Knowledgebase includes UniProtKB/Swiss-Prot and UniProtKB/TrEMBL): The UniProt Knowledgebase (UniProtKB) is the central access point for extensive curated protein information, including function, classification, and cross-references. Search UniProtKB to retrieve "everything that is known" about a particular sequence
  • uniprotkb_swissprot – UniProtKB/Swiss-Prot (The manually annotated section of UniProtKB): The manually curated subsection of the UniProt Knowledgebase
  • uniprotkb_swissprotsv – UniProtKB/Swiss-Prot isoforms (The manually annotated isoforms of UniProtKB/Swiss-Prot): The isoform sequences for the manually curated subsection of the UniProt Knowledgebase
  • uniprotkb_reference_proteomes – UniProtKB Reference Proteomes: Taxonomic subset of the UniProtKB Reference Proteomes
  • uniprotkb_trembl – UniProtKB/TrEMBL (The automatically annotated section of UniProtKB): Subsection of the UniProt Knowledgebase derived from ENA Sequence (formerly EMBL-Bank) coding sequence translations with annotation produced by an automated process
  • uniprotkb_refprotswissprot – UniProtKB Reference Proteomes plus Swiss-Prot: UniProtKB Reference Proteomes plus Swiss-Prot
  • uniprotkb_archaea – UniProtKB Archaea: Taxonomic subset of the UniProt Knowledgebase for archaea
  • uniprotkb_arthropoda – UniProtKB Arthropoda: Taxonomic subset of the UniProt Knowledgebase for arthropoda
  • uniprotkb_bacteria – UniProtKB Bacteria: Taxonomic subset of the UniProt Knowledgebase for bacteria
  • uniprotkb_complete_microbial_proteomes – UniProtKB Complete Microbial Proteomes: Taxonomic subset of the UniProt Knowledgebase for complete microbial proteomes
  • uniprotkb_eukaryota – UniProtKB Eukaryota: Taxonomic subset of the UniProt Knowledgebase for eukaryota
  • uniprotkb_fungi – UniProtKB Fungi: Taxonomic subset of the UniProt Knowledgebase for fungi
  • uniprotkb_human – UniProtKB Human: Taxonomic subset of the UniProt Knowledgebase for human
  • uniprotkb_mammals – UniProtKB Mammals: Taxonomic subset of the UniProt Knowledgebase for mammals
  • uniprotkb_nematoda – UniProtKB Nematoda: Taxonomic subset of the UniProt Knowledgebase for nematoda
  • uniprotkb_rodents – UniProtKB Rodents: Taxonomic subset of the UniProt Knowledgebase for rodents
  • uniprotkb_vertebrates – UniProtKB Vertebrates: Taxonomic subset of the UniProt Knowledgebase for vertebrates
  • uniprotkb_viridiplantae – UniProtKB Viridiplantae: Taxonomic subset of the UniProt Knowledgebase for viridiplantae
  • uniprotkb_viruses – UniProtKB Viruses: Taxonomic subset of the UniProt Knowledgebase for viruses
  • uniprotkb_enzyme – UniProtKB Enzyme: Taxonomic subset of the UniProt Knowledgebase for enzymes
  • uniprotkb_covid19 – UniProtKB COVID-19: Taxonomic subset of the UniProt Knowledgebase for COVID-19
  • uniref100 – UniProt Clusters 100% (UniRef100): The UniProt Reference Clusters (UniRef) containing sequences which are 100% identical.
  • uniref90 – UniProt Clusters 90% (UniRef90): The UniProt Reference Clusters (UniRef) containing sequences which are 90% identical.
  • uniref50 – UniProt Clusters 50% (UniRef50): The UniProt Reference Clusters (UniRef) containing sequences which are 50% identical.
  • pdb – Protein Structure Sequences (PDBe protein structure sequences): Protein sequences from structures described in the Brookhaven Protein Data Bank (PDB)

© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/protein_sequence_similarity_search of google-deepmind/science-skills.

  • SKILL.md
  • references/citation.bib
  • scripts/mmseqs2_search.py
  • scripts/uniprot_blast.py

Open the folder on GitHubat commit 8ab7672

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in google-deepmind/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Protein Sequence Similarity Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Protein Sequence Similarity Search compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Protein Sequence Similarity Search this skillgoogle-deepmind/science-skills3.2k1 repos~2.7kAutomated safety check: NotesApache-2.0
Medical Vector Searchaipoch/medical-research-skills1.9k—~2kAutomated safety check: PassMIT
Drugbank Databasedavila7/claude-code-templates33k10 repos~2.3kAutomated safety check: PassMIT
Scholar RAGjoshzyj/open-scholar-skill168—~7.4kAutomated safety check: NotesCustom licence
Comprehensive Protein AnalysisInternScience/scp1701 repos~2kAutomated safety check: PassMIT
Genimlaipoch/medical-research-skills1.9k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Medical Vector Search

    aipoch/medical-research-skills

    Vector database retrieval and evidence-based answering for medical research topics.

    1.9k GitHub stars~2k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Drugbank Database

    davila7/claude-code-templates

    Access and analyze comprehensive drug information from the DrugBank database including drug properties, interactions, targets, pathways, chemical structures, and pharmacology data.

    33k GitHub starsUsed in 10 repos~2.3k tokens
    Research & ScienceAuto-check passed
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 22 days ago
    Research & ScienceAuto-check: notes
  • Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation.

    170 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Geniml

    aipoch/medical-research-skills

    Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run…

    1.9k GitHub stars~1.9k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Bio Similarity Searching

    FreedomIntelligence/OpenClaw-Medical-Skills

    Performs molecular similarity searches using Tanimoto coefficient on fingerprints via RDKit.

    3.1k GitHub stars~1.7k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed

More from google-deepmind/science-skills

All 40 skills in this repo
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Auto-check: notes
  • Chembl Database

    google-deepmind/science-skills

    Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.

    3.2k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Auto-check: notes
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Auto-check: notes

Questions about Protein Sequence Similarity Search

What does Protein Sequence Similarity Search do?

Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback). Protein Sequence Similarity Search is an agent skill from google-deepmind/science-skills. Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback).

When should I use Protein Sequence Similarity Search?

Protein Sequence Similarity Search fits situations like: this whenever the user provides a protein sequence; FASTA file and asks to find homologues; sequence matches; wants to infer protein function based on sequence similarity.

How do I install Protein Sequence Similarity Search in Claude Code?

Run `npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a claude-code`. Or copy the skill folder (skills/protein_sequence_similarity_search in google-deepmind/science-skills) into .claude/skills/protein-sequence-similarity-search in your project. Claude Code loads it when a task matches its description.

How do I install Protein Sequence Similarity Search in Codex?

Run `npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a codex`. Or copy the skill folder (skills/protein_sequence_similarity_search in google-deepmind/science-skills) into .agents/skills/protein-sequence-similarity-search in your project. Codex loads it when a task matches its description.

Can I use Protein Sequence Similarity Search in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/protein-sequence-similarity-search, .gemini/skills/protein-sequence-similarity-search, .github/skills/protein-sequence-similarity-search and .opencode/skills/protein-sequence-similarity-search in your project.

What does Protein Sequence Similarity Search need to run?

Going by SKILL.md and its folder, Protein Sequence Similarity Search needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Protein Sequence Similarity Search access the network?

SKILL.md names 1 domain. As links in the text: ebi.ac.uk. This is read from the text; nothing was executed.

Is Protein Sequence Similarity Search safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Protein Sequence Similarity Search use?

Protein Sequence Similarity Search is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Protein Sequence Similarity Search use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 532 tokens, read only when the agent opens those files.

What are the alternatives to Protein Sequence Similarity Search?

Skills that share tags, products or a category with Protein Sequence Similarity Search: Medical Vector Search (aipoch/medical-research-skills, 1.9k stars), Drugbank Database (davila7/claude-code-templates, 33k stars), Scholar RAG (joshzyj/open-scholar-skill, 168 stars) and Comprehensive Protein Analysis (InternScience/scp, 170 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Protein Sequence Similarity Search?

google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,233 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on October 9, 2026.

Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.