Medical Vector Search
aipoch/medical-research-skills
Vector database retrieval and evidence-based answering for medical research topics.
Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback).
$ npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google-deepmind/science-skills protein-sequence-similarity-search --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/protein_sequence_similarity_search .claude/skills/protein-sequence-similarity-search && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "protein-sequence-similarity-search" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/protein_sequence_similarity_search into .claude/skills/protein-sequence-similarity-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "protein-sequence-similarity-search", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google-deepmind/science-skills/tree/main/skills/protein_sequence_similarity_searchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google-deepmind/science-skills protein-sequence-similarity-search --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/protein_sequence_similarity_search .agents/skills/protein-sequence-similarity-search && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "protein-sequence-similarity-search" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/protein_sequence_similarity_search into .agents/skills/protein-sequence-similarity-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "protein-sequence-similarity-search", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google-deepmind/science-skills protein-sequence-similarity-search --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/protein_sequence_similarity_search .cursor/skills/protein-sequence-similarity-search && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "protein-sequence-similarity-search" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/protein_sequence_similarity_search into .cursor/skills/protein-sequence-similarity-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "protein-sequence-similarity-search", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google-deepmind/science-skills.git --path skills/protein_sequence_similarity_search--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google-deepmind/science-skills protein-sequence-similarity-search --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/protein_sequence_similarity_search .gemini/skills/protein-sequence-similarity-search && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "protein-sequence-similarity-search" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/protein_sequence_similarity_search into .gemini/skills/protein-sequence-similarity-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "protein-sequence-similarity-search", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google-deepmind/science-skills protein-sequence-similarity-searchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/protein_sequence_similarity_search .github/skills/protein-sequence-similarity-search && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "protein-sequence-similarity-search" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/protein_sequence_similarity_search into .github/skills/protein-sequence-similarity-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "protein-sequence-similarity-search", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google-deepmind/science-skills protein-sequence-similarity-search --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/protein_sequence_similarity_search .opencode/skills/protein-sequence-similarity-search && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "protein-sequence-similarity-search" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/protein_sequence_similarity_search into .opencode/skills/protein-sequence-similarity-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "protein-sequence-similarity-search", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
protein-sequence-similarity-searchSearches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback).
Protein Sequence Similarity Search is an agent skill from google-deepmind/science-skills. Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback). Trigger this whenever the user provides a protein sequence or FASTA file and asks to find homologues, sequence matches, or wants to infer protein function based on sequence similarity, but not when the user wants to infer protein function based on structural similarity.
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `scripts/mmseqs2_search.py` and `scripts/uniprot_blast.py`).
It sits in Research & Science, covering Vector databases. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 8ab7672. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
ebi.ac.ukFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Protein Sequence Similarity Search loads about 2.7k tokens when it runs, and up to ~3.2k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 1,226 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
3. **`.env` file**: Make sure the `.env` file exists in your home directory.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from google-deepmind/science-skills at commit 8ab7672, republished under its Apache-2.0 licence (© google-deepmind). 1,226 words, ~2,669 tokens.
.claude/skills/protein-sequence-similarity-search/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.uv: Read the uv skill and follow its Setup instructions to ensure
uv is installed and on PATH..env file: Make sure the .env file exists in your home directory.
Create one if it does not exist.USER_EMAIL (optional but recommended): Recommended by the EBI for
BLAST job tracking, but the skill works without it. You MUST use the
safe credentials protocol in the credentials skill to check for and
request this credential if this skill looks relevant to the user's request.Take a user-provided amino acid sequence (or a path to a .fasta file), search
for sequence homologues using the fastest available method, generate a
Markdown-formatted table of the top hits, interpret key alignment metrics,
summarize the inferred protein functions, and save results locally for future
programmatic analysis.
.md file for your summary. The JSON
and other outputs are for subsequent tool use only.Choose the search method based on the user's request:
If the user says "quick search" or "fast search", no specific method
requested / general homologue search, of if you are unsure: Run MMseqs2 (fast,
default) using mmseqs2_search.py
If MMseqs2 fails (exit code 2: RATELIMIT or API error) or User explicitly
requests "BLAST" or a specific BLAST database (e.g. uniprotkb_swissprot,
pdb, uniprotkb_human): Run BLAST using uniprot_blast.py
Identify the query from the user. It can be a raw sequence string (e.g., "MKVLY...") or a path to a local file (e.g., "./data/sequence.fasta").
Determine the search method using the list above.
Generate File Names: Generate descriptive output file names based on the
input (e.g., proteinA_mmseqs2.json and proteinA_mmseqs2.md).
Execute the MMseqs2 script:
uv run scripts/mmseqs2_search.py <SEQUENCE_OR_FILE> -o <generated-filename.md> -j <generated-filename.json>uv run scripts/mmseqs2_search.py <SEQUENCE_OR_FILE> -o <generated-filename.md> -j <generated-filename.json> --include-mgnifyThe script will query the ColabFold MMseqs2 API and poll for completion. This is typically fast (under 2 minutes).
If the script exits with code 2 (API failure, rate limit), automatically fall back to BLAST (Path B below). Inform the user: "MMseqs2 search failed, falling back to BLAST."
Read the Results: Open and read the generated .md file.
Database Selection & Validation: Determine the most appropriate database(s) based on the user's prompt.
Database Code (e.g.,
uniprotkb_bacteria).uniprotkb_swissprot.--databases.Generate File Names: (e.g., proteinA_ebi_blast.json and
proteinA_ebi_blast.md).
This API requires the user email address to be set in the USER_EMAIL
environment variable for inclusion in request header. You MUST use the
safe credentials protocol in the credentials skill to check for and
request this credential if this skill looks relevant to the user's request.
Execute the BLAST script:
uv run scripts/uniprot_blast.py <SEQUENCE_OR_FILE> -o <generated-filename.md> -j <generated-filename.json>uv run scripts/uniprot_blast.py <SEQUENCE_OR_FILE> -o <generated-filename.md> -j <generated-filename.json> --databases <db1,db2>The script will query the EBI BLAST API and poll the server. Note: This can take up to 15 minutes; wait patiently.
Read the Results: Open and read the generated .md file.
1e-50) indicate extreme statistical
significance..json and .md) and their
locations.uniprotkb – UniProt Knowledgebase (The UniProt Knowledgebase includes
UniProtKB/Swiss-Prot and UniProtKB/TrEMBL): The UniProt Knowledgebase
(UniProtKB) is the central access point for extensive curated protein
information, including function, classification, and cross-references.
Search UniProtKB to retrieve "everything that is known" about a particular
sequenceuniprotkb_swissprot – UniProtKB/Swiss-Prot (The manually annotated section
of UniProtKB): The manually curated subsection of the UniProt Knowledgebaseuniprotkb_swissprotsv – UniProtKB/Swiss-Prot isoforms (The manually
annotated isoforms of UniProtKB/Swiss-Prot): The isoform sequences for the
manually curated subsection of the UniProt Knowledgebaseuniprotkb_reference_proteomes – UniProtKB Reference Proteomes: Taxonomic
subset of the UniProtKB Reference Proteomesuniprotkb_trembl – UniProtKB/TrEMBL (The automatically annotated section
of UniProtKB): Subsection of the UniProt Knowledgebase derived from ENA
Sequence (formerly EMBL-Bank) coding sequence translations with annotation
produced by an automated processuniprotkb_refprotswissprot – UniProtKB Reference Proteomes plus
Swiss-Prot: UniProtKB Reference Proteomes plus Swiss-Protuniprotkb_archaea – UniProtKB Archaea: Taxonomic subset of the UniProt
Knowledgebase for archaeauniprotkb_arthropoda – UniProtKB Arthropoda: Taxonomic subset of the
UniProt Knowledgebase for arthropodauniprotkb_bacteria – UniProtKB Bacteria: Taxonomic subset of the UniProt
Knowledgebase for bacteriauniprotkb_complete_microbial_proteomes – UniProtKB Complete Microbial
Proteomes: Taxonomic subset of the UniProt Knowledgebase for complete
microbial proteomesuniprotkb_eukaryota – UniProtKB Eukaryota: Taxonomic subset of the UniProt
Knowledgebase for eukaryotauniprotkb_fungi – UniProtKB Fungi: Taxonomic subset of the UniProt
Knowledgebase for fungiuniprotkb_human – UniProtKB Human: Taxonomic subset of the UniProt
Knowledgebase for humanuniprotkb_mammals – UniProtKB Mammals: Taxonomic subset of the UniProt
Knowledgebase for mammalsuniprotkb_nematoda – UniProtKB Nematoda: Taxonomic subset of the UniProt
Knowledgebase for nematodauniprotkb_rodents – UniProtKB Rodents: Taxonomic subset of the UniProt
Knowledgebase for rodentsuniprotkb_vertebrates – UniProtKB Vertebrates: Taxonomic subset of the
UniProt Knowledgebase for vertebratesuniprotkb_viridiplantae – UniProtKB Viridiplantae: Taxonomic subset of the
UniProt Knowledgebase for viridiplantaeuniprotkb_viruses – UniProtKB Viruses: Taxonomic subset of the UniProt
Knowledgebase for virusesuniprotkb_enzyme – UniProtKB Enzyme: Taxonomic subset of the UniProt
Knowledgebase for enzymesuniprotkb_covid19 – UniProtKB COVID-19: Taxonomic subset of the UniProt
Knowledgebase for COVID-19uniref100 – UniProt Clusters 100% (UniRef100): The UniProt Reference
Clusters (UniRef) containing sequences which are 100% identical.uniref90 – UniProt Clusters 90% (UniRef90): The UniProt Reference Clusters
(UniRef) containing sequences which are 90% identical.uniref50 – UniProt Clusters 50% (UniRef50): The UniProt Reference Clusters
(UniRef) containing sequences which are 50% identical.pdb – Protein Structure Sequences (PDBe protein structure sequences):
Protein sequences from structures described in the Brookhaven Protein Data
Bank (PDB)© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts, references) in skills/protein_sequence_similarity_search of google-deepmind/science-skills.
Open the folder on GitHubat commit 8ab7672
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in google-deepmind/science-skills, which our catalogue first saw on October 7, 2026.
Protein Sequence Similarity Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Protein Sequence Similarity Search this skillgoogle-deepmind/science-skills | 3.2k | 1 repos | ~2.7k | Automated safety check: Notes | Apache-2.0 | |
| Medical Vector Searchaipoch/medical-research-skills | 1.9k | — | ~2k | Automated safety check: Pass | MIT | |
| Drugbank Databasedavila7/claude-code-templates | 33k | 10 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Scholar RAGjoshzyj/open-scholar-skill | 168 | — | ~7.4k | Automated safety check: Notes | Custom licence | |
| Comprehensive Protein AnalysisInternScience/scp | 170 | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Genimlaipoch/medical-research-skills | 1.9k | — | ~1.9k | Automated safety check: Pass | MIT |
aipoch/medical-research-skills
Vector database retrieval and evidence-based answering for medical research topics.
davila7/claude-code-templates
Access and analyze comprehensive drug information from the DrugBank database including drug properties, interactions, targets, pathways, chemical structures, and pharmacology data.
joshzyj/open-scholar-skill
Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.
InternScience/scp
Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation.
aipoch/medical-research-skills
Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run…
FreedomIntelligence/OpenClaw-Medical-Skills
Performs molecular similarity searches using Tanimoto coefficient on fingerprints via RDKit.
google-deepmind/science-skills
Retrieve and analyze AlphaFold predicted structures for a protein.
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
google-deepmind/science-skills
Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.
google-deepmind/science-skills
Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
Categories
Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback). Protein Sequence Similarity Search is an agent skill from google-deepmind/science-skills. Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback).
Protein Sequence Similarity Search fits situations like: this whenever the user provides a protein sequence; FASTA file and asks to find homologues; sequence matches; wants to infer protein function based on sequence similarity.
Run `npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a claude-code`. Or copy the skill folder (skills/protein_sequence_similarity_search in google-deepmind/science-skills) into .claude/skills/protein-sequence-similarity-search in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a codex`. Or copy the skill folder (skills/protein_sequence_similarity_search in google-deepmind/science-skills) into .agents/skills/protein-sequence-similarity-search in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill protein-sequence-similarity-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/protein-sequence-similarity-search, .gemini/skills/protein-sequence-similarity-search, .github/skills/protein-sequence-similarity-search and .opencode/skills/protein-sequence-similarity-search in your project.
Going by SKILL.md and its folder, Protein Sequence Similarity Search needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: ebi.ac.uk. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Protein Sequence Similarity Search is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 532 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Protein Sequence Similarity Search: Medical Vector Search (aipoch/medical-research-skills, 1.9k stars), Drugbank Database (davila7/claude-code-templates, 33k stars), Scholar RAG (joshzyj/open-scholar-skill, 168 stars) and Comprehensive Protein Analysis (InternScience/scp, 170 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,233 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on October 9, 2026.
Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.