Agent skill

Ncbi Sequence Fetch

by google-deepmind in google-deepmind/science-skills

Retrieve protein and nucleotide sequences from NCBI databases using E-utilities.

Apache-2.0Auto-check: notesResearch & Science

Install Ncbi Sequence Fetch

skills CLI
$ npx skills add google-deepmind/science-skills --skill ncbi-sequence-fetch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-deepmind/science-skills ncbi-sequence-fetch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ncbi_sequence_fetch .claude/skills/ncbi-sequence-fetch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ncbi-sequence-fetch
GitHub stars
3.2k
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
840 words
Files
3 (incl. scripts, references)
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Retrieve protein and nucleotide sequences from NCBI databases using E-utilities.

  • Works in 10 steps: Fetch Protein by Accession → Fetch Nucleotide by Accession → CDS Translate → …
  • You need to fetch biological sequences by accession
  • SKILL.md covers Prerequisites, Core Rules, Overview and Utility Scripts, plus 2 more sections
  • Runs Python scripts from its folder; calls uv; needs NCBI_API_KEY

What it does

Ncbi Sequence Fetch is an agent skill from google-deepmind/science-skills. Retrieve protein and nucleotide sequences from NCBI databases using E-utilities. Supports direct accession lookup, CDS translation, gene+organism search, locus lookup, PubMed-linked sequences, patent protein extraction, and organism+length fallback search. Use when you need to fetch biological sequences by accession, gene name, locus tag, PubMed ID, or patent number.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `scripts/ncbi_fetch.py`).

It sits in Research & Science, covering Academic paper search, Intellectual property and Translation. It works with NCBI and PubMed. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.

When your agent uses it

  • You need to fetch biological sequences by accession
  • Tasks that involve Academic paper search
  • Tasks that involve Intellectual property

Example prompts

  • “/ncbi-sequence-fetch”

Requirements

  • Python 3
  • A credential in NCBI_API_KEY

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Fetch Protein by Accession
  2. Fetch Nucleotide by Accession
  3. CDS Translate
  4. Search Any Database
  5. Cross-Database Links (elink)
  6. Gene + Organism Search
  7. Locus Tag Search
  8. PubMed-Linked Proteins
  9. Patent Sequence Search
  10. Organism + Length Search

What it can do on your machine

Read from SKILL.md and the folder at commit 6883275. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ncbi.nlm.nih.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NCBI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ncbi Sequence Fetch loads about 2.3k tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 840 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:22
    3.  **`.env` file**: Make sure the `.env` file exists in your home directory.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google-deepmind/science-skills at commit 6883275, republished under its Apache-2.0 licence (© google-deepmind). 840 words, ~2,254 tokens.

Download SKILL.mdSave it as .claude/skills/ncbi-sequence-fetch/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
ncbi-sequence-fetch
description
Retrieve protein and nucleotide sequences from NCBI databases using E-utilities. Supports direct accession lookup, CDS translation, gene+organism search, locus lookup, PubMed-linked sequences, patent protein extraction, and organism+length fallback search. Use when you need to fetch biological sequences by accession, gene name, locus tag, PubMed ID, or patent number.

NCBI Sequence Fetch

Prerequisites

  1. uv: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH.
  2. User Notification: If .licenses/ncbi_sequence_fetch_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://www.ncbi.nlm.nih.gov/ and https://www.ncbi.nlm.nih.gov/home/about/policies/, then (2) create the file recording the notification text and timestamp.
  3. .env file: Make sure the .env file exists in your home directory. Create one if it does not exist.
  4. NCBI_API_KEY (optional): Raises the NCBI rate limit from 3 to 10 requests/second. The skill works without it, but a key is recommended if the user plans many queries or encounters a 429 error. You can register for a key for free at https://www.ncbi.nlm.nih.gov/account/settings/. You MUST use the safe credentials protocol in the credentials skill to check for and request this key if this skill looks relevant to the user's request.

Core Rules

  • Use the Wrapper: ALWAYS execute the provided helper scripts to query the database rather than accessing the database directly. The scripts automatically enforce the required rate limit gracefully.
  • API Key Support: If the user provides an NCBI_API_KEY in their environment, the query speed limits are automatically increased significantly.
  • Notification: If this skill is used, ensure this is mentioned in the output.

Overview

Wraps NCBI's Entrez E-utilities (efetch, esearch, elink, esummary) for retrieving protein and nucleotide sequences. Provides 10 subcommands covering the full range of sequence retrieval workflows:

  • fetch-protein — Direct protein accession lookup (GenPept, RefSeq)
  • fetch-nucleotide — Direct nucleotide accession lookup
  • cds-translate — Fetch CDS and translate to protein (3 methods)
  • search — Free-text search of any NCBI database
  • elink — Follow cross-database links (PubMed→Protein, etc.)
  • gene-protein — Search protein by gene name + organism
  • locus-protein — Search protein by locus tag + organism
  • pubmed-proteins — Find proteins linked to a PubMed article
  • patent-search — Extract protein sequences from patents
  • organism-length — Last-resort search by organism + exact AA length

Utility Scripts

scripts/ncbi_fetch.py — Single script with subcommands.

All subcommands write structured JSON output. Use --output FILE to save to a file, or omit it to print to stdout. A human-readable summary is always printed to stdout.

1. Fetch Protein by Accession

Fetches protein FASTA from NCBI by accession (XP_, NP_, GenPept, etc.)

bash
uv run scripts/ncbi_fetch.py fetch-protein XP_022033624 -o /tmp/result.json
uv run scripts/ncbi_fetch.py fetch-protein NP_001234567 ABC12345.1
2. Fetch Nucleotide by Accession

Fetches nucleotide FASTA from NCBI by accession.

bash
uv run scripts/ncbi_fetch.py fetch-nucleotide MK034466 -o /tmp/result.json
3. CDS Translate

Fetches a CDS/nucleotide accession and translates to protein sequence. Tries three approaches in order: 1. NCBI's pre-translated CDS protein (fasta_cds_aa)

  1. GenBank XML CDS annotation translations 3. Raw nucleotide → 6-frame ORF finding
bash
uv run scripts/ncbi_fetch.py cds-translate MK034466 -o /tmp/result.json
uv run scripts/ncbi_fetch.py cds-translate HQ662330 --target-length 1043

If the accession is a genomic record (not mRNA/CDS), the tool will report is_genomic: true so you can fall back to a homology-based approach instead.

4. Search Any Database

Free-text search using Entrez query syntax. Supports all NCBI databases.

bash
# Search protein database
uv run scripts/ncbi_fetch.py search "WRR4B[Gene Name] AND Arabidopsis[Organism]" \
  --database protein --retmax 5 --fetch-sequences

# Search nucleotide database
uv run scripts/ncbi_fetch.py search "Rz2[Gene Name] AND Beta vulgaris[Organism]" \
  --database nuccore --retmax 10

# Search with patent filter
uv run scripts/ncbi_fetch.py search "disease resistance AND Solanum[Organism] AND patent[Properties]" \
  --database protein --fetch-sequences

# Search by sequence length
uv run scripts/ncbi_fetch.py search '"Oryza sativa"[Organism] AND 1043[SLEN]' \
  --database protein --fetch-sequences --retmax 50

Follow NCBI's cross-database links (e.g., PubMed article → linked proteins).

bash
uv run scripts/ncbi_fetch.py elink 24896089 --dbfrom pubmed --db protein \
  --fetch-sequences -o /tmp/linked.json

Searches for protein sequences by gene name and organism. Searches NCBI Protein with [Gene Name] and [Organism] qualifiers.

bash
uv run scripts/ncbi_fetch.py gene-protein WRR4B --organism "Arabidopsis thaliana"
uv run scripts/ncbi_fetch.py gene-protein Pikh-2 --organism "Oryza sativa" \
  --target-length 1043 -o /tmp/result.json

Searches by locus tag in both NCBI Protein and Nuccore databases. Extracts CDS translations from GenBank XML when direct protein hits aren't available.

bash
uv run scripts/ncbi_fetch.py locus-protein At1g56540 --organism "Arabidopsis thaliana"
uv run scripts/ncbi_fetch.py locus-protein Niben101Scf02422g02015.1 \
  --organism "Nicotiana benthamiana" -o /tmp/result.json
Show full SKILL.md (328 more words)Show less
8. PubMed-Linked Proteins

Finds protein sequences linked to a PubMed article. Searches NCBI Protein by PMID, follows elink PubMed→Protein, and extracts CDS translations from linked Nuccore records.

bash
uv run scripts/ncbi_fetch.py pubmed-proteins 30692254 --identifier WRR4B
uv run scripts/ncbi_fetch.py pubmed-proteins 24896089 --identifier "K2" \
  -o /tmp/result.json

Two modes:

By patent number — fetches all protein sequences from a specific patent: bash uv run scripts/ncbi_fetch.py patent-search --patent-number US10123456 -o /tmp/patent.json

By keywords — searches NCBI Protein with patent[Properties] filter: bash uv run scripts/ncbi_fetch.py patent-search --keywords WRR4B Albugo --organism "Arabidopsis thaliana" -o /tmp/patent.json

[!IMPORTANT] Patent convention: In molecular biology patents, SEQ ID NO: 1 is typically the DNA sequence and SEQ ID NO: 2 is the primary protein. Higher SEQ ID NOs are variants or related sequences. Prefer Sequence 2 when selecting the primary protein of interest.

Last-resort search when only organism and expected protein length are known. Uses NCBI's [SLEN] filter for exact length matching.

bash
uv run scripts/ncbi_fetch.py organism-length \
  --organism "Arabidopsis thaliana" --length 1048 --retmax 50 \
  -o /tmp/result.json

[!NOTE] This often returns multiple candidates. Use the JSON output headers to identify the correct protein.

Workflow

Standard Sequence Retrieval Cascade

When trying to find a protein sequence, follow this priority order:

  1. Direct accession — fetch-protein with GenPept/RefSeq accession
  2. CDS translation — cds-translate with nucleotide/CDS accession
  3. PubMed-linked — pubmed-proteins with PMID + gene name
  4. Locus lookup — locus-protein with locus tag + organism
  5. Gene + organism — gene-protein with gene name + organism
  6. Patent search — patent-search with patent number or keywords
  7. Organism + length — organism-length as last resort
Interpreting Results
  • All subcommands return JSON with a results array
  • Each result has sequence (AA string), length, and header/metadata
  • When multiple results are returned, select by:
    • Closest match to expected length (target_length)
    • Header relevance (matching gene name, "disease resistance" keywords)
    • Source priority (RefSeq > GenPept > patent)

Reference

  • NCBI E-utilities docs: https://www.ncbi.nlm.nih.gov/books/NBK25499/
  • Entrez search syntax: https://www.ncbi.nlm.nih.gov/books/NBK49540/
  • Database list: protein, nuccore, gene, pubmed, pmc, biosample, etc.
  • Common accession formats:
    • XP_ / NP_ — NCBI RefSeq protein
    • AAA to AZZ + digits — GenPept (translated GenBank)
    • MK, MN, HQ, etc. + digits — GenBank nucleotide
    • ENSG, ENST, ENSP — Ensembl (use ensembl-database skill instead)
    • Q, P, O + digits — UniProt (use uniprot-database skill instead)

© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/ncbi_sequence_fetch of google-deepmind/science-skills.

  • SKILL.md
  • references/citation.bib
  • scripts/ncbi_fetch.py

Open the folder on GitHubat commit 6883275

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in google-deepmind/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Ncbi Sequence Fetch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ncbi Sequence Fetch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ncbi Sequence Fetch this skillgoogle-deepmind/science-skills3.2k1 repos~2.3kAutomated safety check: NotesApache-2.0
PubMed REST API Searchdavila7/claude-code-templates32k15 repos~3.9kAutomated safety check: PassMIT
Journal Skillsaipoch/medical-research-skills2k—~1.7kAutomated safety check: PassMIT
Pubmed Databasejaechang-hits/SciAgent-Skills3701 repos~4.4kAutomated safety check: PassCC-BY-4.0
Scientific DB Pubmed Databaseaffaan-m/ECC274k1 repos~1.2kAutomated safety check: PassMIT
Mining Pubmed Literaturemaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • PubMed REST API Search

    davila7/claude-code-templates

    Searches PubMed directly through its E-utilities REST API, with guidance on Boolean and MeSH query syntax, batch retrieval and citation data.

    32k GitHub starsUsed in 15 repos~3.9k tokens
    Research & ScienceAuto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 20 days ago
    Research & ScienceAuto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~4.4k tokens
    Research & ScienceAuto-check passed
  • Direct PubMed and NCBI E-utilities search workflows for biomedical literature, MeSH queries, PMID lookup, citation retrieval, and API-backed literature monitoring.

    274k GitHub starsUsed in 1 repo~1.2k tokens
    Research & ScienceAuto-check passed
  • Mining Pubmed Literature

    maziyarpanahi/openmed

    Searches and fetches PubMed and PMC via NCBI E-utilities (ESearch then EFetch/ESummary) to gather biomedical evidence and build text corpora.

    5.5k GitHub stars~1.7k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Search Lit

    Aperivue/medsci-skills

    A skill your agent uses when finding papers or building a reference list.

    329 GitHub stars~4.7k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed

More from google-deepmind/science-skills

All 40 skills in this repo
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Auto-check: notes
  • Chembl Database

    google-deepmind/science-skills

    Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.

    3.2k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Auto-check: notes
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Auto-check: notes

Works with

Questions about Ncbi Sequence Fetch

What does Ncbi Sequence Fetch do?

Retrieve protein and nucleotide sequences from NCBI databases using E-utilities. Ncbi Sequence Fetch is an agent skill from google-deepmind/science-skills. Retrieve protein and nucleotide sequences from NCBI databases using E-utilities.

When should I use Ncbi Sequence Fetch?

Ncbi Sequence Fetch fits situations like: you need to fetch biological sequences by accession; tasks that involve Academic paper search; tasks that involve Intellectual property.

How do I install Ncbi Sequence Fetch in Claude Code?

Run `npx skills add google-deepmind/science-skills --skill ncbi-sequence-fetch -a claude-code`. Or copy the skill folder (skills/ncbi_sequence_fetch in google-deepmind/science-skills) into .claude/skills/ncbi-sequence-fetch in your project. Claude Code loads it when a task matches its description.

How do I install Ncbi Sequence Fetch in Codex?

Run `npx skills add google-deepmind/science-skills --skill ncbi-sequence-fetch -a codex`. Or copy the skill folder (skills/ncbi_sequence_fetch in google-deepmind/science-skills) into .agents/skills/ncbi-sequence-fetch in your project. Codex loads it when a task matches its description.

Can I use Ncbi Sequence Fetch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill ncbi-sequence-fetch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ncbi-sequence-fetch, .gemini/skills/ncbi-sequence-fetch, .github/skills/ncbi-sequence-fetch and .opencode/skills/ncbi-sequence-fetch in your project.

What does Ncbi Sequence Fetch need to run?

Going by SKILL.md and its folder, Ncbi Sequence Fetch needs Python for the scripts in its folder, the command-line tools its instructions call (uv) and credentials named NCBI_API_KEY. Our summary lists: Python 3; A credential in NCBI_API_KEY.

Does Ncbi Sequence Fetch access the network?

SKILL.md names 1 domain. As links in the text: ncbi.nlm.nih.gov. This is read from the text; nothing was executed.

Is Ncbi Sequence Fetch safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ncbi Sequence Fetch use?

Ncbi Sequence Fetch is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ncbi Sequence Fetch use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 453 tokens, read only when the agent opens those files.

What are the alternatives to Ncbi Sequence Fetch?

Skills that share tags, products or a category with Ncbi Sequence Fetch: PubMed REST API Search (davila7/claude-code-templates, 32k stars), Journal Skills (aipoch/medical-research-skills, 2k stars), Pubmed Database (jaechang-hits/SciAgent-Skills, 370 stars) and Scientific DB Pubmed Database (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ncbi Sequence Fetch?

google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,216 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on September 15, 2026.

Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.