Agent skill

Bio Batch Downloads

by GPTomics in GPTomics/bioSkills

Download large datasets from NCBI efficiently using EPost, history server, batching, rate limiting, and retry logic.

MITAuto-check passedBackend & APIs

Install Bio Batch Downloads

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-batch-downloads -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-batch-downloads --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/database-access/batch-downloads .claude/skills/bio-batch-downloads && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-batch-downloads
GitHub stars
1.2k
Used in
2 other repos
Token cost
~3.9k tokens
SKILL.md length
1,401 words
Files
5
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Download large datasets from NCBI efficiently using EPost, history server, batching, rate limiting, and retry logic.

  • Bulk-fetching tens of thousands of sequences
  • SKILL.md covers Version Compatibility, Required Setup, Decision matrix: which… and Rate-limit math (precise), plus 8 more sections
  • Runs Python scripts from its folder; calls pip
  • Pulling all results of a large ESearch

What it does

Bio Batch Downloads is an agent skill from GPTomics/bioSkills. Download large datasets from NCBI efficiently using EPost, history server, batching, rate limiting, and retry logic. Use when bulk-fetching tens of thousands of sequences, pulling all results of a large ESearch, designing reproducible pipelines, comparing E-utilities to NCBI Datasets v2 CLI, or implementing checksum-validated downloads. Encodes WebEnv TTL (~8h), EPost 200-ID limit, retmax caps, parallelization design, and integrity verification.

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `examples/batch_by_ids.py`, `examples/batch_fasta.py` and `examples/robust_download.py`).

It sits in Backend & APIs, covering Rate limiting, Error handling and Bioinformatics. It works with NCBI and Biopython. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Bulk-fetching tens of thousands of sequences
  • Pulling all results of a large ESearch
  • Designing reproducible pipelines
  • Comparing E-utilities to NCBI Datasets v2 CLI

Example prompts

  • “/bio-batch-downloads”

Requirements

  • Python 3
  • A credential in YOUR_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ncbi.nlm.nih.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Batch Downloads loads about 3.9k tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 1,401 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,401 words, ~3,884 tokens.

Download SKILL.mdSave it as .claude/skills/bio-batch-downloads/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
bio-batch-downloads
description
Download large datasets from NCBI efficiently using EPost, history server, batching, rate limiting, and retry logic. Use when bulk-fetching tens of thousands of sequences, pulling all results of a large ESearch, designing reproducible pipelines, comparing E-utilities to NCBI Datasets v2 CLI, or implementing checksum-validated downloads. Encodes WebEnv TTL (~8h), EPost 200-ID limit, retmax caps, parallelization design, and integrity verification.
tool_type
python
primary_tool
Bio.Entrez

Version Compatibility

Reference examples tested with: BioPython 1.83+, NCBI Datasets CLI 16.0+, Entrez Direct 21.0+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show biopython then help(Bio.Entrez.efetch) to check signatures
  • CLI: datasets --version and efetch -version

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Batch Downloads

"Download N thousand records from NCBI without getting blocked" -> The right answer is rarely "parallelize requests". For >5000 records the answer is the history server: search once, fetch in chunks server-side. For >100,000 records or whole genomes, the modern answer is NCBI Datasets v2 CLI -- the E-utilities are not optimized for bulk genome/gene data anymore.

This skill encodes (a) when to use each retrieval strategy, (b) the precise rate-limit math, (c) WebEnv lifecycle for long-running jobs, (d) how to design retry/resume, and (e) when to defect to Datasets CLI instead.

  • Python: Entrez.esearch(usehistory='y') + chunked Entrez.efetch() (BioPython)
  • CLI: datasets download genome accession ... (NCBI Datasets v2 -- preferred for genome/gene bulk)
  • CLI: epost | efetch -mode webenv (Entrez Direct)

Required Setup

python
from Bio import Entrez
import time
Entrez.email = 'researcher@institution.edu'
Entrez.api_key = 'YOUR_KEY'  # 3 -> 10 req/sec; mandatory for bulk
Entrez.tool = 'project-name'

Decision matrix: which retrieval strategy?

Record countSourceStrategyWhy
< 200 known IDsAny dbEFetch with comma-joined id=Single round-trip; trivial
200-5,000 known IDsAny dbEPost (chunked at 200) -> history -> chunked EFetchURL length limit + chunked retrieval
5,000-100,000 from a queryAny dbESearch with usehistory='y' -> chunked EFetchPush to server once; pull in batches
> 100,000 sequencesnucleotide/proteinConsider FTP mirror or Datasets CLI; chunk if E-utils stillNCBI throttles bulk; offline mirror is faster
Whole genome assembliesAssembly/Datasetsdatasets download genome accession ...Datasets v2 is the modern bulk endpoint
All RefSeq for a speciesDatasetsdatasets download genome taxon ...Replaces assembly_summary.txt scraping
All gene records for a listDatasetsdatasets download gene gene-id ...Cleaner output than EFetch gene XML
Raw sequencing readsSRAprefetch + fasterq-dump (or ENA mirror)See sra-data skill

The Datasets CLI is the right answer for any genome- or gene-centric bulk workflow as of 2023+. The E-utilities remain right for PubMed, ESummary metadata, custom queries, and anything not in the Datasets API. See ncbi-datasets-cli skill.

Rate-limit math (precise)

Authreq/secSleep between callsBulk-friendly notes
Email only30.34 sSingle-threaded only; parallelism violates ToS
Email + API key100.10 sModest parallelism (max ~4 workers) safe
Institutional bulkNegotiatedEmail eutilities@ncbi.nlm.nih.govFor >100K queries; courtesy expected

NCBI's terms ask that heavy automated downloads run outside US weekday business hours (9 AM-5 PM ET). Cron the job for nights/weekends; pipelines that ignore this get IP-throttled.

Critical: parallelizing API calls is the WRONG bulk strategy. One stream with history server + larger batches is faster AND more polite than N parallel streams. The bottleneck is rarely NCBI's throughput at small N -- it's the round-trip count.

History server lifecycle (the long-running-job trap)

PropertyValueFailure mode
TTL8 hours absolute (per NCBI E-utils help)Job started Friday evening dies Saturday morning
Idle eviction~15 min empirically under loadA worker that stalls loses its WebEnv
Per-session isolationOne WebEnv string per sessionDon't share across processes if isolation matters
Expired session behaviorHTTP 200 with <ERROR>WebEnv not found</ERROR>Won't surface as HTTP error -- must parse body
RecoveryRe-run ESearch; resume at retstartNeed to checkpoint progress to disk

Production pattern: checkpoint the retstart cursor after each successful chunk to disk; on restart, re-run ESearch (cheap), pick up retstart from checkpoint, continue.

EPost specifics

EPost pushes a list of UIDs to the history server so downstream EFetch can pull by WebEnv/QueryKey instead of by ID. Two constraints:

  • 200 IDs per EPost call is the hard limit.
  • Chained posts share a WebEnv: pass the WebEnv from the first call into subsequent calls to accumulate IDs under one session; a new QueryKey is issued per call.

To intersect: term=#{key1} AND #{key2} against the WebEnv produces a new key.

Batch size guidelines per rettype

DatabaserettypeOptimal batchPer-record payload
nucleotidefasta500-1000~1 KB
nucleotidegb100-200~10-50 KB
proteinfasta500-1000~0.5 KB
proteingp100-200~5-30 KB
pubmedmedline1000-2000~2 KB
pubmedxml200-500~10-30 KB
anyesummary (docsum)500 per call~1 KB

Smaller batches for GenBank/XML because per-record payload is larger; larger batches for FASTA because the per-call HTTP overhead dominates.

Code patterns

Production batch fetch (history server + retry + checkpoint)

Goal: Download all records matching a query, robust to mid-job failures and session expiry.

Approach: ESearch with history; checkpoint cursor to disk; on error, retry the chunk; on session expiry, re-run ESearch and resume from checkpoint.

Reference (BioPython 1.83+):

python
import json
import time
from pathlib import Path
from urllib.error import HTTPError
from Bio import Entrez


def checkpointed_batch_download(db, term, out_path, ckpt_path, rettype='fasta',
                                 retmode='text', batch_size=500, max_retries=3):
    '''Download all matching records with disk checkpoint for resumability.'''
    delay = 0.1 if Entrez.api_key else 0.34
    ckpt = Path(ckpt_path)
    start = json.loads(ckpt.read_text())['start'] if ckpt.exists() else 0

    h = Entrez.esearch(db=db, term=term, usehistory='y', retmax=0)
    s = Entrez.read(h); h.close()
    webenv, query_key, total = s['WebEnv'], s['QueryKey'], int(s['Count'])
    print(f'{total:,} records matched; resuming at {start:,}')

    mode = 'a' if start else 'w'
    with open(out_path, mode) as out:
        while start < total:
            for attempt in range(max_retries):
                try:
                    h = Entrez.efetch(db=db, rettype=rettype, retmode=retmode,
                                      retstart=start, retmax=batch_size,
                                      webenv=webenv, query_key=query_key)
                    body = h.read(); h.close()
                    if isinstance(body, bytes):
                        body = body.decode('utf-8', errors='replace')
                    if '<ERROR>' in body[:500]:
                        raise RuntimeError(f'Server error in body: {body[:200]}')
                    out.write(body)
                    break
                except HTTPError as e:
                    if e.code == 429:
                        wait = 10 * (attempt + 1)
                        print(f'  Rate-limited; sleeping {wait}s')
                        time.sleep(wait)
                    elif attempt == max_retries - 1:
                        raise
                    else:
                        time.sleep(5 * (attempt + 1))
                except RuntimeError as e:
                    # Likely WebEnv expired; re-run ESearch
                    print(f'  {e}; refreshing WebEnv')
                    h = Entrez.esearch(db=db, term=term, usehistory='y', retmax=0)
                    s = Entrez.read(h); h.close()
                    webenv, query_key = s['WebEnv'], s['QueryKey']

            start += batch_size
            ckpt.write_text(json.dumps({'start': start, 'total': total}))
            time.sleep(delay)
            print(f'  {min(start, total):,}/{total:,}')
    ckpt.unlink(missing_ok=True)
EPost large ID list, then EFetch

Goal: Download by a known list of 5,000 accessions without 414 URI errors.

Approach: EPost in 200-ID chunks; reuse WebEnv across chunks; final fetch reads from history.

Reference (BioPython 1.83+):

python
def epost_and_fetch(db, ids, out_path, rettype='fasta', retmode='text', batch_size=500):
    delay = 0.1 if Entrez.api_key else 0.34
    webenv = None
    posted_keys = []  # (query_key, n_ids) so we iterate each key's actual size
    for i in range(0, len(ids), 200):
        chunk = ids[i:i+200]
        kwargs = {'db': db, 'id': ','.join(chunk)}
        if webenv:
            kwargs['WebEnv'] = webenv
        h = Entrez.epost(**kwargs)
        r = Entrez.read(h); h.close()
        webenv = r['WebEnv']
        posted_keys.append((r['QueryKey'], len(chunk)))
        time.sleep(delay)

    with open(out_path, 'w') as out:
        for qk, n in posted_keys:
            for start in range(0, n, batch_size):
                h = Entrez.efetch(db=db, rettype=rettype, retmode=retmode,
                                  retstart=start, retmax=min(batch_size, n - start),
                                  webenv=webenv, query_key=qk)
                out.write(h.read()); h.close()
                time.sleep(delay)
Integrity check after download

Goal: Confirm downloaded FASTA has the expected record count and no truncation.

Approach: Count expected (from ESearch Count) vs observed (from SeqIO.parse).

python
from Bio import SeqIO

def verify_fasta_count(path, expected):
    observed = sum(1 for _ in SeqIO.parse(path, 'fasta'))
    assert observed == expected, f'Expected {expected:,} records, found {observed:,}'
    return True

For genome assemblies and known-checksum files, NCBI provides MD5 manifests (e.g. md5checksums.txt in FTP genome directories). NCBI Datasets CLI verifies checksums automatically; the FTP-direct route needs explicit md5sum -c.

Show full SKILL.md (554 more words)Show less
Compare E-utils to Datasets CLI cost
python
def estimate_efetch_calls(total, batch_size):
    return -(-total // batch_size)  # ceiling division

For 100,000 nucleotide records at 500/batch with API key: 200 calls * 0.1s = 20s minimum. For the same workflow via datasets download gene gene-id 100000: one CLI invocation, parallel download, automatic checksum. For genome-scale bulk, Datasets wins by an order of magnitude.

Parallelization design (modest)

Goal: Pull from two independent queries concurrently without violating rate limits.

Approach: Async with a global semaphore that enforces the API-key-permitted rate. Max 4 concurrent workers is the polite cap.

python
import asyncio
from asyncio import Semaphore

# Pseudo-pattern; real impl needs aiohttp + Bio.Entrez async wrappers
async def fetch_with_semaphore(sem, db, id_, rettype):
    async with sem:
        # call EFetch
        await asyncio.sleep(0.1)  # rate gate
        # ... actual call

sem = Semaphore(4)

Never exceed 4 concurrent workers with an API key, or 1 without. Above that NCBI throttles by IP and the whole pipeline grinds.

Failure modes

Session expires mid-pipeline
  • Trigger: Job runs >8h or worker idles >15 min.
  • Mechanism: WebEnv evicted; EFetch returns HTTP 200 with <ERROR>WebEnv not found</ERROR> body.
  • Symptom: Silently truncated output mid-file; downstream parsing fails on empty chunks.
  • Fix: Parse body for <ERROR>; re-run ESearch and resume at checkpointed retstart.
URL too long on >200 IDs
  • Trigger: Comma-joined id= to EFetch with 250+ IDs.
  • Mechanism: GET URL exceeds NCBI's ~2000 char limit.
  • Symptom: HTTP 414 URI Too Long, or silent truncation.
  • Fix: EPost in chunks of 200 first, then EFetch by WebEnv/QueryKey.
Rate-limit cascade
  • Trigger: Parallelizing without API key; or >10 req/s with key.
  • Mechanism: NCBI returns 429; aggressive retry triggers IP-level throttle.
  • Symptom: Pipeline gets slower and eventually stops.
  • Fix: Add jittered exponential backoff; reduce concurrency; reach out for institutional access if bulk is the norm.
Datasets / E-utils confusion
  • Trigger: Building a custom assembly_summary.txt scraper instead of using Datasets.
  • Mechanism: Datasets API is the official, supported bulk endpoint for genome/gene data; E-utils is not optimized for it.
  • Symptom: Slow downloads, stale snapshots, missing fields.
  • Fix: Use datasets download genome ... for genomes; datasets download gene ... for gene records. See ncbi-datasets-cli.
Silent retmax cap
  • Trigger: ESearch without usehistory='y'; Count > 9999.
  • Mechanism: Legacy esearch enforces 9999 cap; the rest of the result set is silently dropped.
  • Symptom: Batch loop terminates early; missing thousands of records.
  • Fix: Always set usehistory='y' for any query expected to return >5000.
Checkpoint corruption / partial chunk
  • Trigger: Job crashes mid-chunk; checkpoint hasn't been written.
  • Mechanism: Output file has half a record at the end.
  • Symptom: SeqIO.parse fails on the partial record.
  • Fix: Write checkpoint AFTER successful chunk write + file flush; on resume, truncate the output file at the last newline before continuing.

Common errors

Error / symptomCauseSolution
HTTPError 429Rate limitSleep with backoff; get API key
HTTPError 414URL too longEPost first
<ERROR>WebEnv not found</ERROR> (HTTP 200)Session expiredRe-run ESearch; resume at checkpoint
Output file ends mid-recordCrash mid-chunkTruncate-to-newline on resume
Slow despite API keyToo few records per callIncrease batch_size to 500+ for FASTA
Datasets CLI faster than EFetchWorkflow is genome/gene bulkSwitch to ncbi-datasets-cli

References

  • Sayers EW et al. (2024) Database resources of the National Center for Biotechnology Information in 2024. Nucleic Acids Res 52:D33-D43.
  • Kans J. (2024) Entrez Direct: E-utilities on the Unix Command Line. NCBI Bookshelf NBK179288.
  • NCBI. EPost help and Usage Guidelines. NBK25499.
  • NCBI Datasets documentation: https://www.ncbi.nlm.nih.gov/datasets/docs/v2/
  • entrez-search - Build the query that batch-downloads will fetch
  • entrez-fetch - Single-record EFetch and ESummary
  • entrez-link - Chain ELink with neighbor_history for cross-db bulk
  • ncbi-datasets-cli - Modern bulk endpoint for genome/gene data; preferred over E-utils for that scope
  • sra-data - Raw read downloads via SRA toolkit (not via E-utilities)
  • geo-data - GEO supplementary file downloads

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in database-access/batch-downloads of GPTomics/bioSkills.

  • SKILL.md
  • examples/batch_by_ids.py
  • examples/batch_fasta.py
  • examples/robust_download.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Batch Downloads next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Batch Downloads compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Batch Downloads this skillGPTomics/bioSkills1.2k2 repos~3.9kAutomated safety check: PassMIT
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Azure APIM Policy Authoringthomast1906/github-copilot-agent-skills202—~1.5kAutomated safety check: PassMIT
Biopythondavila7/claude-code-templates33k12 repos~3.4kAutomated safety check: PassMIT
API IntegrationHack23/cia239—~1.9kAutomated safety check: PassApache-2.0
BiopythonK-Dense-AI/scientific-agent-skills48k1 repos~4.3kAutomated safety check: NotesMIT

Similar skills

  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Azure APIM Policy Authoring

    thomast1906/github-copilot-agent-skills

    Generates Azure API Management policy XML for authentication, rate limiting, CORS, error handling and transformations, consulting Azure best-practice and documentation tools first.

    202 GitHub stars~1.5k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • API Integration

    Hack23/cia

    External API integration patterns, retry logic, circuit breakers, caching, rate limiting for government data APIs

    239 GitHub stars~1.9k tokensUpdated today
    Backend & APIsAuto-check passed
  • Biopython

    K-Dense-AI/scientific-agent-skills

    Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).

    48k GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check: notes
  • Resilience

    codewithmukesh/dotnet-claude-kit

    Resilience patterns for .NET 10 applications using Polly v8.

    756 GitHub starsUsed in 1 repo~3.2k tokens
    Backend & APIsAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Categories

Questions about Bio Batch Downloads

What does Bio Batch Downloads do?

Download large datasets from NCBI efficiently using EPost, history server, batching, rate limiting, and retry logic. Bio Batch Downloads is an agent skill from GPTomics/bioSkills. Download large datasets from NCBI efficiently using EPost, history server, batching, rate limiting, and retry logic.

When should I use Bio Batch Downloads?

Bio Batch Downloads fits situations like: bulk-fetching tens of thousands of sequences; pulling all results of a large ESearch; designing reproducible pipelines; comparing E-utilities to NCBI Datasets v2 CLI.

How do I install Bio Batch Downloads in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-batch-downloads -a claude-code`. Or copy the skill folder (database-access/batch-downloads in GPTomics/bioSkills) into .claude/skills/bio-batch-downloads in your project. Claude Code loads it when a task matches its description.

How do I install Bio Batch Downloads in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-batch-downloads -a codex`. Or copy the skill folder (database-access/batch-downloads in GPTomics/bioSkills) into .agents/skills/bio-batch-downloads in your project. Codex loads it when a task matches its description.

Can I use Bio Batch Downloads in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-batch-downloads -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-batch-downloads, .gemini/skills/bio-batch-downloads, .github/skills/bio-batch-downloads and .opencode/skills/bio-batch-downloads in your project.

What does Bio Batch Downloads need to run?

Going by SKILL.md and its folder, Bio Batch Downloads needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A credential in YOUR_KEY.

Does Bio Batch Downloads access the network?

SKILL.md names 1 domain. As links in the text: ncbi.nlm.nih.gov. This is read from the text; nothing was executed.

Is Bio Batch Downloads safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Batch Downloads use?

Bio Batch Downloads is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Batch Downloads use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Batch Downloads?

Skills that share tags, products or a category with Bio Batch Downloads: Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Azure APIM Policy Authoring (thomast1906/github-copilot-agent-skills, 202 stars), Biopython (davila7/claude-code-templates, 33k stars) and API Integration (Hack23/cia, 239 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Batch Downloads?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.