Agent skill

Bio Ncbi Datasets CLI

by GPTomics in GPTomics/bioSkills

Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows).

MITAuto-check passedResearch & Science

Install Bio Ncbi Datasets CLI

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-ncbi-datasets-cli --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/database-access/ncbi-datasets-cli .claude/skills/bio-ncbi-datasets-cli && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-ncbi-datasets-cli
GitHub stars
1.2k
Used in
2 other repos
Token cost
~3.6k tokens
SKILL.md length
1,206 words
Files
5
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows).

  • Works in 3 steps: Discover: datasets download genome taxon… → Inspect: unzip -p human.zip… → Pull: either datasets rehydrate…
  • Bulk-pulling genome assemblies
  • SKILL.md covers Version Compatibility, Installation, What's in scope (use Datasets)… and Subcommand taxonomy, plus 9 more sections
  • Runs Shell scripts from its folder; calls conda and curl; reaches ncbi.nlm.nih.gov and ftp.ncbi.nlm.nih.gov

What it does

Bio Ncbi Datasets CLI is an agent skill from GPTomics/bioSkills. Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows). Use when bulk-pulling genome assemblies, gene metadata across species, ortholog sets, or BLAST databases; when E-utilities are too slow for genome-scale work; or when automatic checksum verification, parallel download, and clean accession-driven retrieval are required. Encodes the JSON-lines output format, dataformat conversion, --dehydrated for…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `examples/bulk_dehydrated.sh`, `examples/download_genome.sh` and `examples/gene_metadata.sh`).

It sits in Research & Science, covering Bioinformatics and Web scraping. It works with NCBI. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Bulk-pulling genome assemblies
  • Gene metadata across species
  • BLAST databases
  • E-utilities are too slow for genome-scale work

Example prompts

  • “/bio-ncbi-datasets-cli”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Discover: datasets download genome taxon human --reference --dehydrated --filename human.zip (fast; ~MB).
  2. Inspect: unzip -p human.zip ncbi_dataset/fetch.txt -- a TSV of all URLs to pull.
  3. Pull: either datasets rehydrate --directory ./human/ or use aria2c --input-file=fetch.txt for parallel pull.

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • conda
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ncbi.nlm.nih.gov
    • ftp.ncbi.nlm.nih.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Ncbi Datasets CLI loads about 3.6k tokens when it runs. Until then it costs about 150 tokens; SKILL.md has 1,206 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~150
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,206 words, ~3,568 tokens.

Download SKILL.mdSave it as .claude/skills/bio-ncbi-datasets-cli/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
bio-ncbi-datasets-cli
description
Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assembly_summary.txt scraping and many EFetch workflows). Use when bulk-pulling genome assemblies, gene metadata across species, ortholog sets, or BLAST databases; when E-utilities are too slow for genome-scale work; or when automatic checksum verification, parallel download, and clean accession-driven retrieval are required. Encodes the JSON-lines output format, dataformat conversion, --dehydrated for cloud workflows, and when Datasets is/isn't the right tool.
tool_type
cli
primary_tool
NCBI Datasets CLI

Version Compatibility

Reference examples tested with: NCBI Datasets CLI 16.0+ (2024), dataformat 16.0+

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: datasets --version, dataformat --version
  • Subcommand help: datasets <subcommand> --help

If a subcommand or flag is unrecognized, run datasets --help and adapt. The CLI is under active development; major releases (v15 -> v16) added subcommands and renamed flags.

NCBI Datasets CLI

"Pull genome / gene / ortholog data from NCBI in 2026" -> The Datasets v2 CLI (launched 2023) is the official, supported bulk endpoint for genome and gene-centric data. It replaces the prior best-practice of scraping assembly_summary.txt + parallel FTP + manual checksum verification. For genome-scale data, it is strictly better than E-utilities (EFetch).

The CLI is not the right answer for everything. PubMed, SRA reads, and custom Entrez queries still belong to E-utilities. The defection rule: if the question is about genome assemblies, gene records, or pre-computed orthologs, use Datasets; otherwise stay with E-utilities.

  • CLI: datasets download genome accession GCF_...
  • CLI: datasets summary gene symbol BRCA1 --taxon human
  • Python: subprocess wrapper; Python client ncbi-datasets-pylib (experimental as of 2024)

Installation

bash
# conda
conda install -c conda-forge ncbi-datasets-cli

# Or direct download (Linux, macOS, Windows binaries)
curl -O https://ftp.ncbi.nlm.nih.gov/pub/datasets/command-line/v2/linux-amd64/datasets

datasets --version    # 16.0+ expected
dataformat --version  # bundled companion tool

What's in scope (use Datasets) vs out of scope (use E-utilities or other tools)

QuestionDatasetsUse instead
Genome assembly downloadyes—
All reference genomes for a taxonyes—
Gene record metadata (multi-species)yes—
Ortholog data for a geneyes (datasets summary gene ... --ortholog)OrthoDB / Compara for tree-aware orthology
Virus data (assemblies, metadata)yes (datasets download virus)—
Annotation files (GFF3, GTF) for a genomeyes—
Protein records (curated, with cross-refs)partialUniProt REST for richer annotation
PubMednoentrez-search / entrez-fetch
SRA readsnosra-data
BLASTnoblast-searches / local-blast
Custom Entrez queriesnoentrez-search
Pre-computed alignments (Compara)noensembl-rest

Subcommand taxonomy

SubcommandPurposeExample
datasets summary genomeMetadata only; JSON outputdatasets summary genome accession GCF_000001405.40
datasets download genomeDownload data filesdatasets download genome accession GCF_...
datasets summary geneGene record metadatadatasets summary gene symbol BRCA1 --taxon human
datasets download geneDownload gene productsdatasets download gene symbol BRCA1 --taxon human
datasets summary taxonomyTaxonomy infodatasets summary taxonomy taxon human
datasets download virusVirus assemblies/proteinsdatasets download virus genome taxon SARS-CoV-2
dataformat tsv / dataformat excelConvert JSON-lines to tabulardataformat tsv gene-summary

datasets summary always returns JSON-lines on stdout (one object per record). datasets download produces a .zip (default) or a "dehydrated" stub for cloud workflows.

Key parameters (download)

FlagEffect
--filename out.zipWhere to write the archive
--include genome,gff3,gtf,protein,cds,rna,seq-reportWhich file types to include
--referenceRestrict to reference assemblies only (one per species)
--annotatedRestrict to annotated assemblies
--assembly-source RefSeq / GenBank / allDatabase source
--assembly-level chromosome,completeAssembly quality level
--released-after 2024-01-01Date filter
--dehydratedSkip data; download just stubs + URL list (for parallel pull)
--api-key XXXOptional API key (raises rate limit)
--no-progressbarFor non-interactive use

For very large pulls (1000+ genomes), --dehydrated is the right choice: download the metadata stubs first, then run datasets rehydrate later or pull URLs in parallel from the manifest.

JSON-lines output + dataformat

datasets summary returns JSON-lines (one JSON object per line) on stdout. Pipe through dataformat tsv for tabular:

bash
datasets summary genome taxon "Escherichia coli" --reference --as-json-lines \
  | dataformat tsv genome --fields accession,organism-name,assembly-level,scaffold-n50 \
  > ecoli_refs.tsv

dataformat subcommands match summary types: genome, gene, virus-genome, etc. The --fields list is documented per type via dataformat tsv <type> --help.

When to use --dehydrated for cloud workflows

The "dehydrated" mode separates data discovery from data transfer:

  1. Discover: datasets download genome taxon human --reference --dehydrated --filename human.zip (fast; ~MB).
  2. Inspect: unzip -p human.zip ncbi_dataset/fetch.txt -- a TSV of all URLs to pull.
  3. Pull: either datasets rehydrate --directory ./human/ or use aria2c --input-file=fetch.txt for parallel pull.

This is essential for HPC / cloud pipelines where inspection of the pending transfer is needed before committing the I/O.

Checksum verification (automatic)

datasets verifies MD5 checksums for every downloaded file automatically. Rehydrate workflows also verify. If a file fails checksum, Datasets retries up to 3 times then errors. This replaces the md5sum -c step that was required with assembly_summary.txt-based scraping.

Code patterns

Download a single reference genome

Goal: Get human reference assembly with genome + GTF + protein + CDS.

Approach: datasets download genome accession ... --include ....

Reference (NCBI Datasets CLI 16.0+):

bash
#!/bin/bash
# Reference: NCBI Datasets CLI 16.0+ | Verify API if version differs

datasets download genome accession GCF_000001405.40 \
    --include genome,gff3,gtf,protein,cds,seq-report \
    --filename human_grch38.zip

unzip -q human_grch38.zip -d human_grch38/
ls -lh human_grch38/ncbi_dataset/data/GCF_000001405.40/
Bulk download all reference bacterial genomes

Goal: Pull every RefSeq reference bacterial assembly with annotation.

Approach: --dehydrated first for inspection; rehydrate with parallel pull.

Reference (NCBI Datasets CLI 16.0+):

bash
#!/bin/bash
# Step 1: dehydrated discovery
datasets download genome taxon Bacteria \
    --reference --annotated --assembly-source RefSeq \
    --include genome,gff3,protein \
    --dehydrated --filename bact_refs.zip

unzip -q bact_refs.zip -d bact_refs/
wc -l bact_refs/ncbi_dataset/fetch.txt   # how many files will be pulled

# Step 2: parallel pull via aria2 (or datasets rehydrate)
aria2c --input-file=bact_refs/ncbi_dataset/fetch.txt \
       --dir=bact_refs/ncbi_dataset/data/ \
       --max-concurrent-downloads=8 \
       --retry-wait=5
Gene metadata across species
bash
datasets summary gene symbol BRCA1 \
    --taxon Mammalia \
    --as-json-lines \
  | dataformat tsv gene --fields gene-id,symbol,taxname,description,nomenclature-authority,chromosomes \
  > brca1_mammals.tsv

head brca1_mammals.tsv
Find orthologs for a gene
bash
datasets summary gene symbol BRCA1 --taxon human --ortholog --as-json-lines \
  | dataformat tsv gene --fields gene-id,symbol,taxname,description \
  > brca1_orthologs.tsv

--ortholog returns NCBI's ortholog set (a single representative per species; tree-aware orthology with multiple co-orthologs is in ortholog-inference / Compara / OMA).

Show full SKILL.md (476 more words)Show less
Filter assemblies by quality and date
bash
datasets summary genome taxon "Salmonella enterica" \
    --assembly-level chromosome,complete \
    --released-after 2024-01-01 \
    --as-json-lines \
  | dataformat tsv genome --fields accession,organism-name,assembly-level,scaffold-n50,submission-date \
  > sal_2024.tsv
Python wrapper with checksum + retry awareness

Reference (NCBI Datasets CLI 16.0+):

python
import subprocess
import json
from pathlib import Path


def datasets_summary(subcommand, *args):
    '''Run `datasets summary` and parse JSON-lines stdout.'''
    cmd = ['datasets', 'summary', subcommand, *args, '--as-json-lines']
    out = subprocess.run(cmd, capture_output=True, text=True, check=True)
    return [json.loads(line) for line in out.stdout.strip().split('\n') if line]


def datasets_download(subcommand, *args, out='dataset.zip', include=None):
    cmd = ['datasets', 'download', subcommand, *args, '--filename', out]
    if include:
        cmd += ['--include', ','.join(include)]
    subprocess.run(cmd, check=True)
    return Path(out)


genomes = datasets_summary('genome', 'taxon', 'Escherichia coli', '--reference')
print(f'{len(genomes)} reference E. coli assemblies')
for g in genomes[:3]:
    acc = g.get('accession')
    n50 = g.get('assemblyStats', {}).get('contigN50')
    print(f'  {acc}  N50={n50}')

datasets_download('genome', 'accession', 'GCF_000005845.2',
                  out='ecoli_k12.zip',
                  include=['genome', 'gff3', 'protein'])
Comparison vs E-utilities
python
# E-utilities path: ESearch in assembly db -> ESummary -> manual FTP pull
#   ~30 API calls + manual md5 + serial download
# Datasets path:
#   datasets download genome accession GCF_...  # one command, automatic md5, parallel inside

For genome workflows, Datasets is 5-50x faster than the equivalent E-utilities pipeline and far more reliable.

Failure modes

Choosing Datasets for the wrong question
  • Trigger: Trying to pull raw SRA reads via Datasets.
  • Mechanism: Datasets covers genome/gene/ortholog, not raw reads.
  • Symptom: Subcommand not found or empty result.
  • Fix: Use sra-data skill (prefetch/fasterq-dump) for raw reads.
--reference filter loses too much
  • Trigger: Bulk pull of "all assemblies for a species"; --reference returns one per species.
  • Mechanism: Reference subset is the canonical single representative.
  • Symptom: Far fewer assemblies than expected for a species with hundreds of submissions.
  • Fix: Drop --reference for full set; add --assembly-level chromosome,complete for quality filter instead.
Dehydrated workflow forgotten
  • Trigger: 1000-genome pull without --dehydrated.
  • Mechanism: Datasets downloads serially within one ZIP; can take hours.
  • Symptom: Slow; no parallelism; one giant ZIP.
  • Fix: Use --dehydrated + aria2c with --max-concurrent-downloads.
dataformat field name guessing
  • Trigger: dataformat tsv genome --fields foo,bar with invented field names.
  • Mechanism: Field names are constrained per summary type.
  • Symptom: "Unknown field" error.
  • Fix: dataformat tsv genome --help lists valid field names; pull JSON-lines and inspect with jq to discover fields.
Old assembly_summary.txt-based scripts still in use
  • Trigger: Legacy pipeline scraping https://ftp.ncbi.nlm.nih.gov/genomes/all/refseq/....
  • Mechanism: Pre-2023 best practice; FTP listing parsing is fragile.
  • Symptom: Slow; brittle; no checksums; broken when NCBI restructures FTP.
  • Fix: Switch to Datasets CLI; the FTP path still works but Datasets is the supported modern path.
API key not used for high-volume
  • Trigger: 1000+ summary calls in a loop without --api-key.
  • Mechanism: NCBI rate-limits unauthenticated bulk traffic.
  • Symptom: Throttling; slow downloads.
  • Fix: Pass --api-key YOUR_KEY to bulk commands; obtain from https://www.ncbi.nlm.nih.gov/account/settings/.
CLI version drift
  • Trigger: Using Datasets v14 with v16 docs.
  • Mechanism: Subcommands and flags renamed between major versions.
  • Symptom: "Unknown flag" or different output structure.
  • Fix: Pin to v16+; conda update ncbi-datasets-cli.

Common errors

Error / symptomCauseSolution
"command not found: datasets"Not installedconda install -c conda-forge ncbi-datasets-cli
Subcommand not foundOld versionUpgrade to v16+
Slow 1000-genome pullSerial downloadUse --dehydrated + aria2c
"Unknown field" in dataformatWrong field nameCheck dataformat <type> --help
Throttled bulk pullNo API keyPass --api-key
--reference returns 1 per speciesBy designDrop the flag or use --assembly-level
MD5 mismatch retriedNetwork issueDatasets retries automatically; persistent failure -> investigate network

References

  • entrez-search - For PubMed, custom queries, and non-genome data
  • entrez-fetch - For single-record fetches outside genome/gene scope
  • batch-downloads - Bulk E-utilities (when not genome-scale)
  • sra-data - Raw sequencing reads (NOT covered by Datasets)
  • ensembl-rest - Ensembl REST as alternative for Ensembl-native species
  • ortholog-inference - Compara/OMA/OrthoDB for tree-aware orthology

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in database-access/ncbi-datasets-cli of GPTomics/bioSkills.

  • SKILL.md
  • examples/bulk_dehydrated.sh
  • examples/download_genome.sh
  • examples/gene_metadata.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Ncbi Datasets CLI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Ncbi Datasets CLI compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Ncbi Datasets CLI this skillGPTomics/bioSkills1.2k2 repos~3.6kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k3 repos~3.4kAutomated safety check: NotesApache-2.0
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates32k12 repos~4.5kAutomated safety check: NotesMIT
Biopythondavila7/claude-code-templates32k13 repos~3.4kAutomated safety check: PassMIT
Clinvar Databasedavila7/claude-code-templates32k11 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 3 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • ETE Toolkit for Phylogenetic Trees

    davila7/claude-code-templates

    Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.

    32k GitHub starsUsed in 12 repos~4.5k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 13 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    davila7/claude-code-templates

    Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.3k tokens
    Research & ScienceAuto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Ncbi Datasets CLI

What does Bio Ncbi Datasets CLI do?

Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows). Bio Ncbi Datasets CLI is an agent skill from GPTomics/bioSkills.txt scraping and many EFetch workflows).

When should I use Bio Ncbi Datasets CLI?

Bio Ncbi Datasets CLI fits situations like: bulk-pulling genome assemblies; gene metadata across species; BLAST databases; E-utilities are too slow for genome-scale work.

How do I install Bio Ncbi Datasets CLI in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a claude-code`. Or copy the skill folder (database-access/ncbi-datasets-cli in GPTomics/bioSkills) into .claude/skills/bio-ncbi-datasets-cli in your project. Claude Code loads it when a task matches its description.

How do I install Bio Ncbi Datasets CLI in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a codex`. Or copy the skill folder (database-access/ncbi-datasets-cli in GPTomics/bioSkills) into .agents/skills/bio-ncbi-datasets-cli in your project. Codex loads it when a task matches its description.

Can I use Bio Ncbi Datasets CLI in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-ncbi-datasets-cli, .gemini/skills/bio-ncbi-datasets-cli, .github/skills/bio-ncbi-datasets-cli and .opencode/skills/bio-ncbi-datasets-cli in your project.

What does Bio Ncbi Datasets CLI need to run?

Going by SKILL.md and its folder, Bio Ncbi Datasets CLI needs a shell for the scripts in its folder and the command-line tools its instructions call (conda and curl). Our summary lists: Python 3; A Bash shell.

Does Bio Ncbi Datasets CLI access the network?

SKILL.md names 2 domains. In commands or code: ncbi.nlm.nih.gov and ftp.ncbi.nlm.nih.gov; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Ncbi Datasets CLI safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Ncbi Datasets CLI use?

Bio Ncbi Datasets CLI is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Ncbi Datasets CLI use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Ncbi Datasets CLI?

Skills that share tags, products or a category with Bio Ncbi Datasets CLI: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 32k stars) and Biopython (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Ncbi Datasets CLI?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.