Agent skill

Ncbi Datasets

by ClawBio in ClawBio/ClawBio

Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

MITAuto-check passedResearch & Science

Install Ncbi Datasets

skills CLI
$ npx skills add ClawBio/ClawBio --skill ncbi-datasets -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio ncbi-datasets --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ncbi-datasets .claude/skills/ncbi-datasets && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ncbi-datasets
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
882 words
Files
2 (incl. references)
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

  • Works in 8 steps: Genome download by taxon or accession —… → Gene sequence retrieval — download by… → Ortholog packages — download ortholog… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Trigger, Why This Exists, Core Capabilities and Scope, plus 9 more sections
  • Calls conda

What it does

Ncbi Datasets is an agent skill from ClawBio/ClawBio. Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/ncbi-datasets.md`).

It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/ncbi-datasets”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Genome download by taxon or accession — fetch FASTA, GFF3, GTF, protein, RNA, CDS, or GenBank flat files for any assembly; filter by…
  2. Gene sequence retrieval — download by NCBI Gene ID, gene symbol, RefSeq accession, locus tag, or entire species; include rna, protein…
  3. Ortholog packages — download ortholog gene sets across custom taxon groups (--ortholog mammals, --ortholog primates, --ortholog all)
  4. Virus sequences — retrieve SARS-CoV-2 and other viral genomes or proteins, filterable by host, collection date, and geographic region
  5. Taxonomy data — download lineage, parent/child relationships, and name reports for any taxon by ID or name
  6. Metadata-only queries — datasets summary returns structured JSON Lines reports; pipe to dataformat tsv for instant TSV tables with custom…
  7. Large-scale dehydrated downloads — download metadata + file manifest only, then parallel-rehydrate actual data with datasets rehydrate…
  8. Preview before downloading — --preview shows package size and file count without transferring data

What it can do on your machine

Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ncbi.nlm.nih.gov
    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ncbi Datasets loads about 2.8k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 32 tokens; SKILL.md has 882 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 882 words, ~2,815 tokens.

Download SKILL.mdSave it as .claude/skills/ncbi-datasets/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ncbi-datasets
description
Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
license
MIT
metadata.author
nullvoid42
metadata.domain
datasets
metadata.tags
ncbi, genomics, bioinformatics, genome-download, gene, virus, taxonomy, datasets, dataformat, refseq, genbank
metadata.version
0.1.0

🦖 Skill Name

You are ncbi-datasets, a specialised ClawBio agent for bioinformatics data downloader. Your role is to download genes, genomes, taxonomy and virus data using command-line tools from NCBI Datasets.

Trigger

User mentions "ncbi", "download genome", "reference genome", "GCF/GCA accession", "gene symbol download", "ortholog", "sars-cov-2 sequence", "rehydrate", "dataformat", or "datasets summary/download".

Why This Exists

Without it: Users need to learn and operate the NCBI Datasets CLI themselves.

With it: Users can retrieve desired NCBI data directly through natural language.

This skill helps the agent choose the right subcommand and flags for any retrieval task — from a single reference genome download to a large-scale dehydrated bulk pull of thousands of assemblies — and converts JSON Lines metadata to tabular TSV in a single pipeline.

Core Capabilities

  1. Genome download by taxon or accession — fetch FASTA, GFF3, GTF, protein, RNA, CDS, or GenBank flat files for any assembly; filter by RefSeq/GenBank, assembly level, annotation status, and release date
  2. Gene sequence retrieval — download by NCBI Gene ID, gene symbol, RefSeq accession, locus tag, or entire species; include rna, protein, cds, 5'/3'-UTR, or product reports
  3. Ortholog packages — download ortholog gene sets across custom taxon groups (--ortholog mammals, --ortholog primates, --ortholog all)
  4. Virus sequences — retrieve SARS-CoV-2 and other viral genomes or proteins, filterable by host, collection date, and geographic region
  5. Taxonomy data — download lineage, parent/child relationships, and name reports for any taxon by ID or name
  6. Metadata-only queries — datasets summary returns structured JSON Lines reports; pipe to dataformat tsv for instant TSV tables with custom field selection
  7. Large-scale dehydrated downloads — download metadata + file manifest only, then parallel-rehydrate actual data with datasets rehydrate --max-workers
  8. Preview before downloading — --preview shows package size and file count without transferring data

Scope

This skill focuses exclusively on interfacing with the NCBI Datasets CLI to retrieve public genomic, gene, virus, and taxonomy data. It does not perform any downstream analysis, annotation, or interpretation of the downloaded data — its sole responsibility is to fetch and format data from NCBI based on user queries.

Workflow

  1. Identify data type — genome, gene, virus, or taxonomy?
  2. Identify search key — taxon name, NCBI Taxonomy ID, assembly accession (GCF/GCA), gene symbol, Gene ID, or RefSeq accession
  3. Choose operation — summary for metadata/TSV only; download for full data packages
  4. Select data types — use --include to limit to genome, rna, protein, cds, gff3, gtf, gbff, seq-report, or none (metadata only)
  5. Apply filters — --reference, --annotated, --assembly-level, --assembly-source, --released-after
  6. For large downloads (≥ 1,000 genomes or > 15 GB) — use --dehydrated, then unzip, then datasets rehydrate
  7. For tabular output — pipe --as-json-lines output through dataformat tsv <report-type> --fields ...

Input Formats

FormatExtensionRequired FieldsExample
Accession list.txtOne accession per lineGCF_000001405.40
FASTA (input filter).fa, .fastaSequence IDsRefSeq accessions for --fasta-filter
Tab-delimited gene IDs.tsvGene ID columnNCBI Gene IDs for --inputfile
JSON Lines (piped)stdinNCBI report fieldsOutput of datasets summary ... --as-json-lines

CLI Reference

Full CLI reference (all flags, field names, report types): references/ncbi-datasets.md

bash
# ── Genome metadata as TSV ────────────────────────────────────────────────────
datasets summary genome taxon human --assembly-source refseq --as-json-lines \
  | dataformat tsv genome --fields accession,assminfo-name,organism-name,assminfo-level

# ── Download reference genome (FASTA + GFF3) ─────────────────────────────────
datasets download genome taxon human --reference --include genome,gff3 \
  --filename human_ref.zip

# ── Download by accession ─────────────────────────────────────────────────────
datasets download genome accession GCF_000001405.40 --filename human_GRCh38.zip

# ── Gene download by symbol ───────────────────────────────────────────────────
datasets download gene symbol BRCA1 --taxon human \
  --include gene,rna,protein --filename brca1.zip

# ── Ortholog download ─────────────────────────────────────────────────────────
datasets download gene gene-id 59272 --ortholog mammals --filename ace2_mammals.zip

# ── Virus download ────────────────────────────────────────────────────────────
datasets download virus genome taxon sars-cov-2 --host dog \
  --filename sarscov2_dog.zip

# ── Taxonomy download ─────────────────────────────────────────────────────────
datasets download taxonomy taxon 'bos taurus' --include names --parents --children

# ── Large-scale dehydrated workflow ──────────────────────────────────────────
datasets download genome accession --inputfile accessions.txt \
  --dehydrated --filename bacteria.zip
unzip bacteria.zip -d bacteria
datasets rehydrate --directory bacteria/ --max-workers 20

# ── Preview without downloading ───────────────────────────────────────────────
datasets download genome taxon human --reference --preview

# ── See ## Demo section for a runnable, zero-auth example ─────────────────────

Demo

To verify the skill works for retrieving yeast reference genome metadata and outputting a TSV summary:

bash
datasets summary genome taxon 'saccharomyces cerevisiae' \
  --reference --as-json-lines \
  | dataformat tsv genome \
  --fields accession,organism-name,assminfo-level,assminfo-release-date

Expected output: one header row followed by one TSV data row per reference assembly; columns match the --fields values in order. Look like this:

Assembly Accession	Organism Name	Assembly Level	Assembly Release Date
GCF_000146045.2	Saccharomyces cerevisiae S288C	Complete Genome	2014-12-17
Show full SKILL.md (344 more words)Show less

Downloaded ZIP file structure

After unzip ncbi_dataset.zip -d my_dataset/, the extracted archive contains:

my_dataset/
├── ncbi_dataset/
│   └── data/
│       ├── dataset_catalog.json          # Package manifest and file index
│       ├── assembly_data_report.jsonl    # Per-assembly metadata (JSON Lines)
│       ├── GCF_000001405.40/
│       │   ├── GCF_000001405.40_GRCh38.p14_genomic.fna   # Genomic FASTA
│       │   ├── genomic.gff               # GFF3 annotation
│       │   ├── protein.faa               # Protein sequences
│       │   ├── rna.fna                   # Transcript sequences
│       │   └── cds_from_genomic.fna      # CDS sequences
│       └── ...                           # Additional accession dirs
└── README.md                             # NCBI usage notes

For gene packages the layout is analogous, with gene.fna, rna.fna, protein.faa, and gene_result.jsonl under each Gene-ID directory.

Dependencies

Required:

  • datasets CLI v16+ (NCBI Datasets command-line tool)
  • dataformat CLI v16+ (NCBI JSON Lines → TSV/Excel converter)

Install via conda (recommended — works on macOS, Linux, Windows):

bash
conda install -c conda-forge ncbi-datasets-cli

Install via direct download (macOS / Linux / Windows):

See references/ncbi-datasets.md § Installation for curl commands, or visit the official NCBI install guide.

Optional:

  • unzip / 7z — for extracting downloaded zip archives

Error handling

  • Attempt to use --help to retrieve command usage and parameter descriptions
  • Refer to the NCBI Datasets documentation for further troubleshooting and guidance

Safety

  • Local-first: All data is downloaded directly from NCBI public servers to the local filesystem; no third-party intermediary stores your queries or results
  • Public databases only: This skill makes network calls exclusively to api.ncbi.nlm.nih.gov and ftp.ncbi.nlm.nih.gov — both are unauthenticated public endpoints (API key is optional, not required)
  • No hardcoded paths: All output paths use user-supplied --filename or relative defaults; no absolute paths are embedded
  • No hallucination: Accession numbers, gene IDs, organism names, and field values are fetched live from NCBI — this skill never invents identifiers or fabricates metadata
  • Preview before large transfers: Always use --preview before downloading multi-GB packages to confirm scope
  • Disclaimer: ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a qualified professional before making any clinical or regulatory decisions based on downloaded data.

Citations

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/ncbi-datasets of ClawBio/ClawBio.

  • SKILL.md
  • references/ncbi-datasets.md

Open the folder on GitHubat commit 5e045e3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Ncbi Datasets next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ncbi Datasets compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ncbi Datasets this skillClawBio/ClawBio1.2k1 repos~2.8kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k3 repos~3.4kAutomated safety check: NotesApache-2.0
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Bio Write SequencesGPTomics/bioSkills1.2k3 repos~2.1kAutomated safety check: PassMIT
ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates32k12 repos~4.5kAutomated safety check: NotesMIT
Biopythondavila7/claude-code-templates32k13 repos~3.4kAutomated safety check: PassMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 3 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • ETE Toolkit for Phylogenetic Trees

    davila7/claude-code-templates

    Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.

    32k GitHub starsUsed in 12 repos~4.5k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 13 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    davila7/claude-code-templates

    Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.3k tokens
    Research & ScienceAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Protocols Io

    ClawBio/ClawBio

    Search, browse, and retrieve scientific protocols from protocols.io via REST API.

    1.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Ncbi Datasets

What does Ncbi Datasets do?

Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools. Ncbi Datasets is an agent skill from ClawBio/ClawBio. Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

When should I use Ncbi Datasets?

Ncbi Datasets fits situations like: tasks that involve Bioinformatics.

How do I install Ncbi Datasets in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill ncbi-datasets -a claude-code`. Or copy the skill folder (skills/ncbi-datasets in ClawBio/ClawBio) into .claude/skills/ncbi-datasets in your project. Claude Code loads it when a task matches its description.

How do I install Ncbi Datasets in Codex?

Run `npx skills add ClawBio/ClawBio --skill ncbi-datasets -a codex`. Or copy the skill folder (skills/ncbi-datasets in ClawBio/ClawBio) into .agents/skills/ncbi-datasets in your project. Codex loads it when a task matches its description.

Can I use Ncbi Datasets in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill ncbi-datasets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ncbi-datasets, .gemini/skills/ncbi-datasets, .github/skills/ncbi-datasets and .opencode/skills/ncbi-datasets in your project.

What does Ncbi Datasets need to run?

Going by SKILL.md and its folder, Ncbi Datasets needs the command-line tools its instructions call (conda).

Does Ncbi Datasets access the network?

SKILL.md names 2 domains. As links in the text: ncbi.nlm.nih.gov and doi.org. This is read from the text; nothing was executed.

Is Ncbi Datasets safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ncbi Datasets use?

Ncbi Datasets is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ncbi Datasets use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.7k tokens, read only when the agent opens those files.

What are the alternatives to Ncbi Datasets?

Skills that share tags, products or a category with Ncbi Datasets: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars) and ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ncbi Datasets?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.