Dbsnp Database
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows).
$ npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-ncbi-datasets-cli --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/database-access/ncbi-datasets-cli .claude/skills/bio-ncbi-datasets-cli && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-ncbi-datasets-cli" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/ncbi-datasets-cli into .claude/skills/bio-ncbi-datasets-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-ncbi-datasets-cli", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/database-access/ncbi-datasets-cliType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-ncbi-datasets-cli --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/database-access/ncbi-datasets-cli .agents/skills/bio-ncbi-datasets-cli && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-ncbi-datasets-cli" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/ncbi-datasets-cli into .agents/skills/bio-ncbi-datasets-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-ncbi-datasets-cli", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-ncbi-datasets-cli --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/database-access/ncbi-datasets-cli .cursor/skills/bio-ncbi-datasets-cli && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-ncbi-datasets-cli" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/ncbi-datasets-cli into .cursor/skills/bio-ncbi-datasets-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-ncbi-datasets-cli", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path database-access/ncbi-datasets-cli--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-ncbi-datasets-cli --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/database-access/ncbi-datasets-cli .gemini/skills/bio-ncbi-datasets-cli && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-ncbi-datasets-cli" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/ncbi-datasets-cli into .gemini/skills/bio-ncbi-datasets-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-ncbi-datasets-cli", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-ncbi-datasets-cliInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/database-access/ncbi-datasets-cli .github/skills/bio-ncbi-datasets-cli && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-ncbi-datasets-cli" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/ncbi-datasets-cli into .github/skills/bio-ncbi-datasets-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-ncbi-datasets-cli", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-ncbi-datasets-cli --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/database-access/ncbi-datasets-cli .opencode/skills/bio-ncbi-datasets-cli && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-ncbi-datasets-cli" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/ncbi-datasets-cli into .opencode/skills/bio-ncbi-datasets-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-ncbi-datasets-cli", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-ncbi-datasets-cliDownload genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows).
Bio Ncbi Datasets CLI is an agent skill from GPTomics/bioSkills. Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows). Use when bulk-pulling genome assemblies, gene metadata across species, ortholog sets, or BLAST databases; when E-utilities are too slow for genome-scale work; or when automatic checksum verification, parallel download, and clean accession-driven retrieval are required. Encodes the JSON-lines output format, dataformat conversion, --dehydrated for…
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `examples/bulk_dehydrated.sh`, `examples/download_genome.sh` and `examples/gene_metadata.sh`).
It sits in Research & Science, covering Bioinformatics and Web scraping. It works with NCBI. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
condacurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
ncbi.nlm.nih.govftp.ncbi.nlm.nih.govFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Ncbi Datasets CLI loads about 3.6k tokens when it runs. Until then it costs about 150 tokens; SKILL.md has 1,206 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,206 words, ~3,568 tokens.
.claude/skills/bio-ncbi-datasets-cli/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Reference examples tested with: NCBI Datasets CLI 16.0+ (2024), dataformat 16.0+
Before using code patterns, verify installed versions match. If versions differ:
datasets --version, dataformat --versiondatasets <subcommand> --helpIf a subcommand or flag is unrecognized, run datasets --help and adapt. The CLI is under active development; major releases (v15 -> v16) added subcommands and renamed flags.
"Pull genome / gene / ortholog data from NCBI in 2026" -> The Datasets v2 CLI (launched 2023) is the official, supported bulk endpoint for genome and gene-centric data. It replaces the prior best-practice of scraping assembly_summary.txt + parallel FTP + manual checksum verification. For genome-scale data, it is strictly better than E-utilities (EFetch).
The CLI is not the right answer for everything. PubMed, SRA reads, and custom Entrez queries still belong to E-utilities. The defection rule: if the question is about genome assemblies, gene records, or pre-computed orthologs, use Datasets; otherwise stay with E-utilities.
datasets download genome accession GCF_...datasets summary gene symbol BRCA1 --taxon humansubprocess wrapper; Python client ncbi-datasets-pylib (experimental as of 2024)# conda
conda install -c conda-forge ncbi-datasets-cli
# Or direct download (Linux, macOS, Windows binaries)
curl -O https://ftp.ncbi.nlm.nih.gov/pub/datasets/command-line/v2/linux-amd64/datasets
datasets --version # 16.0+ expected
dataformat --version # bundled companion tool| Question | Datasets | Use instead |
|---|---|---|
| Genome assembly download | yes | — |
| All reference genomes for a taxon | yes | — |
| Gene record metadata (multi-species) | yes | — |
| Ortholog data for a gene | yes (datasets summary gene ... --ortholog) | OrthoDB / Compara for tree-aware orthology |
| Virus data (assemblies, metadata) | yes (datasets download virus) | — |
| Annotation files (GFF3, GTF) for a genome | yes | — |
| Protein records (curated, with cross-refs) | partial | UniProt REST for richer annotation |
| PubMed | no | entrez-search / entrez-fetch |
| SRA reads | no | sra-data |
| BLAST | no | blast-searches / local-blast |
| Custom Entrez queries | no | entrez-search |
| Pre-computed alignments (Compara) | no | ensembl-rest |
| Subcommand | Purpose | Example |
|---|---|---|
datasets summary genome | Metadata only; JSON output | datasets summary genome accession GCF_000001405.40 |
datasets download genome | Download data files | datasets download genome accession GCF_... |
datasets summary gene | Gene record metadata | datasets summary gene symbol BRCA1 --taxon human |
datasets download gene | Download gene products | datasets download gene symbol BRCA1 --taxon human |
datasets summary taxonomy | Taxonomy info | datasets summary taxonomy taxon human |
datasets download virus | Virus assemblies/proteins | datasets download virus genome taxon SARS-CoV-2 |
dataformat tsv / dataformat excel | Convert JSON-lines to tabular | dataformat tsv gene-summary |
datasets summary always returns JSON-lines on stdout (one object per record). datasets download produces a .zip (default) or a "dehydrated" stub for cloud workflows.
| Flag | Effect |
|---|---|
--filename out.zip | Where to write the archive |
--include genome,gff3,gtf,protein,cds,rna,seq-report | Which file types to include |
--reference | Restrict to reference assemblies only (one per species) |
--annotated | Restrict to annotated assemblies |
--assembly-source RefSeq / GenBank / all | Database source |
--assembly-level chromosome,complete | Assembly quality level |
--released-after 2024-01-01 | Date filter |
--dehydrated | Skip data; download just stubs + URL list (for parallel pull) |
--api-key XXX | Optional API key (raises rate limit) |
--no-progressbar | For non-interactive use |
For very large pulls (1000+ genomes), --dehydrated is the right choice: download the metadata stubs first, then run datasets rehydrate later or pull URLs in parallel from the manifest.
datasets summary returns JSON-lines (one JSON object per line) on stdout. Pipe through dataformat tsv for tabular:
datasets summary genome taxon "Escherichia coli" --reference --as-json-lines \
| dataformat tsv genome --fields accession,organism-name,assembly-level,scaffold-n50 \
> ecoli_refs.tsvdataformat subcommands match summary types: genome, gene, virus-genome, etc. The --fields list is documented per type via dataformat tsv <type> --help.
The "dehydrated" mode separates data discovery from data transfer:
datasets download genome taxon human --reference --dehydrated --filename human.zip (fast; ~MB).unzip -p human.zip ncbi_dataset/fetch.txt -- a TSV of all URLs to pull.datasets rehydrate --directory ./human/ or use aria2c --input-file=fetch.txt for parallel pull.This is essential for HPC / cloud pipelines where inspection of the pending transfer is needed before committing the I/O.
datasets verifies MD5 checksums for every downloaded file automatically. Rehydrate workflows also verify. If a file fails checksum, Datasets retries up to 3 times then errors. This replaces the md5sum -c step that was required with assembly_summary.txt-based scraping.
Goal: Get human reference assembly with genome + GTF + protein + CDS.
Approach: datasets download genome accession ... --include ....
Reference (NCBI Datasets CLI 16.0+):
#!/bin/bash
# Reference: NCBI Datasets CLI 16.0+ | Verify API if version differs
datasets download genome accession GCF_000001405.40 \
--include genome,gff3,gtf,protein,cds,seq-report \
--filename human_grch38.zip
unzip -q human_grch38.zip -d human_grch38/
ls -lh human_grch38/ncbi_dataset/data/GCF_000001405.40/Goal: Pull every RefSeq reference bacterial assembly with annotation.
Approach: --dehydrated first for inspection; rehydrate with parallel pull.
Reference (NCBI Datasets CLI 16.0+):
#!/bin/bash
# Step 1: dehydrated discovery
datasets download genome taxon Bacteria \
--reference --annotated --assembly-source RefSeq \
--include genome,gff3,protein \
--dehydrated --filename bact_refs.zip
unzip -q bact_refs.zip -d bact_refs/
wc -l bact_refs/ncbi_dataset/fetch.txt # how many files will be pulled
# Step 2: parallel pull via aria2 (or datasets rehydrate)
aria2c --input-file=bact_refs/ncbi_dataset/fetch.txt \
--dir=bact_refs/ncbi_dataset/data/ \
--max-concurrent-downloads=8 \
--retry-wait=5datasets summary gene symbol BRCA1 \
--taxon Mammalia \
--as-json-lines \
| dataformat tsv gene --fields gene-id,symbol,taxname,description,nomenclature-authority,chromosomes \
> brca1_mammals.tsv
head brca1_mammals.tsvdatasets summary gene symbol BRCA1 --taxon human --ortholog --as-json-lines \
| dataformat tsv gene --fields gene-id,symbol,taxname,description \
> brca1_orthologs.tsv--ortholog returns NCBI's ortholog set (a single representative per species; tree-aware orthology with multiple co-orthologs is in ortholog-inference / Compara / OMA).
datasets summary genome taxon "Salmonella enterica" \
--assembly-level chromosome,complete \
--released-after 2024-01-01 \
--as-json-lines \
| dataformat tsv genome --fields accession,organism-name,assembly-level,scaffold-n50,submission-date \
> sal_2024.tsvReference (NCBI Datasets CLI 16.0+):
import subprocess
import json
from pathlib import Path
def datasets_summary(subcommand, *args):
'''Run `datasets summary` and parse JSON-lines stdout.'''
cmd = ['datasets', 'summary', subcommand, *args, '--as-json-lines']
out = subprocess.run(cmd, capture_output=True, text=True, check=True)
return [json.loads(line) for line in out.stdout.strip().split('\n') if line]
def datasets_download(subcommand, *args, out='dataset.zip', include=None):
cmd = ['datasets', 'download', subcommand, *args, '--filename', out]
if include:
cmd += ['--include', ','.join(include)]
subprocess.run(cmd, check=True)
return Path(out)
genomes = datasets_summary('genome', 'taxon', 'Escherichia coli', '--reference')
print(f'{len(genomes)} reference E. coli assemblies')
for g in genomes[:3]:
acc = g.get('accession')
n50 = g.get('assemblyStats', {}).get('contigN50')
print(f' {acc} N50={n50}')
datasets_download('genome', 'accession', 'GCF_000005845.2',
out='ecoli_k12.zip',
include=['genome', 'gff3', 'protein'])# E-utilities path: ESearch in assembly db -> ESummary -> manual FTP pull
# ~30 API calls + manual md5 + serial download
# Datasets path:
# datasets download genome accession GCF_... # one command, automatic md5, parallel insideFor genome workflows, Datasets is 5-50x faster than the equivalent E-utilities pipeline and far more reliable.
sra-data skill (prefetch/fasterq-dump) for raw reads.--reference filter loses too much--reference returns one per species.--reference for full set; add --assembly-level chromosome,complete for quality filter instead.--dehydrated.--dehydrated + aria2c with --max-concurrent-downloads.dataformat tsv genome --fields foo,bar with invented field names.dataformat tsv genome --help lists valid field names; pull JSON-lines and inspect with jq to discover fields.https://ftp.ncbi.nlm.nih.gov/genomes/all/refseq/....--api-key.--api-key YOUR_KEY to bulk commands; obtain from https://www.ncbi.nlm.nih.gov/account/settings/.conda update ncbi-datasets-cli.| Error / symptom | Cause | Solution |
|---|---|---|
| "command not found: datasets" | Not installed | conda install -c conda-forge ncbi-datasets-cli |
| Subcommand not found | Old version | Upgrade to v16+ |
| Slow 1000-genome pull | Serial download | Use --dehydrated + aria2c |
| "Unknown field" in dataformat | Wrong field name | Check dataformat <type> --help |
| Throttled bulk pull | No API key | Pass --api-key |
--reference returns 1 per species | By design | Drop the flag or use --assembly-level |
| MD5 mismatch retried | Network issue | Datasets retries automatically; persistent failure -> investigate network |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in database-access/ncbi-datasets-cli of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Ncbi Datasets CLI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Ncbi Datasets CLI this skillGPTomics/bioSkills | 1.2k | 2 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 3 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 | |
| Biopython Bioinformaticsaiming-lab/AutoResearchClaw | 15k | — | ~810 | Automated safety check: Pass | MIT | |
| ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates | 32k | 12 repos | ~4.5k | Automated safety check: Notes | MIT | |
| Biopythondavila7/claude-code-templates | 32k | 13 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Clinvar Databasedavila7/claude-code-templates | 32k | 11 repos | ~3.3k | Automated safety check: Pass | MIT |
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.
davila7/claude-code-templates
Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.
davila7/claude-code-templates
Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.
davila7/claude-code-templates
Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.
ClawBio/ClawBio
Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Works with
Categories
Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows). Bio Ncbi Datasets CLI is an agent skill from GPTomics/bioSkills.txt scraping and many EFetch workflows).
Bio Ncbi Datasets CLI fits situations like: bulk-pulling genome assemblies; gene metadata across species; BLAST databases; E-utilities are too slow for genome-scale work.
Run `npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a claude-code`. Or copy the skill folder (database-access/ncbi-datasets-cli in GPTomics/bioSkills) into .claude/skills/bio-ncbi-datasets-cli in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a codex`. Or copy the skill folder (database-access/ncbi-datasets-cli in GPTomics/bioSkills) into .agents/skills/bio-ncbi-datasets-cli in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-ncbi-datasets-cli -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-ncbi-datasets-cli, .gemini/skills/bio-ncbi-datasets-cli, .github/skills/bio-ncbi-datasets-cli and .opencode/skills/bio-ncbi-datasets-cli in your project.
Going by SKILL.md and its folder, Bio Ncbi Datasets CLI needs a shell for the scripts in its folder and the command-line tools its instructions call (conda and curl). Our summary lists: Python 3; A Bash shell.
SKILL.md names 2 domains. In commands or code: ncbi.nlm.nih.gov and ftp.ncbi.nlm.nih.gov; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Ncbi Datasets CLI is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Ncbi Datasets CLI: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 32k stars) and Biopython (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.