Dbsnp Database
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
$ npx skills add ClawBio/ClawBio --skill ncbi-datasets -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ClawBio/ClawBio ncbi-datasets --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ncbi-datasets .claude/skills/ncbi-datasets && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ncbi-datasets" agent skill from https://github.com/ClawBio/ClawBio/tree/main/skills/ncbi-datasets into .claude/skills/ncbi-datasets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncbi-datasets", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ClawBio/ClawBio/tree/main/skills/ncbi-datasetsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ClawBio/ClawBio --skill ncbi-datasets -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ClawBio/ClawBio ncbi-datasets --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ncbi-datasets .agents/skills/ncbi-datasets && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ncbi-datasets" agent skill from https://github.com/ClawBio/ClawBio/tree/main/skills/ncbi-datasets into .agents/skills/ncbi-datasets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncbi-datasets", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ClawBio/ClawBio --skill ncbi-datasets -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ClawBio/ClawBio ncbi-datasets --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ncbi-datasets .cursor/skills/ncbi-datasets && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ncbi-datasets" agent skill from https://github.com/ClawBio/ClawBio/tree/main/skills/ncbi-datasets into .cursor/skills/ncbi-datasets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncbi-datasets", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ClawBio/ClawBio.git --path skills/ncbi-datasets--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ClawBio/ClawBio --skill ncbi-datasets -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ClawBio/ClawBio ncbi-datasets --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ncbi-datasets .gemini/skills/ncbi-datasets && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ncbi-datasets" agent skill from https://github.com/ClawBio/ClawBio/tree/main/skills/ncbi-datasets into .gemini/skills/ncbi-datasets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncbi-datasets", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ClawBio/ClawBio ncbi-datasetsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ClawBio/ClawBio --skill ncbi-datasets -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ncbi-datasets .github/skills/ncbi-datasets && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ncbi-datasets" agent skill from https://github.com/ClawBio/ClawBio/tree/main/skills/ncbi-datasets into .github/skills/ncbi-datasets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncbi-datasets", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ClawBio/ClawBio --skill ncbi-datasets -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ClawBio/ClawBio ncbi-datasets --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ncbi-datasets .opencode/skills/ncbi-datasets && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ncbi-datasets" agent skill from https://github.com/ClawBio/ClawBio/tree/main/skills/ncbi-datasets into .opencode/skills/ncbi-datasets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncbi-datasets", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ncbi-datasetsDownload genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
Ncbi Datasets is an agent skill from ClawBio/ClawBio. Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/ncbi-datasets.md`).
It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
condaFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
ncbi.nlm.nih.govdoi.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ncbi Datasets loads about 2.8k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 32 tokens; SKILL.md has 882 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 882 words, ~2,815 tokens.
.claude/skills/ncbi-datasets/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.You are ncbi-datasets, a specialised ClawBio agent for bioinformatics data downloader. Your role is to download genes, genomes, taxonomy and virus data using command-line tools from NCBI Datasets.
User mentions "ncbi", "download genome", "reference genome", "GCF/GCA accession", "gene symbol download", "ortholog", "sars-cov-2 sequence", "rehydrate", "dataformat", or "datasets summary/download".
Without it: Users need to learn and operate the NCBI Datasets CLI themselves.
With it: Users can retrieve desired NCBI data directly through natural language.
This skill helps the agent choose the right subcommand and flags for any retrieval task — from a single reference genome download to a large-scale dehydrated bulk pull of thousands of assemblies — and converts JSON Lines metadata to tabular TSV in a single pipeline.
--ortholog mammals, --ortholog primates, --ortholog all)datasets summary returns structured JSON Lines reports; pipe to dataformat tsv for instant TSV tables with custom field selectiondatasets rehydrate --max-workers--preview shows package size and file count without transferring dataThis skill focuses exclusively on interfacing with the NCBI Datasets CLI to retrieve public genomic, gene, virus, and taxonomy data. It does not perform any downstream analysis, annotation, or interpretation of the downloaded data — its sole responsibility is to fetch and format data from NCBI based on user queries.
summary for metadata/TSV only; download for full data packages--include to limit to genome, rna, protein, cds, gff3, gtf, gbff, seq-report, or none (metadata only)--reference, --annotated, --assembly-level, --assembly-source, --released-after--dehydrated, then unzip, then datasets rehydrate--as-json-lines output through dataformat tsv <report-type> --fields ...| Format | Extension | Required Fields | Example |
|---|---|---|---|
| Accession list | .txt | One accession per line | GCF_000001405.40 |
| FASTA (input filter) | .fa, .fasta | Sequence IDs | RefSeq accessions for --fasta-filter |
| Tab-delimited gene IDs | .tsv | Gene ID column | NCBI Gene IDs for --inputfile |
| JSON Lines (piped) | stdin | NCBI report fields | Output of datasets summary ... --as-json-lines |
Full CLI reference (all flags, field names, report types):
references/ncbi-datasets.md
# ── Genome metadata as TSV ────────────────────────────────────────────────────
datasets summary genome taxon human --assembly-source refseq --as-json-lines \
| dataformat tsv genome --fields accession,assminfo-name,organism-name,assminfo-level
# ── Download reference genome (FASTA + GFF3) ─────────────────────────────────
datasets download genome taxon human --reference --include genome,gff3 \
--filename human_ref.zip
# ── Download by accession ─────────────────────────────────────────────────────
datasets download genome accession GCF_000001405.40 --filename human_GRCh38.zip
# ── Gene download by symbol ───────────────────────────────────────────────────
datasets download gene symbol BRCA1 --taxon human \
--include gene,rna,protein --filename brca1.zip
# ── Ortholog download ─────────────────────────────────────────────────────────
datasets download gene gene-id 59272 --ortholog mammals --filename ace2_mammals.zip
# ── Virus download ────────────────────────────────────────────────────────────
datasets download virus genome taxon sars-cov-2 --host dog \
--filename sarscov2_dog.zip
# ── Taxonomy download ─────────────────────────────────────────────────────────
datasets download taxonomy taxon 'bos taurus' --include names --parents --children
# ── Large-scale dehydrated workflow ──────────────────────────────────────────
datasets download genome accession --inputfile accessions.txt \
--dehydrated --filename bacteria.zip
unzip bacteria.zip -d bacteria
datasets rehydrate --directory bacteria/ --max-workers 20
# ── Preview without downloading ───────────────────────────────────────────────
datasets download genome taxon human --reference --preview
# ── See ## Demo section for a runnable, zero-auth example ─────────────────────To verify the skill works for retrieving yeast reference genome metadata and outputting a TSV summary:
datasets summary genome taxon 'saccharomyces cerevisiae' \
--reference --as-json-lines \
| dataformat tsv genome \
--fields accession,organism-name,assminfo-level,assminfo-release-dateExpected output: one header row followed by one TSV data row per reference assembly; columns match the --fields values in order.
Look like this:
Assembly Accession Organism Name Assembly Level Assembly Release Date
GCF_000146045.2 Saccharomyces cerevisiae S288C Complete Genome 2014-12-17After unzip ncbi_dataset.zip -d my_dataset/, the extracted archive contains:
my_dataset/
├── ncbi_dataset/
│ └── data/
│ ├── dataset_catalog.json # Package manifest and file index
│ ├── assembly_data_report.jsonl # Per-assembly metadata (JSON Lines)
│ ├── GCF_000001405.40/
│ │ ├── GCF_000001405.40_GRCh38.p14_genomic.fna # Genomic FASTA
│ │ ├── genomic.gff # GFF3 annotation
│ │ ├── protein.faa # Protein sequences
│ │ ├── rna.fna # Transcript sequences
│ │ └── cds_from_genomic.fna # CDS sequences
│ └── ... # Additional accession dirs
└── README.md # NCBI usage notesFor gene packages the layout is analogous, with gene.fna, rna.fna, protein.faa, and gene_result.jsonl under each Gene-ID directory.
Required:
datasets CLI v16+ (NCBI Datasets command-line tool)dataformat CLI v16+ (NCBI JSON Lines → TSV/Excel converter)Install via conda (recommended — works on macOS, Linux, Windows):
conda install -c conda-forge ncbi-datasets-cliInstall via direct download (macOS / Linux / Windows):
See
references/ncbi-datasets.md § Installationfor curl commands, or visit the official NCBI install guide.
Optional:
unzip / 7z — for extracting downloaded zip archivesapi.ncbi.nlm.nih.gov and ftp.ncbi.nlm.nih.gov — both are unauthenticated public endpoints (API key is optional, not required)--filename or relative defaults; no absolute paths are embedded--preview before downloading multi-GB packages to confirm scope© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/ncbi-datasets of ClawBio/ClawBio.
Open the folder on GitHubat commit 5e045e3
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.
Ncbi Datasets next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ncbi Datasets this skillClawBio/ClawBio | 1.2k | 1 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 3 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 | |
| Biopython Bioinformaticsaiming-lab/AutoResearchClaw | 15k | — | ~810 | Automated safety check: Pass | MIT | |
| Bio Write SequencesGPTomics/bioSkills | 1.2k | 3 repos | ~2.1k | Automated safety check: Pass | MIT | |
| ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates | 32k | 12 repos | ~4.5k | Automated safety check: Notes | MIT | |
| Biopythondavila7/claude-code-templates | 32k | 13 repos | ~3.4k | Automated safety check: Pass | MIT |
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
davila7/claude-code-templates
Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.
davila7/claude-code-templates
Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.
davila7/claude-code-templates
Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.
ClawBio/ClawBio
Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.
ClawBio/ClawBio
Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.
ClawBio/ClawBio
Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.
ClawBio/ClawBio
Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.
ClawBio/ClawBio
Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.
ClawBio/ClawBio
Search, browse, and retrieve scientific protocols from protocols.io via REST API.
Works with
Categories
Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools. Ncbi Datasets is an agent skill from ClawBio/ClawBio. Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
Ncbi Datasets fits situations like: tasks that involve Bioinformatics.
Run `npx skills add ClawBio/ClawBio --skill ncbi-datasets -a claude-code`. Or copy the skill folder (skills/ncbi-datasets in ClawBio/ClawBio) into .claude/skills/ncbi-datasets in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ClawBio/ClawBio --skill ncbi-datasets -a codex`. Or copy the skill folder (skills/ncbi-datasets in ClawBio/ClawBio) into .agents/skills/ncbi-datasets in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill ncbi-datasets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ncbi-datasets, .gemini/skills/ncbi-datasets, .github/skills/ncbi-datasets and .opencode/skills/ncbi-datasets in your project.
Going by SKILL.md and its folder, Ncbi Datasets needs the command-line tools its instructions call (conda).
SKILL.md names 2 domains. As links in the text: ncbi.nlm.nih.gov and doi.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ncbi Datasets is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ncbi Datasets: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars) and ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.
Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.