Dbsnp Database
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC…
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-prokaryotic-annotation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-annotation/prokaryotic-annotation .claude/skills/bio-genome-annotation-prokaryotic-annotation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-genome-annotation-prokaryotic-annotation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/prokaryotic-annotation into .claude/skills/bio-genome-annotation-prokaryotic-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-prokaryotic-annotation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/prokaryotic-annotationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-prokaryotic-annotation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/genome-annotation/prokaryotic-annotation .agents/skills/bio-genome-annotation-prokaryotic-annotation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-genome-annotation-prokaryotic-annotation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/prokaryotic-annotation into .agents/skills/bio-genome-annotation-prokaryotic-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-prokaryotic-annotation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-prokaryotic-annotation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/genome-annotation/prokaryotic-annotation .cursor/skills/bio-genome-annotation-prokaryotic-annotation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-genome-annotation-prokaryotic-annotation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/prokaryotic-annotation into .cursor/skills/bio-genome-annotation-prokaryotic-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-prokaryotic-annotation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path genome-annotation/prokaryotic-annotation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-prokaryotic-annotation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/genome-annotation/prokaryotic-annotation .gemini/skills/bio-genome-annotation-prokaryotic-annotation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-genome-annotation-prokaryotic-annotation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/prokaryotic-annotation into .gemini/skills/bio-genome-annotation-prokaryotic-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-prokaryotic-annotation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-prokaryotic-annotationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/genome-annotation/prokaryotic-annotation .github/skills/bio-genome-annotation-prokaryotic-annotation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-genome-annotation-prokaryotic-annotation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/prokaryotic-annotation into .github/skills/bio-genome-annotation-prokaryotic-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-prokaryotic-annotation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-prokaryotic-annotation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/genome-annotation/prokaryotic-annotation .opencode/skills/bio-genome-annotation-prokaryotic-annotation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-genome-annotation-prokaryotic-annotation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/prokaryotic-annotation into .opencode/skills/bio-genome-annotation-prokaryotic-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-prokaryotic-annotation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-genome-annotation-prokaryotic-annotationAnnotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC…
Bio Genome Annotation Prokaryotic Annotation is an agent skill from GPTomics/bioSkills. Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC locus tags. Covers Bakta-vs-Prokka-vs-PGAP-vs-DFAST choice, light-vs-full database tiers, translation-table selection (11/4/25), archaeal and leaderless-gene caveats, the small-ORF blind spot, pseudogene-vs-phase-variation, the pangenome re-annotation trap, and submission compliance. Use when annotating a newly assembled…
Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/bakta_annotation.sh`, `examples/parse_annotation.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell and Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Genome Annotation Prokaryotic Annotation loads about 4.2k tokens when it runs. Until then it costs about 177 tokens; SKILL.md has 1,848 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,848 words, ~4,205 tokens.
.claude/skills/bio-genome-annotation-prokaryotic-annotation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: Bakta 1.9+, Prokka 1.14.6 (legacy), tRNAscan-SE 2.0+, CheckM2 1.0+, gffutils 0.12+.
Before using code patterns, verify installed versions match. If versions differ:
<tool> --version then <tool> --help to confirm flagspip show <package> then help(module.function) to check signaturesAnnotation content tracks the database version, not just the binary: a Bakta full DB (schema v5+, ~30 GB zipped) and a Bakta light DB give different functional calls, and two runs months apart can differ purely from DB updates. Record the Bakta DB version in methods. Prokka's bundled databases are frozen (~2019-2021). If code throws an error, introspect the installed tool and adapt rather than retrying.
"Annotate my bacterial genome" -> Call protein-coding genes, tRNAs/rRNAs/ncRNAs, and other features, then assign function by database identity, and emit submission-ready files.
bakta --db db/ assembly.fa (default), prokka --outdir out assembly.fa (legacy), NCBI PGAP (submission/RefSeq-grade)For a typical isolate, Prodigal/Pyrodigal recovers >95-99% of true coding genes. The unsolved problems are the start codon (translation initiation site), the small/overlapping/recoded ORFs, and above all the function. Three consequences a postdoc must internalize:
A confidently wrong product name is worse than "hypothetical protein." Names assigned by loose homology transfer across distant lineages and self-amplify through databases - there is no mechanism to retract a correction once it spreads, so annotation accuracy has gone down, not up, as sequencing scaled (Salzberg 2019 Genome Biol 20:92). Bakta's design counters this: gene-level identity only via exact (MD5/UniRef100) match, falling back to cluster level, then "hypothetical" rather than guessing. 25-50% hypothetical is healthy for a non-model organism; near 0% means over-confident transfer.
Comparing gene counts across tools/versions/DBs is invalid. Tool agreement is "highly dependent on the organism of study" and biased toward model organisms (Dimonaco 2022 Bioinformatics 38:1198) - there is no universal best tool. Differences between two annotations are mostly tool artifacts, not biology. For any collection (pangenomics), re-annotate every assembly from FASTA with one pipeline + one pinned DB; merging published annotations inflates the accessory genome ~10× (Tonkin-Hill 2020 Genome Biol 21:180).
Annotation completeness ≠ assembly completeness. A fragmented/contaminated assembly produces truncated partial CDS at every contig break, inflated counts, and missing rRNA operons. Run CheckM2 before trusting any annotation QC number.
| Tool | Maintained | Database approach | Output | When |
|---|---|---|---|---|
| Bakta | Yes (active) | Curated, versioned, alignment-free (UniRef + AMRFinderPlus + expert systems) | GFF3/GBFF/EMBL/FASTA/TSV/JSON + plot | Default for new work; reproducible, MAG-aware |
| Prokka | Frozen ~2021 | BLAST hierarchy vs frozen UniProt/RefSeq + HMM | GFF3/GBK/FAA + tbl | Legacy only; pangenome pipelines (Roary) expect Prokka GFF |
| NCBI PGAP | Yes (NCBI) | RefSeq protein-family models + ProSplign | ASN.1/GenBank, submission-ready | GenBank/RefSeq submission; best pseudogene/frameshift/selenoprotein handling |
| DFAST | Yes (DDBJ) | DFAST reference DBs + swappable refs | INSDC files | DDBJ submission; fast, flexible reference swap |
| RAST / BV-BRC | Yes | SEED subsystems | GenBank/GFF + subsystems | Subsystem/metabolic framing; web; not for direct INSDC submission |
Prokka uses Aragorn (tRNA) + barrnap (rRNA); Bakta uses tRNAscan-SE 2.0 (tRNA, domain-specific) + Infernal/Rfam (rRNA/ncRNA) + PILER-CR (CRISPR arrays) + DeepSig (signal peptides). Bakta is deliberately conservative on naming, so it reports "hypothetical" where Prokka confidently (sometimes wrongly) names a gene - Bakta looking "less annotated" is it being honest.
| Scenario | Recommended | Why |
|---|---|---|
| Routine isolate, reproducible | Bakta, full DB | standardized, versioned, current functional calls |
| Submitting to GenBank/RefSeq | PGAP | RefSeq re-annotates with PGAP regardless; best pseudogene handling |
| Submitting to DDBJ | DFAST | INSDC-ready, DDBJ-aligned |
| MAG / metagenome bin | Bakta --meta + CheckM2 gate | anonymous-mode calling; QC before trust |
| Mycoplasma / Mollicutes | Bakta --translation-table 4 | UGA=Trp; table 11 splits genes at every UGA |
| Archaeon | Bakta + verify tRNAscan-SE archaeal model; PGAP if N-termini matter | leaderless mRNAs degrade Prodigal TIS |
| Inherited Prokka pangenome pipeline | pin Prokka version, or re-call all in Bakta | tool consistency vs current biology |
| Subsystem/metabolic view | RAST / BV-BRC | SEED subsystem categories |
| Genome not yet assembled / poor QC | -> genome-assembly/assembly-qc | fix assembly before annotating |
| AMR for clinical report | -> run AMRFinderPlus/CARD-RGI directly | --organism context, point mutations, partials matter |
# Database (record the version): full ~30 GB for publishable annotation; light for triage/CI
bakta_db download --output db/ --type full
bakta --db db/ --output bakta_out --prefix ecoli_k12 \
--genus Escherichia --species coli --strain K-12 \
--locus-tag ECK12 --gram - --complete --threads 16 \
assembly.fastaKey flags: --translation-table {11,4,25} (default 11), --gram {+,-,?} (gates DeepSig signal-peptide calls; default ?), --complete (all sequences are finished replicons; enables oriC detection - do NOT use on draft contigs), --meta (metagenome/MAG mode), --compliant (enforce INSDC structure), --keep-contig-headers, --proteins <faa> (trusted-protein transfer). Set --genus/--species from a GTDB-Tk classification, not a guess. Primary outputs: .gff3, .gbff, .faa, .ffn, .fna, .tsv, plus .hypotheticals.tsv and .inference.tsv (open the inference column to ask why a product was assigned).
prokka --outdir prokka_out --prefix my_genome --locustag MYORG \
--genus Escherichia --species coli --cpus 8 --rfam assembly.fastaUse only for tool-chain consistency with an existing Prokka-based pangenome workflow, and pin the version. Its bundled databases are frozen: a gene family characterized after ~2019 is "hypothetical" in Prokka but named by current Bakta/PGAP, so the same gene flips core/accessory purely on DB vintage.
Goal: Load Bakta/Prokka GFF3 into a queryable database and compute coding density, the first sanity number.
Approach: Build a gffutils database, sum CDS lengths, divide by genome length; flag values outside the expected band.
import gffutils
CODING_DENSITY_LOW = 0.85 # <0.85 in a free-living bacterium suggests wrong table, fragmentation, or heavy pseudogenization
CODING_DENSITY_HIGH = 0.93 # >0.93 suggests ORF over-calling (spurious short hypotheticals)
def coding_density(gff_file, genome_length):
db = gffutils.create_db(gff_file, ':memory:', merge_strategy='merge')
coding_bp = sum(c.end - c.start + 1 for c in db.features_of_type('CDS'))
density = coding_bp / genome_length
if density < CODING_DENSITY_LOW or density > CODING_DENSITY_HIGH:
print(f'WARNING: coding density {density:.1%} outside expected 85-93%')
return density/pseudo) is the best automated arbiter.prfB/RF2 (a +1 frameshift at a slippery CTTT + internal SD), dnaX (−1), and IS-element transposases encode one protein across two frames; naive callers emit two short ORFs. PGAP recognizes a curated set.Trigger: comparing counts from Bakta vs Prokka vs old RefSeq, or mixed-vintage records. Mechanism: tool/DB differences dominate biological differences. Symptom: "novel genes" or accessory-genome inflation. Fix: re-annotate every genome from FASTA with one pipeline + pinned DB.
Trigger: table 11 on a Mollicute. Mechanism: UGA read as stop. Symptom: low coding density, doubled short gene count, high hypothetical fraction. Fix: --translation-table 4; confirm from GTDB-Tk.
Trigger: running Bakta before CheckM2. Mechanism: contig breaks truncate CDS; contamination mixes organisms. Symptom: partial CDS at ends, inflated/chimeric gene set, missing rRNA operons. Fix: CheckM2 first; contamination >5% -> decontaminate.
Trigger: reading the product column as ground truth. Mechanism: loose-homology transfer below ~40% identity is unreliable, especially for promiscuous folds. Symptom: a "histidine-kinase expansion" that is one over-called domain. Fix: open .inference.tsv; prefer InterPro/Pfam architecture over free-text product when they conflict.
Trigger: Prodigal single mode on a MAG/plasmid/phage. Mechanism: self-training needs a full single-organism genome. Symptom: poor calls on short/mixed input. Fix: Bakta --meta (anonymous mode).
Trigger: Bakta on macOS. Mechanism: DeepSig was dropped from the default mac conda env (~v1.9.4). Symptom: silently absent signal-peptide calls. Fix: verify the run log; run on Linux if SPs matter.
| Threshold | Source | Rationale |
|---|---|---|
| Coding density ~88-90% (band 85-93%) | bacterial genome norm | <85% = wrong table/fragmentation/decay; >93% = over-calling |
| ~850-1,000 genes/Mb (~1 gene/kb) | bacterial gene density | far higher = over-call; far lower = under-call/wrong table |
| Hypothetical 25-50% | annotation norm | ~0% = over-confident transfer; >60-70% = novel lineage or wrong tool/table |
| tRNA count ≥ ~20 (often 40-60) | one isoacceptor set minimum | far fewer = fragmented assembly broke tRNA regions |
| rRNA operons ~1-15 (e.g. ~7 in E. coli) | copy-number norm | zero/fractional in a "complete" genome = short-read repeat collapse |
| Prodigal minimum ~30 aa / 90 nt | caller default | shorter ORFs need a dedicated sORF caller + Ribo-seq |
| CheckM2 contamination ≤5%, completeness ≥90% | MIMAG-aligned | above/below -> annotation QC numbers uninterpretable, fix assembly first |
| Error / symptom | Cause | Solution |
|---|---|---|
| Genes split, low coding density | wrong translation table | --translation-table 4/25; confirm from GTDB-Tk |
| Many hypothetical proteins | novel organism / light DB | full Bakta DB; add InterProScan/eggNOG; normal if 25-50% |
| Low gene / tRNA count, missing rRNA | fragmented assembly | CheckM2; prefer long-read/complete assembly |
| RefSeq record differs from the local GFF | RefSeq is PGAP, GenBank keeps the submitter's annotation | report WP_ accession + locus_tag; expect divergence |
| Pseudogene spike | ONT homopolymer indels vs real decay vs phase variation | check homopolymer context; corroborate with short-read/HiFi |
| Submission rejected on locus tags | invented prefix | register the prefix via BioSample (3-12 alnum, starts with a letter) |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in genome-annotation/prokaryotic-annotation of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Genome Annotation Prokaryotic Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Genome Annotation Prokaryotic Annotation this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.2k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 | |
| Biopython Bioinformaticsaiming-lab/AutoResearchClaw | 15k | — | ~810 | Automated safety check: Pass | MIT | |
| ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates | 33k | 11 repos | ~4.5k | Automated safety check: Notes | MIT | |
| Biopythondavila7/claude-code-templates | 33k | 12 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Clinvar Databasedavila7/claude-code-templates | 33k | 10 repos | ~3.3k | Automated safety check: Pass | MIT |
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.
davila7/claude-code-templates
Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.
davila7/claude-code-templates
Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.
davila7/claude-code-templates
Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.
ClawBio/ClawBio
Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Works with
Categories
Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC…. Bio Genome Annotation Prokaryotic Annotation is an agent skill from GPTomics/bioSkills. Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC locus tags.
Bio Genome Annotation Prokaryotic Annotation fits situations like: annotating a newly assembled prokaryotic genome; choosing an annotation tool; re-annotating a collection for pangenomics; preparing annotations for NCBI/DDBJ submission.
Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a claude-code`. Or copy the skill folder (genome-annotation/prokaryotic-annotation in GPTomics/bioSkills) into .claude/skills/bio-genome-annotation-prokaryotic-annotation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a codex`. Or copy the skill folder (genome-annotation/prokaryotic-annotation in GPTomics/bioSkills) into .agents/skills/bio-genome-annotation-prokaryotic-annotation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-annotation-prokaryotic-annotation, .gemini/skills/bio-genome-annotation-prokaryotic-annotation, .github/skills/bio-genome-annotation-prokaryotic-annotation and .opencode/skills/bio-genome-annotation-prokaryotic-annotation in your project.
Going by SKILL.md and its folder, Bio Genome Annotation Prokaryotic Annotation needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Genome Annotation Prokaryotic Annotation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Genome Annotation Prokaryotic Annotation: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 33k stars) and Biopython (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.