Agent skill

Bio Genome Annotation Prokaryotic Annotation

by GPTomics in GPTomics/bioSkills

Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC…

MITAuto-check passedResearch & Science

Install Bio Genome Annotation Prokaryotic Annotation

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-annotation-prokaryotic-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-annotation/prokaryotic-annotation .claude/skills/bio-genome-annotation-prokaryotic-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-annotation-prokaryotic-annotation
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.2k tokens
SKILL.md length
1,848 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC…

  • Works in 3 steps: A confidently wrong product name is… → Comparing gene counts across… → Annotation completeness ≠ assembly…
  • Annotating a newly assembled prokaryotic genome
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Shell and Python scripts from its folder; calls pip

What it does

Bio Genome Annotation Prokaryotic Annotation is an agent skill from GPTomics/bioSkills. Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC locus tags. Covers Bakta-vs-Prokka-vs-PGAP-vs-DFAST choice, light-vs-full database tiers, translation-table selection (11/4/25), archaeal and leaderless-gene caveats, the small-ORF blind spot, pseudogene-vs-phase-variation, the pangenome re-annotation trap, and submission compliance. Use when annotating a newly assembled…

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/bakta_annotation.sh`, `examples/parse_annotation.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Annotating a newly assembled prokaryotic genome
  • Choosing an annotation tool
  • Re-annotating a collection for pangenomics
  • Preparing annotations for NCBI/DDBJ submission

Example prompts

  • “Use the bio-genome-annotation-prokaryotic-annotation skill to annotate bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active…”
  • “/bio-genome-annotation-prokaryotic-annotation”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A confidently wrong product name is worse than "hypothetical protein." Names assigned by loose homology transfer across distant lineages…
  2. Comparing gene counts across tools/versions/DBs is invalid. Tool agreement is "highly dependent on the organism of study" and biased…
  3. Annotation completeness ≠ assembly completeness. A fragmented/contaminated assembly produces truncated partial CDS at every contig break…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Annotation Prokaryotic Annotation loads about 4.2k tokens when it runs. Until then it costs about 177 tokens; SKILL.md has 1,848 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~177
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,848 words, ~4,205 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-annotation-prokaryotic-annotation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-genome-annotation-prokaryotic-annotation
description
Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC locus tags. Covers Bakta-vs-Prokka-vs-PGAP-vs-DFAST choice, light-vs-full database tiers, translation-table selection (11/4/25), archaeal and leaderless-gene caveats, the small-ORF blind spot, pseudogene-vs-phase-variation, the pangenome re-annotation trap, and submission compliance. Use when annotating a newly assembled prokaryotic genome, choosing an annotation tool, re-annotating a collection for pangenomics, or preparing annotations for NCBI/DDBJ submission.
tool_type
cli
primary_tool
Bakta

Version Compatibility

Reference examples tested with: Bakta 1.9+, Prokka 1.14.6 (legacy), tRNAscan-SE 2.0+, CheckM2 1.0+, gffutils 0.12+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures

Annotation content tracks the database version, not just the binary: a Bakta full DB (schema v5+, ~30 GB zipped) and a Bakta light DB give different functional calls, and two runs months apart can differ purely from DB updates. Record the Bakta DB version in methods. Prokka's bundled databases are frozen (~2019-2021). If code throws an error, introspect the installed tool and adapt rather than retrying.

Prokaryotic Genome Annotation

"Annotate my bacterial genome" -> Call protein-coding genes, tRNAs/rRNAs/ncRNAs, and other features, then assign function by database identity, and emit submission-ready files.

  • CLI: bakta --db db/ assembly.fa (default), prokka --outdir out assembly.fa (legacy), NCBI PGAP (submission/RefSeq-grade)

The Single Most Important Modern Insight -- Gene Calling Is Near-Solved; Function Is Not, and Wrong Labels Propagate

For a typical isolate, Prodigal/Pyrodigal recovers >95-99% of true coding genes. The unsolved problems are the start codon (translation initiation site), the small/overlapping/recoded ORFs, and above all the function. Three consequences a postdoc must internalize:

  1. A confidently wrong product name is worse than "hypothetical protein." Names assigned by loose homology transfer across distant lineages and self-amplify through databases - there is no mechanism to retract a correction once it spreads, so annotation accuracy has gone down, not up, as sequencing scaled (Salzberg 2019 Genome Biol 20:92). Bakta's design counters this: gene-level identity only via exact (MD5/UniRef100) match, falling back to cluster level, then "hypothetical" rather than guessing. 25-50% hypothetical is healthy for a non-model organism; near 0% means over-confident transfer.

  2. Comparing gene counts across tools/versions/DBs is invalid. Tool agreement is "highly dependent on the organism of study" and biased toward model organisms (Dimonaco 2022 Bioinformatics 38:1198) - there is no universal best tool. Differences between two annotations are mostly tool artifacts, not biology. For any collection (pangenomics), re-annotate every assembly from FASTA with one pipeline + one pinned DB; merging published annotations inflates the accessory genome ~10× (Tonkin-Hill 2020 Genome Biol 21:180).

  3. Annotation completeness ≠ assembly completeness. A fragmented/contaminated assembly produces truncated partial CDS at every contig break, inflated counts, and missing rRNA operons. Run CheckM2 before trusting any annotation QC number.

Tool Taxonomy

ToolMaintainedDatabase approachOutputWhen
BaktaYes (active)Curated, versioned, alignment-free (UniRef + AMRFinderPlus + expert systems)GFF3/GBFF/EMBL/FASTA/TSV/JSON + plotDefault for new work; reproducible, MAG-aware
ProkkaFrozen ~2021BLAST hierarchy vs frozen UniProt/RefSeq + HMMGFF3/GBK/FAA + tblLegacy only; pangenome pipelines (Roary) expect Prokka GFF
NCBI PGAPYes (NCBI)RefSeq protein-family models + ProSplignASN.1/GenBank, submission-readyGenBank/RefSeq submission; best pseudogene/frameshift/selenoprotein handling
DFASTYes (DDBJ)DFAST reference DBs + swappable refsINSDC filesDDBJ submission; fast, flexible reference swap
RAST / BV-BRCYesSEED subsystemsGenBank/GFF + subsystemsSubsystem/metabolic framing; web; not for direct INSDC submission

Prokka uses Aragorn (tRNA) + barrnap (rRNA); Bakta uses tRNAscan-SE 2.0 (tRNA, domain-specific) + Infernal/Rfam (rRNA/ncRNA) + PILER-CR (CRISPR arrays) + DeepSig (signal peptides). Bakta is deliberately conservative on naming, so it reports "hypothetical" where Prokka confidently (sometimes wrongly) names a gene - Bakta looking "less annotated" is it being honest.

Decision Tree by Scenario

ScenarioRecommendedWhy
Routine isolate, reproducibleBakta, full DBstandardized, versioned, current functional calls
Submitting to GenBank/RefSeqPGAPRefSeq re-annotates with PGAP regardless; best pseudogene handling
Submitting to DDBJDFASTINSDC-ready, DDBJ-aligned
MAG / metagenome binBakta --meta + CheckM2 gateanonymous-mode calling; QC before trust
Mycoplasma / MollicutesBakta --translation-table 4UGA=Trp; table 11 splits genes at every UGA
ArchaeonBakta + verify tRNAscan-SE archaeal model; PGAP if N-termini matterleaderless mRNAs degrade Prodigal TIS
Inherited Prokka pangenome pipelinepin Prokka version, or re-call all in Baktatool consistency vs current biology
Subsystem/metabolic viewRAST / BV-BRCSEED subsystem categories
Genome not yet assembled / poor QC-> genome-assembly/assembly-qcfix assembly before annotating
AMR for clinical report-> run AMRFinderPlus/CARD-RGI directly--organism context, point mutations, partials matter

Bakta (Default)

bash
# Database (record the version): full ~30 GB for publishable annotation; light for triage/CI
bakta_db download --output db/ --type full

bakta --db db/ --output bakta_out --prefix ecoli_k12 \
    --genus Escherichia --species coli --strain K-12 \
    --locus-tag ECK12 --gram - --complete --threads 16 \
    assembly.fasta

Key flags: --translation-table {11,4,25} (default 11), --gram {+,-,?} (gates DeepSig signal-peptide calls; default ?), --complete (all sequences are finished replicons; enables oriC detection - do NOT use on draft contigs), --meta (metagenome/MAG mode), --compliant (enforce INSDC structure), --keep-contig-headers, --proteins <faa> (trusted-protein transfer). Set --genus/--species from a GTDB-Tk classification, not a guess. Primary outputs: .gff3, .gbff, .faa, .ffn, .fna, .tsv, plus .hypotheticals.tsv and .inference.tsv (open the inference column to ask why a product was assigned).

Prokka (Legacy)

bash
prokka --outdir prokka_out --prefix my_genome --locustag MYORG \
    --genus Escherichia --species coli --cpus 8 --rfam assembly.fasta

Use only for tool-chain consistency with an existing Prokka-based pangenome workflow, and pin the version. Its bundled databases are frozen: a gene family characterized after ~2019 is "hypothetical" in Prokka but named by current Bakta/PGAP, so the same gene flips core/accessory purely on DB vintage.

Coding Density and CDS Extraction with Python

Goal: Load Bakta/Prokka GFF3 into a queryable database and compute coding density, the first sanity number.

Approach: Build a gffutils database, sum CDS lengths, divide by genome length; flag values outside the expected band.

python
import gffutils

CODING_DENSITY_LOW = 0.85   # <0.85 in a free-living bacterium suggests wrong table, fragmentation, or heavy pseudogenization
CODING_DENSITY_HIGH = 0.93  # >0.93 suggests ORF over-calling (spurious short hypotheticals)

def coding_density(gff_file, genome_length):
    db = gffutils.create_db(gff_file, ':memory:', merge_strategy='merge')
    coding_bp = sum(c.end - c.start + 1 for c in db.features_of_type('CDS'))
    density = coding_bp / genome_length
    if density < CODING_DENSITY_LOW or density > CODING_DENSITY_HIGH:
        print(f'WARNING: coding density {density:.1%} outside expected 85-93%')
    return density

Hard Cases the Caller Gets Wrong

  • Wrong genetic code is silent and looks like fragmentation. A Mycoplasma run under table 11 yields anomalously low coding density + many short "hypothetical" fragments because every internal UGA split a gene. Confirm the table from taxonomy (GTDB-Tk), never a guess. Table 4 (UGA=Trp, Mollicutes), table 25 (UGA=Gly, Gracilibacteria/SR1).
  • Pseudogene over-calling in reductive genomes is real biology. Mycobacterium leprae has ~1,116 pseudogenes vs ~1,604 intact CDS (Cole 2001 Nature 409:1007) - high pseudogene fraction in a symbiont/host-restricted pathogen (Rickettsia, Sodalis) is a lifestyle signal, not a defect. PGAP (ProSplign aligns through frameshifts, emits /pseudo) is the best automated arbiter.
  • Phase variation is not a pseudogene. A frameshift inside a homopolymer/SSR tract in a contingency locus (Neisseria ~65 candidate loci, Haemophilus, Campylobacter poly-G tracts) means the assembled cell was in the OFF phase - do NOT "correct" or polish it away. This confounds with ONT homopolymer indel error: a pseudogene spike is ambiguous between reductive biology, phase variation, and basecalling error; disambiguate with orthogonal (short-read/HiFi) data.
  • Programmed frameshifts make one gene look like two. prfB/RF2 (a +1 frameshift at a slippery CTTT + internal SD), dnaX (−1), and IS-element transposases encode one protein across two frames; naive callers emit two short ORFs. PGAP recognizes a curated set.
  • Small ORFs (<~50 aa) are systematically missed - callers impose a ~30 aa minimum because short ORFs arise by chance and coding statistics are unreliable there. E. coli K-12 was missing dozens of 16-50 aa proteins. Treat the gene count as a lower bound on the small proteome; for sORF biology use a dedicated caller (smORFer, smORFinder) and, ideally, Ribo-seq (the experimental arbiter of translation).
  • Archaea and Actinobacteria use leaderless mRNAs (no 5' UTR, no Shine-Dalgarno) - Prodigal's RBS-first scoring degrades TIS placement; GeneMarkS-2 (and PGAP) model leaderless transcription. Use tRNAscan-SE's archaeal model for archaeal intron-containing tRNAs.
Show full SKILL.md (735 more words)Show less

Per-Method Failure Modes

Cross-tool gene-count comparison

Trigger: comparing counts from Bakta vs Prokka vs old RefSeq, or mixed-vintage records. Mechanism: tool/DB differences dominate biological differences. Symptom: "novel genes" or accessory-genome inflation. Fix: re-annotate every genome from FASTA with one pipeline + pinned DB.

Wrong translation table

Trigger: table 11 on a Mollicute. Mechanism: UGA read as stop. Symptom: low coding density, doubled short gene count, high hypothetical fraction. Fix: --translation-table 4; confirm from GTDB-Tk.

Annotating an unvetted assembly

Trigger: running Bakta before CheckM2. Mechanism: contig breaks truncate CDS; contamination mixes organisms. Symptom: partial CDS at ends, inflated/chimeric gene set, missing rRNA operons. Fix: CheckM2 first; contamination >5% -> decontaminate.

Trusting an over-specific product name

Trigger: reading the product column as ground truth. Mechanism: loose-homology transfer below ~40% identity is unreliable, especially for promiscuous folds. Symptom: a "histidine-kinase expansion" that is one over-called domain. Fix: open .inference.tsv; prefer InterPro/Pfam architecture over free-text product when they conflict.

Single-mode calling on a metagenome or tiny replicon

Trigger: Prodigal single mode on a MAG/plasmid/phage. Mechanism: self-training needs a full single-organism genome. Symptom: poor calls on short/mixed input. Fix: Bakta --meta (anonymous mode).

Missing signal peptides on macOS

Trigger: Bakta on macOS. Mechanism: DeepSig was dropped from the default mac conda env (~v1.9.4). Symptom: silently absent signal-peptide calls. Fix: verify the run log; run on Linux if SPs matter.

Quantitative Thresholds

ThresholdSourceRationale
Coding density ~88-90% (band 85-93%)bacterial genome norm<85% = wrong table/fragmentation/decay; >93% = over-calling
~850-1,000 genes/Mb (~1 gene/kb)bacterial gene densityfar higher = over-call; far lower = under-call/wrong table
Hypothetical 25-50%annotation norm~0% = over-confident transfer; >60-70% = novel lineage or wrong tool/table
tRNA count ≥ ~20 (often 40-60)one isoacceptor set minimumfar fewer = fragmented assembly broke tRNA regions
rRNA operons ~1-15 (e.g. ~7 in E. coli)copy-number normzero/fractional in a "complete" genome = short-read repeat collapse
Prodigal minimum ~30 aa / 90 ntcaller defaultshorter ORFs need a dedicated sORF caller + Ribo-seq
CheckM2 contamination ≤5%, completeness ≥90%MIMAG-alignedabove/below -> annotation QC numbers uninterpretable, fix assembly first

Common Errors

Error / symptomCauseSolution
Genes split, low coding densitywrong translation table--translation-table 4/25; confirm from GTDB-Tk
Many hypothetical proteinsnovel organism / light DBfull Bakta DB; add InterProScan/eggNOG; normal if 25-50%
Low gene / tRNA count, missing rRNAfragmented assemblyCheckM2; prefer long-read/complete assembly
RefSeq record differs from the local GFFRefSeq is PGAP, GenBank keeps the submitter's annotationreport WP_ accession + locus_tag; expect divergence
Pseudogene spikeONT homopolymer indels vs real decay vs phase variationcheck homopolymer context; corroborate with short-read/HiFi
Submission rejected on locus tagsinvented prefixregister the prefix via BioSample (3-12 alnum, starts with a letter)

References

  • Schwengers O, et al. 2021. Bakta: rapid and standardized annotation of bacterial genomes via alignment-free sequence identification. Microb Genom 7:000685.
  • Seemann T. 2014. Prokka: rapid prokaryotic genome annotation. Bioinformatics 30:2068-2069.
  • Hyatt D, et al. 2010. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 11:119.
  • Larralde M. 2022. Pyrodigal: Python bindings and interface to Prodigal. J Open Source Softw 7:4296.
  • Lomsadze A, et al. 2018. Modeling leaderless transcription and atypical genes results in more accurate gene prediction in prokaryotes (GeneMarkS-2). Genome Res 28:1079-1089.
  • Tatusova T, et al. 2016. NCBI prokaryotic genome annotation pipeline. Nucleic Acids Res 44:6614-6624.
  • Li W, et al. 2021. RefSeq: expanding the Prokaryotic Genome Annotation Pipeline reach with protein family model curation. Nucleic Acids Res 49:D1020-D1028.
  • Tanizawa Y, et al. 2018. DFAST: a flexible prokaryotic genome annotation pipeline for faster genome publication. Bioinformatics 34:1037-1039.
  • Chan PP, et al. 2021. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res 49:9077-9096.
  • Chklovski A, et al. 2023. CheckM2: a rapid, scalable and accurate tool for assessing microbial genome quality using machine learning. Nat Methods 20:1203-1212.
  • Dimonaco NJ, et al. 2022. No one tool to rule them all: prokaryotic gene prediction tool annotations are highly dependent on the organism of study. Bioinformatics 38:1198-1207.
  • Tonkin-Hill G, et al. 2020. Producing polished prokaryotic pangenomes with the Panaroo pipeline. Genome Biol 21:180.
  • Salzberg SL. 2019. Next-generation genome annotation: we still struggle to get it right. Genome Biol 20:92.
  • Cole ST, et al. 2001. Massive gene decay in the leprosy bacillus. Nature 409:1007-1011.
  • functional-annotation - Add GO/KEGG/Pfam to hypothetical proteins with eggNOG-mapper/InterProScan
  • ncrna-annotation - Detailed ncRNA identification with Infernal/Rfam beyond the built-in callers
  • annotation-qc - CheckM2 completeness/contamination gate and gene-set sanity metrics
  • genome-assembly/assembly-qc - Assess assembly quality before annotation
  • genome-intervals/gtf-gff-handling - Parse and manipulate GFF3 output
  • comparative-genomics/pangenome-analysis - Uniform re-annotation before pangenome clustering

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in genome-annotation/prokaryotic-annotation of GPTomics/bioSkills.

  • SKILL.md
  • examples/bakta_annotation.sh
  • examples/parse_annotation.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Annotation Prokaryotic Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Annotation Prokaryotic Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Annotation Prokaryotic Annotation this skillGPTomics/bioSkills1.2k1 repos~4.2kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates33k11 repos~4.5kAutomated safety check: NotesMIT
Biopythondavila7/claude-code-templates33k12 repos~3.4kAutomated safety check: PassMIT
Clinvar Databasedavila7/claude-code-templates33k10 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • ETE Toolkit for Phylogenetic Trees

    davila7/claude-code-templates

    Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.

    33k GitHub starsUsed in 11 repos~4.5k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    davila7/claude-code-templates

    Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 10 repos~3.3k tokens
    Research & ScienceAuto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Genome Annotation Prokaryotic Annotation

What does Bio Genome Annotation Prokaryotic Annotation do?

Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC…. Bio Genome Annotation Prokaryotic Annotation is an agent skill from GPTomics/bioSkills. Annotates bacterial and archaeal genomes (isolates, MAGs, plasmids) with Bakta (active versioned databases, NCBI-compliant output) or Prokka (legacy), producing GFF3/GenBank/EMBL/FASTA with INSDC locus tags.

When should I use Bio Genome Annotation Prokaryotic Annotation?

Bio Genome Annotation Prokaryotic Annotation fits situations like: annotating a newly assembled prokaryotic genome; choosing an annotation tool; re-annotating a collection for pangenomics; preparing annotations for NCBI/DDBJ submission.

How do I install Bio Genome Annotation Prokaryotic Annotation in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a claude-code`. Or copy the skill folder (genome-annotation/prokaryotic-annotation in GPTomics/bioSkills) into .claude/skills/bio-genome-annotation-prokaryotic-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Annotation Prokaryotic Annotation in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a codex`. Or copy the skill folder (genome-annotation/prokaryotic-annotation in GPTomics/bioSkills) into .agents/skills/bio-genome-annotation-prokaryotic-annotation in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Annotation Prokaryotic Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-prokaryotic-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-annotation-prokaryotic-annotation, .gemini/skills/bio-genome-annotation-prokaryotic-annotation, .github/skills/bio-genome-annotation-prokaryotic-annotation and .opencode/skills/bio-genome-annotation-prokaryotic-annotation in your project.

What does Bio Genome Annotation Prokaryotic Annotation need to run?

Going by SKILL.md and its folder, Bio Genome Annotation Prokaryotic Annotation needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Genome Annotation Prokaryotic Annotation access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Genome Annotation Prokaryotic Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Annotation Prokaryotic Annotation use?

Bio Genome Annotation Prokaryotic Annotation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Annotation Prokaryotic Annotation use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Annotation Prokaryotic Annotation?

Skills that share tags, products or a category with Bio Genome Annotation Prokaryotic Annotation: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 33k stars) and Biopython (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Annotation Prokaryotic Annotation?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.