Agent skill

Bakta Genome Annotation

by jaechang-hits in jaechang-hits/SciAgent-Skills

Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline.

GPL-3.0Auto-check passedResearch & Science

Install Bakta Genome Annotation

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills bakta-genome-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/bakta-genome-annotation .claude/skills/bakta-genome-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bakta-genome-annotation
GitHub stars
374
Used in
1 other repo
Token cost
~5.9k tokens
SKILL.md length
1,195 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
GPL-3.0

At a glance

Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline.

  • Works in 8 steps: Install Bakta and Download the Database → Prepare the Input Assembly → Run Standard Bakta Annotation → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 7 more sections
  • Calls mamba, pip and python

What it does

Bakta Genome Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline. Identifies CDS, ncRNA, tRNA, rRNA, tmRNA, sORFs, CRISPR arrays, oriC/oriV/oriT, and gaps against a curated UniRef-derived database. Produces NCBI-compatible GFF3, GenBank, EMBL, JSON, FASTA, TSV, and a circular genome plot. Use Prokka for legacy pipelines or non-bacterial kingdoms; PGAP for NCBI GenBank submission.

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/bakta-genome-annotation”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Install Bakta and Download the Database
  2. Prepare the Input Assembly
  3. Run Standard Bakta Annotation
  4. Parse the JSON Summary
  5. Parse the TSV Feature Table
  6. Render the Circular Genome Plot
  7. Compute Annotation Quality Statistics
  8. Batch Annotation Across Multiple Genomes

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • mamba
    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • doi.org
    • biopython.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bakta Genome Annotation loads about 5.9k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 1,195 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its GPL-3.0 licence (© jaechang-hits). 1,195 words, ~5,871 tokens.

Download SKILL.mdSave it as .claude/skills/bakta-genome-annotation/SKILL.md (or your agent's skills folder).
name
bakta-genome-annotation
description
Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline. Identifies CDS, ncRNA, tRNA, rRNA, tmRNA, sORFs, CRISPR arrays, oriC/oriV/oriT, and gaps against a curated UniRef-derived database. Produces NCBI-compatible GFF3, GenBank, EMBL, JSON, FASTA, TSV, and a circular genome plot. Use Prokka for legacy pipelines or non-bacterial kingdoms; PGAP for NCBI GenBank submission.
license
GPL-3.0

Bakta Genome Annotation

Overview

Bakta is a command-line pipeline for rapid, standardized annotation of bacterial and archaeal genomes and plasmids. It combines Prodigal for CDS prediction, tRNAscan-SE/Aragorn/Barrnap/Infernal for non-coding RNA, PILER-CR/PILERCR for CRISPR detection, and a tiered DIAMOND/HMM search against a curated UniRef100 + IPS/UPS database to assign gene names, EC numbers, GO terms, and COG categories. Bakta produces NCBI-compatible outputs (GFF3, GenBank, EMBL, INSDC-formatted FASTA, plus a JSON summary and a circular Circos plot) for a typical 5 Mb genome in 5–15 minutes on 8 CPUs.

When to Use

  • Annotating bacterial or archaeal genome assemblies (Illumina, PacBio, Nanopore) with NCBI-compatible locus tags and product names
  • Annotating plasmids and other circular replicons separately with --plasmid and --complete flags
  • Producing JSON-structured annotation outputs that can be parsed without GenBank or GFF3 detours
  • Generating a publication-ready circular genome plot via the bundled bakta_plot command
  • Annotating MAGs (metagenome-assembled genomes) with --meta to disable Prodigal training
  • Use Prokka instead when you need viral/mitochondrial kingdoms or when you must reproduce a legacy Prokka pipeline exactly
  • Use PGAP instead when submitting to NCBI GenBank with full standards compliance
  • Use Bakta when you want faster runs, regularly updated UniRef-derived databases, AMRFinderPlus integration, and a JSON summary out of the box

Prerequisites

  • Software: Bakta ≥ 1.9, Python 3.8+, Prodigal, tRNAscan-SE, Aragorn, Barrnap, Infernal, DIAMOND, HMMER3, PILER-CR, BLAST+, AMRFinderPlus
  • Database: Bakta DB (full ~70 GB, or light ~3 GB) downloaded once with bakta_db download
  • Python packages (for output parsing): biopython, pandas, matplotlib
  • Input: assembled genome in FASTA format (one or more contigs)
  • Hardware: ≥ 16 GB RAM for full DB, ≥ 4 GB RAM for light DB; ≥ 8 CPUs recommended

Check before installing: The tool may already be available in the current environment (e.g., inside a pixi / conda env). Run command -v bakta first and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool via pixi run bakta rather than bare bakta.

bash
# Install Bakta via conda/mamba (recommended)
mamba install -c conda-forge -c bioconda bakta

# Verify installation
bakta --version
# bakta 1.9.4

# Download the light database (~3 GB, faster, fewer functional hits)
bakta_db download --output db/ --type light

# Or full database (~70 GB, comprehensive UniRef100 coverage)
# bakta_db download --output db/ --type full

# Install Python parsing dependencies
pip install biopython pandas matplotlib

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: database
    kind: required
    source: user
    ask: "Annotate against the light database or the full one? The light database leaves more genes described only generically."
    default: null

  - id: D2
    param: organismMetadata
    kind: required
    source: user
    ask: "Which genus, species, and strain should be recorded in the output?"
    default: null

  - id: D3
    param: repliconTopology
    kind: required
    source: data
    ask: "Are these contigs complete circular replicons, plasmids, or draft fragments?"
    default: "draft fragments - no circularity assumed"

  - id: D4
    param: assemblyOrigin
    kind: required
    source: data
    ask: "Is this an isolate genome, or a bin recovered from a metagenome?"
    default: "isolate - gene prediction trains on the assembly itself"

  - id: D5
    param: translationTable
    kind: required
    source: user
    ask: "Which genetic code does this organism use?"
    default: "11, the bacterial and archaeal code"

  - id: D6
    param: locusTagPrefix
    kind: optional
    source: user
    ask: "Which prefix should locus tags carry, if they will be submitted or cross-referenced?"
    default: "generated automatically"

  - id: D7
    param: minContigLength
    kind: optional_conditional
    source: data
    ask: "Should short contigs be skipped?"
    default: "no minimum"

  - id: D8
    param: threadsAndPlotting
    kind: never_ask
    source: data
    reason: "Thread count and the circular plot affect runtime and output files, not the annotation"
    default: "available cores, plot enabled"

D5 is asked rather than defaulted because the exceptions are real and silent: Mycoplasma and relatives reassign a stop codon, so annotating them with the standard bacterial code truncates a large fraction of their genes into fragments that still look like ordinary short CDS calls.

Quick Start

bash
# Annotate a bacterial genome — results in results/ directory
bakta genome.fasta \
    --db db/bakta_db_light \
    --output results/ \
    --prefix sample1 \
    --threads 8

# Inspect the JSON summary for feature counts
python -c "
import json
with open('results/sample1.json') as f:
    d = json.load(f)
print('Genus:', d['genome'].get('genus'))
print('Length:', d['genome']['size'], 'bp')
print('CDS:', sum(1 for f in d['features'] if f['type'] == 'cds'))
print('tRNA:', sum(1 for f in d['features'] if f['type'] == 'tRNA'))
"

Workflow

Step 1: Install Bakta and Download the Database

Install Bakta and prepare the reference database. The database download is one-time and reused across runs.

bash
# Create a dedicated conda environment (avoids dependency conflicts)
mamba create -n bakta_env -c conda-forge -c bioconda bakta python=3.11 -y
mamba activate bakta_env

# Verify Bakta and its dependencies
bakta --version
# bakta 1.9.4

bakta --help | head -20

# Download the light database (sufficient for routine annotation)
mkdir -p db/
bakta_db download --output db/ --type light
# Downloads ~3 GB; expands to ~5 GB on disk

# Verify the database was extracted correctly
ls db/bakta_db_light/
# antifam.h3f  bakta.db  expert  oric.fna  pfam.h3f  rfam-go.tsv  ...

# (Optional) Update AMRFinderPlus DB used by Bakta for AMR gene calling
amrfinder -u

# Install Python parsing tools
pip install biopython pandas matplotlib
Step 2: Prepare the Input Assembly

Bakta requires clean FASTA headers without spaces or special characters. Pre-clean and optionally filter short contigs.

python
from Bio import SeqIO
import re

input_fasta = "genome.fasta"
records = list(SeqIO.parse(input_fasta, "fasta"))
print(f"Input assembly: {len(records)} contigs")
total_bases = sum(len(r) for r in records)
print(f"Total bases: {total_bases:,}")
print(f"Largest contig: {max(len(r) for r in records):,} bp")

# Bakta preferred: short, alphanumeric, unique IDs
cleaned = []
for i, rec in enumerate(records, 1):
    new_id = f"contig_{i:04d}"
    new_rec = rec.__class__(rec.seq, id=new_id, description="")
    cleaned.append(new_rec)

SeqIO.write(cleaned, "genome_clean.fasta", "fasta")
print(f"Wrote genome_clean.fasta with {len(cleaned)} contigs")
bash
# Filter out short contigs (<200 bp) which contribute little to annotation
awk 'BEGIN{RS=">"; ORS=""} NR>1 {n=split($0, a, "\n"); seq=""; for(i=2;i<=n;i++) seq=seq a[i]; if (length(seq) >= 200) print ">" $0}' \
    genome_clean.fasta > genome_filtered.fasta

echo "Filtered assembly: $(grep -c '>' genome_filtered.fasta) contigs"
Step 3: Run Standard Bakta Annotation

Run Bakta with genus/species hints. Locus tags are auto-generated from the strain field.

bash
# Standard annotation for a draft bacterial genome
bakta genome_clean.fasta \
    --db db/bakta_db_light \
    --output annotation/ \
    --prefix E_coli_K12 \
    --genus Escherichia \
    --species coli \
    --strain K12 \
    --locus-tag ECOLI \
    --threads 8 \
    --keep-contig-headers

# Expected runtime: 5–15 min for ~5 Mb genome on 8 CPUs (light DB)

echo "Bakta annotation outputs:"
ls annotation/
# E_coli_K12.embl   E_coli_K12.faa     E_coli_K12.ffn
# E_coli_K12.fna    E_coli_K12.gbff    E_coli_K12.gff3
# E_coli_K12.hypotheticals.faa  E_coli_K12.hypotheticals.tsv
# E_coli_K12.json   E_coli_K12.log     E_coli_K12.png
# E_coli_K12.svg    E_coli_K12.tsv     E_coli_K12.txt
Step 4: Parse the JSON Summary

Bakta's JSON output is the canonical, machine-readable annotation. Parse it directly for downstream pipelines.

python
import json
import pandas as pd
from collections import Counter

with open("annotation/E_coli_K12.json") as f:
    bakta = json.load(f)

# Genome-level metadata
genome = bakta["genome"]
print(f"Organism: {genome.get('genus')} {genome.get('species')} {genome.get('strain')}")
print(f"Size: {genome['size']:,} bp across {len(bakta['sequences'])} sequences")
print(f"GC content: {genome['gc']:.2%}")

# Feature type counts
features = bakta["features"]
type_counts = Counter(f["type"] for f in features)
print("\nFeature counts:")
for ftype, n in sorted(type_counts.items(), key=lambda x: -x[1]):
    print(f"  {ftype:>10}: {n}")

# Build a tidy CDS DataFrame
cds_rows = []
for f in features:
    if f["type"] != "cds":
        continue
    cds_rows.append({
        "locus_tag": f.get("locus", ""),
        "contig":    f.get("contig", ""),
        "start":     f.get("start"),
        "stop":      f.get("stop"),
        "strand":    f.get("strand"),
        "gene":      f.get("gene", ""),
        "product":   f.get("product", ""),
        "length_aa": len(f.get("aa", "")),
    })

cds_df = pd.DataFrame(cds_rows)
print(f"\nTotal CDS: {len(cds_df)}")
print(cds_df.head(5).to_string(index=False))
Step 5: Parse the TSV Feature Table

The TSV output is convenient for spreadsheet workflows and quick filtering.

python
import pandas as pd

# Bakta TSV begins with comment lines starting with '#'
df = pd.read_csv("annotation/E_coli_K12.tsv", sep="\t", comment="#",
                 names=["sequence_id", "type", "start", "stop", "strand",
                        "locus_tag", "gene", "product", "dbxrefs"])
print(f"Total features: {len(df)}")
print(f"Feature types: {df['type'].value_counts().to_dict()}")

# Hypothetical vs annotated CDS
cds = df[df["type"] == "cds"].copy()
hypothetical = cds["product"].str.contains("hypothetical", case=False, na=True)
print(f"\nCDS with assigned function: {(~hypothetical).sum()} / {len(cds)}")
print(f"Hypothetical proteins: {hypothetical.sum()}")

# Cross-references (UniRef, KEGG, EC, GO, etc.) parsed from the dbxrefs column
def split_xrefs(xref_str):
    if not isinstance(xref_str, str) or xref_str in ("", "-"):
        return []
    return [x.strip() for x in xref_str.split(",")]

cds["dbxref_list"] = cds["dbxrefs"].apply(split_xrefs)
ec_hits = cds[cds["dbxref_list"].apply(lambda xs: any(x.startswith("EC:") for x in xs))]
print(f"CDS with EC numbers: {len(ec_hits)}")
print(ec_hits[["locus_tag", "gene", "product"]].head(5).to_string(index=False))
Step 6: Render the Circular Genome Plot

Bakta emits a Circos-style PNG/SVG by default. Regenerate with custom styling using bakta_plot.

bash
# Re-render the plot from the existing JSON with a different style
bakta_plot --output annotation/ \
           --prefix E_coli_K12_replot \
           --type cog \
           --dpi 300 \
           annotation/E_coli_K12.json

ls annotation/E_coli_K12_replot.*
# E_coli_K12_replot.png  E_coli_K12_replot.svg

# Display the rendered PNG inline in a Jupyter notebook
python <<'PY'
from pathlib import Path
import matplotlib.pyplot as plt
import matplotlib.image as mpimg

img = mpimg.imread("annotation/E_coli_K12.png")
fig, ax = plt.subplots(figsize=(6, 6))
ax.imshow(img)
ax.axis("off")
ax.set_title("Bakta circular genome annotation")
plt.savefig("bakta_circular_thumbnail.png", dpi=120, bbox_inches="tight")
print("Saved bakta_circular_thumbnail.png")
PY
Step 7: Compute Annotation Quality Statistics

Summarize hypothetical-protein rate, gene density, and feature coverage to assess annotation quality.

python
import json
import pandas as pd
import matplotlib.pyplot as plt

with open("annotation/E_coli_K12.json") as f:
    bakta = json.load(f)

genome_size = bakta["genome"]["size"]
features = bakta["features"]

# Per-feature-type counts and density (per Mb)
counts = {}
for f in features:
    counts[f["type"]] = counts.get(f["type"], 0) + 1
density = {k: v / (genome_size / 1e6) for k, v in counts.items()}

cds = [f for f in features if f["type"] == "cds"]
n_cds = len(cds)
n_hypo = sum(1 for f in cds if "hypothetical" in (f.get("product") or "").lower())
n_known = n_cds - n_hypo
coding_density = sum(abs(f["stop"] - f["start"] + 1) for f in cds) / genome_size

print(f"Genome: {genome_size:,} bp")
print(f"CDS: {n_cds}  ({density.get('cds', 0):.1f} per Mb)")
print(f"  Known function: {n_known}  Hypothetical: {n_hypo}")
print(f"Coding density: {coding_density:.1%}")
print(f"tRNA: {counts.get('tRNA', 0)}   rRNA: {counts.get('rRNA', 0)}")
print(f"ncRNA: {counts.get('ncRNA', 0)}  CRISPR: {counts.get('crispr', 0)}")

# Bar plot of feature type counts
fig, ax = plt.subplots(figsize=(8, 4))
items = sorted(counts.items(), key=lambda x: -x[1])
labels = [k for k, _ in items]
values = [v for _, v in items]
bars = ax.bar(labels, values, color="#2980B9", edgecolor="white")
for bar, v in zip(bars, values):
    ax.text(bar.get_x() + bar.get_width() / 2, v + max(values) * 0.01,
            str(v), ha="center", fontsize=9)
ax.set_ylabel("Count")
ax.set_title("Bakta feature counts by type")
plt.xticks(rotation=20)
plt.tight_layout()
plt.savefig("bakta_feature_counts.png", dpi=150, bbox_inches="tight")
print("Saved bakta_feature_counts.png")
Step 8: Batch Annotation Across Multiple Genomes

Run Bakta over a directory of assemblies and aggregate per-sample summary statistics.

bash
#!/bin/bash
# batch_bakta.sh — annotate all FASTA files in a directory
INPUT_DIR="genomes/"
OUTPUT_DIR="annotations/"
DB="db/bakta_db_light"
mkdir -p "$OUTPUT_DIR"

for FASTA in "$INPUT_DIR"/*.fasta; do
    SAMPLE=$(basename "$FASTA" .fasta)
    echo "Annotating: $SAMPLE"
    bakta "$FASTA" \
        --db "$DB" \
        --output "${OUTPUT_DIR}/${SAMPLE}" \
        --prefix "$SAMPLE" \
        --threads 4 \
        --skip-plot \
        --force
done

echo "Batch annotation complete."
python
# Aggregate per-genome summaries from each sample's JSON file
from pathlib import Path
import json
import pandas as pd

annotation_dir = Path("annotations/")
rows = []
for json_file in sorted(annotation_dir.glob("*/*.json")):
    with open(json_file) as f:
        bakta = json.load(f)
    sample = json_file.stem
    counts = {}
    for feat in bakta["features"]:
        counts[feat["type"]] = counts.get(feat["type"], 0) + 1
    rows.append({
        "sample":  sample,
        "size":    bakta["genome"]["size"],
        "gc":      bakta["genome"]["gc"],
        "CDS":     counts.get("cds", 0),
        "tRNA":    counts.get("tRNA", 0),
        "rRNA":    counts.get("rRNA", 0),
        "ncRNA":   counts.get("ncRNA", 0),
        "crispr":  counts.get("crispr", 0),
    })

summary_df = pd.DataFrame(rows)
print(f"Annotated genomes: {len(summary_df)}")
print(summary_df.to_string(index=False))
summary_df.to_csv("batch_bakta_summary.csv", index=False)
print("Saved: batch_bakta_summary.csv")

Key Parameters

ParameterDefaultRange / OptionsEffect
--db—path to Bakta DB directoryRequired; selects light or full reference DB
--genus—any genus name stringSets organism metadata in GenBank/EMBL output
--species—any species name stringCombined with --genus for organism qualifier
--strain—any stringAdds strain qualifier to organism metadata
--locus-tagauto3–24 alphanumeric charsPrefix for locus tags (e.g., ECOLI_00001)
--completeoffflagTreat all input contigs as complete circular replicons
--plasmidoffflagTreat all input contigs as plasmids (circular)
--metaoffflagDisable Prodigal training (use for MAGs / metagenomic bins)
--translation-table11NCBI table IDs (1, 4, 11, 25, …)Genetic code used by Prodigal for CDS translation
--min-contig-length1any integer (bp)Skip contigs shorter than this length
--threads11–CPU countParallel threads for DIAMOND/HMMER searches
--skip-plotoffflagSkip the slow circular plot rendering step
--keep-contig-headersoffflagPreserve original contig IDs instead of renaming
--proteins—path to GenBank or FASTA fileCustom expert protein DB used before UniRef search

Common Recipes

Recipe: Plasmid-only Annotation

When to use: A finished circular plasmid sequence that should be annotated without chromosome assumptions.

bash
bakta plasmid.fasta \
    --db db/bakta_db_light \
    --output plasmid_annotation/ \
    --prefix pBR322 \
    --plasmid \
    --complete \
    --threads 4

# Inspect plasmid-typing results (incompatibility group, replication initiator)
grep -E "rep|inc" plasmid_annotation/pBR322.tsv | head -10
Show full SKILL.md (470 more words)Show less
Recipe: MAG (Metagenome-Assembled Genome) Annotation

When to use: Annotating a bin from metagenomic assembly where Prodigal cannot train on the full sequence.

bash
bakta MAG_bin_42.fasta \
    --db db/bakta_db_light \
    --output mag_annotation/ \
    --prefix MAG_bin_42 \
    --meta \
    --min-contig-length 500 \
    --skip-plot \
    --threads 8

cat mag_annotation/MAG_bin_42.txt | head -20
Recipe: Annotation with Custom Expert Protein Database

When to use: You have curated reference proteins (e.g., a virulence factor catalog) that should take priority over UniRef.

bash
# Custom proteins are searched before the bundled UniRef DB
bakta genome.fasta \
    --db db/bakta_db_light \
    --proteins virulence_factors.faa \
    --output annotation_custom/ \
    --prefix novel_strain \
    --threads 8

echo "Annotations referencing custom DB:"
grep "User-provided" annotation_custom/novel_strain.tsv | head
Recipe: Convert Bakta GFF3 to a Pandas DataFrame

When to use: You need feature coordinates and qualifiers for downstream coordinate-overlap analyses.

python
import pandas as pd

def parse_gff3(gff_file):
    rows = []
    with open(gff_file) as f:
        for line in f:
            if line.startswith("##FASTA"):
                break
            if line.startswith("#") or not line.strip():
                continue
            parts = line.rstrip("\n").split("\t")
            if len(parts) < 9:
                continue
            attrs = {}
            for item in parts[8].split(";"):
                if "=" in item:
                    k, v = item.split("=", 1)
                    attrs[k] = v
            rows.append({
                "seqname":   parts[0],
                "feature":   parts[2],
                "start":     int(parts[3]),
                "end":       int(parts[4]),
                "strand":    parts[6],
                "locus_tag": attrs.get("locus_tag", attrs.get("ID", "")),
                "gene":      attrs.get("gene", ""),
                "product":   attrs.get("product", ""),
            })
    return pd.DataFrame(rows)

gff_df = parse_gff3("annotation/E_coli_K12.gff3")
print(f"Features parsed: {len(gff_df)}")
print(gff_df[gff_df["feature"] == "CDS"].head(5).to_string(index=False))

Expected Outputs

Output FileFormatDescription
{prefix}.gff3GFF3Genome annotation with feature coordinates and attributes; INSDC-compliant
{prefix}.gbffGenBank Flat FileAnnotated GenBank record for use in Geneious, BioPython, NCBI submission prep
{prefix}.emblEMBLEMBL-format annotation for ENA submission
{prefix}.fnaFASTANucleotide sequences of the input contigs
{prefix}.ffnFASTANucleotide sequences of all annotated features
{prefix}.faaFASTAProtein sequences for all CDS features
{prefix}.hypotheticals.faaFASTAProtein sequences flagged as hypothetical (for further investigation)
{prefix}.hypotheticals.tsvTSVDetailed table for hypothetical CDS with low-confidence hits
{prefix}.tsvTSVFull feature table: seqid, type, start, stop, strand, locus_tag, gene, product, dbxrefs
{prefix}.jsonJSONMachine-readable summary with genome metadata + every feature with full qualifiers
{prefix}.png / .svgImageCircos circular genome plot colored by COG category
{prefix}.txtTextPlain-text summary of feature counts and per-replicon stats
{prefix}.logTextFull Bakta run log including timing per step

Troubleshooting

ProblemCauseSolution
Error: database not found at <path>DB path incorrect, or DB not extractedRe-run bakta_db download --output db/ --type light; pass full path to --db
KILLED during DIAMOND stepOut of memory on full DBSwitch to --type light, reduce --threads, or run on a host with ≥ 32 GB RAM
Very high hypothetical-protein rate (>60 %)Divergent strain or light DB without close hitsRe-run with the full DB or supply a curated --proteins reference set
tRNAscan-SE: Error opening fileMissing tRNAscan-SE installation in conda envmamba install -c bioconda trnascan-se and rerun
Bakta complains about contig headersSpaces or special characters in FASTA IDsPre-clean headers with awk '/^>/{print ">contig_"++i; next}{print}'
Plot generation hangs or failsCircos / matplotlib backend issueRe-run with --skip-plot; render later via bakta_plot from the JSON file
AMRFinderPlus: database not up-to-date warningAMRFinderPlus DB staleamrfinder -u to refresh, or pass --skip-amr to bypass
Different locus tag prefix than expected--locus-tag not set, default uses random prefixExplicitly pass --locus-tag MYORG to control prefix
Run is much slower than reportedDefault --threads 1Set --threads to physical core count; use --skip-plot for batch jobs

References

© jaechang-hits, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/annotation/bakta-genome-annotation of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bakta Genome Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bakta Genome Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bakta Genome Annotation this skilljaechang-hits/SciAgent-Skills3741 repos~5.9kAutomated safety check: PassGPL-3.0
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Bio Write SequencesGPTomics/bioSkills1.2k3 repos~2.1kAutomated safety check: PassMIT
ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates33k11 repos~4.5kAutomated safety check: NotesMIT
Biopythondavila7/claude-code-templates33k12 repos~3.4kAutomated safety check: PassMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • ETE Toolkit for Phylogenetic Trees

    davila7/claude-code-templates

    Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.

    33k GitHub starsUsed in 11 repos~4.5k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    davila7/claude-code-templates

    Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 10 repos~3.3k tokens
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Works with

Questions about Bakta Genome Annotation

What does Bakta Genome Annotation do?

Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline. Bakta Genome Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline.

When should I use Bakta Genome Annotation?

Bakta Genome Annotation fits situations like: tasks that involve Bioinformatics.

How do I install Bakta Genome Annotation in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/bakta-genome-annotation in jaechang-hits/SciAgent-Skills) into .claude/skills/bakta-genome-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Bakta Genome Annotation in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/bakta-genome-annotation in jaechang-hits/SciAgent-Skills) into .agents/skills/bakta-genome-annotation in your project. Codex loads it when a task matches its description.

Can I use Bakta Genome Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bakta-genome-annotation, .gemini/skills/bakta-genome-annotation, .github/skills/bakta-genome-annotation and .opencode/skills/bakta-genome-annotation in your project.

What does Bakta Genome Annotation need to run?

Going by SKILL.md and its folder, Bakta Genome Annotation needs the command-line tools its instructions call (mamba, pip and python). Our summary lists: Python 3.

Does Bakta Genome Annotation access the network?

SKILL.md names 3 domains. As links in the text: github.com, doi.org and biopython.org. This is read from the text; nothing was executed.

Is Bakta Genome Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bakta Genome Annotation use?

Bakta Genome Annotation is published under the GPL-3.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bakta Genome Annotation use?

About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bakta Genome Annotation?

Skills that share tags, products or a category with Bakta Genome Annotation: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars) and ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bakta Genome Annotation?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.