Dbsnp Database
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bakta-genome-annotation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/bakta-genome-annotation .claude/skills/bakta-genome-annotation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bakta-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/bakta-genome-annotation into .claude/skills/bakta-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bakta-genome-annotation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/bakta-genome-annotationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bakta-genome-annotation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/bakta-genome-annotation .agents/skills/bakta-genome-annotation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bakta-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/bakta-genome-annotation into .agents/skills/bakta-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bakta-genome-annotation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bakta-genome-annotation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/bakta-genome-annotation .cursor/skills/bakta-genome-annotation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bakta-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/bakta-genome-annotation into .cursor/skills/bakta-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bakta-genome-annotation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/annotation/bakta-genome-annotation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bakta-genome-annotation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/bakta-genome-annotation .gemini/skills/bakta-genome-annotation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bakta-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/bakta-genome-annotation into .gemini/skills/bakta-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bakta-genome-annotation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills bakta-genome-annotationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/bakta-genome-annotation .github/skills/bakta-genome-annotation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bakta-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/bakta-genome-annotation into .github/skills/bakta-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bakta-genome-annotation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bakta-genome-annotation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/bakta-genome-annotation .opencode/skills/bakta-genome-annotation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bakta-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/bakta-genome-annotation into .opencode/skills/bakta-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bakta-genome-annotation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bakta-genome-annotationAnnotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline.
Bakta Genome Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline. Identifies CDS, ncRNA, tRNA, rRNA, tmRNA, sORFs, CRISPR arrays, oriC/oriV/oriT, and gaps against a curated UniRef-derived database. Produces NCBI-compatible GFF3, GenBank, EMBL, JSON, FASTA, TSV, and a circular genome plot. Use Prokka for legacy pipelines or non-bacterial kingdoms; PGAP for NCBI GenBank submission.
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is GPL-3.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
mambapippythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comdoi.orgbiopython.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bakta Genome Annotation loads about 5.9k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 1,195 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its GPL-3.0 licence (© jaechang-hits). 1,195 words, ~5,871 tokens.
.claude/skills/bakta-genome-annotation/SKILL.md (or your agent's skills folder).Bakta is a command-line pipeline for rapid, standardized annotation of bacterial and archaeal genomes and plasmids. It combines Prodigal for CDS prediction, tRNAscan-SE/Aragorn/Barrnap/Infernal for non-coding RNA, PILER-CR/PILERCR for CRISPR detection, and a tiered DIAMOND/HMM search against a curated UniRef100 + IPS/UPS database to assign gene names, EC numbers, GO terms, and COG categories. Bakta produces NCBI-compatible outputs (GFF3, GenBank, EMBL, INSDC-formatted FASTA, plus a JSON summary and a circular Circos plot) for a typical 5 Mb genome in 5–15 minutes on 8 CPUs.
--plasmid and --complete flagsbakta_plot command--meta to disable Prodigal trainingbakta_db downloadbiopython, pandas, matplotlibCheck before installing: The tool may already be available in the current environment (e.g., inside a
pixi/condaenv). Runcommand -v baktafirst and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool viapixi run baktarather than barebakta.
# Install Bakta via conda/mamba (recommended)
mamba install -c conda-forge -c bioconda bakta
# Verify installation
bakta --version
# bakta 1.9.4
# Download the light database (~3 GB, faster, fewer functional hits)
bakta_db download --output db/ --type light
# Or full database (~70 GB, comprehensive UniRef100 coverage)
# bakta_db download --output db/ --type full
# Install Python parsing dependencies
pip install biopython pandas matplotlibSettle these with the user before writing any analysis code.
decisions:
- id: D1
param: database
kind: required
source: user
ask: "Annotate against the light database or the full one? The light database leaves more genes described only generically."
default: null
- id: D2
param: organismMetadata
kind: required
source: user
ask: "Which genus, species, and strain should be recorded in the output?"
default: null
- id: D3
param: repliconTopology
kind: required
source: data
ask: "Are these contigs complete circular replicons, plasmids, or draft fragments?"
default: "draft fragments - no circularity assumed"
- id: D4
param: assemblyOrigin
kind: required
source: data
ask: "Is this an isolate genome, or a bin recovered from a metagenome?"
default: "isolate - gene prediction trains on the assembly itself"
- id: D5
param: translationTable
kind: required
source: user
ask: "Which genetic code does this organism use?"
default: "11, the bacterial and archaeal code"
- id: D6
param: locusTagPrefix
kind: optional
source: user
ask: "Which prefix should locus tags carry, if they will be submitted or cross-referenced?"
default: "generated automatically"
- id: D7
param: minContigLength
kind: optional_conditional
source: data
ask: "Should short contigs be skipped?"
default: "no minimum"
- id: D8
param: threadsAndPlotting
kind: never_ask
source: data
reason: "Thread count and the circular plot affect runtime and output files, not the annotation"
default: "available cores, plot enabled"D5 is asked rather than defaulted because the exceptions are real and silent: Mycoplasma and relatives reassign a stop codon, so annotating them with the standard bacterial code truncates a large fraction of their genes into fragments that still look like ordinary short CDS calls.
# Annotate a bacterial genome — results in results/ directory
bakta genome.fasta \
--db db/bakta_db_light \
--output results/ \
--prefix sample1 \
--threads 8
# Inspect the JSON summary for feature counts
python -c "
import json
with open('results/sample1.json') as f:
d = json.load(f)
print('Genus:', d['genome'].get('genus'))
print('Length:', d['genome']['size'], 'bp')
print('CDS:', sum(1 for f in d['features'] if f['type'] == 'cds'))
print('tRNA:', sum(1 for f in d['features'] if f['type'] == 'tRNA'))
"Install Bakta and prepare the reference database. The database download is one-time and reused across runs.
# Create a dedicated conda environment (avoids dependency conflicts)
mamba create -n bakta_env -c conda-forge -c bioconda bakta python=3.11 -y
mamba activate bakta_env
# Verify Bakta and its dependencies
bakta --version
# bakta 1.9.4
bakta --help | head -20
# Download the light database (sufficient for routine annotation)
mkdir -p db/
bakta_db download --output db/ --type light
# Downloads ~3 GB; expands to ~5 GB on disk
# Verify the database was extracted correctly
ls db/bakta_db_light/
# antifam.h3f bakta.db expert oric.fna pfam.h3f rfam-go.tsv ...
# (Optional) Update AMRFinderPlus DB used by Bakta for AMR gene calling
amrfinder -u
# Install Python parsing tools
pip install biopython pandas matplotlibBakta requires clean FASTA headers without spaces or special characters. Pre-clean and optionally filter short contigs.
from Bio import SeqIO
import re
input_fasta = "genome.fasta"
records = list(SeqIO.parse(input_fasta, "fasta"))
print(f"Input assembly: {len(records)} contigs")
total_bases = sum(len(r) for r in records)
print(f"Total bases: {total_bases:,}")
print(f"Largest contig: {max(len(r) for r in records):,} bp")
# Bakta preferred: short, alphanumeric, unique IDs
cleaned = []
for i, rec in enumerate(records, 1):
new_id = f"contig_{i:04d}"
new_rec = rec.__class__(rec.seq, id=new_id, description="")
cleaned.append(new_rec)
SeqIO.write(cleaned, "genome_clean.fasta", "fasta")
print(f"Wrote genome_clean.fasta with {len(cleaned)} contigs")# Filter out short contigs (<200 bp) which contribute little to annotation
awk 'BEGIN{RS=">"; ORS=""} NR>1 {n=split($0, a, "\n"); seq=""; for(i=2;i<=n;i++) seq=seq a[i]; if (length(seq) >= 200) print ">" $0}' \
genome_clean.fasta > genome_filtered.fasta
echo "Filtered assembly: $(grep -c '>' genome_filtered.fasta) contigs"Run Bakta with genus/species hints. Locus tags are auto-generated from the strain field.
# Standard annotation for a draft bacterial genome
bakta genome_clean.fasta \
--db db/bakta_db_light \
--output annotation/ \
--prefix E_coli_K12 \
--genus Escherichia \
--species coli \
--strain K12 \
--locus-tag ECOLI \
--threads 8 \
--keep-contig-headers
# Expected runtime: 5–15 min for ~5 Mb genome on 8 CPUs (light DB)
echo "Bakta annotation outputs:"
ls annotation/
# E_coli_K12.embl E_coli_K12.faa E_coli_K12.ffn
# E_coli_K12.fna E_coli_K12.gbff E_coli_K12.gff3
# E_coli_K12.hypotheticals.faa E_coli_K12.hypotheticals.tsv
# E_coli_K12.json E_coli_K12.log E_coli_K12.png
# E_coli_K12.svg E_coli_K12.tsv E_coli_K12.txtBakta's JSON output is the canonical, machine-readable annotation. Parse it directly for downstream pipelines.
import json
import pandas as pd
from collections import Counter
with open("annotation/E_coli_K12.json") as f:
bakta = json.load(f)
# Genome-level metadata
genome = bakta["genome"]
print(f"Organism: {genome.get('genus')} {genome.get('species')} {genome.get('strain')}")
print(f"Size: {genome['size']:,} bp across {len(bakta['sequences'])} sequences")
print(f"GC content: {genome['gc']:.2%}")
# Feature type counts
features = bakta["features"]
type_counts = Counter(f["type"] for f in features)
print("\nFeature counts:")
for ftype, n in sorted(type_counts.items(), key=lambda x: -x[1]):
print(f" {ftype:>10}: {n}")
# Build a tidy CDS DataFrame
cds_rows = []
for f in features:
if f["type"] != "cds":
continue
cds_rows.append({
"locus_tag": f.get("locus", ""),
"contig": f.get("contig", ""),
"start": f.get("start"),
"stop": f.get("stop"),
"strand": f.get("strand"),
"gene": f.get("gene", ""),
"product": f.get("product", ""),
"length_aa": len(f.get("aa", "")),
})
cds_df = pd.DataFrame(cds_rows)
print(f"\nTotal CDS: {len(cds_df)}")
print(cds_df.head(5).to_string(index=False))The TSV output is convenient for spreadsheet workflows and quick filtering.
import pandas as pd
# Bakta TSV begins with comment lines starting with '#'
df = pd.read_csv("annotation/E_coli_K12.tsv", sep="\t", comment="#",
names=["sequence_id", "type", "start", "stop", "strand",
"locus_tag", "gene", "product", "dbxrefs"])
print(f"Total features: {len(df)}")
print(f"Feature types: {df['type'].value_counts().to_dict()}")
# Hypothetical vs annotated CDS
cds = df[df["type"] == "cds"].copy()
hypothetical = cds["product"].str.contains("hypothetical", case=False, na=True)
print(f"\nCDS with assigned function: {(~hypothetical).sum()} / {len(cds)}")
print(f"Hypothetical proteins: {hypothetical.sum()}")
# Cross-references (UniRef, KEGG, EC, GO, etc.) parsed from the dbxrefs column
def split_xrefs(xref_str):
if not isinstance(xref_str, str) or xref_str in ("", "-"):
return []
return [x.strip() for x in xref_str.split(",")]
cds["dbxref_list"] = cds["dbxrefs"].apply(split_xrefs)
ec_hits = cds[cds["dbxref_list"].apply(lambda xs: any(x.startswith("EC:") for x in xs))]
print(f"CDS with EC numbers: {len(ec_hits)}")
print(ec_hits[["locus_tag", "gene", "product"]].head(5).to_string(index=False))Bakta emits a Circos-style PNG/SVG by default. Regenerate with custom styling using bakta_plot.
# Re-render the plot from the existing JSON with a different style
bakta_plot --output annotation/ \
--prefix E_coli_K12_replot \
--type cog \
--dpi 300 \
annotation/E_coli_K12.json
ls annotation/E_coli_K12_replot.*
# E_coli_K12_replot.png E_coli_K12_replot.svg
# Display the rendered PNG inline in a Jupyter notebook
python <<'PY'
from pathlib import Path
import matplotlib.pyplot as plt
import matplotlib.image as mpimg
img = mpimg.imread("annotation/E_coli_K12.png")
fig, ax = plt.subplots(figsize=(6, 6))
ax.imshow(img)
ax.axis("off")
ax.set_title("Bakta circular genome annotation")
plt.savefig("bakta_circular_thumbnail.png", dpi=120, bbox_inches="tight")
print("Saved bakta_circular_thumbnail.png")
PYSummarize hypothetical-protein rate, gene density, and feature coverage to assess annotation quality.
import json
import pandas as pd
import matplotlib.pyplot as plt
with open("annotation/E_coli_K12.json") as f:
bakta = json.load(f)
genome_size = bakta["genome"]["size"]
features = bakta["features"]
# Per-feature-type counts and density (per Mb)
counts = {}
for f in features:
counts[f["type"]] = counts.get(f["type"], 0) + 1
density = {k: v / (genome_size / 1e6) for k, v in counts.items()}
cds = [f for f in features if f["type"] == "cds"]
n_cds = len(cds)
n_hypo = sum(1 for f in cds if "hypothetical" in (f.get("product") or "").lower())
n_known = n_cds - n_hypo
coding_density = sum(abs(f["stop"] - f["start"] + 1) for f in cds) / genome_size
print(f"Genome: {genome_size:,} bp")
print(f"CDS: {n_cds} ({density.get('cds', 0):.1f} per Mb)")
print(f" Known function: {n_known} Hypothetical: {n_hypo}")
print(f"Coding density: {coding_density:.1%}")
print(f"tRNA: {counts.get('tRNA', 0)} rRNA: {counts.get('rRNA', 0)}")
print(f"ncRNA: {counts.get('ncRNA', 0)} CRISPR: {counts.get('crispr', 0)}")
# Bar plot of feature type counts
fig, ax = plt.subplots(figsize=(8, 4))
items = sorted(counts.items(), key=lambda x: -x[1])
labels = [k for k, _ in items]
values = [v for _, v in items]
bars = ax.bar(labels, values, color="#2980B9", edgecolor="white")
for bar, v in zip(bars, values):
ax.text(bar.get_x() + bar.get_width() / 2, v + max(values) * 0.01,
str(v), ha="center", fontsize=9)
ax.set_ylabel("Count")
ax.set_title("Bakta feature counts by type")
plt.xticks(rotation=20)
plt.tight_layout()
plt.savefig("bakta_feature_counts.png", dpi=150, bbox_inches="tight")
print("Saved bakta_feature_counts.png")Run Bakta over a directory of assemblies and aggregate per-sample summary statistics.
#!/bin/bash
# batch_bakta.sh — annotate all FASTA files in a directory
INPUT_DIR="genomes/"
OUTPUT_DIR="annotations/"
DB="db/bakta_db_light"
mkdir -p "$OUTPUT_DIR"
for FASTA in "$INPUT_DIR"/*.fasta; do
SAMPLE=$(basename "$FASTA" .fasta)
echo "Annotating: $SAMPLE"
bakta "$FASTA" \
--db "$DB" \
--output "${OUTPUT_DIR}/${SAMPLE}" \
--prefix "$SAMPLE" \
--threads 4 \
--skip-plot \
--force
done
echo "Batch annotation complete."# Aggregate per-genome summaries from each sample's JSON file
from pathlib import Path
import json
import pandas as pd
annotation_dir = Path("annotations/")
rows = []
for json_file in sorted(annotation_dir.glob("*/*.json")):
with open(json_file) as f:
bakta = json.load(f)
sample = json_file.stem
counts = {}
for feat in bakta["features"]:
counts[feat["type"]] = counts.get(feat["type"], 0) + 1
rows.append({
"sample": sample,
"size": bakta["genome"]["size"],
"gc": bakta["genome"]["gc"],
"CDS": counts.get("cds", 0),
"tRNA": counts.get("tRNA", 0),
"rRNA": counts.get("rRNA", 0),
"ncRNA": counts.get("ncRNA", 0),
"crispr": counts.get("crispr", 0),
})
summary_df = pd.DataFrame(rows)
print(f"Annotated genomes: {len(summary_df)}")
print(summary_df.to_string(index=False))
summary_df.to_csv("batch_bakta_summary.csv", index=False)
print("Saved: batch_bakta_summary.csv")| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
--db | — | path to Bakta DB directory | Required; selects light or full reference DB |
--genus | — | any genus name string | Sets organism metadata in GenBank/EMBL output |
--species | — | any species name string | Combined with --genus for organism qualifier |
--strain | — | any string | Adds strain qualifier to organism metadata |
--locus-tag | auto | 3–24 alphanumeric chars | Prefix for locus tags (e.g., ECOLI_00001) |
--complete | off | flag | Treat all input contigs as complete circular replicons |
--plasmid | off | flag | Treat all input contigs as plasmids (circular) |
--meta | off | flag | Disable Prodigal training (use for MAGs / metagenomic bins) |
--translation-table | 11 | NCBI table IDs (1, 4, 11, 25, …) | Genetic code used by Prodigal for CDS translation |
--min-contig-length | 1 | any integer (bp) | Skip contigs shorter than this length |
--threads | 1 | 1–CPU count | Parallel threads for DIAMOND/HMMER searches |
--skip-plot | off | flag | Skip the slow circular plot rendering step |
--keep-contig-headers | off | flag | Preserve original contig IDs instead of renaming |
--proteins | — | path to GenBank or FASTA file | Custom expert protein DB used before UniRef search |
When to use: A finished circular plasmid sequence that should be annotated without chromosome assumptions.
bakta plasmid.fasta \
--db db/bakta_db_light \
--output plasmid_annotation/ \
--prefix pBR322 \
--plasmid \
--complete \
--threads 4
# Inspect plasmid-typing results (incompatibility group, replication initiator)
grep -E "rep|inc" plasmid_annotation/pBR322.tsv | head -10When to use: Annotating a bin from metagenomic assembly where Prodigal cannot train on the full sequence.
bakta MAG_bin_42.fasta \
--db db/bakta_db_light \
--output mag_annotation/ \
--prefix MAG_bin_42 \
--meta \
--min-contig-length 500 \
--skip-plot \
--threads 8
cat mag_annotation/MAG_bin_42.txt | head -20When to use: You have curated reference proteins (e.g., a virulence factor catalog) that should take priority over UniRef.
# Custom proteins are searched before the bundled UniRef DB
bakta genome.fasta \
--db db/bakta_db_light \
--proteins virulence_factors.faa \
--output annotation_custom/ \
--prefix novel_strain \
--threads 8
echo "Annotations referencing custom DB:"
grep "User-provided" annotation_custom/novel_strain.tsv | headWhen to use: You need feature coordinates and qualifiers for downstream coordinate-overlap analyses.
import pandas as pd
def parse_gff3(gff_file):
rows = []
with open(gff_file) as f:
for line in f:
if line.startswith("##FASTA"):
break
if line.startswith("#") or not line.strip():
continue
parts = line.rstrip("\n").split("\t")
if len(parts) < 9:
continue
attrs = {}
for item in parts[8].split(";"):
if "=" in item:
k, v = item.split("=", 1)
attrs[k] = v
rows.append({
"seqname": parts[0],
"feature": parts[2],
"start": int(parts[3]),
"end": int(parts[4]),
"strand": parts[6],
"locus_tag": attrs.get("locus_tag", attrs.get("ID", "")),
"gene": attrs.get("gene", ""),
"product": attrs.get("product", ""),
})
return pd.DataFrame(rows)
gff_df = parse_gff3("annotation/E_coli_K12.gff3")
print(f"Features parsed: {len(gff_df)}")
print(gff_df[gff_df["feature"] == "CDS"].head(5).to_string(index=False))| Output File | Format | Description |
|---|---|---|
{prefix}.gff3 | GFF3 | Genome annotation with feature coordinates and attributes; INSDC-compliant |
{prefix}.gbff | GenBank Flat File | Annotated GenBank record for use in Geneious, BioPython, NCBI submission prep |
{prefix}.embl | EMBL | EMBL-format annotation for ENA submission |
{prefix}.fna | FASTA | Nucleotide sequences of the input contigs |
{prefix}.ffn | FASTA | Nucleotide sequences of all annotated features |
{prefix}.faa | FASTA | Protein sequences for all CDS features |
{prefix}.hypotheticals.faa | FASTA | Protein sequences flagged as hypothetical (for further investigation) |
{prefix}.hypotheticals.tsv | TSV | Detailed table for hypothetical CDS with low-confidence hits |
{prefix}.tsv | TSV | Full feature table: seqid, type, start, stop, strand, locus_tag, gene, product, dbxrefs |
{prefix}.json | JSON | Machine-readable summary with genome metadata + every feature with full qualifiers |
{prefix}.png / .svg | Image | Circos circular genome plot colored by COG category |
{prefix}.txt | Text | Plain-text summary of feature counts and per-replicon stats |
{prefix}.log | Text | Full Bakta run log including timing per step |
| Problem | Cause | Solution |
|---|---|---|
Error: database not found at <path> | DB path incorrect, or DB not extracted | Re-run bakta_db download --output db/ --type light; pass full path to --db |
KILLED during DIAMOND step | Out of memory on full DB | Switch to --type light, reduce --threads, or run on a host with ≥ 32 GB RAM |
| Very high hypothetical-protein rate (>60 %) | Divergent strain or light DB without close hits | Re-run with the full DB or supply a curated --proteins reference set |
tRNAscan-SE: Error opening file | Missing tRNAscan-SE installation in conda env | mamba install -c bioconda trnascan-se and rerun |
| Bakta complains about contig headers | Spaces or special characters in FASTA IDs | Pre-clean headers with awk '/^>/{print ">contig_"++i; next}{print}' |
| Plot generation hangs or fails | Circos / matplotlib backend issue | Re-run with --skip-plot; render later via bakta_plot from the JSON file |
AMRFinderPlus: database not up-to-date warning | AMRFinderPlus DB stale | amrfinder -u to refresh, or pass --skip-amr to bypass |
| Different locus tag prefix than expected | --locus-tag not set, default uses random prefix | Explicitly pass --locus-tag MYORG to control prefix |
| Run is much slower than reported | Default --threads 1 | Set --threads to physical core count; use --skip-plot for batch jobs |
.gbff output in Python© jaechang-hits, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/annotation/bakta-genome-annotation of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Bakta Genome Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bakta Genome Annotation this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~5.9k | Automated safety check: Pass | GPL-3.0 | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 | |
| Biopython Bioinformaticsaiming-lab/AutoResearchClaw | 15k | — | ~810 | Automated safety check: Pass | MIT | |
| Bio Write SequencesGPTomics/bioSkills | 1.2k | 3 repos | ~2.1k | Automated safety check: Pass | MIT | |
| ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates | 33k | 11 repos | ~4.5k | Automated safety check: Notes | MIT | |
| Biopythondavila7/claude-code-templates | 33k | 12 repos | ~3.4k | Automated safety check: Pass | MIT |
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
davila7/claude-code-templates
Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.
davila7/claude-code-templates
Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.
davila7/claude-code-templates
Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline. Bakta Genome Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate bacterial and archaeal genomes and plasmids with Bakta's Prodigal/HMM/diamond pipeline.
Bakta Genome Annotation fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/bakta-genome-annotation in jaechang-hits/SciAgent-Skills) into .claude/skills/bakta-genome-annotation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/bakta-genome-annotation in jaechang-hits/SciAgent-Skills) into .agents/skills/bakta-genome-annotation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill bakta-genome-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bakta-genome-annotation, .gemini/skills/bakta-genome-annotation, .github/skills/bakta-genome-annotation and .opencode/skills/bakta-genome-annotation in your project.
Going by SKILL.md and its folder, Bakta Genome Annotation needs the command-line tools its instructions call (mamba, pip and python). Our summary lists: Python 3.
SKILL.md names 3 domains. As links in the text: github.com, doi.org and biopython.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bakta Genome Annotation is published under the GPL-3.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bakta Genome Annotation: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars) and ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.