Dbsnp Database
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline.
$ npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills prokka-genome-annotation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/prokka-genome-annotation .claude/skills/prokka-genome-annotation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "prokka-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/prokka-genome-annotation into .claude/skills/prokka-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prokka-genome-annotation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/prokka-genome-annotationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills prokka-genome-annotation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/prokka-genome-annotation .agents/skills/prokka-genome-annotation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "prokka-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/prokka-genome-annotation into .agents/skills/prokka-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prokka-genome-annotation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills prokka-genome-annotation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/prokka-genome-annotation .cursor/skills/prokka-genome-annotation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "prokka-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/prokka-genome-annotation into .cursor/skills/prokka-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prokka-genome-annotation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/annotation/prokka-genome-annotation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills prokka-genome-annotation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/prokka-genome-annotation .gemini/skills/prokka-genome-annotation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "prokka-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/prokka-genome-annotation into .gemini/skills/prokka-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prokka-genome-annotation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills prokka-genome-annotationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/prokka-genome-annotation .github/skills/prokka-genome-annotation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "prokka-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/prokka-genome-annotation into .github/skills/prokka-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prokka-genome-annotation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills prokka-genome-annotation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/prokka-genome-annotation .opencode/skills/prokka-genome-annotation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "prokka-genome-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/prokka-genome-annotation into .opencode/skills/prokka-genome-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prokka-genome-annotation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
prokka-genome-annotationAnnotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline.
Prokka Genome Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline. Identifies CDS, rRNA, tRNA, tmRNA, signal peptides against Pfam, TIGRFAMs, RefSeq. Outputs GFF3, GenBank, FASTA, TSV. Use PGAP for NCBI GenBank submission; Bakta for faster NCBI-compatible annotation.
Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is GPL-3.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
condapipmambaFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comdoi.orgbiopython.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Prokka Genome Annotation loads about 6.1k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 1,099 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its GPL-3.0 licence (© jaechang-hits). 1,099 words, ~6,096 tokens.
.claude/skills/prokka-genome-annotation/SKILL.md (or your agent's skills folder).Prokka is a command-line pipeline for rapid annotation of prokaryotic genomes (bacteria, archaea, and viruses). It uses a tiered search strategy: protein-coding genes (CDS) are predicted with Prodigal and searched first against a genus-specific database, then RefSeq proteins, then Pfam/TIGRFAMs HMMs. Non-coding RNA genes (rRNA, tRNA, tmRNA) are identified with Barrnap, Aragorn, and Infernal. Prokka processes a single FASTA assembly in minutes and outputs a comprehensive annotation in GFF3, GenBank, FASTA, and tabular formats.
--metagenome flagbiopython, pandas, matplotlibCheck before installing: The tool may already be available in the current environment (e.g., inside a
pixi/condaenv). Runcommand -v prokkafirst and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool viapixi run prokkarather than bareprokka.
# Install Prokka via conda/mamba (recommended)
conda install -c conda-forge -c bioconda prokka
# Or with mamba (faster)
mamba install -c conda-forge -c bioconda prokka
# Verify installation and database setup
prokka --version
# prokka 1.14.6
# Check that required tools are on PATH
prokka --depends
# prokka needs: awk, sed, grep, makeblastdb, blastp, hmmscan, ...
# Install Python parsing dependencies
pip install biopython pandas matplotlibSettle these with the user before writing any analysis code.
decisions:
- id: D1
param: kingdom
kind: required
source: user
ask: "Is this a bacterial, archaeal, viral, or mitochondrial assembly?"
default: "Bacteria"
- id: D2
param: organismMetadata
kind: required
source: user
ask: "Which genus, species, and strain should be recorded, and used to prioritise a genus-specific protein database?"
default: null
- id: D3
param: assemblyOrigin
kind: required
source: data
ask: "Is this an isolate genome, or a bin recovered from a metagenome?"
default: "isolate - gene prediction trains on the assembly itself"
- id: D4
param: rnaPrediction
kind: optional
source: user
ask: "Should rRNA, tRNA, and other non-coding RNA be predicted alongside coding genes?"
default: "tRNA and rRNA predicted; Rfam search off as it is slow"
- id: D5
param: customProteins
kind: optional
source: user
ask: "Is there a curated protein set that should be searched before the general reference database?"
default: "none"
- id: D6
param: hitEvalue
kind: optional
source: user
ask: "How confident must a database hit be before it names a gene?"
default: "1e-6"
- id: D7
param: minContigLength
kind: optional_conditional
source: data
ask: "Should short contigs be skipped - and does the assembly have enough of them to matter?"
default: "no minimum"
- id: D8
param: cpus
kind: never_ask
source: data
reason: "Affects runtime only, not the annotation"
default: "min(8, available_cores)"D3 changes how genes are predicted rather than how they are named. Prodigal trains its model on the input; a metagenomic bin holding more than one organism trains a blended model, and the resulting gene boundaries are wrong in a way that no downstream step flags.
# Annotate a bacterial genome assembly — results in results/ directory
prokka genome.fasta \
--outdir results/ \
--prefix sample1 \
--kingdom Bacteria \
--cpus 4
# Check output summary
cat results/sample1.txt
# Organism: Genus species strain
# Contigs: 1
# Bases: 4639675
# CDS: 4140
# rRNA: 22
# tRNA: 86
echo "Annotation complete. Key output files:"
ls results/sample1.{gff,gbk,faa,ffn,tsv}Install Prokka and confirm all dependent tools are accessible in the current environment.
# Create a dedicated conda environment
conda create -n prokka_env -c conda-forge -c bioconda prokka python=3.10 -y
conda activate prokka_env
# Verify Prokka version and all tool dependencies
prokka --version
# prokka 1.14.6
prokka --depends
# Checking that required tools are installed...
# OK: makeblastdb is installed (2.13.0+)
# OK: blastp is installed (2.13.0+)
# OK: hmmscan is installed (3.3.2)
# OK: prodigal is installed (2.6.3)
# OK: barrnap is installed (0.9)
# Check available genus-specific databases bundled with Prokka
ls $(conda info --base)/envs/prokka_env/db/genus/
# Archaea Bacteria Mitochondria Viruses
# Install Python parsing tools
pip install biopython pandas matplotlibClean and rename contigs to comply with Prokka's header requirements before annotation.
from Bio import SeqIO
import re
# Load and inspect assembly
input_fasta = "genome.fasta"
records = list(SeqIO.parse(input_fasta, "fasta"))
print(f"Input assembly: {len(records)} contigs")
total_bases = sum(len(r) for r in records)
print(f"Total bases: {total_bases:,}")
print(f"Largest contig: {max(len(r) for r in records):,} bp")
print(f"N50 approx: see assembly stats tool")
# Rename contigs to short IDs compatible with Prokka (max 37 chars)
# Prokka requires: no spaces, no special characters in header
cleaned = []
for i, rec in enumerate(records, 1):
new_id = f"contig_{i:04d}"
new_rec = rec.__class__(rec.seq, id=new_id, description=f"len={len(rec.seq)}")
cleaned.append(new_rec)
SeqIO.write(cleaned, "genome_clean.fasta", "fasta")
print(f"\nWrote genome_clean.fasta with {len(cleaned)} renamed contigs")
# genome_clean.fasta: contig_0001 through contig_NNNN# Alternatively, clean headers with a simple bash one-liner
awk '/^>/{print ">contig_" ++i; next}{print}' genome.fasta > genome_clean.fasta
# Filter out short contigs (< 200 bp) to reduce annotation noise
awk '/^>/{header=$0; next} length($0) >= 200 {print header; print}' \
genome_clean.fasta > genome_filtered.fasta
echo "Filtered assembly ready: $(grep -c '>' genome_filtered.fasta) contigs"Run Prokka with standard options for a bacterial genome, specifying genus/species for database selection.
# Basic annotation with genus/species hint (uses genus-specific protein database first)
prokka genome_clean.fasta \
--outdir annotation/ \
--prefix E_coli_K12 \
--kingdom Bacteria \
--genus Escherichia \
--species coli \
--strain K12 \
--cpus 8 \
--mincontiglen 200
# Expected runtime: 2–10 minutes for a typical 4–6 Mb bacterial genome
echo "Prokka annotation output files:"
ls annotation/
# E_coli_K12.err E_coli_K12.faa E_coli_K12.ffn
# E_coli_K12.fna E_coli_K12.gbk E_coli_K12.gff
# E_coli_K12.log E_coli_K12.sqn E_coli_K12.tbl
# E_coli_K12.tsv E_coli_K12.txtLoad the TSV output for a quick overview of annotated features and their functional assignments.
import pandas as pd
# Load the annotation TSV (tab-delimited feature table)
tsv_file = "annotation/E_coli_K12.tsv"
df = pd.read_csv(tsv_file, sep="\t")
print(f"Total features: {len(df)}")
print(f"Columns: {list(df.columns)}")
# Columns: [locus_tag, ftype, length_bp, gene, EC_number, COG, product]
# Feature type summary
print("\nFeature type counts:")
print(df["ftype"].value_counts().to_string())
# CDS 4140
# tRNA 86
# rRNA 22
# tmRNA 1
# Functional gene annotations (non-hypothetical CDS)
cds_df = df[df["ftype"] == "CDS"].copy()
hypothetical = cds_df["product"].str.contains("hypothetical", case=False, na=True)
print(f"\nCDS with known function: {(~hypothetical).sum()}")
print(f"Hypothetical proteins: {hypothetical.sum()}")
# Genes with EC numbers (enzymes)
ec_annotated = cds_df[cds_df["EC_number"].notna() & (cds_df["EC_number"] != "")]
print(f"CDS with EC numbers: {len(ec_annotated)}")
print(ec_annotated[["locus_tag", "gene", "EC_number", "product"]].head(5).to_string(index=False))Read the GenBank file to access per-gene sequences, qualifiers, and feature coordinates.
from Bio import SeqIO
import pandas as pd
# Parse GenBank file
gbk_file = "annotation/E_coli_K12.gbk"
records = list(SeqIO.parse(gbk_file, "genbank"))
print(f"Contigs in GenBank: {len(records)}")
# Iterate over CDS features and extract details
rows = []
for rec in records:
for feat in rec.features:
if feat.type != "CDS":
continue
qualifiers = feat.qualifiers
rows.append({
"contig": rec.id,
"locus_tag": qualifiers.get("locus_tag", ["?"])[0],
"gene": qualifiers.get("gene", [""])[0],
"product": qualifiers.get("product", ["hypothetical protein"])[0],
"EC_number": qualifiers.get("EC_number", [""])[0],
"protein_id": qualifiers.get("protein_id", [""])[0],
"start": int(feat.location.start),
"end": int(feat.location.end),
"strand": feat.location.strand,
"aa_length": len(qualifiers.get("translation", [""])[0]),
})
features_df = pd.DataFrame(rows)
print(f"CDS features extracted: {len(features_df)}")
print(features_df.head(3).to_string(index=False))
# Retrieve protein sequence for a specific gene
gene_name = "dnaA"
gene_feat = features_df[features_df["gene"] == gene_name]
if not gene_feat.empty:
idx = gene_feat.index[0]
print(f"\n{gene_name}: {features_df.loc[idx, 'aa_length']} aa")Generate a summary barplot of feature types and functional annotation coverage.
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
tsv_file = "annotation/E_coli_K12.tsv"
df = pd.read_csv(tsv_file, sep="\t")
cds = df[df["ftype"] == "CDS"].copy()
n_cds = len(cds)
n_known = (~cds["product"].str.contains("hypothetical", case=False, na=True)).sum()
n_hypo = n_cds - n_known
n_ec = cds["EC_number"].notna().sum()
n_rrna = (df["ftype"] == "rRNA").sum()
n_trna = (df["ftype"] == "tRNA").sum()
fig, axes = plt.subplots(1, 2, figsize=(11, 4))
# Left panel: feature type counts
types = df["ftype"].value_counts()
colors_left = ["#2980B9", "#27AE60", "#E74C3C", "#F39C12"] + ["#95A5A6"] * len(types)
axes[0].bar(types.index, types.values, color=colors_left[:len(types)], edgecolor="white")
axes[0].set_title("Annotated Feature Counts")
axes[0].set_xlabel("Feature Type")
axes[0].set_ylabel("Count")
for i, (label, val) in enumerate(types.items()):
axes[0].text(i, val + 5, str(val), ha="center", fontsize=9)
# Right panel: CDS functional annotation breakdown
labels = ["Known function", "Hypothetical protein", "Enzyme (EC number)"]
sizes = [n_known, n_hypo, n_ec]
colors_right = ["#2980B9", "#BDC3C7", "#E74C3C"]
bars = axes[1].bar(labels, sizes, color=colors_right, edgecolor="white")
axes[1].set_title(f"CDS Functional Annotation\n(n={n_cds} total CDS)")
axes[1].set_ylabel("Count")
axes[1].tick_params(axis="x", rotation=15)
for bar, val in zip(bars, sizes):
axes[1].text(bar.get_x() + bar.get_width() / 2,
val + 10, str(val), ha="center", fontsize=9)
plt.suptitle("Prokka Genome Annotation Summary", fontsize=12, fontweight="bold")
plt.tight_layout()
plt.savefig("prokka_annotation_summary.png", dpi=150, bbox_inches="tight")
print(f"Saved prokka_annotation_summary.png")
print(f"CDS: {n_cds} | Known function: {n_known} | Hypothetical: {n_hypo}")
print(f"rRNA: {n_rrna} | tRNA: {n_trna} | EC-annotated CDS: {n_ec}")Annotate multiple genome assemblies sequentially and collect summary statistics.
#!/bin/bash
# batch_prokka.sh — annotate all FASTA files in a directory
INPUT_DIR="genomes/"
OUTPUT_DIR="annotations/"
mkdir -p "$OUTPUT_DIR"
for FASTA in "$INPUT_DIR"/*.fasta; do
SAMPLE=$(basename "$FASTA" .fasta)
echo "Annotating: $SAMPLE"
prokka "$FASTA" \
--outdir "${OUTPUT_DIR}/${SAMPLE}" \
--prefix "$SAMPLE" \
--kingdom Bacteria \
--cpus 4 \
--mincontiglen 200 \
--quiet
echo " Done: ${OUTPUT_DIR}/${SAMPLE}/${SAMPLE}.txt"
done
echo "Batch annotation complete."# Collect summary statistics from all Prokka .txt files
from pathlib import Path
import pandas as pd
annotation_dir = Path("annotations/")
rows = []
for txt_file in sorted(annotation_dir.glob("*/*.txt")):
sample = txt_file.stem
stats = {}
with open(txt_file) as f:
for line in f:
line = line.strip()
if ": " in line:
key, val = line.split(": ", 1)
stats[key.strip()] = val.strip()
rows.append({
"sample": sample,
"contigs": int(stats.get("Contigs", 0)),
"bases": int(stats.get("Bases", 0)),
"CDS": int(stats.get("CDS", 0)),
"rRNA": int(stats.get("rRNA", 0)),
"tRNA": int(stats.get("tRNA", 0)),
})
summary_df = pd.DataFrame(rows)
print(f"Annotated genomes: {len(summary_df)}")
print(summary_df.to_string(index=False))
summary_df.to_csv("batch_annotation_summary.csv", index=False)
print("\nSaved: batch_annotation_summary.csv")Use the protein FASTA outputs for cross-strain comparison with identity-based clustering.
from Bio import SeqIO, pairwise2
import pandas as pd
# Load protein sequences from two strains
def load_proteins(faa_file):
"""Return dict of locus_tag -> protein sequence."""
return {rec.id: str(rec.seq)
for rec in SeqIO.parse(faa_file, "fasta")}
strain_a = load_proteins("annotation_A/strain_A.faa")
strain_b = load_proteins("annotation_B/strain_B.faa")
print(f"Strain A proteins: {len(strain_a)}")
print(f"Strain B proteins: {len(strain_b)}")
# Compare gene counts and size distributions
import matplotlib.pyplot as plt
import numpy as np
len_a = [len(seq) for seq in strain_a.values()]
len_b = [len(seq) for seq in strain_b.values()]
fig, ax = plt.subplots(figsize=(8, 4))
bins = np.linspace(0, 1500, 50)
ax.hist(len_a, bins=bins, alpha=0.6, label=f"Strain A (n={len(len_a)})", color="#2980B9")
ax.hist(len_b, bins=bins, alpha=0.6, label=f"Strain B (n={len(len_b)})", color="#E74C3C")
ax.set_xlabel("Protein Length (aa)")
ax.set_ylabel("Count")
ax.set_title("Protein Length Distribution by Strain")
ax.legend()
plt.tight_layout()
plt.savefig("strain_protein_length_comparison.png", dpi=150, bbox_inches="tight")
print("Saved strain_protein_length_comparison.png")
# Find proteins unique to each strain by size/count difference
print(f"\nSize difference: {abs(len(strain_a) - len(strain_b))} proteins")
print("→ Use Roary or OrthoFinder for formal pan-genome analysis")| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
--kingdom | Bacteria | Bacteria, Archaea, Viruses, Mitochondria | Selects gene prediction model and default databases |
--genus | — | any genus name string | Prioritizes genus-specific protein database for CDS annotation |
--species | — | any species name string | Combined with --genus for locus_tag prefix and organism metadata |
--strain | — | any string | Added to organism metadata in GenBank output |
--proteins | — | path to FASTA file | Custom protein database prepended before RefSeq search |
--hmms | — | path to HMM file | Additional custom HMM database for specialized annotation |
--evalue | 1e-6 | 1e-4–1e-9 | E-value cutoff for BLAST and HMMER hits |
--cpus | 8 | 1–CPU count | Parallel processes for BLAST and HMMER searches |
--mincontiglen | 1 | any integer | Skip contigs shorter than this length (bp) |
--metagenome | off | flag | Disables Prodigal training on this genome (for MAGs) |
--rfam | off | flag | Enable Infernal rRNA/ncRNA search against Rfam (slower) |
--norrna | off | flag | Skip rRNA prediction (use when assembly has no rRNA genes) |
--notrna | off | flag | Skip tRNA prediction |
When to use: You have closely related reference proteins (e.g., characterized isolate) to improve annotation accuracy.
# Use a custom protein database to enhance annotation of a novel strain
# Custom proteins are searched first, before internal Prokka databases
prokka genome.fasta \
--proteins reference_proteins.faa \
--outdir custom_annotation/ \
--prefix novel_strain \
--kingdom Bacteria \
--cpus 4
echo "Custom DB annotation complete:"
grep "CDS" custom_annotation/novel_strain.txtWhen to use: Annotating a MAG where Prodigal cannot train on the full genome sequence.
# --metagenome disables Prodigal model training (uses meta mode)
# --mincontiglen 500 discards short, potentially chimeric contigs
prokka MAG_bin_42.fasta \
--metagenome \
--kingdom Bacteria \
--outdir mag_annotation/ \
--prefix MAG_bin_42 \
--mincontiglen 500 \
--cpus 4
echo "MAG annotation complete:"
cat mag_annotation/MAG_bin_42.txtWhen to use: Retrieve nucleotide or protein sequences for a target gene or pathway from the annotation.
from Bio import SeqIO
# Extract all sequences for a specific gene name from protein FASTA
faa_file = "annotation/E_coli_K12.faa"
target_gene = "dnaA"
matches = []
for rec in SeqIO.parse(faa_file, "fasta"):
# Prokka FASTA header: >locus_tag gene product
if target_gene in rec.description:
matches.append(rec)
print(f"Proteins matching '{target_gene}': {len(matches)}")
for rec in matches:
print(f" {rec.id}: {len(rec.seq)} aa — {rec.description}")
# Save matching sequences
if matches:
SeqIO.write(matches, f"{target_gene}_proteins.faa", "fasta")
print(f"Saved: {target_gene}_proteins.faa")When to use: Work with genomic coordinates for feature overlap analysis or visualization.
import pandas as pd
def parse_gff(gff_file):
"""Parse Prokka GFF3 file into a DataFrame (skips sequence section)."""
rows = []
with open(gff_file) as f:
for line in f:
if line.startswith("##FASTA"):
break
if line.startswith("#") or not line.strip():
continue
parts = line.rstrip("\n").split("\t")
if len(parts) < 9:
continue
attr_dict = {}
for item in parts[8].split(";"):
if "=" in item:
k, v = item.split("=", 1)
attr_dict[k] = v
rows.append({
"seqname": parts[0],
"source": parts[1],
"feature": parts[2],
"start": int(parts[3]),
"end": int(parts[4]),
"score": parts[5],
"strand": parts[6],
"frame": parts[7],
"locus_tag": attr_dict.get("ID", ""),
"gene": attr_dict.get("gene", ""),
"product": attr_dict.get("product", ""),
})
return pd.DataFrame(rows)
gff_df = parse_gff("annotation/E_coli_K12.gff")
print(f"GFF features: {len(gff_df)}")
print(gff_df[gff_df["feature"] == "CDS"].head(5)[
["seqname", "start", "end", "strand", "gene", "product"]
].to_string(index=False))| Output File | Format | Description |
|---|---|---|
{prefix}.gff | GFF3 | Genome annotation with feature coordinates and attributes; includes FASTA sequence at the end |
{prefix}.gbk | GenBank | Full GenBank-format annotation for use in Geneious, BioPython, and submission prep |
{prefix}.faa | FASTA | Predicted protein sequences for all CDS features |
{prefix}.ffn | FASTA | Nucleotide sequences for all annotated features (CDS, rRNA, tRNA) |
{prefix}.fna | FASTA | Nucleotide FASTA of the complete assembly (contigs) |
{prefix}.tsv | TSV | Tab-delimited summary table: locus_tag, ftype, length_bp, gene, EC_number, COG, product |
{prefix}.txt | Text | One-line summary counts: Contigs, Bases, CDS, rRNA, tRNA, tmRNA |
{prefix}.tbl | TBL | Feature table format for tbl2asn GenBank submission |
{prefix}.sqn | SQN | ASN.1 format for NCBI submission (generated by tbl2asn) |
{prefix}.err | Text | Warnings and errors from tbl2asn validation step |
{prefix}.log | Text | Full Prokka run log with timing per step |
| Problem | Cause | Solution |
|---|---|---|
FATAL: Can't find any tRNA genes | No tRNA detected on short or highly fragmented assembly | Add --notrna flag; check if assembly is too fragmented (N50 < 1 kb) |
Can't exec "makeblastdb" | BLAST+ not on PATH | conda install -c bioconda blast; ensure prokka env is activated |
Contig ID too long (>37 chars) | Assembler produced long contig headers | Pre-process FASTA to shorten headers: awk '/^>/{print ">contig_"++i; next}{print}' |
| Very high hypothetical protein rate (>60%) | Divergent organism with few database matches | Add --proteins with closely related strain FAA; consider --genus flag |
ERROR: Argument --proteins: file does not exist | Path to custom protein file is incorrect | Use absolute path; verify file exists with ls -la custom.faa |
tbl2asn error: multiple /product | Duplicate product qualifiers in annotation | Ignore if exporting for local use; for NCBI submission, use PGAP instead |
| Annotation is very slow (>30 min for ~5 Mb) | --cpus not set or set to 1; --rfam enabled | Set --cpus to available thread count; disable --rfam for faster runs |
| rRNA count is 0 for complete genome | --norrna flag was set, or barrnap threshold too strict | Remove --norrna; check barrnap is installed with barrnap --version |
© jaechang-hits, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/annotation/prokka-genome-annotation of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Prokka Genome Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Prokka Genome Annotation this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~6.1k | Automated safety check: Pass | GPL-3.0 | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 | |
| Biopython Bioinformaticsaiming-lab/AutoResearchClaw | 15k | — | ~810 | Automated safety check: Pass | MIT | |
| Bio Write SequencesGPTomics/bioSkills | 1.2k | 3 repos | ~2.1k | Automated safety check: Pass | MIT | |
| ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates | 33k | 11 repos | ~4.5k | Automated safety check: Notes | MIT | |
| Biopythondavila7/claude-code-templates | 33k | 12 repos | ~3.4k | Automated safety check: Pass | MIT |
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
davila7/claude-code-templates
Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.
davila7/claude-code-templates
Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.
davila7/claude-code-templates
Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline. Prokka Genome Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline.
Prokka Genome Annotation fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/prokka-genome-annotation in jaechang-hits/SciAgent-Skills) into .claude/skills/prokka-genome-annotation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/prokka-genome-annotation in jaechang-hits/SciAgent-Skills) into .agents/skills/prokka-genome-annotation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prokka-genome-annotation, .gemini/skills/prokka-genome-annotation, .github/skills/prokka-genome-annotation and .opencode/skills/prokka-genome-annotation in your project.
Going by SKILL.md and its folder, Prokka Genome Annotation needs the command-line tools its instructions call (conda, pip and mamba). Our summary lists: Python 3.
SKILL.md names 3 domains. As links in the text: github.com, doi.org and biopython.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Prokka Genome Annotation is published under the GPL-3.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Prokka Genome Annotation: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars) and ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.