Agent skill

Prokka Genome Annotation

by jaechang-hits in jaechang-hits/SciAgent-Skills

Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline.

GPL-3.0Auto-check passedResearch & Science

Install Prokka Genome Annotation

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills prokka-genome-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/prokka-genome-annotation .claude/skills/prokka-genome-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prokka-genome-annotation
GitHub stars
374
Used in
1 other repo
Token cost
~6.1k tokens
SKILL.md length
1,099 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
GPL-3.0

At a glance

Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline.

  • Works in 8 steps: Install and Verify Prokka → Prepare the Input Genome → Run Basic Prokka Annotation → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 7 more sections
  • Calls conda, pip and mamba

What it does

Prokka Genome Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline. Identifies CDS, rRNA, tRNA, tmRNA, signal peptides against Pfam, TIGRFAMs, RefSeq. Outputs GFF3, GenBank, FASTA, TSV. Use PGAP for NCBI GenBank submission; Bakta for faster NCBI-compatible annotation.

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/prokka-genome-annotation”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Install and Verify Prokka
  2. Prepare the Input Genome
  3. Run Basic Prokka Annotation
  4. Parse Annotation Summary (TSV)
  5. Parse GenBank Output with BioPython
  6. Visualize Annotation Statistics
  7. Batch Annotation Across Multiple Genomes
  8. Compare Annotations Between Strains

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • conda
    • pip
    • mamba

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • doi.org
    • biopython.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prokka Genome Annotation loads about 6.1k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 1,099 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its GPL-3.0 licence (© jaechang-hits). 1,099 words, ~6,096 tokens.

Download SKILL.mdSave it as .claude/skills/prokka-genome-annotation/SKILL.md (or your agent's skills folder).
name
prokka-genome-annotation
description
Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline. Identifies CDS, rRNA, tRNA, tmRNA, signal peptides against Pfam, TIGRFAMs, RefSeq. Outputs GFF3, GenBank, FASTA, TSV. Use PGAP for NCBI GenBank submission; Bakta for faster NCBI-compatible annotation.
license
GPL-3.0

Prokka Genome Annotation

Overview

Prokka is a command-line pipeline for rapid annotation of prokaryotic genomes (bacteria, archaea, and viruses). It uses a tiered search strategy: protein-coding genes (CDS) are predicted with Prodigal and searched first against a genus-specific database, then RefSeq proteins, then Pfam/TIGRFAMs HMMs. Non-coding RNA genes (rRNA, tRNA, tmRNA) are identified with Barrnap, Aragorn, and Infernal. Prokka processes a single FASTA assembly in minutes and outputs a comprehensive annotation in GFF3, GenBank, FASTA, and tabular formats.

When to Use

  • Annotating a newly assembled bacterial or archaeal genome from Illumina, PacBio, or Nanopore assemblies
  • Getting functional protein annotations (CDS with product names, EC numbers, GO terms) from a draft or complete genome
  • Preparing annotation files for downstream comparative genomics (Roary pan-genome, OrthoFinder)
  • Annotating viral or phage genomes when kingdom-specific databases are important
  • Performing metagenome-assembled genome (MAG) annotation with the --metagenome flag
  • Parsing annotated outputs in Python with BioPython for downstream sequence or feature analysis
  • Use PGAP (NCBI Prokaryotic Genome Annotation Pipeline) instead when the goal is NCBI GenBank submission with standards compliance
  • Use Bakta instead for faster annotation with built-in NCBI-compatible outputs and a more regularly updated database

Prerequisites

  • Software: Prokka ≥ 1.14, Perl 5, Prodigal, Barrnap, HMMER3, BLAST+, Aragorn, Infernal, tbl2asn
  • Python packages (for output parsing): biopython, pandas, matplotlib
  • Input: assembled genome in FASTA format (complete or draft with multiple contigs)
  • Environment: conda strongly recommended to handle the Perl and C dependency stack

Check before installing: The tool may already be available in the current environment (e.g., inside a pixi / conda env). Run command -v prokka first and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool via pixi run prokka rather than bare prokka.

bash
# Install Prokka via conda/mamba (recommended)
conda install -c conda-forge -c bioconda prokka

# Or with mamba (faster)
mamba install -c conda-forge -c bioconda prokka

# Verify installation and database setup
prokka --version
# prokka 1.14.6

# Check that required tools are on PATH
prokka --depends
# prokka needs: awk, sed, grep, makeblastdb, blastp, hmmscan, ...

# Install Python parsing dependencies
pip install biopython pandas matplotlib

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: kingdom
    kind: required
    source: user
    ask: "Is this a bacterial, archaeal, viral, or mitochondrial assembly?"
    default: "Bacteria"

  - id: D2
    param: organismMetadata
    kind: required
    source: user
    ask: "Which genus, species, and strain should be recorded, and used to prioritise a genus-specific protein database?"
    default: null

  - id: D3
    param: assemblyOrigin
    kind: required
    source: data
    ask: "Is this an isolate genome, or a bin recovered from a metagenome?"
    default: "isolate - gene prediction trains on the assembly itself"

  - id: D4
    param: rnaPrediction
    kind: optional
    source: user
    ask: "Should rRNA, tRNA, and other non-coding RNA be predicted alongside coding genes?"
    default: "tRNA and rRNA predicted; Rfam search off as it is slow"

  - id: D5
    param: customProteins
    kind: optional
    source: user
    ask: "Is there a curated protein set that should be searched before the general reference database?"
    default: "none"

  - id: D6
    param: hitEvalue
    kind: optional
    source: user
    ask: "How confident must a database hit be before it names a gene?"
    default: "1e-6"

  - id: D7
    param: minContigLength
    kind: optional_conditional
    source: data
    ask: "Should short contigs be skipped - and does the assembly have enough of them to matter?"
    default: "no minimum"

  - id: D8
    param: cpus
    kind: never_ask
    source: data
    reason: "Affects runtime only, not the annotation"
    default: "min(8, available_cores)"

D3 changes how genes are predicted rather than how they are named. Prodigal trains its model on the input; a metagenomic bin holding more than one organism trains a blended model, and the resulting gene boundaries are wrong in a way that no downstream step flags.

Quick Start

bash
# Annotate a bacterial genome assembly — results in results/ directory
prokka genome.fasta \
    --outdir results/ \
    --prefix sample1 \
    --kingdom Bacteria \
    --cpus 4

# Check output summary
cat results/sample1.txt
# Organism: Genus species strain
# Contigs: 1
# Bases: 4639675
# CDS: 4140
# rRNA: 22
# tRNA: 86

echo "Annotation complete. Key output files:"
ls results/sample1.{gff,gbk,faa,ffn,tsv}

Workflow

Step 1: Install and Verify Prokka

Install Prokka and confirm all dependent tools are accessible in the current environment.

bash
# Create a dedicated conda environment
conda create -n prokka_env -c conda-forge -c bioconda prokka python=3.10 -y
conda activate prokka_env

# Verify Prokka version and all tool dependencies
prokka --version
# prokka 1.14.6

prokka --depends
# Checking that required tools are installed...
# OK: makeblastdb is installed (2.13.0+)
# OK: blastp is installed (2.13.0+)
# OK: hmmscan is installed (3.3.2)
# OK: prodigal is installed (2.6.3)
# OK: barrnap is installed (0.9)

# Check available genus-specific databases bundled with Prokka
ls $(conda info --base)/envs/prokka_env/db/genus/
# Archaea  Bacteria  Mitochondria  Viruses

# Install Python parsing tools
pip install biopython pandas matplotlib
Step 2: Prepare the Input Genome

Clean and rename contigs to comply with Prokka's header requirements before annotation.

python
from Bio import SeqIO
import re

# Load and inspect assembly
input_fasta = "genome.fasta"
records = list(SeqIO.parse(input_fasta, "fasta"))
print(f"Input assembly: {len(records)} contigs")
total_bases = sum(len(r) for r in records)
print(f"Total bases: {total_bases:,}")
print(f"Largest contig: {max(len(r) for r in records):,} bp")
print(f"N50 approx: see assembly stats tool")

# Rename contigs to short IDs compatible with Prokka (max 37 chars)
# Prokka requires: no spaces, no special characters in header
cleaned = []
for i, rec in enumerate(records, 1):
    new_id = f"contig_{i:04d}"
    new_rec = rec.__class__(rec.seq, id=new_id, description=f"len={len(rec.seq)}")
    cleaned.append(new_rec)

SeqIO.write(cleaned, "genome_clean.fasta", "fasta")
print(f"\nWrote genome_clean.fasta with {len(cleaned)} renamed contigs")
# genome_clean.fasta: contig_0001 through contig_NNNN
bash
# Alternatively, clean headers with a simple bash one-liner
awk '/^>/{print ">contig_" ++i; next}{print}' genome.fasta > genome_clean.fasta

# Filter out short contigs (< 200 bp) to reduce annotation noise
awk '/^>/{header=$0; next} length($0) >= 200 {print header; print}' \
    genome_clean.fasta > genome_filtered.fasta

echo "Filtered assembly ready: $(grep -c '>' genome_filtered.fasta) contigs"
Step 3: Run Basic Prokka Annotation

Run Prokka with standard options for a bacterial genome, specifying genus/species for database selection.

bash
# Basic annotation with genus/species hint (uses genus-specific protein database first)
prokka genome_clean.fasta \
    --outdir annotation/ \
    --prefix E_coli_K12 \
    --kingdom Bacteria \
    --genus Escherichia \
    --species coli \
    --strain K12 \
    --cpus 8 \
    --mincontiglen 200

# Expected runtime: 2–10 minutes for a typical 4–6 Mb bacterial genome

echo "Prokka annotation output files:"
ls annotation/
# E_coli_K12.err   E_coli_K12.faa   E_coli_K12.ffn
# E_coli_K12.fna   E_coli_K12.gbk   E_coli_K12.gff
# E_coli_K12.log   E_coli_K12.sqn   E_coli_K12.tbl
# E_coli_K12.tsv   E_coli_K12.txt
Step 4: Parse Annotation Summary (TSV)

Load the TSV output for a quick overview of annotated features and their functional assignments.

python
import pandas as pd

# Load the annotation TSV (tab-delimited feature table)
tsv_file = "annotation/E_coli_K12.tsv"
df = pd.read_csv(tsv_file, sep="\t")
print(f"Total features: {len(df)}")
print(f"Columns: {list(df.columns)}")
# Columns: [locus_tag, ftype, length_bp, gene, EC_number, COG, product]

# Feature type summary
print("\nFeature type counts:")
print(df["ftype"].value_counts().to_string())
# CDS     4140
# tRNA      86
# rRNA      22
# tmRNA      1

# Functional gene annotations (non-hypothetical CDS)
cds_df = df[df["ftype"] == "CDS"].copy()
hypothetical = cds_df["product"].str.contains("hypothetical", case=False, na=True)
print(f"\nCDS with known function: {(~hypothetical).sum()}")
print(f"Hypothetical proteins: {hypothetical.sum()}")

# Genes with EC numbers (enzymes)
ec_annotated = cds_df[cds_df["EC_number"].notna() & (cds_df["EC_number"] != "")]
print(f"CDS with EC numbers: {len(ec_annotated)}")
print(ec_annotated[["locus_tag", "gene", "EC_number", "product"]].head(5).to_string(index=False))
Step 5: Parse GenBank Output with BioPython

Read the GenBank file to access per-gene sequences, qualifiers, and feature coordinates.

python
from Bio import SeqIO
import pandas as pd

# Parse GenBank file
gbk_file = "annotation/E_coli_K12.gbk"
records = list(SeqIO.parse(gbk_file, "genbank"))
print(f"Contigs in GenBank: {len(records)}")

# Iterate over CDS features and extract details
rows = []
for rec in records:
    for feat in rec.features:
        if feat.type != "CDS":
            continue
        qualifiers = feat.qualifiers
        rows.append({
            "contig":      rec.id,
            "locus_tag":   qualifiers.get("locus_tag", ["?"])[0],
            "gene":        qualifiers.get("gene", [""])[0],
            "product":     qualifiers.get("product", ["hypothetical protein"])[0],
            "EC_number":   qualifiers.get("EC_number", [""])[0],
            "protein_id":  qualifiers.get("protein_id", [""])[0],
            "start":       int(feat.location.start),
            "end":         int(feat.location.end),
            "strand":      feat.location.strand,
            "aa_length":   len(qualifiers.get("translation", [""])[0]),
        })

features_df = pd.DataFrame(rows)
print(f"CDS features extracted: {len(features_df)}")
print(features_df.head(3).to_string(index=False))

# Retrieve protein sequence for a specific gene
gene_name = "dnaA"
gene_feat = features_df[features_df["gene"] == gene_name]
if not gene_feat.empty:
    idx = gene_feat.index[0]
    print(f"\n{gene_name}: {features_df.loc[idx, 'aa_length']} aa")
Step 6: Visualize Annotation Statistics

Generate a summary barplot of feature types and functional annotation coverage.

python
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches

tsv_file = "annotation/E_coli_K12.tsv"
df = pd.read_csv(tsv_file, sep="\t")

cds = df[df["ftype"] == "CDS"].copy()
n_cds = len(cds)
n_known = (~cds["product"].str.contains("hypothetical", case=False, na=True)).sum()
n_hypo = n_cds - n_known
n_ec = cds["EC_number"].notna().sum()
n_rrna = (df["ftype"] == "rRNA").sum()
n_trna = (df["ftype"] == "tRNA").sum()

fig, axes = plt.subplots(1, 2, figsize=(11, 4))

# Left panel: feature type counts
types = df["ftype"].value_counts()
colors_left = ["#2980B9", "#27AE60", "#E74C3C", "#F39C12"] + ["#95A5A6"] * len(types)
axes[0].bar(types.index, types.values, color=colors_left[:len(types)], edgecolor="white")
axes[0].set_title("Annotated Feature Counts")
axes[0].set_xlabel("Feature Type")
axes[0].set_ylabel("Count")
for i, (label, val) in enumerate(types.items()):
    axes[0].text(i, val + 5, str(val), ha="center", fontsize=9)

# Right panel: CDS functional annotation breakdown
labels = ["Known function", "Hypothetical protein", "Enzyme (EC number)"]
sizes  = [n_known, n_hypo, n_ec]
colors_right = ["#2980B9", "#BDC3C7", "#E74C3C"]
bars = axes[1].bar(labels, sizes, color=colors_right, edgecolor="white")
axes[1].set_title(f"CDS Functional Annotation\n(n={n_cds} total CDS)")
axes[1].set_ylabel("Count")
axes[1].tick_params(axis="x", rotation=15)
for bar, val in zip(bars, sizes):
    axes[1].text(bar.get_x() + bar.get_width() / 2,
                 val + 10, str(val), ha="center", fontsize=9)

plt.suptitle("Prokka Genome Annotation Summary", fontsize=12, fontweight="bold")
plt.tight_layout()
plt.savefig("prokka_annotation_summary.png", dpi=150, bbox_inches="tight")
print(f"Saved prokka_annotation_summary.png")
print(f"CDS: {n_cds}  |  Known function: {n_known}  |  Hypothetical: {n_hypo}")
print(f"rRNA: {n_rrna}  |  tRNA: {n_trna}  |  EC-annotated CDS: {n_ec}")
Step 7: Batch Annotation Across Multiple Genomes

Annotate multiple genome assemblies sequentially and collect summary statistics.

bash
#!/bin/bash
# batch_prokka.sh — annotate all FASTA files in a directory

INPUT_DIR="genomes/"
OUTPUT_DIR="annotations/"
mkdir -p "$OUTPUT_DIR"

for FASTA in "$INPUT_DIR"/*.fasta; do
    SAMPLE=$(basename "$FASTA" .fasta)
    echo "Annotating: $SAMPLE"
    prokka "$FASTA" \
        --outdir "${OUTPUT_DIR}/${SAMPLE}" \
        --prefix "$SAMPLE" \
        --kingdom Bacteria \
        --cpus 4 \
        --mincontiglen 200 \
        --quiet
    echo "  Done: ${OUTPUT_DIR}/${SAMPLE}/${SAMPLE}.txt"
done

echo "Batch annotation complete."
python
# Collect summary statistics from all Prokka .txt files
from pathlib import Path
import pandas as pd

annotation_dir = Path("annotations/")
rows = []

for txt_file in sorted(annotation_dir.glob("*/*.txt")):
    sample = txt_file.stem
    stats = {}
    with open(txt_file) as f:
        for line in f:
            line = line.strip()
            if ": " in line:
                key, val = line.split(": ", 1)
                stats[key.strip()] = val.strip()
    rows.append({
        "sample":   sample,
        "contigs":  int(stats.get("Contigs", 0)),
        "bases":    int(stats.get("Bases", 0)),
        "CDS":      int(stats.get("CDS", 0)),
        "rRNA":     int(stats.get("rRNA", 0)),
        "tRNA":     int(stats.get("tRNA", 0)),
    })

summary_df = pd.DataFrame(rows)
print(f"Annotated genomes: {len(summary_df)}")
print(summary_df.to_string(index=False))
summary_df.to_csv("batch_annotation_summary.csv", index=False)
print("\nSaved: batch_annotation_summary.csv")
Step 8: Compare Annotations Between Strains

Use the protein FASTA outputs for cross-strain comparison with identity-based clustering.

python
from Bio import SeqIO, pairwise2
import pandas as pd

# Load protein sequences from two strains
def load_proteins(faa_file):
    """Return dict of locus_tag -> protein sequence."""
    return {rec.id: str(rec.seq)
            for rec in SeqIO.parse(faa_file, "fasta")}

strain_a = load_proteins("annotation_A/strain_A.faa")
strain_b = load_proteins("annotation_B/strain_B.faa")

print(f"Strain A proteins: {len(strain_a)}")
print(f"Strain B proteins: {len(strain_b)}")

# Compare gene counts and size distributions
import matplotlib.pyplot as plt
import numpy as np

len_a = [len(seq) for seq in strain_a.values()]
len_b = [len(seq) for seq in strain_b.values()]

fig, ax = plt.subplots(figsize=(8, 4))
bins = np.linspace(0, 1500, 50)
ax.hist(len_a, bins=bins, alpha=0.6, label=f"Strain A (n={len(len_a)})", color="#2980B9")
ax.hist(len_b, bins=bins, alpha=0.6, label=f"Strain B (n={len(len_b)})", color="#E74C3C")
ax.set_xlabel("Protein Length (aa)")
ax.set_ylabel("Count")
ax.set_title("Protein Length Distribution by Strain")
ax.legend()
plt.tight_layout()
plt.savefig("strain_protein_length_comparison.png", dpi=150, bbox_inches="tight")
print("Saved strain_protein_length_comparison.png")

# Find proteins unique to each strain by size/count difference
print(f"\nSize difference: {abs(len(strain_a) - len(strain_b))} proteins")
print("→ Use Roary or OrthoFinder for formal pan-genome analysis")

Key Parameters

ParameterDefaultRange / OptionsEffect
--kingdomBacteriaBacteria, Archaea, Viruses, MitochondriaSelects gene prediction model and default databases
--genus—any genus name stringPrioritizes genus-specific protein database for CDS annotation
--species—any species name stringCombined with --genus for locus_tag prefix and organism metadata
--strain—any stringAdded to organism metadata in GenBank output
--proteins—path to FASTA fileCustom protein database prepended before RefSeq search
--hmms—path to HMM fileAdditional custom HMM database for specialized annotation
--evalue1e-61e-4–1e-9E-value cutoff for BLAST and HMMER hits
--cpus81–CPU countParallel processes for BLAST and HMMER searches
--mincontiglen1any integerSkip contigs shorter than this length (bp)
--metagenomeoffflagDisables Prodigal training on this genome (for MAGs)
--rfamoffflagEnable Infernal rRNA/ncRNA search against Rfam (slower)
--norrnaoffflagSkip rRNA prediction (use when assembly has no rRNA genes)
--notrnaoffflagSkip tRNA prediction

Common Recipes

Recipe: Annotation with Custom Protein Database

When to use: You have closely related reference proteins (e.g., characterized isolate) to improve annotation accuracy.

bash
# Use a custom protein database to enhance annotation of a novel strain
# Custom proteins are searched first, before internal Prokka databases
prokka genome.fasta \
    --proteins reference_proteins.faa \
    --outdir custom_annotation/ \
    --prefix novel_strain \
    --kingdom Bacteria \
    --cpus 4

echo "Custom DB annotation complete:"
grep "CDS" custom_annotation/novel_strain.txt
Show full SKILL.md (430 more words)Show less
Recipe: Metagenome-Assembled Genome (MAG) Annotation

When to use: Annotating a MAG where Prodigal cannot train on the full genome sequence.

bash
# --metagenome disables Prodigal model training (uses meta mode)
# --mincontiglen 500 discards short, potentially chimeric contigs
prokka MAG_bin_42.fasta \
    --metagenome \
    --kingdom Bacteria \
    --outdir mag_annotation/ \
    --prefix MAG_bin_42 \
    --mincontiglen 500 \
    --cpus 4

echo "MAG annotation complete:"
cat mag_annotation/MAG_bin_42.txt
Recipe: Extract Specific Gene Sequences

When to use: Retrieve nucleotide or protein sequences for a target gene or pathway from the annotation.

python
from Bio import SeqIO

# Extract all sequences for a specific gene name from protein FASTA
faa_file = "annotation/E_coli_K12.faa"
target_gene = "dnaA"

matches = []
for rec in SeqIO.parse(faa_file, "fasta"):
    # Prokka FASTA header: >locus_tag gene product
    if target_gene in rec.description:
        matches.append(rec)

print(f"Proteins matching '{target_gene}': {len(matches)}")
for rec in matches:
    print(f"  {rec.id}: {len(rec.seq)} aa — {rec.description}")

# Save matching sequences
if matches:
    SeqIO.write(matches, f"{target_gene}_proteins.faa", "fasta")
    print(f"Saved: {target_gene}_proteins.faa")
Recipe: Convert GFF to Pandas DataFrame

When to use: Work with genomic coordinates for feature overlap analysis or visualization.

python
import pandas as pd

def parse_gff(gff_file):
    """Parse Prokka GFF3 file into a DataFrame (skips sequence section)."""
    rows = []
    with open(gff_file) as f:
        for line in f:
            if line.startswith("##FASTA"):
                break
            if line.startswith("#") or not line.strip():
                continue
            parts = line.rstrip("\n").split("\t")
            if len(parts) < 9:
                continue
            attr_dict = {}
            for item in parts[8].split(";"):
                if "=" in item:
                    k, v = item.split("=", 1)
                    attr_dict[k] = v
            rows.append({
                "seqname":  parts[0],
                "source":   parts[1],
                "feature":  parts[2],
                "start":    int(parts[3]),
                "end":      int(parts[4]),
                "score":    parts[5],
                "strand":   parts[6],
                "frame":    parts[7],
                "locus_tag": attr_dict.get("ID", ""),
                "gene":     attr_dict.get("gene", ""),
                "product":  attr_dict.get("product", ""),
            })
    return pd.DataFrame(rows)

gff_df = parse_gff("annotation/E_coli_K12.gff")
print(f"GFF features: {len(gff_df)}")
print(gff_df[gff_df["feature"] == "CDS"].head(5)[
    ["seqname", "start", "end", "strand", "gene", "product"]
].to_string(index=False))

Expected Outputs

Output FileFormatDescription
{prefix}.gffGFF3Genome annotation with feature coordinates and attributes; includes FASTA sequence at the end
{prefix}.gbkGenBankFull GenBank-format annotation for use in Geneious, BioPython, and submission prep
{prefix}.faaFASTAPredicted protein sequences for all CDS features
{prefix}.ffnFASTANucleotide sequences for all annotated features (CDS, rRNA, tRNA)
{prefix}.fnaFASTANucleotide FASTA of the complete assembly (contigs)
{prefix}.tsvTSVTab-delimited summary table: locus_tag, ftype, length_bp, gene, EC_number, COG, product
{prefix}.txtTextOne-line summary counts: Contigs, Bases, CDS, rRNA, tRNA, tmRNA
{prefix}.tblTBLFeature table format for tbl2asn GenBank submission
{prefix}.sqnSQNASN.1 format for NCBI submission (generated by tbl2asn)
{prefix}.errTextWarnings and errors from tbl2asn validation step
{prefix}.logTextFull Prokka run log with timing per step

Troubleshooting

ProblemCauseSolution
FATAL: Can't find any tRNA genesNo tRNA detected on short or highly fragmented assemblyAdd --notrna flag; check if assembly is too fragmented (N50 < 1 kb)
Can't exec "makeblastdb"BLAST+ not on PATHconda install -c bioconda blast; ensure prokka env is activated
Contig ID too long (>37 chars)Assembler produced long contig headersPre-process FASTA to shorten headers: awk '/^>/{print ">contig_"++i; next}{print}'
Very high hypothetical protein rate (>60%)Divergent organism with few database matchesAdd --proteins with closely related strain FAA; consider --genus flag
ERROR: Argument --proteins: file does not existPath to custom protein file is incorrectUse absolute path; verify file exists with ls -la custom.faa
tbl2asn error: multiple /productDuplicate product qualifiers in annotationIgnore if exporting for local use; for NCBI submission, use PGAP instead
Annotation is very slow (>30 min for ~5 Mb)--cpus not set or set to 1; --rfam enabledSet --cpus to available thread count; disable --rfam for faster runs
rRNA count is 0 for complete genome--norrna flag was set, or barrnap threshold too strictRemove --norrna; check barrnap is installed with barrnap --version

References

© jaechang-hits, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/annotation/prokka-genome-annotation of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Prokka Genome Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prokka Genome Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prokka Genome Annotation this skilljaechang-hits/SciAgent-Skills3741 repos~6.1kAutomated safety check: PassGPL-3.0
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Bio Write SequencesGPTomics/bioSkills1.2k3 repos~2.1kAutomated safety check: PassMIT
ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates33k11 repos~4.5kAutomated safety check: NotesMIT
Biopythondavila7/claude-code-templates33k12 repos~3.4kAutomated safety check: PassMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • ETE Toolkit for Phylogenetic Trees

    davila7/claude-code-templates

    Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.

    33k GitHub starsUsed in 11 repos~4.5k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    davila7/claude-code-templates

    Query NCBI ClinVar for variant clinical significance. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 10 repos~3.3k tokens
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Works with

Questions about Prokka Genome Annotation

What does Prokka Genome Annotation do?

Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline. Prokka Genome Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate prokaryotic genomes (bacteria, archaea, viruses) via Prokka's BLAST/HMM pipeline.

When should I use Prokka Genome Annotation?

Prokka Genome Annotation fits situations like: tasks that involve Bioinformatics.

How do I install Prokka Genome Annotation in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/prokka-genome-annotation in jaechang-hits/SciAgent-Skills) into .claude/skills/prokka-genome-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Prokka Genome Annotation in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/prokka-genome-annotation in jaechang-hits/SciAgent-Skills) into .agents/skills/prokka-genome-annotation in your project. Codex loads it when a task matches its description.

Can I use Prokka Genome Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill prokka-genome-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prokka-genome-annotation, .gemini/skills/prokka-genome-annotation, .github/skills/prokka-genome-annotation and .opencode/skills/prokka-genome-annotation in your project.

What does Prokka Genome Annotation need to run?

Going by SKILL.md and its folder, Prokka Genome Annotation needs the command-line tools its instructions call (conda, pip and mamba). Our summary lists: Python 3.

Does Prokka Genome Annotation access the network?

SKILL.md names 3 domains. As links in the text: github.com, doi.org and biopython.org. This is read from the text; nothing was executed.

Is Prokka Genome Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prokka Genome Annotation use?

Prokka Genome Annotation is published under the GPL-3.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prokka Genome Annotation use?

About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prokka Genome Annotation?

Skills that share tags, products or a category with Prokka Genome Annotation: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars) and ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prokka Genome Annotation?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.