Agent skill

Roary Pangenome

by jaechang-hits in jaechang-hits/SciAgent-Skills

Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline.

GPL-3.0Auto-check passedResearch & Science

Install Roary Pangenome

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills roary-pangenome --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/roary-pangenome .claude/skills/roary-pangenome && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
roary-pangenome
GitHub stars
374
Used in
1 other repo
Token cost
~5.8k tokens
SKILL.md length
1,229 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
GPL-3.0

At a glance

Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline.

  • Works in 8 steps: Install Roary and Verify Dependencies → Prepare Per-Sample GFF3 Files from… → Run Roary with Core Gene Alignment → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 7 more sections
  • Calls mamba, pip and python; reaches github.com

What it does

Roary Pangenome is an agent skill from jaechang-hits/SciAgent-Skills. Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline. Builds gene presence/absence matrices, core/soft-core/shell/cloud partitions, multi-FASTA core gene alignments (with -e), and a pan-genome reference. Use Panaroo for higher-accuracy pan-genomes from highly fragmented assemblies, PIRATE for paralog-aware clustering, or PPanGGOLiN for graph-based partitioning.

Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/roary-pangenome”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Install Roary and Verify Dependencies
  2. Prepare Per-Sample GFF3 Files from Prokka or Bakta
  3. Run Roary with Core Gene Alignment
  4. Parse the Gene Presence/Absence Matrix
  5. Visualize the Pan-Genome Frequency Distribution
  6. Build a Phylogenetic Tree from the Core Gene Alignment
  7. Produce Roary's Built-in Summary Plots
  8. Compute Per-Genome Accessory Gene Counts

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • mamba
    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • doi.org
    • sanger-pathogens.github.io
    • microbesonline.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Roary Pangenome loads about 5.8k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 1,229 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~5.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its GPL-3.0 licence (© jaechang-hits). 1,229 words, ~5,770 tokens.

Download SKILL.mdSave it as .claude/skills/roary-pangenome/SKILL.md (or your agent's skills folder).
name
roary-pangenome
description
Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline. Builds gene presence/absence matrices, core/soft-core/shell/cloud partitions, multi-FASTA core gene alignments (with `-e`), and a pan-genome reference. Use Panaroo for higher-accuracy pan-genomes from highly fragmented assemblies, PIRATE for paralog-aware clustering, or PPanGGOLiN for graph-based partitioning.
license
GPL-3.0

Roary Pan-Genome Pipeline

Overview

Roary is a high-throughput pan-genome pipeline for prokaryotes that takes per-sample GFF3 annotations (typically from Prokka or Bakta) and produces a clustered gene presence/absence matrix across the entire input set. It first reduces redundancy with CD-HIT iterative clustering, then performs an all-vs-all BLASTP within each pre-cluster, and finally applies MCL graph clustering to define orthologous gene families. The output partitions the gene space into core (≥ 99 %), soft-core (95–99 %), shell (15–95 %), and cloud (< 15 %) genes and optionally builds a concatenated core-gene alignment suitable for phylogenetic inference.

When to Use

  • Computing a pan-genome from a set of bacterial isolate annotations (10–10,000 genomes)
  • Producing a gene_presence_absence.csv matrix for downstream GWAS, accessory-gene mining, or core-gene phylogenetics
  • Building a concatenated core-gene multi-FASTA alignment for ML/Bayesian phylogenetic trees
  • Generating a pan-genome reference FASTA to use as a non-redundant gene catalog
  • Comparative genomics across closely related strains where >95 % nucleotide identity is expected
  • Use Panaroo instead when assemblies are highly fragmented or annotations are noisy (Panaroo aggressively cleans annotation errors)
  • Use PIRATE instead when paralog-aware clustering with multiple identity thresholds is needed
  • Use PPanGGOLiN instead when graph-based, statistically grounded gene-family partitioning is preferred over fixed-frequency cutoffs

Prerequisites

  • Software: Roary ≥ 3.13, Perl 5, BLAST+, CD-HIT, MCL, BEDTools, PRANK or MAFFT (for -e core alignment), FastTree (optional)
  • Python packages (for output parsing): pandas, matplotlib, seaborn, biopython, dendropy
  • Input: per-sample GFF3 files with embedded FASTA at the end (Prokka/Bakta default output)
  • Hardware: ≥ 8 GB RAM for ~50 genomes; 32 GB+ recommended for ~500 genomes
  • Naming: each GFF3 filename becomes the sample column header in the output matrix; use sample_id.gff style

Check before installing: The tool may already be available in the current environment (e.g., inside a pixi / conda env). Run command -v roary first and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool via pixi run roary rather than bare roary.

bash
# Install Roary via conda/mamba (recommended)
mamba install -c conda-forge -c bioconda roary

# Verify installation
roary --version
# 3.13.0

# Verify dependent tools
which cd-hit blastp mcl bedtools mafft
# /opt/conda/bin/cd-hit
# /opt/conda/bin/blastp
# /opt/conda/bin/mcl
# /opt/conda/bin/bedtools
# /opt/conda/bin/mafft

# Install Python parsing dependencies
pip install pandas matplotlib seaborn biopython dendropy

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: inputAnnotations
    kind: derived
    source: upstream
    ask: "Which GFF3 files, and were they all produced by the same annotation tool and version?"
    default: "the annotation stage output"

  - id: D2
    param: identityThreshold
    kind: required
    source: user
    ask: "How similar must two proteins be to count as the same gene? Lower values merge diverged orthologs; higher values split them."
    default: "95%"

  - id: D3
    param: coreDefinition
    kind: required
    source: user
    ask: "In what fraction of isolates must a gene appear before it counts as core?"
    default: "99%"

  - id: D4
    param: paralogSplitting
    kind: required
    source: user
    ask: "Should duplicated genes within a genome be split into separate families, or kept together?"
    default: "split"

  - id: D5
    param: coreAlignment
    kind: required
    source: upstream
    ask: "Is a concatenated core-gene alignment needed for a downstream phylogeny?"
    default: "not produced - it is the slowest stage by a wide margin"

  - id: D6
    param: alignmentSpeed
    kind: optional
    source: user
    depends_on: [D5]
    ask: "Align every gene accurately, or use the fast route sufficient for tree building?"
    default: "fast alignment"
    skip_if: "no core alignment requested"

  - id: D7
    param: clusterInflation
    kind: optional
    source: user
    ask: "How tightly should the clustering step group families?"
    default: 1.5

  - id: D8
    param: maxClusters
    kind: optional_conditional
    source: data
    ask: "Is the expected number of gene families above the default cap?"
    default: 50000

  - id: D9
    param: processes
    kind: never_ask
    source: data
    reason: "Affects runtime only, not the gene families"
    default: "min(8, available_cores)"

D1 is derived and carries a consistency requirement rather than a choice: Roary compares annotations, so GFF3 files from different annotation tools or database versions produce apparent presence/absence differences that are annotation artefacts, not biology. D2 and D3 together decide the size of the core genome, which is usually the headline number of the analysis.

Quick Start

bash
# Run Roary on all GFF3 files in current directory; emit core gene alignment
roary -e --mafft -p 8 -o pangenome -f roary_out/ *.gff

# Inspect summary statistics
cat roary_out/summary_statistics.txt
# Core genes        (99% <= strains <= 100%) 2823
# Soft core genes   (95% <= strains <  99%)   78
# Shell genes       (15% <= strains <  95%)  1542
# Cloud genes       ( 0% <= strains <  15%)  3104
# Total genes                                 7547

Workflow

Step 1: Install Roary and Verify Dependencies

Install Roary in a dedicated environment to avoid Perl module conflicts.

bash
# Create a dedicated conda environment
mamba create -n roary_env -c conda-forge -c bioconda roary mafft fasttree python=3.11 -y
mamba activate roary_env

# Verify Roary
roary --version
# 3.13.0

# Confirm pipeline dependencies are reachable
for tool in cd-hit blastp mcl bedtools mafft FastTree; do
    if command -v $tool >/dev/null; then
        echo "OK: $tool"
    else
        echo "MISSING: $tool"
    fi
done

# Install Python parsing tools
pip install pandas matplotlib seaborn biopython dendropy
Step 2: Prepare Per-Sample GFF3 Files from Prokka or Bakta

Roary requires GFF3 files with an embedded FASTA section (the ##FASTA block). Both Prokka .gff and Bakta .gff3 outputs satisfy this format.

bash
# Verify GFF3 files include the embedded FASTA section
mkdir -p roary_input/
cp annotations/*/sample*.gff roary_input/

for GFF in roary_input/*.gff; do
    if grep -q "^##FASTA" "$GFF"; then
        echo "OK: $GFF"
    else
        echo "MISSING ##FASTA in: $GFF"
    fi
done

# Roary uses GFF filename (without .gff) as the sample column name.
# Rename or symlink to ensure unique, descriptive names.
cd roary_input/
ls *.gff | head
# strain_A.gff  strain_B.gff  strain_C.gff
python
# Sanity-check sample size and unique names
from pathlib import Path

gff_files = sorted(Path("roary_input/").glob("*.gff"))
print(f"GFF files: {len(gff_files)}")

names = [g.stem for g in gff_files]
duplicates = [n for n in names if names.count(n) > 1]
assert not duplicates, f"Duplicate sample names: {set(duplicates)}"

# Show first 5
for g in gff_files[:5]:
    size_mb = g.stat().st_size / 1e6
    print(f"  {g.name}: {size_mb:.1f} MB")
Step 3: Run Roary with Core Gene Alignment

The -e --mafft combination produces the concatenated core-gene alignment used downstream for phylogenetics.

bash
# Run Roary on the prepared GFF directory
# -e: extract core gene alignment per gene, then concatenate
# --mafft: use MAFFT for alignment (faster than default PRANK)
# -p 8: 8 parallel processes
# -i 95: min BLASTP percentage identity (default 95)
# -cd 99: % isolates a gene must be in to be "core" (default 99)
roary -e --mafft \
    -p 8 \
    -i 95 \
    -cd 99 \
    -o pangenome \
    -f roary_out/ \
    roary_input/*.gff

# Expected runtime:
#   ~50 genomes:  10–20 min
#   ~500 genomes: 4–8 hours

ls roary_out/
# accessory_binary_genes.fa            gene_presence_absence.csv
# accessory_binary_genes.fa.newick     gene_presence_absence.Rtab
# accessory.header.embl                number_of_conserved_genes.Rtab
# accessory.tab                        number_of_genes_in_pan_genome.Rtab
# blast_identity_frequency.Rtab        number_of_new_genes.Rtab
# clustered_proteins                   number_of_unique_genes.Rtab
# core_accessory.header.embl           pan_genome_reference.fa
# core_accessory.tab                   summary_statistics.txt
# core_gene_alignment.aln
Step 4: Parse the Gene Presence/Absence Matrix

The CSV output is the canonical pan-genome matrix: rows are gene families, columns include metadata and one column per sample (filled with locus tags or empty).

python
import pandas as pd
import numpy as np

df = pd.read_csv("roary_out/gene_presence_absence.csv", low_memory=False)
print(f"Gene families: {len(df)}")
print(f"Columns: {len(df.columns)}")
# Standard metadata columns end with 'Inference', then sample columns follow
metadata_cols = list(df.columns[:14])
sample_cols = list(df.columns[14:])
print(f"Samples: {len(sample_cols)}")

# Build a binary presence/absence matrix (1 = locus tag present, 0 = empty)
binary = df[sample_cols].notna().astype(int)
binary.index = df["Gene"]
print(f"Binary matrix shape: {binary.shape}")

# Pan-genome partition by frequency
n_samples = len(sample_cols)
freq = binary.sum(axis=1) / n_samples
core      = freq[freq >= 0.99]
soft_core = freq[(freq >= 0.95) & (freq < 0.99)]
shell     = freq[(freq >= 0.15) & (freq < 0.95)]
cloud     = freq[freq < 0.15]

print(f"\nPan-genome partition (n={n_samples} genomes):")
print(f"  Core      (>=99 %): {len(core):>5}")
print(f"  Soft-core (95–99 %): {len(soft_core):>5}")
print(f"  Shell    (15–95 %): {len(shell):>5}")
print(f"  Cloud      (<15 %): {len(cloud):>5}")
print(f"  Total            : {len(freq):>5}")
Step 5: Visualize the Pan-Genome Frequency Distribution

A frequency histogram exposes whether the input set has a U-shaped (open) or core-skewed (closed) pan-genome.

python
import pandas as pd
import matplotlib.pyplot as plt
import numpy as np

df = pd.read_csv("roary_out/gene_presence_absence.csv", low_memory=False)
sample_cols = list(df.columns[14:])
n = len(sample_cols)
binary = df[sample_cols].notna().astype(int)
freq = binary.sum(axis=1)

fig, axes = plt.subplots(1, 2, figsize=(11, 4))

# Histogram of gene frequencies
axes[0].hist(freq, bins=range(1, n + 2), color="#2980B9", edgecolor="white")
axes[0].axvline(0.99 * n, color="red", linestyle="--", label="99 % core cutoff")
axes[0].axvline(0.15 * n, color="orange", linestyle="--", label="15 % shell cutoff")
axes[0].set_xlabel("Number of genomes containing the gene")
axes[0].set_ylabel("Number of gene families")
axes[0].set_title("Pan-genome frequency distribution")
axes[0].legend()

# Pie chart of pan-genome partition
freqp = freq / n
core_n      = (freqp >= 0.99).sum()
soft_n      = ((freqp >= 0.95) & (freqp < 0.99)).sum()
shell_n     = ((freqp >= 0.15) & (freqp < 0.95)).sum()
cloud_n     = (freqp < 0.15).sum()
axes[1].pie([core_n, soft_n, shell_n, cloud_n],
            labels=[f"Core ({core_n})", f"Soft-core ({soft_n})",
                    f"Shell ({shell_n})", f"Cloud ({cloud_n})"],
            colors=["#27AE60", "#3498DB", "#F39C12", "#95A5A6"],
            autopct="%1.1f%%", startangle=90)
axes[1].set_title(f"Pan-genome partition (n={n} genomes)")

plt.tight_layout()
plt.savefig("pangenome_distribution.png", dpi=150, bbox_inches="tight")
print(f"Saved pangenome_distribution.png  (total families: {len(freq)})")
Step 6: Build a Phylogenetic Tree from the Core Gene Alignment

Use FastTree on the concatenated core-gene alignment to recover strain relationships.

bash
# FastTree expects an aligned multi-FASTA; Roary produces this with `-e`
FastTree -nt -gtr -nosupport \
    < roary_out/core_gene_alignment.aln \
    > roary_out/core_gene_tree.nwk

echo "Tree built: roary_out/core_gene_tree.nwk"
# (Optional) IQ-TREE for support values:
# iqtree -s core_gene_alignment.aln -m GTR+G -bb 1000 -nt 8
python
# Visualize the FastTree result alongside the gene presence/absence matrix
import dendropy
import matplotlib.pyplot as plt
import pandas as pd
import numpy as np

tree = dendropy.Tree.get(path="roary_out/core_gene_tree.nwk", schema="newick")
leaves = [leaf.taxon.label for leaf in tree.leaf_node_iter()]
print(f"Tree leaves: {len(leaves)}")
print(f"Tree length: {tree.length():.4f} substitutions/site")

# Print first 10 leaves with their root-to-tip distances
for leaf in list(tree.leaf_node_iter())[:10]:
    dist = leaf.distance_from_root()
    print(f"  {leaf.taxon.label:<25}  d={dist:.4f}")
Step 7: Produce Roary's Built-in Summary Plots

Roary ships with a roary_plots.py companion that builds a tree-anchored matrix figure.

bash
# Use the bundled python helper from the Roary install
# (download from: https://github.com/sanger-pathogens/Roary/blob/master/contrib/roary_plots/roary_plots.py)
python roary_plots.py \
    roary_out/core_gene_tree.nwk \
    roary_out/gene_presence_absence.csv \
    --format png \
    --labels

ls *.png
# pangenome_frequency.png
# pangenome_matrix.png
# pangenome_pie.png
Step 8: Compute Per-Genome Accessory Gene Counts

Identify which genomes carry the most unique (cloud) genes — useful for strain-specific gene investigation.

python
import pandas as pd

df = pd.read_csv("roary_out/gene_presence_absence.csv", low_memory=False)
sample_cols = list(df.columns[14:])
n = len(sample_cols)
binary = df[sample_cols].notna().astype(int)
binary.index = df["Gene"]

freq = binary.sum(axis=1)
cloud_mask = freq < (0.15 * n)
shell_mask = (freq >= 0.15 * n) & (freq < 0.95 * n)

per_genome = pd.DataFrame({
    "genes_total":   binary.sum(axis=0),
    "cloud_genes":   binary.loc[cloud_mask].sum(axis=0),
    "shell_genes":   binary.loc[shell_mask].sum(axis=0),
})
per_genome = per_genome.sort_values("cloud_genes", ascending=False)
print(per_genome.head(10).to_string())

per_genome.to_csv("per_genome_accessory_counts.csv")
print("\nSaved per_genome_accessory_counts.csv")

Key Parameters

ParameterDefaultRange / OptionsEffect
-i9570–99 (% identity)Minimum BLASTP percentage identity for clustering
-cd9990–100 (%)Minimum % of isolates a gene must be in to be called "core"
-p11–CPU countNumber of parallel processes for BLAST and MCL stages
-eoffflagExtract concatenated core gene multi-FASTA alignment (per-gene MSA + concat)
--mafftoffflag (with -e)Use MAFFT for alignment (faster than default PRANK)
-soffflagDo not split paralogs into separate gene families
-noffflagUse fast core gene alignment with MAFFT (skips per-gene PRANK)
-g50000any integerMaximum number of clusters expected (cap on output size)
-iv1.51.2–5.0MCL inflation value; higher = tighter clusters
-roffflagGenerate R plots (presence/absence heatmap, pan-genome curves)
-oclustered_proteinsstringOutput prefix for clustered proteins file
-f.pathOutput directory
-zoffflagDon't delete intermediate files (useful for debugging)

Common Recipes

Show full SKILL.md (503 more words)Show less
Recipe: Core Gene Alignment from Existing Roary Output

When to use: You already ran Roary without -e and now need a phylogenetic alignment.

bash
# Re-run only the alignment step using the extracted core genes
# query_pan_genome from the Roary suite extracts gene-family multi-FASTAs
query_pan_genome -a intersection \
    -g roary_out/clustered_proteins \
    -o core_gene_list.txt \
    roary_input/*.gff

echo "Core genes: $(wc -l < core_gene_list.txt)"
# Then re-run Roary with -e to materialize the alignment:
# roary -e --mafft -p 8 -f roary_out2/ roary_input/*.gff
Recipe: Lower Identity Threshold for Diverse Genus-Level Sets

When to use: Comparing genomes across multiple closely related species (e.g., genus-level pan-genome).

bash
# Drop identity to 70 % to capture orthologs across species boundaries
roary -e --mafft \
    -p 8 \
    -i 70 \
    -cd 99 \
    -f roary_genus/ \
    roary_input/*.gff

cat roary_genus/summary_statistics.txt
# Expect a smaller core and a much larger shell + cloud
Recipe: Roary on a Subset of Strains

When to use: You want a pan-genome of a specific subclade without re-annotating all genomes.

bash
# Stage GFF files for the subset
mkdir -p subset/
for SAMPLE in strain_A strain_B strain_E strain_K; do
    cp roary_input/${SAMPLE}.gff subset/
done

roary -e --mafft -p 4 -f roary_subset/ subset/*.gff
echo "Subset pan-genome:"
cat roary_subset/summary_statistics.txt
Recipe: Cross-Tabulate Accessory Genes Against Strain Metadata

When to use: You want to find genes enriched in a specific phenotype or geographic group.

python
import pandas as pd
from scipy.stats import fisher_exact

# Load presence/absence and strain metadata
pa = pd.read_csv("roary_out/gene_presence_absence.csv", low_memory=False)
sample_cols = list(pa.columns[14:])
binary = pa[sample_cols].notna().astype(int)
binary.index = pa["Gene"]
binary.columns = sample_cols

# Metadata: sample_id, phenotype (e.g., "resistant" / "susceptible")
meta = pd.read_csv("strain_metadata.csv").set_index("sample_id")
group_resistant = meta[meta["phenotype"] == "resistant"].index
group_susceptible = meta[meta["phenotype"] == "susceptible"].index
group_resistant = [g for g in group_resistant if g in binary.columns]
group_susceptible = [g for g in group_susceptible if g in binary.columns]

# Fisher's exact test per gene family
results = []
for gene, row in binary.iterrows():
    a = int(row[group_resistant].sum())
    b = len(group_resistant) - a
    c = int(row[group_susceptible].sum())
    d = len(group_susceptible) - c
    if a + c == 0 or a + c == len(group_resistant) + len(group_susceptible):
        continue  # skip invariant genes
    odds, pval = fisher_exact([[a, b], [c, d]])
    results.append({"gene": gene, "n_resistant": a, "n_susceptible": c,
                    "odds_ratio": odds, "pvalue": pval})

results_df = pd.DataFrame(results).sort_values("pvalue")
print(f"Gene-phenotype associations (top 10):")
print(results_df.head(10).to_string(index=False))
results_df.to_csv("phenotype_associated_genes.csv", index=False)

Expected Outputs

Output FileFormatDescription
gene_presence_absence.csvCSVPan-genome matrix; rows = gene families, columns = sample locus tags
gene_presence_absence.RtabTSVBinary 0/1 matrix in R-friendly format
summary_statistics.txtTextCounts of core / soft-core / shell / cloud / total gene families
pan_genome_reference.faFASTANon-redundant nucleotide reference of one representative per gene family
clustered_proteinsTextCluster definitions: cluster name → constituent locus tags
core_gene_alignment.alnFASTA (aligned)Concatenated multi-FASTA of core gene alignments (only with -e)
accessory_binary_genes.faFASTABinary 0/1 sequences of accessory presence (input to FastTree)
accessory_binary_genes.fa.newickNewickAccessory gene phylogeny (built from binary FASTA)
number_of_conserved_genes.RtabTSVRarefaction curve of core genes vs. genome count
number_of_genes_in_pan_genome.RtabTSVRarefaction curve of pan-genome growth
number_of_new_genes.RtabTSVNew genes added per genome
number_of_unique_genes.RtabTSVPer-genome unique gene counts
blast_identity_frequency.RtabTSVDistribution of BLAST identities used during clustering

Troubleshooting

ProblemCauseSolution
GFF file does not contain ##FASTAAnnotation file is not a Prokka/Bakta-style GFF3 with embedded sequencesRe-export from Prokka/Bakta; do not strip the FASTA tail
Could not run blastpBLAST+ missing from the conda envmamba install -c bioconda blast and re-activate the env
Roary produces 0 core genesOne sample has very different annotations or is misidentifiedInspect gene_presence_absence.csv; remove outlier with roary --check_input style sanity checks
Job runs out of memory on > 200 genomesDefault identity (95 %) creates many BLAST jobsIncrease -i to 90 if accuracy permits; split into batches; use Panaroo on RAM-tight machines
Paralogs split into many tiny clustersMCL inflation too high or duplicate locus tags across samplesAdd -s to keep paralogs together, or rerun Prokka/Bakta with unique --locus-tag per sample
FastTree: bad alignment on core_gene_alignment.alnAlignment built with PRANK was truncatedRe-run with --mafft for MAFFT-based alignment
Output column names look mangled or missingGFF filenames had spaces or special charactersRename input GFFs to [a-zA-Z0-9_]+.gff before running
Same genome counted twiceDuplicate GFF filenames (e.g., copied symlinks)Run `ls roary_input/*.gff
Pan-genome size grows linearly without saturationTruly open pan-genome (expected for many bacteria)This is biological, not a bug; report rarefaction curves and use Heaps' law for confirmation

References

© jaechang-hits, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/annotation/roary-pangenome of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Roary Pangenome next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Roary Pangenome compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Roary Pangenome this skilljaechang-hits/SciAgent-Skills3741 repos~5.8kAutomated safety check: PassGPL-3.0
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Roary Pangenome

What does Roary Pangenome do?

Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline. Roary Pangenome is an agent skill from jaechang-hits/SciAgent-Skills. Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline.

When should I use Roary Pangenome?

Roary Pangenome fits situations like: tasks that involve Bioinformatics.

How do I install Roary Pangenome in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/roary-pangenome in jaechang-hits/SciAgent-Skills) into .claude/skills/roary-pangenome in your project. Claude Code loads it when a task matches its description.

How do I install Roary Pangenome in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/roary-pangenome in jaechang-hits/SciAgent-Skills) into .agents/skills/roary-pangenome in your project. Codex loads it when a task matches its description.

Can I use Roary Pangenome in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/roary-pangenome, .gemini/skills/roary-pangenome, .github/skills/roary-pangenome and .opencode/skills/roary-pangenome in your project.

What does Roary Pangenome need to run?

Going by SKILL.md and its folder, Roary Pangenome needs the command-line tools its instructions call (mamba, pip and python). Our summary lists: Python 3.

Does Roary Pangenome access the network?

SKILL.md names 4 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: doi.org, sanger-pathogens.github.io and microbesonline.org. This is read from the text; nothing was executed.

Is Roary Pangenome safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Roary Pangenome use?

Roary Pangenome is published under the GPL-3.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Roary Pangenome use?

About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Roary Pangenome?

Skills that share tags, products or a category with Roary Pangenome: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Roary Pangenome?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.