Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline.
$ npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills roary-pangenome --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/roary-pangenome .claude/skills/roary-pangenome && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "roary-pangenome" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/roary-pangenome into .claude/skills/roary-pangenome/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "roary-pangenome", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/roary-pangenomeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills roary-pangenome --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/roary-pangenome .agents/skills/roary-pangenome && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "roary-pangenome" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/roary-pangenome into .agents/skills/roary-pangenome/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "roary-pangenome", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills roary-pangenome --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/roary-pangenome .cursor/skills/roary-pangenome && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "roary-pangenome" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/roary-pangenome into .cursor/skills/roary-pangenome/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "roary-pangenome", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/annotation/roary-pangenome--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills roary-pangenome --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/roary-pangenome .gemini/skills/roary-pangenome && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "roary-pangenome" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/roary-pangenome into .gemini/skills/roary-pangenome/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "roary-pangenome", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills roary-pangenomeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/roary-pangenome .github/skills/roary-pangenome && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "roary-pangenome" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/roary-pangenome into .github/skills/roary-pangenome/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "roary-pangenome", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills roary-pangenome --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/annotation/roary-pangenome .opencode/skills/roary-pangenome && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "roary-pangenome" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/annotation/roary-pangenome into .opencode/skills/roary-pangenome/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "roary-pangenome", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
roary-pangenomeCompute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline.
Roary Pangenome is an agent skill from jaechang-hits/SciAgent-Skills. Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline. Builds gene presence/absence matrices, core/soft-core/shell/cloud partitions, multi-FASTA core gene alignments (with -e), and a pan-genome reference. Use Panaroo for higher-accuracy pan-genomes from highly fragmented assemblies, PIRATE for paralog-aware clustering, or PPanGGOLiN for graph-based partitioning.
Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is GPL-3.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
mambapippythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comAlso links to:
doi.orgsanger-pathogens.github.iomicrobesonline.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Roary Pangenome loads about 5.8k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 1,229 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its GPL-3.0 licence (© jaechang-hits). 1,229 words, ~5,770 tokens.
.claude/skills/roary-pangenome/SKILL.md (or your agent's skills folder).Roary is a high-throughput pan-genome pipeline for prokaryotes that takes per-sample GFF3 annotations (typically from Prokka or Bakta) and produces a clustered gene presence/absence matrix across the entire input set. It first reduces redundancy with CD-HIT iterative clustering, then performs an all-vs-all BLASTP within each pre-cluster, and finally applies MCL graph clustering to define orthologous gene families. The output partitions the gene space into core (≥ 99 %), soft-core (95–99 %), shell (15–95 %), and cloud (< 15 %) genes and optionally builds a concatenated core-gene alignment suitable for phylogenetic inference.
gene_presence_absence.csv matrix for downstream GWAS, accessory-gene mining, or core-gene phylogenetics-e core alignment), FastTree (optional)pandas, matplotlib, seaborn, biopython, dendropysample_id.gff styleCheck before installing: The tool may already be available in the current environment (e.g., inside a
pixi/condaenv). Runcommand -v roaryfirst and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool viapixi run roaryrather than bareroary.
# Install Roary via conda/mamba (recommended)
mamba install -c conda-forge -c bioconda roary
# Verify installation
roary --version
# 3.13.0
# Verify dependent tools
which cd-hit blastp mcl bedtools mafft
# /opt/conda/bin/cd-hit
# /opt/conda/bin/blastp
# /opt/conda/bin/mcl
# /opt/conda/bin/bedtools
# /opt/conda/bin/mafft
# Install Python parsing dependencies
pip install pandas matplotlib seaborn biopython dendropySettle these with the user before writing any analysis code.
decisions:
- id: D1
param: inputAnnotations
kind: derived
source: upstream
ask: "Which GFF3 files, and were they all produced by the same annotation tool and version?"
default: "the annotation stage output"
- id: D2
param: identityThreshold
kind: required
source: user
ask: "How similar must two proteins be to count as the same gene? Lower values merge diverged orthologs; higher values split them."
default: "95%"
- id: D3
param: coreDefinition
kind: required
source: user
ask: "In what fraction of isolates must a gene appear before it counts as core?"
default: "99%"
- id: D4
param: paralogSplitting
kind: required
source: user
ask: "Should duplicated genes within a genome be split into separate families, or kept together?"
default: "split"
- id: D5
param: coreAlignment
kind: required
source: upstream
ask: "Is a concatenated core-gene alignment needed for a downstream phylogeny?"
default: "not produced - it is the slowest stage by a wide margin"
- id: D6
param: alignmentSpeed
kind: optional
source: user
depends_on: [D5]
ask: "Align every gene accurately, or use the fast route sufficient for tree building?"
default: "fast alignment"
skip_if: "no core alignment requested"
- id: D7
param: clusterInflation
kind: optional
source: user
ask: "How tightly should the clustering step group families?"
default: 1.5
- id: D8
param: maxClusters
kind: optional_conditional
source: data
ask: "Is the expected number of gene families above the default cap?"
default: 50000
- id: D9
param: processes
kind: never_ask
source: data
reason: "Affects runtime only, not the gene families"
default: "min(8, available_cores)"D1 is derived and carries a consistency requirement rather than a choice: Roary compares annotations, so GFF3 files from different annotation tools or database versions produce apparent presence/absence differences that are annotation artefacts, not biology. D2 and D3 together decide the size of the core genome, which is usually the headline number of the analysis.
# Run Roary on all GFF3 files in current directory; emit core gene alignment
roary -e --mafft -p 8 -o pangenome -f roary_out/ *.gff
# Inspect summary statistics
cat roary_out/summary_statistics.txt
# Core genes (99% <= strains <= 100%) 2823
# Soft core genes (95% <= strains < 99%) 78
# Shell genes (15% <= strains < 95%) 1542
# Cloud genes ( 0% <= strains < 15%) 3104
# Total genes 7547Install Roary in a dedicated environment to avoid Perl module conflicts.
# Create a dedicated conda environment
mamba create -n roary_env -c conda-forge -c bioconda roary mafft fasttree python=3.11 -y
mamba activate roary_env
# Verify Roary
roary --version
# 3.13.0
# Confirm pipeline dependencies are reachable
for tool in cd-hit blastp mcl bedtools mafft FastTree; do
if command -v $tool >/dev/null; then
echo "OK: $tool"
else
echo "MISSING: $tool"
fi
done
# Install Python parsing tools
pip install pandas matplotlib seaborn biopython dendropyRoary requires GFF3 files with an embedded FASTA section (the ##FASTA block). Both Prokka .gff and Bakta .gff3 outputs satisfy this format.
# Verify GFF3 files include the embedded FASTA section
mkdir -p roary_input/
cp annotations/*/sample*.gff roary_input/
for GFF in roary_input/*.gff; do
if grep -q "^##FASTA" "$GFF"; then
echo "OK: $GFF"
else
echo "MISSING ##FASTA in: $GFF"
fi
done
# Roary uses GFF filename (without .gff) as the sample column name.
# Rename or symlink to ensure unique, descriptive names.
cd roary_input/
ls *.gff | head
# strain_A.gff strain_B.gff strain_C.gff# Sanity-check sample size and unique names
from pathlib import Path
gff_files = sorted(Path("roary_input/").glob("*.gff"))
print(f"GFF files: {len(gff_files)}")
names = [g.stem for g in gff_files]
duplicates = [n for n in names if names.count(n) > 1]
assert not duplicates, f"Duplicate sample names: {set(duplicates)}"
# Show first 5
for g in gff_files[:5]:
size_mb = g.stat().st_size / 1e6
print(f" {g.name}: {size_mb:.1f} MB")The -e --mafft combination produces the concatenated core-gene alignment used downstream for phylogenetics.
# Run Roary on the prepared GFF directory
# -e: extract core gene alignment per gene, then concatenate
# --mafft: use MAFFT for alignment (faster than default PRANK)
# -p 8: 8 parallel processes
# -i 95: min BLASTP percentage identity (default 95)
# -cd 99: % isolates a gene must be in to be "core" (default 99)
roary -e --mafft \
-p 8 \
-i 95 \
-cd 99 \
-o pangenome \
-f roary_out/ \
roary_input/*.gff
# Expected runtime:
# ~50 genomes: 10–20 min
# ~500 genomes: 4–8 hours
ls roary_out/
# accessory_binary_genes.fa gene_presence_absence.csv
# accessory_binary_genes.fa.newick gene_presence_absence.Rtab
# accessory.header.embl number_of_conserved_genes.Rtab
# accessory.tab number_of_genes_in_pan_genome.Rtab
# blast_identity_frequency.Rtab number_of_new_genes.Rtab
# clustered_proteins number_of_unique_genes.Rtab
# core_accessory.header.embl pan_genome_reference.fa
# core_accessory.tab summary_statistics.txt
# core_gene_alignment.alnThe CSV output is the canonical pan-genome matrix: rows are gene families, columns include metadata and one column per sample (filled with locus tags or empty).
import pandas as pd
import numpy as np
df = pd.read_csv("roary_out/gene_presence_absence.csv", low_memory=False)
print(f"Gene families: {len(df)}")
print(f"Columns: {len(df.columns)}")
# Standard metadata columns end with 'Inference', then sample columns follow
metadata_cols = list(df.columns[:14])
sample_cols = list(df.columns[14:])
print(f"Samples: {len(sample_cols)}")
# Build a binary presence/absence matrix (1 = locus tag present, 0 = empty)
binary = df[sample_cols].notna().astype(int)
binary.index = df["Gene"]
print(f"Binary matrix shape: {binary.shape}")
# Pan-genome partition by frequency
n_samples = len(sample_cols)
freq = binary.sum(axis=1) / n_samples
core = freq[freq >= 0.99]
soft_core = freq[(freq >= 0.95) & (freq < 0.99)]
shell = freq[(freq >= 0.15) & (freq < 0.95)]
cloud = freq[freq < 0.15]
print(f"\nPan-genome partition (n={n_samples} genomes):")
print(f" Core (>=99 %): {len(core):>5}")
print(f" Soft-core (95–99 %): {len(soft_core):>5}")
print(f" Shell (15–95 %): {len(shell):>5}")
print(f" Cloud (<15 %): {len(cloud):>5}")
print(f" Total : {len(freq):>5}")A frequency histogram exposes whether the input set has a U-shaped (open) or core-skewed (closed) pan-genome.
import pandas as pd
import matplotlib.pyplot as plt
import numpy as np
df = pd.read_csv("roary_out/gene_presence_absence.csv", low_memory=False)
sample_cols = list(df.columns[14:])
n = len(sample_cols)
binary = df[sample_cols].notna().astype(int)
freq = binary.sum(axis=1)
fig, axes = plt.subplots(1, 2, figsize=(11, 4))
# Histogram of gene frequencies
axes[0].hist(freq, bins=range(1, n + 2), color="#2980B9", edgecolor="white")
axes[0].axvline(0.99 * n, color="red", linestyle="--", label="99 % core cutoff")
axes[0].axvline(0.15 * n, color="orange", linestyle="--", label="15 % shell cutoff")
axes[0].set_xlabel("Number of genomes containing the gene")
axes[0].set_ylabel("Number of gene families")
axes[0].set_title("Pan-genome frequency distribution")
axes[0].legend()
# Pie chart of pan-genome partition
freqp = freq / n
core_n = (freqp >= 0.99).sum()
soft_n = ((freqp >= 0.95) & (freqp < 0.99)).sum()
shell_n = ((freqp >= 0.15) & (freqp < 0.95)).sum()
cloud_n = (freqp < 0.15).sum()
axes[1].pie([core_n, soft_n, shell_n, cloud_n],
labels=[f"Core ({core_n})", f"Soft-core ({soft_n})",
f"Shell ({shell_n})", f"Cloud ({cloud_n})"],
colors=["#27AE60", "#3498DB", "#F39C12", "#95A5A6"],
autopct="%1.1f%%", startangle=90)
axes[1].set_title(f"Pan-genome partition (n={n} genomes)")
plt.tight_layout()
plt.savefig("pangenome_distribution.png", dpi=150, bbox_inches="tight")
print(f"Saved pangenome_distribution.png (total families: {len(freq)})")Use FastTree on the concatenated core-gene alignment to recover strain relationships.
# FastTree expects an aligned multi-FASTA; Roary produces this with `-e`
FastTree -nt -gtr -nosupport \
< roary_out/core_gene_alignment.aln \
> roary_out/core_gene_tree.nwk
echo "Tree built: roary_out/core_gene_tree.nwk"
# (Optional) IQ-TREE for support values:
# iqtree -s core_gene_alignment.aln -m GTR+G -bb 1000 -nt 8# Visualize the FastTree result alongside the gene presence/absence matrix
import dendropy
import matplotlib.pyplot as plt
import pandas as pd
import numpy as np
tree = dendropy.Tree.get(path="roary_out/core_gene_tree.nwk", schema="newick")
leaves = [leaf.taxon.label for leaf in tree.leaf_node_iter()]
print(f"Tree leaves: {len(leaves)}")
print(f"Tree length: {tree.length():.4f} substitutions/site")
# Print first 10 leaves with their root-to-tip distances
for leaf in list(tree.leaf_node_iter())[:10]:
dist = leaf.distance_from_root()
print(f" {leaf.taxon.label:<25} d={dist:.4f}")Roary ships with a roary_plots.py companion that builds a tree-anchored matrix figure.
# Use the bundled python helper from the Roary install
# (download from: https://github.com/sanger-pathogens/Roary/blob/master/contrib/roary_plots/roary_plots.py)
python roary_plots.py \
roary_out/core_gene_tree.nwk \
roary_out/gene_presence_absence.csv \
--format png \
--labels
ls *.png
# pangenome_frequency.png
# pangenome_matrix.png
# pangenome_pie.pngIdentify which genomes carry the most unique (cloud) genes — useful for strain-specific gene investigation.
import pandas as pd
df = pd.read_csv("roary_out/gene_presence_absence.csv", low_memory=False)
sample_cols = list(df.columns[14:])
n = len(sample_cols)
binary = df[sample_cols].notna().astype(int)
binary.index = df["Gene"]
freq = binary.sum(axis=1)
cloud_mask = freq < (0.15 * n)
shell_mask = (freq >= 0.15 * n) & (freq < 0.95 * n)
per_genome = pd.DataFrame({
"genes_total": binary.sum(axis=0),
"cloud_genes": binary.loc[cloud_mask].sum(axis=0),
"shell_genes": binary.loc[shell_mask].sum(axis=0),
})
per_genome = per_genome.sort_values("cloud_genes", ascending=False)
print(per_genome.head(10).to_string())
per_genome.to_csv("per_genome_accessory_counts.csv")
print("\nSaved per_genome_accessory_counts.csv")| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
-i | 95 | 70–99 (% identity) | Minimum BLASTP percentage identity for clustering |
-cd | 99 | 90–100 (%) | Minimum % of isolates a gene must be in to be called "core" |
-p | 1 | 1–CPU count | Number of parallel processes for BLAST and MCL stages |
-e | off | flag | Extract concatenated core gene multi-FASTA alignment (per-gene MSA + concat) |
--mafft | off | flag (with -e) | Use MAFFT for alignment (faster than default PRANK) |
-s | off | flag | Do not split paralogs into separate gene families |
-n | off | flag | Use fast core gene alignment with MAFFT (skips per-gene PRANK) |
-g | 50000 | any integer | Maximum number of clusters expected (cap on output size) |
-iv | 1.5 | 1.2–5.0 | MCL inflation value; higher = tighter clusters |
-r | off | flag | Generate R plots (presence/absence heatmap, pan-genome curves) |
-o | clustered_proteins | string | Output prefix for clustered proteins file |
-f | . | path | Output directory |
-z | off | flag | Don't delete intermediate files (useful for debugging) |
When to use: You already ran Roary without -e and now need a phylogenetic alignment.
# Re-run only the alignment step using the extracted core genes
# query_pan_genome from the Roary suite extracts gene-family multi-FASTAs
query_pan_genome -a intersection \
-g roary_out/clustered_proteins \
-o core_gene_list.txt \
roary_input/*.gff
echo "Core genes: $(wc -l < core_gene_list.txt)"
# Then re-run Roary with -e to materialize the alignment:
# roary -e --mafft -p 8 -f roary_out2/ roary_input/*.gffWhen to use: Comparing genomes across multiple closely related species (e.g., genus-level pan-genome).
# Drop identity to 70 % to capture orthologs across species boundaries
roary -e --mafft \
-p 8 \
-i 70 \
-cd 99 \
-f roary_genus/ \
roary_input/*.gff
cat roary_genus/summary_statistics.txt
# Expect a smaller core and a much larger shell + cloudWhen to use: You want a pan-genome of a specific subclade without re-annotating all genomes.
# Stage GFF files for the subset
mkdir -p subset/
for SAMPLE in strain_A strain_B strain_E strain_K; do
cp roary_input/${SAMPLE}.gff subset/
done
roary -e --mafft -p 4 -f roary_subset/ subset/*.gff
echo "Subset pan-genome:"
cat roary_subset/summary_statistics.txtWhen to use: You want to find genes enriched in a specific phenotype or geographic group.
import pandas as pd
from scipy.stats import fisher_exact
# Load presence/absence and strain metadata
pa = pd.read_csv("roary_out/gene_presence_absence.csv", low_memory=False)
sample_cols = list(pa.columns[14:])
binary = pa[sample_cols].notna().astype(int)
binary.index = pa["Gene"]
binary.columns = sample_cols
# Metadata: sample_id, phenotype (e.g., "resistant" / "susceptible")
meta = pd.read_csv("strain_metadata.csv").set_index("sample_id")
group_resistant = meta[meta["phenotype"] == "resistant"].index
group_susceptible = meta[meta["phenotype"] == "susceptible"].index
group_resistant = [g for g in group_resistant if g in binary.columns]
group_susceptible = [g for g in group_susceptible if g in binary.columns]
# Fisher's exact test per gene family
results = []
for gene, row in binary.iterrows():
a = int(row[group_resistant].sum())
b = len(group_resistant) - a
c = int(row[group_susceptible].sum())
d = len(group_susceptible) - c
if a + c == 0 or a + c == len(group_resistant) + len(group_susceptible):
continue # skip invariant genes
odds, pval = fisher_exact([[a, b], [c, d]])
results.append({"gene": gene, "n_resistant": a, "n_susceptible": c,
"odds_ratio": odds, "pvalue": pval})
results_df = pd.DataFrame(results).sort_values("pvalue")
print(f"Gene-phenotype associations (top 10):")
print(results_df.head(10).to_string(index=False))
results_df.to_csv("phenotype_associated_genes.csv", index=False)| Output File | Format | Description |
|---|---|---|
gene_presence_absence.csv | CSV | Pan-genome matrix; rows = gene families, columns = sample locus tags |
gene_presence_absence.Rtab | TSV | Binary 0/1 matrix in R-friendly format |
summary_statistics.txt | Text | Counts of core / soft-core / shell / cloud / total gene families |
pan_genome_reference.fa | FASTA | Non-redundant nucleotide reference of one representative per gene family |
clustered_proteins | Text | Cluster definitions: cluster name → constituent locus tags |
core_gene_alignment.aln | FASTA (aligned) | Concatenated multi-FASTA of core gene alignments (only with -e) |
accessory_binary_genes.fa | FASTA | Binary 0/1 sequences of accessory presence (input to FastTree) |
accessory_binary_genes.fa.newick | Newick | Accessory gene phylogeny (built from binary FASTA) |
number_of_conserved_genes.Rtab | TSV | Rarefaction curve of core genes vs. genome count |
number_of_genes_in_pan_genome.Rtab | TSV | Rarefaction curve of pan-genome growth |
number_of_new_genes.Rtab | TSV | New genes added per genome |
number_of_unique_genes.Rtab | TSV | Per-genome unique gene counts |
blast_identity_frequency.Rtab | TSV | Distribution of BLAST identities used during clustering |
| Problem | Cause | Solution |
|---|---|---|
GFF file does not contain ##FASTA | Annotation file is not a Prokka/Bakta-style GFF3 with embedded sequences | Re-export from Prokka/Bakta; do not strip the FASTA tail |
Could not run blastp | BLAST+ missing from the conda env | mamba install -c bioconda blast and re-activate the env |
| Roary produces 0 core genes | One sample has very different annotations or is misidentified | Inspect gene_presence_absence.csv; remove outlier with roary --check_input style sanity checks |
| Job runs out of memory on > 200 genomes | Default identity (95 %) creates many BLAST jobs | Increase -i to 90 if accuracy permits; split into batches; use Panaroo on RAM-tight machines |
| Paralogs split into many tiny clusters | MCL inflation too high or duplicate locus tags across samples | Add -s to keep paralogs together, or rerun Prokka/Bakta with unique --locus-tag per sample |
FastTree: bad alignment on core_gene_alignment.aln | Alignment built with PRANK was truncated | Re-run with --mafft for MAFFT-based alignment |
| Output column names look mangled or missing | GFF filenames had spaces or special characters | Rename input GFFs to [a-zA-Z0-9_]+.gff before running |
| Same genome counted twice | Duplicate GFF filenames (e.g., copied symlinks) | Run `ls roary_input/*.gff |
| Pan-genome size grows linearly without saturation | Truly open pan-genome (expected for many bacteria) | This is biological, not a bug; report rarefaction curves and use Heaps' law for confirmation |
© jaechang-hits, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/annotation/roary-pangenome of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Roary Pangenome next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Roary Pangenome this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~5.8k | Automated safety check: Pass | GPL-3.0 | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline. Roary Pangenome is an agent skill from jaechang-hits/SciAgent-Skills. Compute the bacterial pan-genome from Prokka/Bakta GFF3 annotations with Roary's CD-HIT + BLAST + MCL clustering pipeline.
Roary Pangenome fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/roary-pangenome in jaechang-hits/SciAgent-Skills) into .claude/skills/roary-pangenome in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/annotation/roary-pangenome in jaechang-hits/SciAgent-Skills) into .agents/skills/roary-pangenome in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill roary-pangenome -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/roary-pangenome, .gemini/skills/roary-pangenome, .github/skills/roary-pangenome and .opencode/skills/roary-pangenome in your project.
Going by SKILL.md and its folder, Roary Pangenome needs the command-line tools its instructions call (mamba, pip and python). Our summary lists: Python 3.
SKILL.md names 4 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: doi.org, sanger-pathogens.github.io and microbesonline.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Roary Pangenome is published under the GPL-3.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Roary Pangenome: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.