Agent skill

Bio Comparative Genomics Gene Family Evolution

by GPTomics in GPTomics/bioSkills

Model gene-family birth-death dynamics across a species tree using CAFE5 (Mendes et al 2020 Bioinformatics 36:5516 gamma-distributed rate categories), CAFE5-error (annotation-error-aware), Count…

MITAuto-check passedResearch & Science

Install Bio Comparative Genomics Gene Family Evolution

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-gene-family-evolution -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-comparative-genomics-gene-family-evolution --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/comparative-genomics/gene-family-evolution .claude/skills/bio-comparative-genomics-gene-family-evolution && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-comparative-genomics-gene-family-evolution
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.6k tokens
SKILL.md length
2,431 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Model gene-family birth-death dynamics across a species tree using CAFE5 (Mendes et al 2020 Bioinformatics 36:5516 gamma-distributed rate categories), CAFE5-error (annotation-error-aware), Count…

  • Correlating gene-family changes with phenotype evolution
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Decision Tree by Experimental… and Per-Tool Failure Modes, plus 12 more sections
  • Runs Shell scripts from its folder; calls conda, java and python3; reaches github.com and iro.umontreal.ca
  • Ranking lineages by adaptive gene-family-rate shifts

What it does

Bio Comparative Genomics Gene Family Evolution is an agent skill from GPTomics/bioSkills. Model gene-family birth-death dynamics across a species tree using CAFE5 (Mendes et al 2020 Bioinformatics 36:5516 gamma-distributed rate categories), CAFE5-error (annotation-error-aware), Count (Csurös 2010 ancestral state reconstruction), BadiRate (Librado 2012 likelihood + parsimony), DupliPHY-Family, and ALE/AleRax (for per-family DTL; see [[gene-tree-species-tree-reconciliation]]). Test lineage-specific gene-family expansions and contractions, distinguish biological dynamics from annotation artifacts…

Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/cafe5_birth_death_analysis.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Accounting and bookkeeping. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Correlating gene-family changes with phenotype evolution
  • Ranking lineages by adaptive gene-family-rate shifts
  • Post-WGD dosage-balance analysis
  • Building Birth-death models from OrthoFinder presence/absence matrices

Example prompts

  • “/bio-comparative-genomics-gene-family-evolution”

Requirements

  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • conda
    • java
    • python3
    • wget
    • git
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • iro.umontreal.ca

    Also links to:

    • timetree.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Comparative Genomics Gene Family Evolution loads about 6.6k tokens when it runs. Until then it costs about 223 tokens; SKILL.md has 2,431 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~223
When it runs · the whole SKILL.md, loaded when a task matches
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,431 words, ~6,648 tokens.

Download SKILL.mdSave it as .claude/skills/bio-comparative-genomics-gene-family-evolution/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-comparative-genomics-gene-family-evolution
description
Model gene-family birth-death dynamics across a species tree using CAFE5 (Mendes et al 2020 Bioinformatics 36:5516 gamma-distributed rate categories), CAFE5-error (annotation-error-aware), Count (Csurös 2010 ancestral state reconstruction), BadiRate (Librado 2012 likelihood + parsimony), DupliPHY-Family, and ALE/AleRax (for per-family DTL; see [[gene-tree-species-tree-reconciliation]]). Test lineage-specific gene-family expansions and contractions, distinguish biological dynamics from annotation artifacts, account for assembly fragmentation, identify functional enrichment in expanded / contracted families. Use when correlating gene-family changes with phenotype evolution, ranking lineages by adaptive gene-family-rate shifts, post-WGD dosage-balance analysis, or building Birth-death models from OrthoFinder presence/absence matrices.
tool_type
cli
primary_tool
CAFE5

Version Compatibility

Reference examples tested with: CAFE5 5.1.0+ (Mendes et al 2020 Bioinformatics 36(22-23):5516-5518), Count 11.0319+ (Csurös 2010 Bioinformatics 26:1910), BadiRate 1.35+ (Librado 2012 Bioinformatics 28:279), DupliPHY-Family (Ames et al 2012), CAFExp (legacy CAFE 4.2 -- DEPRECATED; use CAFE5), OrthoFinder 3.0+ for HOG input, R 4.4+, mclust 6.1+, phytools 2.3+, ETE4 4.1.0+ for tree manipulation. ALE/GeneRax/AleRax in companion skill [[gene-tree-species-tree-reconciliation]].

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: cafe5 --help; Count.exe (Java); badirate --help
  • R: packageVersion('phytools')
  • Python: pip show ete4

If code throws CAFE5: lambda did not converge, Count negative branch length, BadiRate gamma not initialized, the most common causes are: (1) annotation heterogeneity inflating family sizes, (2) saturated families (CAFE5 needs reasonable rate variation), (3) negative branch lengths in input tree (Count requires ultrametric). Pre-process: filter OG matrix to families present in >= 50% of species; resolve polytomies; ultrametricize tree.

Gene Family Evolution

"Which gene families expanded or contracted in which lineages?" -> Birth-death models on phylogeny (Hahn 2005; Csurös 2010) treat each orthogroup's per-species count as evolving under a stochastic birth-death process; lineage-specific rate shifts are detected as departures from a global rate. Annotation heterogeneity is the single largest confounder: different annotation pipelines predict different numbers of genes per family, producing apparent lineage-specific expansions that are artifacts of annotation choice (Tonkin-Hill 2020 demonstrated this for bacterial pangenomes). Consistent annotation + BUSCO/Compleasm completeness filtering are mandatory before any birth-death model interpretation. CAFE5 (Mendes et al 2020 Bioinformatics 36:5516) replaces older CAFE versions with gamma-distributed rate categories for more biologically realistic modeling.

  • CLI: cafe5 -i orthogroup_counts.tsv -t species_tree.nwk -p -- main CAFE5 workflow
  • CLI: cafe5 -e -- error-aware mode for annotation uncertainty
  • CLI: Count -- Java GUI / CLI for parsimony + likelihood ASR
  • CLI: badirate -t tree.nwk -d counts.tsv -- likelihood birth-death + branch parsimony

Algorithmic Taxonomy

ToolApproachOutputStrengthFails when
CAFE5 (Mendes et al 2020 Bioinformatics 36(22-23):5516-5518)Birth-death with gamma rate categoriesGlobal / per-family lambda + significant rate shiftsModern standard; handles rate heterogeneity; explicit Type-I controlAnnotation heterogeneity confounds; needs > 100 families
CAFE5-errorAnnotation-error-aware extensionSame plus error estimatesCritical for noisy annotationsManual error-rate specification or estimation
Count (Csurös 2010 Bioinformatics 26:1910)Both ML and parsimony ASRBranch event counts (D, L) per familyComprehensive output; GUISlower than CAFE5; less modern UX
BadiRate (Librado 2012 Bioinformatics 28:279)Likelihood birth-death + branch parsimonyLineage-specific rate shiftsCombines stochastic + parsimonyLess commonly used; older
DupliPHY-Family (Ames et al 2012)Per-family birth-deathAncestral counts per familyFamily-level granularityOlder; less integrated with modern OrthoFinder
ALE / GeneRax / AleRax (Szöllősi 2013; Morel 2024)Per-family DTL reconciliationPer-family D/T/L event countsDirect integration with [[gene-tree-species-tree-reconciliation]]Slower; per-family rather than across-family
CAFExp / CAFE 4.2 (DEPRECATED)Earlier CAFE--HistoricalUse CAFE5
Whale.jl with WGD (Zwaenepoel 2019)Bayesian DL+WGDWGD-aware family dynamicsNative WGD integrationJulia ecosystem
Functional enrichment downstreamclusterProfiler / topGO on expanded/contractedGO/KEGG enrichmentStandardMultiple testing across families

Methodology evolves; CAFE5 is the modern standard; verify the current CAFE5 manual (hahnlab/CAFE5) and Hahn lab papers before locking on a single approach.

Decision Tree by Experimental Scenario

ScenarioRecommended approachWhy
Standard CAFE-style birth-death analysisCAFE5 with gamma rate categoriesModern standard; handles rate variation
Annotation-pipeline-heterogeneityCAFE5-error modeExplicit error modeling
Post-WGD retention biasCAFE5 + DupGen_finder classification + functional enrichmentCombine birth-death with WGD-specific analysis
Lineage-specific gene-family-rate shifts correlated with phenotypeCAFE5 with binary phenotypeStandard CAFE workflow
Ancestral gene-family counts at internal nodesCount ASRPer-node count posteriors
Test for "fast-evolving" family on specific lineageCAFE5 lambda-tree (per-clade lambda)Compares lambdas across clades
Functional enrichment in expanded familiesclusterProfiler / topGO on expansion listsStandard
HGT-affected families (prokaryotes)ALE / GeneRax per-family DTL (see [[gene-tree-species-tree-reconciliation]])DTL framework explicit
Test if all families share single lambdaCAFE5 global lambda hypothesisRestricted model
Specific family analysis (e.g. immune gene family)ALE / AleRax per-familyPer-family detail
Distinguish gain from loss as primary driverCount separate D and L countsStandard parsimony
Convergent gene-family-rate shiftsRERconverge on family-count vectorsTrait-correlated rate shifts
Identify ancestral pan-clade family complementCAFE5 ASR at MRCAPre-radiation family complement
Sub-clade-specific expansions in plant genomesCAFE5 with clade-specific lambdaCompare angiosperm to gymnosperm

Per-Tool Failure Modes

Annotation heterogeneity inflating expansions / contractions

Trigger: Running CAFE5 on counts from genomes annotated by different pipelines (Augustus, MAKER, BRAKER, NCBI).

Mechanism: Different annotation tools predict different numbers of genes per family; the same biological gene family may be annotated with 5 genes in BRAKER and 8 in MAKER. CAFE5 sees the difference as an "expansion" in the MAKER-annotated species (Tonkin-Hill 2020 documented this for bacterial pangenomes; same principle in eukaryotes).

Symptom: "Most expanded families" cluster in species annotated by a single pipeline (often more permissive tool); per-species "expansion rate" correlates with annotation pipeline rather than biology.

Fix: Re-annotate all genomes with a single pipeline (currently BRAKER3 or Funannotate for eukaryotes; Bakta for bacteria) before CAFE5. Alternatively use CAFE5-error mode with explicit error rates per species. Document annotation pipeline + version per species.

Assembly fragmentation creating false contractions

Trigger: Including draft assemblies with low N50 in CAFE5 analysis.

Mechanism: Fragmented assemblies miss genes; the same family appears with fewer genes than expected. CAFE5 reports this as a contraction in the affected species.

Symptom: "Contracted" families in fragmented assemblies; correlation between BUSCO completeness and CAFE5 "contraction" rate; per-species missing-gene count varies 5-10x.

Fix: Require >= 90% BUSCO/Compleasm completeness for inclusion. Exclude species with > 5% lower BUSCO than median. Document N50 + BUSCO per assembly. For unavoidably fragmented assemblies, exclude from CAFE5 or use CAFE5-error with empirical error estimates.

CAFE5 lambda non-convergence

Trigger: CAFE5 reports "lambda did not converge"; lambda jumps between values across runs.

Mechanism: Insufficient data (< 100 families); strong rate heterogeneity not captured; or input tree non-ultrametric.

Symptom: lambda estimate unstable; AIC of model selection variable.

Fix: Require >= 100 orthogroups (preferably > 1000) in input. Ensure tree is ultrametric (ape::chronos() or treePL); CAFE5 expects time-scaled tree. Use gamma rate categories (CAFE5 default -k 4) for rate heterogeneity. If still non-convergent, restrict to single-copy or small families; check tree branch lengths for negative values.

Gamma rate-category misinterpretation

Trigger: Reporting "gamma rate categories" as biological gene-family clusters.

Mechanism: CAFE5 gamma categories are statistical buckets representing rate heterogeneity across families; they're not "fast-evolving family clusters" with biological meaning per se.

Symptom: Confusion in interpretation; "category-1 families are special."

Fix: Treat gamma categories as a statistical device. Report per-family lambda (estimated under per-family rate model) or per-clade lambda. Functional interpretation comes from family-level statistical tests, not category assignment.

Multiple-testing across many families

Trigger: Reporting "significant" expansions / contractions without correction.

Mechanism: With ~10,000 families tested, ~500 will be significant at p=0.05 under H0. Without correction, false-discovery rate is high.

Symptom: Long lists of "expanded" families; functional enrichment dominated by chance hits.

Fix: Apply FDR (Benjamini-Hochberg) across families; or restrict to a priori hypothesized families. CAFE5 reports per-family p-values; FDR-correct downstream.

Tree non-ultrametric / negative branches

Trigger: Using ML tree directly (substitutions per site) as input to CAFE5.

Mechanism: CAFE5 expects ultrametric (time-calibrated) tree; substitution-based branch lengths are not time-calibrated and may even have negative branches after rate variation.

Symptom: CAFE5 errors with "negative branch length"; or produces unreliable lambda estimates.

Fix: Time-calibrate the tree: use ape::chronos(), treePL (Smith & O'Meara 2012 Bioinformatics 28:2689), or LSD2 (To et al 2016 Syst Biol 65:82; gascuel-lab/LSD2) for fast NP-like calibration. Or use existing time-calibrated tree from TimeTree database.

Outlier-family-driven lambda estimate

Trigger: Including extreme-size families (e.g., NLR resistance gene clusters in plants with 500+ members) in CAFE5.

Mechanism: Birth-death model's likelihood is dominated by large families; one or two outlier families can overwhelm the global lambda estimate.

Symptom: Excluding the top-5 largest families changes lambda by > 30%; lambda confidence intervals huge.

Fix: Robust analysis: report lambda with and without outliers; consider per-family lambda for largest families. Functional annotation of outliers reveals if they're biologically expected expansions (rapidly evolving gene families).

Convergent gene-family-rate shifts not captured

Trigger: Phylogenomic question is whether multiple independent lineages show similar gene-family-rate shifts.

Mechanism: CAFE5 doesn't natively test for convergent rate shifts; per-clade lambda is the closest analog.

Symptom: Manual inspection of expansions across independent lineages shows pattern; CAFE5 doesn't formalize it.

Fix: Combine CAFE5 per-family lambdas with RERconverge (Redlich et al 2024 MBE 41:msae210) for trait-correlated rate shifts; use CSUBST (Fukushima 2023 Nat Eco Evo 7:155) for convergent substitution patterns.

CAFE5 vs ALE for HGT-affected family

Trigger: Applying CAFE5 to bacterial families where transfer is common.

Mechanism: CAFE5 models birth-death (D, L) only; ignores transfer (T). For HGT-affected families, the dynamics include T events that CAFE5 cannot capture.

Symptom: CAFE5 lambda for bacterial families is implausibly high; family counts don't fit birth-death model.

Fix: Use ALE / GeneRax / AleRax for HGT-affected families ([[gene-tree-species-tree-reconciliation]]); CAFE5 is appropriate for vertical-inheritance-dominated families. Combine: CAFE5 for the bulk; ALE for HGT-confirmed families.

Show full SKILL.md (999 more words)Show less
Family classification (single-copy / multi-copy) affecting interpretation

Trigger: Treating single-copy orthogroups (always = 1 per species) and multi-copy uniformly.

Mechanism: Single-copy OGs have no count variation; including them dilutes the analysis. Multi-copy families show meaningful variation.

Symptom: Including single-copy OGs lowers lambda; per-family results dominated by them.

Fix: Restrict CAFE5 input to multi-copy orthogroups (max count >= 2 in any species). Filter via OrthoFinder HOG-classification single-copy column.

Quantitative Thresholds

QuantityThresholdSource / Rationale
Minimum orthogroups for CAFE5>= 100; preferably > 1000Mendes et al 2020 Bioinformatics 36:5516
Maximum gene-family countdepends on tree depth; typical < 1000 per familyPractical
BUSCO/Compleasm completeness>= 90% for inclusionStandard
FDR for expansions / contractionsq < 0.05 (Benjamini-Hochberg)Standard
Significant lambda shiftLRT p < 0.05 / corrected for clade testsCAFE5 LRT
Gamma categories4 default (-k 4)CAFE5 docs
Tree ultrametricityrequired; calibrate with ape::chronos or treePLCAFE5 requirement
Per-clade lambda significanceLRT vs global lambdaCAFE5 docs
Annotation pipelineone pipeline across all species; document versionBest practice
Functional enrichment q-valueq < 0.05 (BH-corrected)Standard
Multi-copy family filtermax count across species >= 2Practical
Tree time-calibration sourceTimeTree, treePL, LSD2Standard alternatives
Excluded singletonsyes, in some workflowsStandard
Outlier family exclusion thresholdtop 5% by max count, exclude or per-family lambdaRobust
Convergent shift detectionRERconverge or CSUBSTCompanion methods

CAFE5 Standard Workflow

Goal: Identify gene families with lineage-specific rate shifts (expansions / contractions).

Approach: Prepare ortholog count matrix -> ultrametricize tree -> run CAFE5 with gamma categories -> FDR-correct -> annotate.

bash
# 1. Generate ortholog matrix from OrthoFinder v3 HOG file (inline Python; no external script needed)
python3 <<'PY'
import pandas as pd
hog = pd.read_csv('Phylogenetic_Hierarchical_Orthogroups/N0.tsv', sep='\t')
sp_cols = [c for c in hog.columns if c not in ('HOG', 'OG', 'Gene Tree Parent Clade')]
counts = pd.DataFrame({sp: hog[sp].fillna('').apply(
    lambda x: len(str(x).split(',')) if x and ',' in str(x) else (1 if str(x).strip() else 0)
) for sp in sp_cols})
counts.insert(0, 'family_id', hog['HOG'])
counts.insert(0, 'Description', '(null)')
counts.to_csv('cafe_input.tsv', sep='\t', index=False)
PY

# Output format: Description  family_id  Species1  Species2  ... (counts per species)

# 2. Ultrametricize tree
Rscript -e "
library(ape)
tree <- read.tree('SpeciesTree_rooted.txt')
tree_ultra <- chronos(tree)
write.tree(tree_ultra, 'tree_ultrametric.nwk')
"

# 3. Run CAFE5 with gamma categories
cafe5 \
    -i cafe_input.tsv \
    -t tree_ultrametric.nwk \
    -p \
    -k 4 \
    -e \
    -o cafe_output
# `-e` (no argument) lets CAFE5 estimate a global error model. To supply a pre-built
# error model file, use `-eerror.txt` (CAFE5 concatenates the flag and argument).
# Verify with `cafe5 --help`.

# Output:
#   cafe_output/Base_results.txt        Per-family results
#   cafe_output/Base_clade_results.txt   Per-clade lambda
#   cafe_output/Base_asr.tre            ASR tree
#   cafe_output/Base_clade_results.txt   LRT against null

# 4. FDR-correct across families (inline Python; see filter_significant_families() below)
python
'''Filter expanded / contracted families with FDR-correction.'''
import pandas as pd
from statsmodels.stats.multitest import multipletests


def filter_significant_families(cafe_results_path, change_table_path, fdr_threshold=0.05):
    '''Identify families with significant rate shifts and their per-branch expansion/contraction direction.

    CAFE5 outputs (Base_family_results.txt or similar) include a per-family p-value (vs the null
    of a single global lambda). The per-branch expansion/contraction comes from a separate
    `Base_change.tab` file (per-family x per-branch count change). Per-family lambda is NOT
    a column in Base_family_results.txt; that's a global parameter.
    '''
    df = pd.read_csv(cafe_results_path, sep='\t')
    pcol = 'p-value' if 'p-value' in df.columns else 'pvalue'
    df['fdr'] = multipletests(df[pcol].fillna(1.0), method='fdr_bh')[1]
    significant = df[df['fdr'] < fdr_threshold]

    # Pair with per-branch change table; positive cell = expansion on that branch
    changes = pd.read_csv(change_table_path, sep='\t')
    return {'significant_families': significant, 'per_branch_changes': changes}

CAFE5-error for Annotation Heterogeneity

bash
# Prepare per-species error rates (from BUSCO completeness or empirical)
cat > error_rates.tsv << 'EOF'
species_id    error_rate
Species_A     0.05
Species_B     0.08
Species_C     0.03
EOF

# Run with error-aware mode (supply pre-built error model file).
# CAFE5 concatenates the -e flag with its argument (no space): -e<error_model_file>
cafe5 \
    -i cafe_input.tsv \
    -t tree_ultrametric.nwk \
    -p \
    -k 4 \
    -eerror_model.txt \
    -o cafe_error_output

Count for Per-Branch Ancestral State

bash
# Count requires Java
java -jar Count.jar -i count_format.tsv -t species_tree.nwk \
    -o count_output --ancestral_counts

# Per-branch D and L counts in count_output/branch_events.tsv

Functional Enrichment Downstream

r
library(clusterProfiler)
library(org.Hs.eg.db)  # or appropriate species DB

# Enrichment of expanded families
expanded_genes <- read.csv('expanded_genes.tsv', stringsAsFactors = FALSE)$gene_id

go_enrich <- enrichGO(gene = expanded_genes,
                      OrgDb = org.Hs.eg.db,
                      ont = 'BP',
                      pAdjustMethod = 'BH',
                      pvalueCutoff = 0.05)

kegg_enrich <- enrichKEGG(gene = expanded_genes,
                          organism = 'hsa',
                          pvalueCutoff = 0.05)

Reconciliation: When Methods Disagree

PatternLikely causeAction
CAFE5 significant; ALE shows no DTL signalCAFE5 detects count change; ALE detects eventsBoth can be true; CAFE5 is count-level, ALE is event-level
CAFE5 significant in bacterial familyPossibly HGT-driven; CAFE5 ignores TRe-run with ALE for HGT-affected families
Per-clade lambda differs in CAFE5 vs uniform lambdaReal rate heterogeneityTrust per-clade
Annotation heterogeneity hypothesisRe-annotation eliminates "expansion"Confirm annotation artifact; report as such
CAFE5 + RERconverge agree on expansion + trait shiftConvergent biological mechanismStrong evidence
Outlier-family-driven global lambdaTop-5 families distort estimateReport robust lambda; manually flag outliers
CAFE5 expansion in fragmented speciesFalse; assembly fragmentationRe-assess BUSCO; exclude or correct
Single-copy family flagged as significantStatistical artifact; no variationExclude single-copy OGs; restrict to multi-copy
Whale.jl WGD branch shows D burst, CAFE5 shows expansionSame event; WGD modeling preferredWhale.jl is more biological for known WGD
Count parsimony vs CAFE5 likelihood disagreeParsimony underestimates lossesTrust CAFE5 likelihood

Operational rule for publication: CAFE5 with gamma rate categories + annotation pipeline normalized + BUSCO completeness > 90% + FDR-corrected significance + functional enrichment of expanded/contracted + (for bacteria) ALE complement = publication-grade gene-family evolution analysis.

Cohort Gotchas

  • WGD lineages: post-WGD retention bias; gene balance hypothesis (Birchler & Veitia 2007); analyze with [[whole-genome-duplication]] context
  • Plant gene families: NLR clusters (resistance) and ribosomal proteins are inherently large; expect lineage variation
  • Mammalian gene families: olfactory receptors are highly variable; expect lineage-specific changes
  • Bacterial gene families: HGT-driven dynamics; use ALE/AleRax instead of CAFE5
  • Polyploid species: subgenome assignment first; analyze each subgenome separately
  • Rapidly evolving lineages: higher branch-specific rates; per-clade lambda model
  • Conserved species (e.g., extant cyanobacteria): lambda may be very low; few changes to detect
  • Recent radiations: insufficient time for divergence; CAFE5 may have low power
  • Highly fragmented MAGs: include only high-quality MAGs

Anticipated Reviewer Pushback

PushbackStandard response
"Annotation pipeline?"Single pipeline (BRAKER3 / Funannotate / Bakta) across all species; version pinned
"BUSCO completeness?">= 90% required; per-species reported
"Assembly fragmentation?"N50 reported; species with > 5% lower BUSCO than median excluded
"Multiple testing?"FDR (Benjamini-Hochberg) applied across families
"Lambda non-convergence?"CAFE5 with -k 4 gamma categories; tree ultrametricized
"Outlier families?"Robust analysis with and without; outliers individually annotated
"HGT in bacteria?"ALE / GeneRax cross-checked for HGT-affected families
"Functional enrichment?"clusterProfiler / topGO with FDR; pathway-level interpretation
"Tree time-calibration?"TimeTree-based or treePL/LSD2; documented
"WGD effects?"DupGen_finder + WGD-specific analysis; subgenome-aware
"Convergent shifts?"RERconverge complementary analysis

Common Errors

Error / symptomCauseSolution
CAFE5 "lambda did not converge"Insufficient data or non-ultrametric treeUltrametricize; increase families; check tree
CAFE5 "negative branch length"ML tree inputTime-calibrate with chronos / treePL / LSD2
CAFE5 errors on input formatWrong column orderCheck OrthoFinder HOG format; species in column 2 onward
Count GUI hangsJava memoryIncrease Java heap: java -Xmx16G -jar Count.jar
BadiRate gamma errorsInitialization issueUse default -g 1 or specify per family
Per-clade lambda implausibly highOutlier family or tree issueExclude top families; re-run
All families significantNo FDR correctionApply BH correction
Annotation heterogeneity not addressedMixed pipelinesRe-annotate consistently
Family with 0 count for all speciesAnnotation issueFilter rows with all zeros
Highly variable family count (e.g. 500-5000)Real biological variation or annotationAnnotate manually; consider exclusion

Tool Installation Notes

bash
# CAFE5
conda install -c bioconda cafe
# Or: git clone https://github.com/hahnlab/CAFE5

# Count
wget http://www.iro.umontreal.ca/~csuros/gene_content/count.tar.gz
tar xf count.tar.gz

# BadiRate
git clone https://github.com/PauloRoldan/badirate

# DupliPHY-Family
# Web only; no public CLI

# Whale.jl (Julia) -- see [[gene-tree-species-tree-reconciliation]]
julia -e 'using Pkg; Pkg.add("Whale")'

# R packages
install.packages(c('ape', 'phytools', 'clusterProfiler', 'org.Hs.eg.db'))

# Time calibration
conda install -c bioconda treepl
# Or use ape::chronos (R)

# For OrthoFinder input
conda install -c bioconda orthofinder

For Funannotate / BRAKER3 reannotation (essential pre-CAFE5):

bash
conda install -c bioconda funannotate braker3

References

  • Hahn MW et al 2005 Genome Res 15:1153 (CAFE original framework)
  • Mendes FK et al 2020 Bioinformatics 36:5516 (CAFE5)
  • Csurös M 2010 Bioinformatics 26:1910 (Count)
  • Librado P et al 2012 Bioinformatics 28:279 (BadiRate)
  • Ames RM et al 2012 Bioinformatics 28:48 (DupliPHY-Family)
  • Tonkin-Hill G et al 2020 Genome Biol 21:180 (Panaroo; annotation heterogeneity)
  • Smith SA & Dunn CW 2008 Bioinformatics 24:715 (Phyutility)
  • Smith SA & O'Meara BC 2012 Bioinformatics 28:2689 (treePL)
  • To T-H, Jung M, Lycett S & Gascuel O 2016 Syst Biol 65:82 (LSD2)
  • Birchler JA & Veitia RA 2007 Plant Cell 19:395 (gene balance)
  • Redlich R et al 2024 MBE 41:msae210 (RERconverge categorical)
  • Fukushima K & Pollock DD 2023 Nat Eco Evo 7:155 (CSUBST)
  • Lynch M & Conery JS 2000 Science 290:1151 (gene duplication mechanism)
  • Force A et al 1999 Genetics 151:1531 (subfunctionalization)
  • De Bie T et al 2006 Bioinformatics 22:1269 (CAFE original software)
  • Han MV et al 2013 MBE 30:1987 (CAFE 3)
  • Otto SP & Whitton J 2000 Annu Rev Genet 34:401 (polyploidy mechanisms)
  • TimeTree (database, http://www.timetree.org)
  • comparative-genomics/ortholog-inference - OrthoFinder HOG matrix is CAFE5 input
  • comparative-genomics/gene-tree-species-tree-reconciliation - ALE per-family DTL, complement to CAFE5
  • comparative-genomics/whole-genome-duplication - Post-WGD retention bias context
  • comparative-genomics/positive-selection - Selection within expanded families
  • comparative-genomics/ancestral-reconstruction - Ancestral count reconstruction
  • phylogenetics/divergence-dating - Time-calibrated tree for CAFE5
  • phylogenetics/modern-tree-inference - Species tree input
  • pathway-analysis/go-enrichment - Functional enrichment of expanded families
  • pathway-analysis/gsea - GSEA on family expansions
  • single-cell/cell-annotation - Cell-type-specific gene-family expansions

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in comparative-genomics/gene-family-evolution of GPTomics/bioSkills.

  • SKILL.md
  • examples/cafe5_birth_death_analysis.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Comparative Genomics Gene Family Evolution next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Comparative Genomics Gene Family Evolution compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Comparative Genomics Gene Family Evolution this skillGPTomics/bioSkills1.2k2 repos~6.6kAutomated safety check: PassMIT
EtetoolkitK-Dense-AI/scientific-agent-skills48k1 repos~3.3kAutomated safety check: NotesGPL-3.0-or-later
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Etetoolkit

    K-Dense-AI/scientific-agent-skills

    Analyzes, manipulates, compares, annotates, and visualizes phylogenetic or other hierarchical trees with ETE 4.

    48k GitHub starsUsed in 1 repo~3.3k tokens
    Research & ScienceAuto-check: notes
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Comparative Genomics Gene Family Evolution

What does Bio Comparative Genomics Gene Family Evolution do?

Model gene-family birth-death dynamics across a species tree using CAFE5 (Mendes et al 2020 Bioinformatics 36:5516 gamma-distributed rate categories), CAFE5-error (annotation-error-aware), Count…. Bio Comparative Genomics Gene Family Evolution is an agent skill from GPTomics/bioSkills. Model gene-family birth-death dynamics across a species tree using CAFE5 (Mendes et al 2020 Bioinformatics 36:5516 gamma-distributed rate categories), CAFE5-error (annotation-error-aware), Count (Csurös 2010 ancestral state reconstruction), BadiRate (Librado 2012 likelihood + parsimony), DupliPHY-Family, and ALE/AleRax (for per-family DTL; see [[gene-tree-species-tree-reconciliation]]).

When should I use Bio Comparative Genomics Gene Family Evolution?

Bio Comparative Genomics Gene Family Evolution fits situations like: correlating gene-family changes with phenotype evolution; ranking lineages by adaptive gene-family-rate shifts; post-WGD dosage-balance analysis; building Birth-death models from OrthoFinder presence/absence matrices.

How do I install Bio Comparative Genomics Gene Family Evolution in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-gene-family-evolution -a claude-code`. Or copy the skill folder (comparative-genomics/gene-family-evolution in GPTomics/bioSkills) into .claude/skills/bio-comparative-genomics-gene-family-evolution in your project. Claude Code loads it when a task matches its description.

How do I install Bio Comparative Genomics Gene Family Evolution in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-gene-family-evolution -a codex`. Or copy the skill folder (comparative-genomics/gene-family-evolution in GPTomics/bioSkills) into .agents/skills/bio-comparative-genomics-gene-family-evolution in your project. Codex loads it when a task matches its description.

Can I use Bio Comparative Genomics Gene Family Evolution in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-gene-family-evolution -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-comparative-genomics-gene-family-evolution, .gemini/skills/bio-comparative-genomics-gene-family-evolution, .github/skills/bio-comparative-genomics-gene-family-evolution and .opencode/skills/bio-comparative-genomics-gene-family-evolution in your project.

What does Bio Comparative Genomics Gene Family Evolution need to run?

Going by SKILL.md and its folder, Bio Comparative Genomics Gene Family Evolution needs a shell for the scripts in its folder and the command-line tools its instructions call (conda, java, python3, wget, git and pip). Our summary lists: A Bash shell.

Does Bio Comparative Genomics Gene Family Evolution access the network?

SKILL.md names 3 domains. In commands or code: github.com and iro.umontreal.ca; the agent is likely to contact these when it follows the instructions. As links in the text: timetree.org. This is read from the text; nothing was executed.

Is Bio Comparative Genomics Gene Family Evolution safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Comparative Genomics Gene Family Evolution use?

Bio Comparative Genomics Gene Family Evolution is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Comparative Genomics Gene Family Evolution use?

About 6.6k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Comparative Genomics Gene Family Evolution?

Skills that share tags, products or a category with Bio Comparative Genomics Gene Family Evolution: Etetoolkit (K-Dense-AI/scientific-agent-skills, 48k stars), Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars) and Clinvar Database (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Comparative Genomics Gene Family Evolution?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.