Agent skill

Bio Comparative Genomics Whole Genome Duplication

by GPTomics in GPTomics/bioSkills

Detect, date, and contextualize whole-genome duplication (WGD / paleopolyploidy) events using wgd v2 (Chen et al 2024), KsRates (Sensalari 2022 substitution-rate-corrected Ks dating), DupGenfinder…

MITAuto-check passedResearch & Science

Install Bio Comparative Genomics Whole Genome Duplication

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-whole-genome-duplication -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-comparative-genomics-whole-genome-duplication --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/comparative-genomics/whole-genome-duplication .claude/skills/bio-comparative-genomics-whole-genome-duplication && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-comparative-genomics-whole-genome-duplication
GitHub stars
1.2k
Used in
2 other repos
Token cost
~7.4k tokens
SKILL.md length
3,080 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Detect, date, and contextualize whole-genome duplication (WGD / paleopolyploidy) events using wgd v2 (Chen et al 2024), KsRates (Sensalari 2022 substitution-rate-corrected Ks dating), DupGenfinder…

  • Identifying ancient polyploidy from Ks distributions and synteny block analysis
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Decision Tree by Experimental… and Per-Tool Failure Modes, plus 12 more sections
  • Runs Shell scripts from its folder; calls git, pip and conda; reaches github.com and bitbucket.org
  • Positioning WGD events relative to speciation

What it does

Bio Comparative Genomics Whole Genome Duplication is an agent skill from GPTomics/bioSkills. Detect, date, and contextualize whole-genome duplication (WGD / paleopolyploidy) events using wgd v2 (Chen et al 2024), KsRates (Sensalari 2022 substitution-rate-corrected Ks dating), DupGenfinder (Qiao 2019), MAPS (Li 2018 phylogenomic), POInT (Conant 2008 ordered-block), SLEDGe (2024 ML-based), Whale.jl (Bayesian DL+WGD), and synteny-anchored paranome construction. Use when identifying ancient polyploidy from Ks distributions and synteny block analysis, positioning WGD events relative to speciation…

Its SKILL.md is about 7.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/wgd_v2_ks_pipeline.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Identifying ancient polyploidy from Ks distributions and synteny block analysis
  • Positioning WGD events relative to speciation
  • Distinguishing tandem from segmental from WGD duplications
  • Dating the 2R/3R vertebrate / fish / salmonid WGDs

Example prompts

  • “/bio-comparative-genomics-whole-genome-duplication”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • pip
    • conda
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • bitbucket.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Comparative Genomics Whole Genome Duplication loads about 7.4k tokens when it runs. Until then it costs about 214 tokens; SKILL.md has 3,080 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~214
When it runs · the whole SKILL.md, loaded when a task matches
~7.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,080 words, ~7,423 tokens.

Download SKILL.mdSave it as .claude/skills/bio-comparative-genomics-whole-genome-duplication/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-comparative-genomics-whole-genome-duplication
description
Detect, date, and contextualize whole-genome duplication (WGD / paleopolyploidy) events using wgd v2 (Chen et al 2024), KsRates (Sensalari 2022 substitution-rate-corrected Ks dating), DupGen_finder (Qiao 2019), MAPS (Li 2018 phylogenomic), POInT (Conant 2008 ordered-block), SLEDGe (2024 ML-based), Whale.jl (Bayesian DL+WGD), and synteny-anchored paranome construction. Use when identifying ancient polyploidy from Ks distributions and synteny block analysis, positioning WGD events relative to speciation, distinguishing tandem from segmental from WGD duplications, dating the 2R/3R vertebrate / fish / salmonid WGDs, building paranome and Ks-age mixture models, applying KsRates substitution-rate correction across lineages, or testing alternative biased-fractionation / dosage-balance models post-WGD.
tool_type
mixed
primary_tool
wgd

Version Compatibility

Reference examples tested with: wgd v2.0.31+ (heche-psb/wgd; Chen et al 2024 Bioinformatics 40:btae272), KsRates 1.1.3+ (VIB-PSB/ksrates; Sensalari 2022 Bioinformatics 38:530), DupGen_finder (Qiao 2019 Genome Biol 20:38), MAPS 1.0 (Li 2018), POInT (Conant lab), SLEDGe (bioRxiv 2024.01.17.574559), Whale.jl 2.0+, ksrates pip 1.1+, MCScanX 1.0+, PAML 4.10+ (yn00/codeml for Ks), BLAT 36+, DIAMOND 2.1+, R 4.4+, mclust 6.1+ (for mixture models). Python 3.10+ required for wgd v2.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: wgd --version; ksrates --version; wgd ksd --help
  • Python: pip show wgd ksrates
  • R: packageVersion('mclust')

If code throws wgd ksd: cannot find PAML output, KsRates: insufficient sister species, MAPS: missing tree, these tools have specific input expectations: wgd needs codon-aware MAFFT/MUSCLE alignment; KsRates needs configured config_ksrates.txt; MAPS needs nucleotide tree. The deprecated arzwa/wgd v1 is replaced by heche-psb/wgd v2.

Whole Genome Duplication Analysis

"Are there WGD events in this lineage and when did they occur?" -> WGD detection combines Ks distributions (synonymous-substitution rates between gene paralog pairs, showing peaks at past polyploidy events) and synteny block analysis (parallel collinear blocks within a genome). Modern best practice uses wgd v2 (Chen et al 2024 Bioinformatics 40:btae272) as an integrated pipeline. KsRates (Sensalari 2022 Bioinformatics 38:530) is mandatory for cross-lineage comparison because substitution rates vary across the tree -- ignoring this places WGDs incorrectly relative to speciation events. The fundamental tradeoff: Ks plot peaks are visually obvious but biologically ambiguous between (1) small-scale tandem duplications, (2) segmental duplications, and (3) true WGD; combining Ks with synteny anchors disambiguates these.

  • CLI: wgd dmd and wgd ksd -- paranome construction + Ks distribution
  • CLI: wgd syn -- synteny-anchored WGD signal extraction
  • CLI: ksrates init then ksrates wgd-paralogs ortho -- substitution-rate-corrected positioning
  • CLI: MCScanX -h then dupgen_finder -- duplication-class assignment
  • CLI: mapsR -- gene-tree-based WGD phylogenetic placement

Algorithmic Taxonomy

ToolApproachOutputStrengthFails when
wgd v2 (Chen et al 2024 Bioinformatics 40:btae272)Integrated paranome + Ks + synteny + WGD datingKs distributions, collinearity plots, GMM/ELMM mixture fits, datingStandard 2024 pipeline; replaces deprecated arzwa/wgd v1Saturation at Ks > 2; single-lineage substitution rate variation
KsRates (Sensalari 2022 Bioinformatics 38:530)Substitution-rate correction via outgroup pairsAdjusted Ks ages of focal-species paralogs vs orthologsMANDATORY when comparing WGDs across lineages with different ratesRequires at least 2 outgroups for rate calibration
DupGen_finder (Qiao 2019 Genome Biol 20:38)Classifies duplications by genomic contextPer-gene class: tandem, proximal, transposed, dispersed, WGDDisambiguates duplication typeClass assignment depends on intervening-gene-count windows
MAPS (Li 2018)Phylogenomic placement of WGD via gene-tree topology mappingWGD position on species treeDetects WGDs from gene-tree-species-tree discordanceComputationally heavy; requires many gene trees
POInT (Conant & Wolfe 2008 Genetics 179:1681)Order-aware reconstruction of WGD chromosomesReconstructed ancestral WGD genomeStrong inference for syntenic-block agesLineage-specific tuning required
SLEDGe (bioRxiv 2024.01.17.574559)ML classifier on Ks plot featuresWGD vs no-WGD binary call + confidenceReduces visual-peak-fitting subjectivityNewer; less validated
wgd v1 (arzwa/wgd; DEPRECATED)Predecessor of v2--HistoricalUse v2 (heche-psb/wgd)
Whale.jl (Zwaenepoel & Van de Peer 2019 MBE 36:1384)Bayesian DL + WGD reconciliationWGD posterior at species-tree nodesNative WGD modeling; integrates with [[gene-tree-species-tree-reconciliation]]Julia ecosystem
WGDexploreR (legacy)Visualize Ks plotsPlots onlyVisualization aidNot for inference
McLuster-WGD (custom workflows)Mixture-model fitting on KsGMM / ELMM componentsFor custom peak fittingNot a standard tool
ksrates web (sensalari 2022)Web interface to ksratesSame as CLIUser-friendlyManual config; not scriptable for genome-wide
wgs2pep (Vandepoele/wgs)Pep-to-WGD ortholog identificationPer-paralog WGD assignmentPlant-focusedLess general

Methodology evolves; verify the current wgd v2 manual and the Chen & Zwaenepoel 2023 review chapter (in Polyploidy: Methods and Protocols) before locking on a single workflow. The 2R vertebrate / 3R teleost / Ss4R salmonid WGDs are well-established; novel WGD claims require concordance across Ks, synteny, and phylogenomic placement.

Decision Tree by Experimental Scenario

ScenarioRecommended approachWhy
Plant comparative genomics; suspect WGDwgd v2 + KsRatesModern standard pipeline; plants are WGD-prone
Vertebrate 2R or 3R WGD analysiswgd v2 + MAPS phylogenomic placementVertebrate-specific; combines Ks + tree topology
Salmonid Ss4R WGDWhale.jl with WGD node in species treeNative WGD modeling in Bayesian framework
Distinguish WGD from sequential small duplicationsKs GMM fit + synteny block analysis (wgd syn)Ks peak alone insufficient; synteny confirms
Date WGD relative to a speciation eventKsRates with outgroup speciation calibrationRequired when substitution rates differ
Identify duplicates retained vs lost after WGDDupGen_finder + Ks distributionClassification + age estimation
Detect novel WGD in non-model organismwgd dmd -> wgd ksd -> visual GMM/ELMM mixtureStandard discovery workflow
Reconstruct ancestral WGD genome architecturePOInTOrder-aware ancestral reconstruction
Test post-WGD bias in fractionation (dosage-balance)DupGen_finder + per-gene functional annotationCompare retained vs lost genes
Distinguish recent vs ancient WGDKsRates + Ks peak location after correctionKs peak < 0.1 recent; 0.5-1.5 ancient
WGD in clade with rapid evolution (e.g. fish)KsRates mandatoryWithout rate correction, Ks peaks displaced
ML classification of WGD signalSLEDGeNewer ML-based alternative to manual fitting
Integrate WGD with DTL inferenceWhale.jl OR ALE with WGD branchSee [[gene-tree-species-tree-reconciliation]]
Pan-clade WGD survey (e.g. all angiosperms)wgd v2 batch + ksrates aggregatedStandard for sustained phylogenomic surveys
Polyploid genome with subgenomesDupGen_finder for tandem/transposed + AnchorWave proali for subgenome-aware syntenySubgenome assignment first
Recently diverged species pair, possible recent WGDminimap2 -x asm5 + SyRI + Ks distributionHigh-resolution recent-WGD detection

Per-Tool Failure Modes

Saturation at Ks > 2

Trigger: Computing Ks for ancient WGD candidates from distantly related taxa.

Mechanism: Synonymous substitutions saturate after Ks ~2; each site has undergone multiple substitutions. Observed Ks underestimates true Ks; the relationship between Ks and time becomes non-monotonic above 1.5. WGD peaks at age 200+ Myr in vertebrates / 100+ Myr in plants are at Ks > 2 and unreliable (Vanneste 2013 MBE 30:177).

Symptom: wgd ksd output shows peak at Ks > 1.5 with broad distribution; KsRates rate-corrected Ks even more uncertain; visual fitting yields ambiguous components.

Fix: Restrict Ks-based WGD inference to Ks < 1.5; for older WGDs, use phylogenomic methods (MAPS, POInT, Whale.jl) which don't rely on Ks alone. Report Ks-based dating with explicit saturation caveat. For 2R vertebrate WGD (~500 Myr), MAPS / Whale.jl is required because Ks is saturated.

Substitution-rate variation across lineages

Trigger: Comparing WGD position in two lineages with different molecular evolutionary rates.

Mechanism: If lineage A evolves twice as fast as lineage B, the same biological time is at twice the Ks in lineage A. A WGD at Ks = 0.5 in lineage A and Ks = 0.25 in lineage B might be the same event biologically.

Symptom: WGD claimed at "different times" in different lineages; speciation-vs-WGD relative timing flips depending on which lineage is the focal species.

Fix: Always use KsRates (Sensalari 2022) when comparing WGDs across lineages. KsRates uses outgroup-species speciation events to calibrate the relative substitution rates, then rescales Ks to a common scale. Single-lineage Ks dating is unreliable for inter-lineage comparison.

Tandem duplication masquerading as WGD peak

Trigger: Ks distribution shows a peak at low Ks (< 0.5); user concludes "recent WGD."

Mechanism: Tandem duplications create paralog pairs with low Ks; these accumulate in clusters (especially in NLR genes in plants, OR genes in mammals). The Ks peak reflects clustered tandem origins, not WGD.

Symptom: "WGD" peak Ks-distribution is dominated by paralogs from tandem clusters; synteny analysis shows few cross-chromosome parallel blocks.

Fix: Use DupGen_finder to classify each paralog pair as tandem / proximal / transposed / dispersed / WGD. The true WGD signal comes from the WGD-classified pairs. Recompute Ks distribution restricted to WGD-classified pairs (or anchor-pair restricted to synteny blocks). wgd v2 integrates this filtering.

Synteny block age inconsistency

Trigger: Different synteny blocks from the same suspected WGD have different Ks distributions.

Mechanism: Real WGD blocks should share a synchronized Ks distribution centered at the WGD age; if blocks have very different ages, the inferred "WGD" was probably a series of segmental duplications.

Symptom: Per-block Ks median varies widely; some blocks show Ks ~ 0.3 and others Ks ~ 1.0; not consistent with single WGD.

Fix: Compute per-block Ks distribution; require synchronization (interquartile range overlapping across blocks). wgd v2 reports per-block age dispersion. If dispersion is high, downgrade to "potentially WGD-like" or attribute to ancient segmental duplication history.

Mixture model under/overfitting components

Trigger: Fitting GMM (gaussian mixture model) or ELMM (exponential-lognormal mixture model) to Ks distribution.

Mechanism: Mixture models with too few components miss real WGD signal; too many components find spurious peaks. BIC-based model selection is standard but sensitive to bin choice in histograms.

Symptom: Different runs with different bin counts give different component numbers; visual peaks don't match component means.

Fix: Use BIC-based component selection with 5-fold cross-validation; report uncertainty in component count. wgd v2 fits both GMM and ELMM and reports BIC for each model up to 5 components. Visually validate against synteny block analysis.

Comparing wgd v1 (deprecated) vs v2 outputs

Trigger: Mixing scripts written for arzwa/wgd v1 with heche-psb/wgd v2.

Mechanism: wgd v1 was deprecated 2023; v2 has different command-line interface, default parameters, and output file structure. Scripts from v1 fail or produce different results in v2.

Symptom: Older scripts using wgd ksd v1 syntax don't work in v2; output paths differ.

Fix: Update all scripts to wgd v2 syntax. The v1 wgd is no longer maintained; v2 is the current standard.

KsRates failing with insufficient species

Trigger: Running KsRates with one focal species and one outgroup.

Mechanism: KsRates calibrates substitution-rate variation via at least two outgroup speciation events; one outgroup species provides only one rate point and cannot resolve rate heterogeneity.

Symptom: KsRates errors or produces unreliable rate-corrected Ks.

Fix: Include >= 2 outgroup species at different phylogenetic distances; for plants, a model angiosperm + gymnosperm pair; for vertebrates, lamprey + invertebrate outgroup. Document outgroup sampling.

Polyploid genome confusion

Trigger: Running wgd v2 on an unassigned-subgenome polyploid (e.g. wheat hexaploid without A/B/D subgenome labels).

Mechanism: WGD signal is muddled when subgenomes are not separated; orthologs across subgenomes appear as paralogs, intermixing tandem / segmental / WGD / homeolog signal.

Symptom: Ks distribution shows multiple overlapping peaks at varying intensities; component fits are unstable.

Fix: Assign subgenomes first using k-mer methods (KMC2 + SubPhaser; Jia 2022 New Phytol 235:801) or synteny-based assignment (GENESPACE; Lovell 2022). Then run wgd v2 on each subgenome separately. Document subgenome assignment.

Confusion of homeologous (WGD) vs orthologous (speciation) pairs

Trigger: Computing Ks across recent WGD species without explicit homeolog labeling.

Mechanism: In recent WGD lineages (e.g. salmonid Ss4R), the "ortholog" between species and the "homeolog within species" can have similar Ks; standard ortholog detection (OrthoFinder) lumps both.

Symptom: wgd v2 reports Ks distributions but doesn't distinguish homeologs from orthologs; downstream dating confused.

Fix: Use synteny-anchored ortholog detection (GENESPACE, ProteinOrtho-synteny) to separate cross-species orthologs from within-species homeologs. Run wgd v2 separately on each class. Report each Ks distribution explicitly.

Show full SKILL.md (1,331 more words)Show less

Quantitative Thresholds

QuantityThresholdSource / Rationale
Ks saturation upper limitKs < 1.5 for reliable inference; >= 2 saturatedVanneste 2013 MBE 30:177
Recent WGD Ks rangeKs 0.1-0.5Standard convention
Ancient WGD Ks rangeKs 0.5-1.5Standard convention
WGD synteny block minimum>= 5 anchors per blockwgd v2 / GENESPACE default
Tandem duplicate window for DupGen_finder5 genes default; species-tunableQiao 2019
Proximal duplicate window5-25 genesStandard
Number of mixture components to consider1-5 with BIC selectionwgd v2 default
Cross-block age synchrony for WGDper-block Ks IQR overlapVisual + statistical
KsRates minimum outgroups>= 2 species at different distancesSensalari 2022
KsRates substitution rate correction valid rangeKs < 1.5 in original; correction can extend slightlyKsRates docs
MAPS minimum gene trees>= 1000 single-copy ortholog treesLi 2018
Whale.jl WGD detection powerdepends on retention rate; > 0.3 retention typically detectableZwaenepoel 2019
Vertebrate 2R Kssaturated; estimated 500-700 MyrDehal 2005
Teleost 3R Kssaturated; estimated 250-350 MyrGlasauer 2014
Salmonid Ss4R Ks80-100 Myr; Ks ~0.1-0.2Lien 2016
Plant 1R / 2Rvaries clade; core-eudicot gamma paleohexaploidy ~120 MyrJiao 2012 (dating); Soltis 2009 (placement)
Synonymous codon site countdS reliable when >= 30 synonymous sites per pairYang 2007 PAML
Per-pair gene length minimum>= 300 bp CDSwgd v2 default
Codon-aware MSA alignerMAFFT --auto or MUSCLE; for distant pairs PRANKwgd v2 supports all
Bootstrap support for MAPS>= 80 on gene-tree branchesLi 2018

wgd v2 Standard Workflow

Goal: Detect and date WGD events from a focal-species proteome and CDS, with synteny-anchored Ks distribution.

Approach: Build paranome -> compute Ks distribution -> identify synteny anchors -> mixture model fit -> visualize.

bash
# Install (Python 3.10+ required)
pip install wgd

# 1. Build paranome (all-vs-all paralog identification)
wgd dmd cds.fasta -o output/paranome.tsv -t 16

# 2. Compute Ks distribution
# Verify exact flags with `wgd ksd --help` (the wgd v2 CLI evolves; --pairwise / --ks-method spelling differs across versions).
wgd ksd output/paranome.tsv cds.fasta -o output/ksd \
    --aligner mafft --n-threads 16

# 3. Synteny anchors (intra-genome)
wgd syn output/paranome.tsv gff.bed cds.fasta -o output/syn \
    --feature gene --gene-attribute Name --min-block-size 5

# 4. Mixture model on synteny-anchored Ks
wgd mix output/ksd/ks_distributions.tsv -o output/mix \
    --model both --components 1 2 3 4 5

# 5. Visualize
wgd viz output/mix/mix_results.tsv -o output/viz \
    --type histogram --bins 50
python
'''Parse wgd v2 outputs and identify WGD peak from Ks distribution.'''
import pandas as pd
from scipy import stats


def load_ks_distribution(ksd_file):
    '''Load wgd ksd output: pair  Ks  Ka  ...'''
    df = pd.read_csv(ksd_file, sep='\t')
    # Filter saturated
    df_filt = df[(df['Ks'] > 0) & (df['Ks'] < 2)]
    return df_filt


def fit_mixture(ks_values, n_components=3):
    '''Fit GMM (gaussian mixture model) on Ks via sklearn.'''
    from sklearn.mixture import GaussianMixture
    gmm = GaussianMixture(n_components=n_components, random_state=42)
    gmm.fit(ks_values.reshape(-1, 1))
    return {
        'means': gmm.means_.flatten(),
        'variances': gmm.covariances_.flatten(),
        'weights': gmm.weights_,
        'bic': gmm.bic(ks_values.reshape(-1, 1))
    }


def identify_wgd_peaks(ks_values, min_components=1, max_components=5):
    '''BIC-based model selection.'''
    results = {}
    for n in range(min_components, max_components + 1):
        results[n] = fit_mixture(ks_values, n)
    best_n = min(results.keys(), key=lambda k: results[k]['bic'])
    return best_n, results[best_n]

KsRates Substitution-Rate-Corrected Workflow

Goal: Position a focal-species WGD relative to speciation events with substitution-rate correction.

Approach: Define focal species + 2+ outgroups -> compute orthologous Ks (focal vs outgroup) -> compute paralogous Ks (focal-internal) -> rescale via outgroup calibration.

bash
# KsRates is best driven through its Nextflow pipeline (which orchestrates ortholog Ks,
# paralog Ks, rate correction, and plotting). Subcommand naming differs across releases;
# always verify with `ksrates --help` against the installed version.

# 1. Generate a working config (subcommand spelling varies by release; e.g. `init` in some,
#    `generate-config` in others). Inspect `ksrates --help`.
ksrates init config_ksrates.txt        # OR: ksrates generate-config config_ksrates.txt

# 2. Edit config_ksrates.txt to set focal_species, outgroups, FASTA + GFF paths, tree.

# 3. Run the full pipeline (Nextflow-driven):
ksrates --config config_ksrates.txt --n-threads 16
# Or invoke individual stages (subcommand names: see `ksrates --help`).

The output plot shows the rate-corrected focal-species paralog Ks distribution with vertical lines indicating outgroup-species speciation events; WGD peaks before vs after speciation events can be distinguished.

DupGen_finder for Duplication Class Assignment

Goal: Classify each paralog pair as tandem / proximal / transposed / dispersed / WGD by genomic context.

Approach: MCScanX collinearity + intervening-gene count -> classify each duplicate by class.

bash
# Input: MCScanX collinearity file + GFF
git clone https://github.com/qiao-xin/DupGen_finder
cd DupGen_finder

# Run with intervening-gene count
./DupGen_finder.pl \
    -i input_collinearity_file \
    -t species_name \
    -c gene_count_file \
    -o output

# Output: tandem, proximal, transposed, dispersed, wgd duplicate lists
python
'''Aggregate DupGen_finder output for downstream Ks distribution per class.'''
def load_dupgen(class_dir):
    classes = {}
    for cls in ('tandem', 'proximal', 'transposed', 'dispersed', 'wgd'):
        path = f'{class_dir}/{cls}.pairs'
        try:
            classes[cls] = pd.read_csv(path, sep='\t', header=None,
                                       names=['gene1', 'gene2'])
        except FileNotFoundError:
            classes[cls] = pd.DataFrame(columns=['gene1', 'gene2'])
    return classes

MAPS Phylogenomic WGD Placement

Goal: Place a WGD event on a species tree using gene-tree-species-tree mapping.

Approach: Build many single-copy ortholog gene trees -> map each tree's topology to the species tree -> identify branches with topology consistent with a WGD event.

MAPS is heavyweight; requires CRG database setup and significant compute. See https://bitbucket.org/barker-lab/maps/src for the current pipeline.

Reconciliation: When Methods Disagree

PatternLikely causeAction
wgd v2 peak at Ks 0.5, MAPS supports WGD on different branchRate variation across lineagesKsRates with outgroups; trust rate-corrected position
Ks peak prominent, synteny anchors sparseTandem-driven peakDupGen_finder classification; verify WGD class
Synteny blocks present, Ks peak diffuseOld WGD; saturated KsRestrict to recent paralogs; expect Ks 0.5-1.5
One outgroup's KsRates says WGD before speciation, other afterOutgroup substitution rate variationUse multiple outgroups; report ranges; explicit caveat
GMM fits 2 components, ELMM fits 3Model class differsVisual inspection; report both; choose by BIC
DupGen_finder calls "WGD" but Ks > 2SaturationKs unreliable; verify via synteny + phylogenomic placement
Two recently sequenced sister species disagree on WGD dateDifferent annotation pipelinesRe-annotate consistently; verify orthology
wgd v2 detects WGD; Whale.jl posterior doesn'tDifferent inference frameworksWhale.jl is Bayesian and gene-tree-based; trust if MAPS / phylogeny supports
Salmonid Ss4R detectable by wgd v2 in young salmonidsRecent WGD; clearer signalConfirmation, not contradiction

Operational rule for publication: Concordance across Ks distribution (wgd v2), synteny-anchored block analysis (wgd syn / GENESPACE), KsRates rate-corrected positioning, and at least one phylogenomic method (MAPS / Whale.jl / ALE with WGD node) = publication-grade WGD claim. Single-Ks-peak claims should be downgraded to "Ks evidence for possible WGD."

Cohort Gotchas

  • Plant WGD legacy: All angiosperms share at least one ancestral WGD (the epsilon event, ~192 Myr); the core-eudicot gamma paleohexaploidy is younger (~120 Myr; Jiao 2012 Genome Biol 13:R3), core-eudicot-specific (Soltis 2009 Am J Bot 96:336); confirm consensus against published plant lineages
  • Vertebrate 2R: ~500-600 Myr; Ks saturated; use MAPS phylogenomic placement (Dehal & Boore 2005)
  • Teleost 3R: ~320 Myr; in addition to 2R; doubles ohnologs in fish
  • Salmonid Ss4R: ~80-100 Myr; recent enough that Ks is informative; ohnologs identifiable
  • Plant polyploids (wheat, Brassica): recent; subgenomes assignable; subgenome-stratify before wgd v2
  • Allopolyploids vs autopolyploids: allopolyploids have higher inter-subgenome divergence (easier to assign subgenomes); autopolyploids have similar subgenomes (chimeric assembly risk)
  • Lineage-specific extra duplications post-WGD: make Ks peak broader; differentiate from sequential WGDs

Anticipated Reviewer Pushback

PushbackStandard response
"Saturation?"Ks < 1.5 reported; for older WGDs, MAPS or Whale.jl used
"Rate variation?"KsRates applied with 2+ outgroups; rate-corrected positioning
"Tandem vs WGD?"DupGen_finder classification; WGD-class subset reported
"Synteny confirmation?"wgd syn or GENESPACE; per-block Ks synchrony reported
"Phylogenomic support?"MAPS, Whale.jl, or ALE-with-WGD agrees
"Mixture model uncertainty?"BIC-based model selection; 1-5 components tested
"Subgenome assignment (polyploid)?"k-mer-based or synteny-based subgenome assignment first
"Why wgd v2 over wgd v1?"v1 deprecated 2023; v2 (heche-psb/wgd) is current
"Comparison to published WGD dates?"Match published vertebrate 2R / teleost 3R / salmonid Ss4R; cite primary literature

Common Errors

Error / symptomCauseSolution
wgd ksd reports few Ks valuesCDS sequences not in-frame or shortVerify CDS quality; filter short sequences
KsRates "no orthologs found"Outgroup species names mismatchedNormalize species labels exactly
GMM fits with 5 components, all similar meansOverfittingUse BIC; restrict max components to 3-4
Per-block Ks IQR > 0.5Block age dispersion; likely segmental not WGDTrust dispersion; reclassify as segmental
wgd v2 ELMM converges to single componentInsufficient pairsReduce filtering; include more paralog pairs
MAPS slow / OOMMany gene trees; large datasetCluster; reduce to representative single-copy orthologs
Whale.jl Turing fails to mixNUTS step size; tree dimensionHMC manual tuning; or reduce species tree
DupGen_finder produces empty tandem setDefault window inappropriateAdjust window for genome density
Ks distribution has secondary peak at 1.8Saturation artifact (not WGD)Recompute with PRANK or PAML codeml; restrict Ks < 1.5
Polyploid "WGD" appears at Ks 0.05Recent polyploidy is ohnologous WGD, expectedVerify with subgenome assignment
Plant analysis lacks reference outgroup gymnospermKsRates needs gymnosperm + angiospermAdd Selaginella or moss outgroup

Tool Installation Notes

bash
# wgd v2
pip install wgd
# Or: git clone https://github.com/heche-psb/wgd && cd wgd && pip install .

# KsRates
pip install ksrates

# DupGen_finder (Perl + MCScanX dependency)
git clone https://github.com/qiao-xin/DupGen_finder
# Install MCScanX (Wang 2012): git clone https://github.com/wyp1125/MCScanX && cd MCScanX && make

# MAPS (Python; requires CRG, ete3)
git clone https://bitbucket.org/barker-lab/maps

# Whale.jl (Julia)
julia -e 'using Pkg; Pkg.add("Whale")'

# SLEDGe
git clone https://github.com/SLEDGe-team/SLEDGe

# POInT (C++)
git clone https://github.com/gconant0/PoInT
cd PoInT && make

# PAML for Ks computation (yn00 method)
conda install -c bioconda paml

# DIAMOND / BLAST for paranome
conda install -c bioconda diamond blast

# R packages
install.packages(c('mclust', 'rmixmod'))

For new analyses, default to wgd v2 + KsRates as the primary pipeline; MAPS or Whale.jl for confirmation on phylogenomic placement.

References

  • Chen H et al 2024 Bioinformatics 40:btae272 (wgd v2)
  • Sensalari C et al 2022 Bioinformatics 38:530 (KsRates)
  • Qiao X et al 2019 Genome Biol 20:38 (DupGen_finder)
  • Li Z et al 2018 PNAS 115:4713 (MAPS phylogenomic WGD placement)
  • Conant GC & Wolfe KH 2008 Genetics 179:1681 (POInT)
  • Zwaenepoel A & Van de Peer Y 2019 MBE 36:1384 (Whale.jl)
  • Sutherland BL et al 2024 bioRxiv 2024.01.17.574559 (SLEDGe ML classifier)
  • Vanneste K et al 2013 MBE 30:177 (Ks saturation)
  • Holland PWH et al 1994 Development Suppl:125 (2R hypothesis)
  • Dehal P & Boore JL 2005 PLoS Biol 3:e314 (vertebrate 2R confirmation)
  • Glasauer SMK & Neuhauss SCF 2014 Mol Genet Genomics 289:1045 (teleost 3R)
  • Lien S et al 2016 Nature 533:200 (salmonid Ss4R)
  • Soltis DE et al 2009 Am J Bot 96:336 (polyploidy and angiosperm diversification)
  • Jiao Y et al 2012 Genome Biol 13:R3 (dates the core-eudicot gamma triplication)
  • Birchler JA & Veitia RA 2007 Plant Cell 19:395 (gene balance hypothesis)
  • Force A et al 1999 Genetics 151:1531 (subfunctionalization)
  • Maere S et al 2005 PNAS 102:5454 (post-WGD retention bias)
  • Tang H et al 2008 GR 18:1944 (synteny / MCScan)
  • Lovell JT et al 2022 eLife 11:e78526 (GENESPACE)
  • Jia K-H et al 2022 New Phytol 235:801 (SubPhaser subgenome assignment)
  • Yang Z 2007 PAML manual (yn00 codon Ks)
  • Chen H & Zwaenepoel A 2023 Methods Mol Biol 2545:3 (inference of ancient polyploidy from genomic data)
  • comparative-genomics/synteny-analysis - Synteny-anchored Ks (wgd syn / GENESPACE)
  • comparative-genomics/ortholog-inference - Ortholog detection feeds DupGen / wgd
  • comparative-genomics/gene-tree-species-tree-reconciliation - Whale.jl native WGD modeling; ALE with WGD branch
  • comparative-genomics/gene-family-evolution - CAFE5 birth-death often shows post-WGD retention bias
  • comparative-genomics/positive-selection - Selection on retained ohnologs (post-WGD diversification)
  • comparative-genomics/ancestral-reconstruction - Ancestral pre-WGD gene state
  • phylogenetics/divergence-dating - Time-calibrated WGD positioning
  • alignment/multiple-alignment - Codon-aware MSA for Ks computation
  • alignment/pairwise-alignment - Pairwise Ks via PAML yn00 / codeml
  • genome-assembly/assembly-qc - BUSCO / Compleasm before WGD analysis

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in comparative-genomics/whole-genome-duplication of GPTomics/bioSkills.

  • SKILL.md
  • examples/wgd_v2_ks_pipeline.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Comparative Genomics Whole Genome Duplication next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Comparative Genomics Whole Genome Duplication compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Comparative Genomics Whole Genome Duplication this skillGPTomics/bioSkills1.2k2 repos~7.4kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Comparative Genomics Whole Genome Duplication

What does Bio Comparative Genomics Whole Genome Duplication do?

Detect, date, and contextualize whole-genome duplication (WGD / paleopolyploidy) events using wgd v2 (Chen et al 2024), KsRates (Sensalari 2022 substitution-rate-corrected Ks dating), DupGenfinder…. Bio Comparative Genomics Whole Genome Duplication is an agent skill from GPTomics/bioSkills.jl (Bayesian DL+WGD), and synteny-anchored paranome construction.

When should I use Bio Comparative Genomics Whole Genome Duplication?

Bio Comparative Genomics Whole Genome Duplication fits situations like: identifying ancient polyploidy from Ks distributions and synteny block analysis; positioning WGD events relative to speciation; distinguishing tandem from segmental from WGD duplications; dating the 2R/3R vertebrate / fish / salmonid WGDs.

How do I install Bio Comparative Genomics Whole Genome Duplication in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-whole-genome-duplication -a claude-code`. Or copy the skill folder (comparative-genomics/whole-genome-duplication in GPTomics/bioSkills) into .claude/skills/bio-comparative-genomics-whole-genome-duplication in your project. Claude Code loads it when a task matches its description.

How do I install Bio Comparative Genomics Whole Genome Duplication in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-whole-genome-duplication -a codex`. Or copy the skill folder (comparative-genomics/whole-genome-duplication in GPTomics/bioSkills) into .agents/skills/bio-comparative-genomics-whole-genome-duplication in your project. Codex loads it when a task matches its description.

Can I use Bio Comparative Genomics Whole Genome Duplication in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-whole-genome-duplication -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-comparative-genomics-whole-genome-duplication, .gemini/skills/bio-comparative-genomics-whole-genome-duplication, .github/skills/bio-comparative-genomics-whole-genome-duplication and .opencode/skills/bio-comparative-genomics-whole-genome-duplication in your project.

What does Bio Comparative Genomics Whole Genome Duplication need to run?

Going by SKILL.md and its folder, Bio Comparative Genomics Whole Genome Duplication needs a shell for the scripts in its folder and the command-line tools its instructions call (git, pip, conda and make). Our summary lists: Python 3; A Bash shell.

Does Bio Comparative Genomics Whole Genome Duplication access the network?

SKILL.md names 2 domains. In commands or code: github.com and bitbucket.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Comparative Genomics Whole Genome Duplication safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Comparative Genomics Whole Genome Duplication use?

Bio Comparative Genomics Whole Genome Duplication is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Comparative Genomics Whole Genome Duplication use?

About 7.4k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Comparative Genomics Whole Genome Duplication?

Skills that share tags, products or a category with Bio Comparative Genomics Whole Genome Duplication: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Comparative Genomics Whole Genome Duplication?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.