Agent skill

Bio Comparative Genomics Introgression Detection

by GPTomics in GPTomics/bioSkills

Detect introgression and admixture between species or populations using Dsuite (Malinsky 2021 fast D-statistics), Patterson's D / ABBA-BABA test (Green 2010; Durand 2011), f4-ratio and f-branch…

MITAuto-check passedResearch & Science

Install Bio Comparative Genomics Introgression Detection

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-introgression-detection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-comparative-genomics-introgression-detection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/comparative-genomics/introgression-detection .claude/skills/bio-comparative-genomics-introgression-detection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-comparative-genomics-introgression-detection
GitHub stars
1.2k
Used in
2 other repos
Token cost
~8k tokens
SKILL.md length
3,203 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Detect introgression and admixture between species or populations using Dsuite (Malinsky 2021 fast D-statistics), Patterson's D / ABBA-BABA test (Green 2010; Durand 2011), f4-ratio and f-branch…

  • Testing inter-species gene flow
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Decision Tree by Experimental… and Per-Method Failure Modes, plus 12 more sections
  • Runs Shell scripts from its folder; calls git, python and pip; reaches github.com
  • Dating admixture events

What it does

Bio Comparative Genomics Introgression Detection is an agent skill from GPTomics/bioSkills. Detect introgression and admixture between species or populations using Dsuite (Malinsky 2021 fast D-statistics), Patterson's D / ABBA-BABA test (Green 2010; Durand 2011), f4-ratio and f-branch statistic (Malinsky 2018), TreeMix (Pickrell & Pritchard 2012), HyDe (Blischak 2018), QuIBL (Edelman 2019), sprime (Browning 2018), Twisst (Martin 2017), PhyloNet (Than 2008) for explicit phylogenetic networks, and qpAdm / qpGraph (Patterson 2012). Distinguish introgression from incomplete lineage sorting (ILS), ancestral…

Its SKILL.md is about 8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/dsuite_abba_baba_workflow.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Statistics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Testing inter-species gene flow
  • Dating admixture events
  • Identifying introgressed segments
  • Building phylogenetic networks for reticulate evolution

Example prompts

  • “/bio-comparative-genomics-introgression-detection”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • python
    • pip
    • make
    • conda
    • mvn

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Comparative Genomics Introgression Detection loads about 8k tokens when it runs. Until then it costs about 223 tokens; SKILL.md has 3,203 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~223
When it runs · the whole SKILL.md, loaded when a task matches
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,203 words, ~8,018 tokens.

Download SKILL.mdSave it as .claude/skills/bio-comparative-genomics-introgression-detection/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-comparative-genomics-introgression-detection
description
Detect introgression and admixture between species or populations using Dsuite (Malinsky 2021 fast D-statistics), Patterson's D / ABBA-BABA test (Green 2010; Durand 2011), f4-ratio and f-branch statistic (Malinsky 2018), TreeMix (Pickrell & Pritchard 2012), HyDe (Blischak 2018), QuIBL (Edelman 2019), sprime (Browning 2018), Twisst (Martin 2017), PhyloNet (Than 2008) for explicit phylogenetic networks, and qpAdm / qpGraph (Patterson 2012). Distinguish introgression from incomplete lineage sorting (ILS), ancestral structure, ghost-lineage admixture, and rate variation. Use when testing inter-species gene flow, dating admixture events, identifying introgressed segments, building phylogenetic networks for reticulate evolution, or applying the ABBAclustering (Koppetsch-Malinsky-Matschiner 2024) framework for divergent-species gene flow.
tool_type
cli
primary_tool
Dsuite

Version Compatibility

Reference examples tested with: Dsuite 0.5+ (millanek/Dsuite; Malinsky 2021 Mol Ecol Res 21:584; ABBAclustering option from Koppetsch-Malinsky-Matschiner 2024 Syst Biol), HyDe 0.4.3+ (Blischak 2018 Syst Biol 67:821), QuIBL (Edelman 2019 Science 366:594), TreeMix 1.13+ (Pickrell & Pritchard 2012 PLoS Genet 8:e1002967), sprime (Browning 2018 Cell 173:53), Twisst (Martin & Van Belleghem 2017 Genetics 206:429), PhyloNet 3.8.2+ (NakhlehLab/PhyloNet; Than-Ruths-Nakhleh 2008 BMC Bioinf 9:322) and PhyloNetworks 0.16+ (JuliaPhylo/PhyloNetworks; Solis-Lemus, Bastide & Ane 2017 MBE 34:3292), qpAdm / qpGraph (AdmixTools v2.0+; Maier 2023), ADMIXTOOLS2 R wrapper (Maier 2023 eLife 12:e85492), MaCS-like simulators (msprime 1.3+ for testing), bcftools 1.21+, samtools 1.21+, vcftools 0.1.16+, R 4.4+. See upstream Dsuite docs for visualization helpers.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: Dsuite --version; treemix --help; qpDstat --help (AdmixTools)
  • Python: pip show msprime hyde
  • R: packageVersion('admixtools')

If code throws Dsuite SETS file format error, TreeMix matrix singular, qpAdm rotation failed, these tools share strict input requirements: Dsuite needs a SETS file mapping samples to populations + an outgroup; TreeMix needs allele-frequency matrix; qpAdm needs ind / snp / geno trio (EIGENSTRAT format).

Introgression and Admixture Detection

"Has there been gene flow between these species / populations?" -> Tests for inter-population admixture span site-frequency (ABBA-BABA, f4), tree-topology (Dsuite f-branch, QuIBL, Twisst), explicit-network (PhyloNet), and haplotype-tract (sprime) approaches. The fundamental confounder is incomplete lineage sorting (ILS): under a symmetric tree with no gene flow, the ABBA and BABA patterns occur with equal frequency from ancestral polymorphism. A significant D-statistic indicates EITHER (a) introgression, OR (b) ancestral structure, OR (c) sampling from a ghost lineage -- additional evidence is required to distinguish (Green 2010 Science 328:710; Durand 2011 MBE 28:2239; Eriksson & Manica 2012 PNAS 109:13956). For high-confidence claims, combine D-statistic with f-branch mapping (assigns admixture to specific branches), Twisst / QuIBL topology-weighting, and TreeMix migration edges.

  • CLI: Dsuite Dtrios -- standard D-statistic for trios of populations
  • CLI: Dsuite Fbranch -- assign admixture signal to specific tree branches
  • CLI: treemix -i freq.gz -k 1000 -m 0 -o out -- migration edges
  • CLI: qpAdm (AdmixTools) -- rotation-based admixture proportion testing
  • R: admixtools::qpadm() modern wrapper
  • CLI: HyDe -i sequences -t taxa.txt --hyptest -- hybridization detection at site level
  • CLI: python QuIBL.py inputfile.txt (config file with treefile path + parameters) -- topology weighting for ILS vs introgression

Algorithmic Taxonomy

Tool / StatisticApproachOutputStrengthFails when
Patterson's D / ABBA-BABA (Green 2010 Science 328:710; Durand 2011 MBE 28:2239)Counts ABBA vs BABA site patterns in fixed tree (((P1, P2), P3), Out)D = (ABBA - BABA) / (ABBA + BABA)Standard introgression test; tractableD > 0 means EITHER admixture OR ancestral structure OR ghost lineage
Dsuite Dtrios (Malinsky 2021 Mol Ecol Res 21:584)Fast D for all population triosD + jackknife SE + p-valueScales to many populations; integrated with phylogenetic treeInherits ABBA-BABA assumptions; symmetric tree only
f4-ratio statistic (Patterson 2012 Genetics 192:1065)Quartet-based admixture proportionAdmixture proportion alphaCleaner than D for inferring admixture amountRequires correct phylogeny; ILS confound
Dsuite Fbranch (Malinsky 2018 Nat Eco Evo 2:1940)Tree-aware f4-ratio mappingBranch-specific admixture signalBest at distinguishing admixture among related lineagesTree must be reliable; ghost branches not detectable
ABBAclustering (Koppetsch-Malinsky-Matschiner 2024 Syst Biol)Dsuite option for divergent-species gene flowCluster-based admixtureDesigned for cases where standard ABBA-BABA failsNewer; specific use case
TreeMix (Pickrell & Pritchard 2012 PLoS Genet 8:e1002967)ML tree with migration edges from drift covarianceTree + migration edges with weightsVisualizes complex admixtureMigration-edge selection subjective; allele-frequency-based
qpAdm (Patterson 2012; AdmixTools v2.0; Maier 2023 eLife 12:e85492)Tests if target derives from sources via rotationPass/fail per source set + admixture proportionRobust statistical framework; rotation controlSource population sampling critical
qpGraph (Patterson 2012)ML phylogenetic graph with admixtureBest-fit graph with admixture nodesQuantitative; explicit tree+admixture modelComputationally heavy; manual graph topology design
HyDe (Blischak 2018 Syst Biol 67:821)Site-level hybridization detectionPer-site / per-locus hybrid evidenceSensitive to recent hybridizationLess specific to ancient admixture
QuIBL (Edelman 2019 Science 366:594)Topology weighting on gene treesPer-tree-weight relative to alternativeDistinguishes ILS from introgression at locus levelPer-locus inference; cross-validation needed
Twisst (Martin & Van Belleghem 2017 Genetics 206:429)Topology weighting on phylogenetic treesTree topology weights per genomic windowVisualize topology variation across genomeComputational cost; tree-window decisions
PhyloNetworks (Solis-Lemus, Bastide & Ane 2017 MBE 34:3292; Julia, JuliaPhylo/PhyloNetworks)SNaQ pseudo-likelihood network inference from gene trees / concordance factorsExplicit reticulation networkModern Julia ecosystem; SNaQ scales wellDifferent software than the Java "PhyloNet"
PhyloNet (Than-Ruths-Nakhleh 2008 BMC Bioinf 9:322; Java, NakhlehLab/PhyloNet)ML / MP / Bayesian network inferenceExplicit reticulation networkMature Java tool; many inference modesComputationally heavy; Maven build
sprime (Browning 2018 Cell 173:53)Per-individual archaic introgression detectionIntrogressed haplotype tractsDesigned for archaic-human-like casesSpecific to closely related introgression source
Relate (Speidel 2019 Nat Genet 51:1321)Phasing-aware genome-wide genealogy reconstructionCoalescence times dating introgression segmentsModern haplotype-aware method; times admixture via tract coalescenceRequires phased data; genealogy timing indirect for admixture
F3 statistic (Patterson 2012)Three-population test for admixtureF3 with SESpecifically detects admixture (vs negative drift)F3 < 0 indicates admixture; mixed signals
F4-statistic (Patterson 2012)Four-population testF4 + p-valueGeneralizes D; allows variable outgroup distanceSymmetric assumption

Methodology evolves; verify the Dsuite documentation and Malinsky 2024 review (eLife) before locking on a single approach. The combination of (1) Dsuite D + Fbranch + (2) Twisst / QuIBL + (3) network method (PhyloNet) is the modern best practice for publication-grade introgression claims.

Decision Tree by Experimental Scenario

ScenarioRecommended approachWhy
Test for introgression between two species, outgroup availableDsuite Dtrios + FbranchStandard ABBA-BABA + tree-aware branch mapping
Multi-population complex admixture inferenceTreeMix + qpAdmMigration edges + rotation tests
Distinguish introgression from ILSQuIBL + Twisst across many lociTopology weighting at locus level
Date the admixture eventRelate (Speidel 2019) + Twisst window analysisGenealogy reconstruction + haplotype-length-distribution dating + topology
Quantify admixture proportion alphaf4-ratio (Patterson 2012) or qpAdmStandard framework
Identify introgressed genomic regionssprime + TwisstPer-individual + per-window
Explicit phylogenetic networkPhyloNet with model selection (DLRS-NL or InferNetwork_MP)Quantitative reticulation
Test for ghost-lineage admixtureqpGraph manual topology + AdmixTools rotationCompare alternative graphs
Recent hybridizationHyDe + per-individual analysisSite-level hybrid evidence
Ancient introgressionDsuite + qpAdm; Patterson f-statistics for archaicTree-aware methods preferred
Plant-plant hybridization with parental polyploidyHyDe + sprime + chromosome-level analysisPolyploid context
Human-archaic introgression (Neanderthal, Denisovan)sprime + qpAdm + ChromoPainterSample-specific haplotype analysis
Cichlid radiationDsuite + TreeMix + QuIBL (Malinsky 2018 Nat Eco Evo 2:1940)Established workflow
Drosophila species groupDsuite + Twisst + qpAdmStandard workflow
Genomic regions of introgression vs vertical inheritanceTwisst per-window + chromosome paintingVisualization of variation across genome
Test for introgression direction (P1 -> P2 vs P2 -> P1)qpAdm rotation; Twisst directionality from phasingqpAdm handles direction better than D
Symmetric phylogeny violation (no clear outgroup)f3 statistic + qpGraph topologyAvoid D-statistic if outgroup unclear
Suspect ABBA-BABA ILS-confoundedSwitch to QuIBL or Twisst for locus-level discriminationThese distinguish ILS from introgression

Per-Method Failure Modes

ILS confounded with introgression in D-statistic

Trigger: D-statistic significantly different from zero for a population trio.

Mechanism: Under symmetric phylogeny with no gene flow, ABBA and BABA patterns occur with equal frequency from ancestral polymorphism (ILS). Random allele sorting in descendant lineages produces the same site patterns as introgression. D > 0 is therefore consistent with introgression OR with asymmetric ILS OR with ancestral structure (Eriksson & Manica 2012 PNAS 109:13956).

Symptom: Many tested trios show D > 0 with low effect size; Twisst / QuIBL on the same loci shows mixed topologies; phylogenetic distance to outgroup is large.

Fix: Use Dsuite f-branch (Malinsky 2018) to assign admixture to specific branches; verify with QuIBL or Twisst topology weighting at the locus level. ABBAclustering (Koppetsch-Malinsky-Matschiner 2024) specifically handles divergent-species gene flow.

Ghost lineage admixture mimicking signal

Trigger: Inferring P1 -> P2 admixture without considering that the admixed lineage could come from an unsampled extinct sister to P3.

Mechanism: D > 0 with P3 as donor is identical to D > 0 with a ghost sister of P3 as donor. The data cannot distinguish without explicit modeling.

Symptom: Hypothesized donor population doesn't match expected geography or ecology; qpGraph cannot fit a parsimonious admixture model without invoking ghost lineages.

Fix: Build qpGraph with potential ghost lineages explicitly; report multiple compatible scenarios; consult known paleontology / biogeography to constrain possibilities. The Neanderthal-Denisovan-anatomically-modern-human framework is heavily ghost-lineage-aware (Slon 2018; Mafessoni 2020).

Ancestral structure (population subdivision before admixture)

Trigger: Significant D-statistic in a recently diverged radiation.

Mechanism: Ancestral subdivision in the ancestor of (P1, P2, P3) produces D > 0 from biased ancestral allele inheritance, not from gene flow between descendants.

Symptom: D positive across multiple trios but pattern not consistent with simple admixture; effect size differs from f-statistic predictions; rapid radiation context.

Fix: Examine F-statistics framework (f3, f4); use qpGraph with structured ancestry; consider PhyloNet for explicit reticulate model. Eriksson & Manica 2012 PNAS 109:13956 documents this confound.

Outgroup-distance effect on D-statistic

Trigger: Using a distantly related outgroup (>200 Myr) for the D-test.

Mechanism: Long branches to the outgroup accumulate convergent substitutions; some sites called ABBA / BABA are actually convergent across branches, not true ABBA / BABA patterns.

Symptom: D-statistic at convergent-evolution-prone sites is elevated; D varies with outgroup choice.

Fix: Use a closer outgroup (< 100 Myr); avoid extremely deep outgroups. Mask sites with extreme conservation across the four taxa. Standard practice in animal species groups uses Drosophila simulans / yakuba for D. melanogaster trios.

Sample-size bias in D-statistic

Trigger: P1 and P2 have very different sample sizes (e.g. 1 vs 50).

Mechanism: Allele frequencies in small samples are more extreme; counting ABBA/BABA across genome is dominated by the population with larger sample.

Symptom: D-statistic dominated by the larger-sample population; jackknife SE underestimated.

Fix: Use one individual per population (Dsuite default for Dtrios); or downsample to consistent N; report effective per-population sample size.

TreeMix migration-edge selection

Trigger: Choosing number of migration edges via likelihood ratio or AIC alone.

Mechanism: TreeMix likelihood increases with each migration edge added (more parameters); the "right" number requires statistical justification, but published criteria vary.

Symptom: Different selection criteria yield different optimal migration edge counts; results sensitive to k-block size (in jackknife).

Fix: Use OptM (Fitak 2021 Biol Methods Protoc 6:bpab017) for principled migration-edge selection; report cross-validation across k-block sizes. Combine with qpAdm rotation tests as independent corroboration.

qpAdm rotation failures

Trigger: qpAdm reporting "rotation rejected" for proposed source population set.

Mechanism: qpAdm tests whether target population can be modeled as admixture of source populations; rotation tests check that swapping sources for the outgroup set doesn't break the model. Rotation failure means the proposed sources don't capture the actual admixture.

Symptom: qpAdm fits with feasible-sounding sources, but rotation indicates poor fit; admixture proportions sum to weird values.

Fix: Reconsider source populations; include "right" populations representing ancestral structures; document rotation results. AdmixTools v2.0 (Maier 2023) provides better diagnostics.

Multiple-testing correction across population trios

Trigger: Running Dsuite Dtrios across all population trios genome-wide and reporting significant ones.

Mechanism: With N populations, there are N choose 4 quartets to test; with 50 populations, ~1.5M tests. Naive 5% Type-I gives 75,000 false positives.

Symptom: Many "significant" trios reported but with low effect sizes; biological interpretation impossible.

Fix: Apply FDR correction across trios; or restrict to a priori hypothesized trios. Dsuite computes Bonferroni-corrected p-values via -c flag.

Genomic-window choice in Twisst / QuIBL

Trigger: Running Twisst on 50 kb windows vs 100 kb windows giving different results.

Mechanism: Smaller windows have more topology stochasticity; larger windows average across recombination events. The "correct" window depends on local recombination rate.

Symptom: Twisst conclusions reverse with window size; spatial pattern of introgression unstable.

Fix: Choose windows matching local LD decay (typically 10-100 kb for vertebrates; 5-20 kb for Drosophila). Report sensitivity to window choice. Combine with Twisst-explore.py for spatial visualization.

Show full SKILL.md (1,306 more words)Show less
Phylogenetic network not unique

Trigger: PhyloNet inferring different reticulation networks with similar likelihood.

Mechanism: Phylogenetic networks are non-unique; multiple networks can have similar likelihoods. PhyloNet's MP / ML criterion doesn't disambiguate beyond a tolerance.

Symptom: PhyloNet returns several networks within delta-LRT = 5; biological interpretation differs.

Fix: Report top-N networks; combine with f-statistics (qpAdm / qpGraph) as independent confirmation; use biological knowledge (geography, ecology) to constrain.

sprime contamination from non-archaic introgression source

Trigger: sprime applied to populations with both ancient and recent admixture sources.

Mechanism: sprime detects archaic-like haplotypes by comparing to outgroup; haplotype tracts from any non-outgroup-resembling source produce false positives.

Symptom: sprime calls many haplotype tracts in populations with known recent introgression that's not from the "archaic" target.

Fix: Use sprime only when sources are well-characterized; cross-validate with paired analysis using known archaic vs known recent sources separately.

Quantitative Thresholds

QuantityThresholdSource / Rationale
Patterson's D significancejackknife z-scorez
D-statistic interpretationD
f4-ratio admixture proportion0 <= alpha <= 1Patterson 2012
Dsuite Fbranch significancejackknife z > 3 per branchMalinsky 2018
TreeMix migration weighttypically 0.05-0.30 for biologically relevantPickrell 2012
TreeMix k-block size500-1000 SNPs defaultEmpirical
qpAdm sources passing rotationrotation-corrected p > 0.05Patterson 2012; Maier 2023
qpAdm admixture proportion0 <= alpha <= 1 with SEStandard
QuIBL trio testbest-trio topology weight > 0.5 + significant zEdelman 2019
Twisst window size10-100 kb (vertebrate); 5-20 kb (Drosophila)Empirical
Per-window Twisst topology weightdominant topology > 50%Standard
HyDe per-individual significanceLRT p < 0.05 vs no-hybridizationBlischak 2018
sprime archaic haplotype length> 50 kb typical archaic; < 20 kb questionableBrowning 2018
Sample size per population (Dsuite)1+ (defaults to one per pop); 5+ preferredEmpirical
Outgroup distance for D-stat< 100 Myr (close); 50-200 Myr (typical)Empirical
FDR threshold across triosq < 0.05 (BH)Standard
Bonferroni for N choose 4 triosadjust per number testedStandard
f3 statistic for admixturef3 < 0 indicates admixturePatterson 2012
ABBAclustering thresholddepends on clustering algorithmKoppetsch 2024
Phylogenetic network tolerancedelta-LRT < 2-5 for "similar"Standard

Dsuite Standard Workflow

Goal: Compute Patterson's D and Fbranch across all population trios from a VCF.

Approach: Prepare SETS file mapping samples to populations -> Dsuite Dtrios -> Dsuite Fbranch -> visualize.

bash
# 1. Prepare SETS file (samples to populations)
cat > SETS.tsv << 'EOF'
sample_id    population
A1           Population_A
A2           Population_A
B1           Population_B
B2           Population_B
C1           Population_C
D1           Outgroup
EOF

# 2. Run Dsuite Dtrios
# Real CLI flags (verify with `Dsuite Dtrios --help` against installed version):
#   -t/--tree=FILE      species tree (Newick)
#   -o/--out-prefix=    output prefix
#   -k/--no-f4-ratio    skip f4-ratio computation
#   -c/--no-combine     do not write the _combine.txt file
# Fbranch is its own subcommand (`Dsuite Fbranch`); no Dtrios flag toggles it.
Dsuite Dtrios \
    --tree=species_tree.nwk \
    -o trios_run \
    population_genotypes.vcf.gz \
    SETS.tsv

# Output: trios_run_BBAA.txt, trios_run_Dmin.txt, trios_run_tree.txt, trios_run_combine.txt

# 3. Run Fbranch for tree-aware admixture mapping
Dsuite Fbranch species_tree.nwk trios_run_tree.txt > trios_run_fbranch.txt

# 4. ABBAclustering test (Koppetsch-Malinsky-Matschiner 2024).
# Implementation specifics depend on the Dsuite branch / fork distributing the ABBAclustering option;
# consult the Koppetsch 2024 supplement and `Dsuite --help` for current invocation. The option may
# be exposed as a subcommand or per-trio flag rather than a global Dtrios switch.
python
'''Parse Dsuite output for per-trio admixture signal.'''
import pandas as pd


def load_dsuite_trios(path):
    '''SETS.tsv_BBAA.txt columns: P1, P2, P3, D, p-value, Z, BBAA, ABBA, BABA, f_d, f_dM, df, ...'''
    df = pd.read_csv(path, sep='\t')
    df['admixture_signal'] = (df['Z'] > 3) & (df['p_value'] < 0.05)
    return df


def filter_top_admixture(df, by='Z', top=20):
    return df.sort_values(by, ascending=False).head(top)

TreeMix for Population Tree with Migration Edges

Goal: Build a population tree with migration edges showing admixture.

Approach: Convert VCF to allele-frequency matrix -> run TreeMix -> select migration count via OptM.

bash
# 1. Convert VCF to TreeMix input
python tree_mix_input.py --vcf population.vcf.gz --popmap popmap.tsv \
    --output input.frq.gz

# 2. Run TreeMix
treemix -i input.frq.gz -m 0 -bootstrap -k 1000 -o output_m0
for m in 1 2 3 4 5; do
    treemix -i input.frq.gz -m $m -bootstrap -k 1000 -o output_m${m}
done

# 3. Select optimal m via OptM (R)
Rscript -e "
library(OptM)
# optM(folder, method, ...): folder is the TreeMix output directory (not the input frq.gz).
# Accepts method='Evanno' (delta-m), 'linear', or 'SiZer'. See ?OptM::optM for details.
optM_result <- optM(folder='.', method='Evanno', tsv=NULL)
"

QuIBL for Locus-Level ILS vs Introgression

Goal: Distinguish ILS from introgression at the locus level via topology weighting.

Approach: Build per-locus phylogenetic trees -> QuIBL on tree list.

bash
# Build per-locus gene trees
for region in regions/*.fa; do
    iqtree2 -s $region -m GTR+G -B 1000 -nt 2 --prefix gene_trees/$(basename $region .fa)
done

# Combine into tree list
cat gene_trees/*.treefile > combined_trees.txt

# Run QuIBL: input is a config file pointing at the tree list + parameters
cat > quibl_input.txt << 'EOF'
treefile: combined_trees.txt
outputfile: quibl_results.txt
likelihoodthresh: 0.01
gradascentscalar: 0.5
totaloutgroup: outgroup_taxon
overallnumlambda: 2
EOF
python QuIBL.py quibl_input.txt

qpAdm Rotation Test

Goal: Test if target population is admixed from candidate sources.

Approach: Define source set + outgroups; AdmixTools qpAdm rotation tests source population set.

bash
# AdmixTools formatted input: ind / snp / geno (EIGENSTRAT)
# Or use admixtools R wrapper

# In R:
Rscript -e "
library(admixtools)
# Setup left (sources) + right (outgroups)
left <- c('SourceA', 'SourceB')
right <- c('Outgroup1', 'Outgroup2', 'Outgroup3')
target <- 'AdmixedPop'

results <- qpadm(prefix='eigenstrat_data', target=target,
                 left=left, right=right)
results$rankdrop  # rotation test rejection
results$weights   # admixture proportions
"

Reconciliation: When Methods Disagree

PatternLikely causeAction
D-statistic significant, Twisst dominant topology supports species treeILS confoundTrust Twisst; D was ILS-driven
D significant, Fbranch attributes to specific branchTrue introgressionConfirmed admixture
qpAdm fails rotation; TreeMix doesn't show migration edgeInconsistent admixture scenarios; insufficient dataReconsider sources; report exploratory
QuIBL supports introgression, D-statistic nullD-statistic underpowered or sample-size biasedTrust QuIBL locus-level
TreeMix migration edge with TreeMix likelihood gain only with high mOverfittingUse OptM; restrict to robust edges
PhyloNet network is non-uniqueMultiple valid reticulation modelsReport all; use prior knowledge to constrain
sprime calls many tracts in non-archaic-source populationFalse positive from non-archaicRestrict sprime to populations with known archaic source
HyDe site-level signal at conserved genesConvergent evolution masquerading as hybridizationFilter conserved sites; verify with non-coding regions
D-statistic stable; Fbranch shows nothingAll admixture attributable to a single ghost branchGhost-lineage candidate

Operational rule for publication: Patterson D + Fbranch + Twisst/QuIBL window-level + at least one network method (PhyloNet or qpGraph) agreeing on admixture; multiple-testing correction across trios; ILS quantified via Twisst or coalescent simulation; explicit acknowledgement of ghost-lineage / ancestral-structure alternatives.

Cohort Gotchas

  • Recent radiations: ILS confounded with introgression; require additional evidence beyond D-statistic
  • Hybridization vs introgression: hybridization is contemporary; introgression is ancestral admixture
  • Plant polyploids: subgenome assignment first; cross-subgenome "introgression" is typically homeologous
  • Reduced-genome organisms: low marker density; ABBA-BABA underpowered
  • Sex chromosomes: non-recombining; restrict introgression tests to autosomal data
  • Distantly related introgression source (>200 Myr): sample more carefully; convergent substitutions confounding
  • Cytonuclear discordance: mtDNA introgression may differ from autosomal; report separately
  • Migration vs introgression: D-statistic detects allele introgression; demographic models needed for migration history
  • Rapid radiation example: cichlids (Malinsky 2018 Nat Eco Evo 2:1940); standard workflow documented

Anticipated Reviewer Pushback

PushbackStandard response
"ILS vs introgression?"Multi-method approach: D + Fbranch + Twisst + QuIBL; ILS-confounded loci identified
"Ghost lineage?"qpGraph models tested with potential ghost lineages; biological context constrains; report alternatives
"Ancestral structure?"qpAdm rotation rules out simple admixture from sampled sources; Eriksson & Manica 2012 alternative explicitly considered
"Outgroup choice?"Outgroup distance reported; sensitivity tested with multiple outgroups
"Multiple-testing correction?"FDR across population trios; or restricted to a priori hypothesized
"TreeMix migration count?"OptM Evanno method; cross-validated across k-block sizes
"qpAdm sources?"Rotation tests passed; multiple alternative source sets evaluated
"Twisst window size?"Matched local LD decay; sensitivity tested
"Direction of introgression?"qpAdm directional + Twisst phasing data
"Date of admixture?"Relate genealogy + tract-length-distribution-based dating; CIs reported

Common Errors

Error / symptomCauseSolution
Dsuite SETS file emptyMissing samples in VCFVerify VCF + SETS sample names
TreeMix matrix singularInsufficient SNPs or perfect collinearityIncrease k-block size; remove duplicates
qpDstat AdmixTools failsEIGENSTRAT format issueRe-convert with convertf
HyDe per-individual no resultDefault thresholds too strictRelax --signal threshold
QuIBL slowMany treesParallelize; reduce per-tree size
Twisst output emptyWindow-tree mismatchVerify gene trees for each window
PhyloNet OOMMany taxa; network complexityReduce taxon set; restrict to candidate reticulations
sprime no haplotypesPhasing or outgroup issueVerify input phased; check outgroup
D-statistic sign confusedAllele orientation issueUse derived alleles consistently
qpAdm convergence failureSource overlap or rare allelesFilter MAF > 0.05; check source overlap

Tool Installation Notes

bash
# Dsuite
git clone https://github.com/millanek/Dsuite && cd Dsuite && make
# Or: conda install -c bioconda dsuite

# TreeMix
conda install -c bioconda treemix

# AdmixTools (qpDstat, qpAdm, qpGraph)
git clone https://github.com/DReichLab/AdmixTools && cd AdmixTools && make

# Modern AdmixTools v2 (R)
remotes::install_github('uqrmaie1/admixtools')

# HyDe
git clone https://github.com/pblischak/HyDe && cd HyDe && pip install -e .

# QuIBL
git clone https://github.com/miriamtnzr/QuIBL

# Twisst
git clone https://github.com/simonhmartin/twisst && pip install -e .

# PhyloNet
git clone https://github.com/NakhlehLab/PhyloNet && cd PhyloNet && mvn package

# sprime
conda install -c bioconda sprime

For population-genetic analyses, the Dsuite + AdmixTools v2 (R) + TreeMix combination is standard. PhyloNet is heavier and requires Maven build.

References

  • Green RE et al 2010 Science 328:710 (Patterson D / ABBA-BABA, Neanderthal admixture)
  • Durand EY et al 2011 MBE 28:2239 (D-statistic theoretical framework)
  • Patterson N et al 2012 Genetics 192:1065 (f-statistics framework)
  • Maier R et al 2023 eLife 12:e85492 (AdmixTools v2 / admixtools R)
  • Pickrell JK & Pritchard JK 2012 PLoS Genet 8:e1002967 (TreeMix)
  • Malinsky M et al 2021 Mol Ecol Resources 21:584 (Dsuite)
  • Malinsky M et al 2018 Nat Eco Evo 2:1940 (f-branch statistic; Lake Malawi cichlid radiation)
  • Koppetsch T et al 2024 Syst Biol (ABBAclustering)
  • Edelman NB et al 2019 Science 366:594 (QuIBL)
  • Martin SH & Van Belleghem SM 2017 Genetics 206:429 (Twisst)
  • Blischak PD, Chifman J, Wolfe AD & Kubatko LS 2018 Syst Biol 67:821 (HyDe)
  • Solis-Lemus C, Bastide P & Ane C 2017 MBE 34:3292 (PhyloNetworks / SNaQ)
  • Than C, Ruths D & Nakhleh L 2008 BMC Bioinf 9:322 (PhyloNet)
  • Browning SR et al 2018 Cell 173:53 (sprime archaic introgression)
  • Eriksson A & Manica A 2012 PNAS 109:13956 (ancestral structure produces D-stat without admixture); Soraggi S et al 2018 G3 8:551 (D-statistic with low-coverage data)
  • Slon V et al 2018 Nature 561:113 (Denisovan-Neanderthal hybrid)
  • Mafessoni F et al 2020 PNAS 117:15132 (high-coverage Chagyrskaya Neanderthal genome)
  • Speidel L et al 2019 Nat Genet 51:1321 (Relate genealogy)
  • Fitak RR 2021 Biol Methods Protoc 6:bpab017 (OptM)
  • Lawson DJ et al 2012 PLoS Genet 8:e1002453 (ChromoPainter; haplotype-based admixture)
  • Patin E et al 2017 Science 356:543 (sub-Saharan admixture example)
  • Frantz LAF et al 2019 PNAS 116:17231 (animal domestication admixture)
  • comparative-genomics/hgt-detection - Distinguish HGT from hybridization in microbes
  • comparative-genomics/gene-tree-species-tree-reconciliation - Reconciliation alternative for introgression
  • comparative-genomics/synteny-analysis - Per-window topology analysis
  • population-genetics/population-structure - PCA / ADMIXTURE precedes introgression testing
  • population-genetics/selection-statistics - Selection on introgressed regions
  • population-genetics/linkage-disequilibrium - LD haplotype context for sprime
  • phylogenetics/modern-tree-inference - Gene-tree inference for QuIBL / Twisst
  • phylogenetics/species-trees - Coalescent species tree under ILS
  • causal-genomics/heritability-partitioning - Inheritance partitioning by genomic region
  • variant-calling/joint-calling - Multi-sample VCF for introgression analysis
  • read-alignment/bwa-alignment - Alignment underlies variant calling for D-statistic

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in comparative-genomics/introgression-detection of GPTomics/bioSkills.

  • SKILL.md
  • examples/dsuite_abba_baba_workflow.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Comparative Genomics Introgression Detection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Comparative Genomics Introgression Detection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Comparative Genomics Introgression Detection this skillGPTomics/bioSkills1.2k2 repos~8kAutomated safety check: PassMIT
PyDESeq2 Differential Expressiondavila7/claude-code-templates32k11 repos~4kAutomated safety check: PassMIT
Ukb Ppp Region FetchClawBio/ClawBio1.2k—~4.6kAutomated safety check: PassMIT
TiledbvcfK-Dense-AI/scientific-agent-skills48k1 repos~3.5kAutomated safety check: PassMIT
Volcano Plot Scriptaipoch/medical-research-skills2k—~2.5kAutomated safety check: PassMIT
Tooluniverse Epigenomicswu-yc/LabClaw1.1k2 repos~14kAutomated safety check: PassNone

Similar skills

  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    32k GitHub starsUsed in 11 repos~4k tokens
    Research & ScienceAuto-check passed
  • Ukb Ppp Region Fetch

    ClawBio/ClawBio

    Fetch a regional slice of plasma pQTL summary statistics from the UK Biobank Pharma Proteomics Project (UKB-PPP; Sun 2023 Nature) for a specific (protein, ancestry) measurement.

    1.2k GitHub stars~4.6k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Tiledbvcf

    K-Dense-AI/scientific-agent-skills

    Stores and retrieves genomic variant calls with TileDB-VCF. An agent skill from K-Dense-AI/scientific-agent-skills.

    48k GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed
  • Volcano Plot Script

    aipoch/medical-research-skills

    Generate R/Python code for volcano plots from DEG (Differentially Expressed Genes) analysis results.

    2k GitHub stars~2.5k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • Analyze metabolomics data including metabolite identification, quantification, pathway analysis, and metabolic flux.

    1.1k GitHub starsUsed in 2 repos~5.9k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Comparative Genomics Introgression Detection

What does Bio Comparative Genomics Introgression Detection do?

Detect introgression and admixture between species or populations using Dsuite (Malinsky 2021 fast D-statistics), Patterson's D / ABBA-BABA test (Green 2010; Durand 2011), f4-ratio and f-branch…. Bio Comparative Genomics Introgression Detection is an agent skill from GPTomics/bioSkills. Detect introgression and admixture between species or populations using Dsuite (Malinsky 2021 fast D-statistics), Patterson's D / ABBA-BABA test (Green 2010; Durand 2011), f4-ratio and f-branch statistic (Malinsky 2018), TreeMix (Pickrell & Pritchard 2012), HyDe (Blischak 2018), QuIBL (Edelman 2019), sprime (Browning 2018), Twisst (Martin 2017), PhyloNet (Than 2008) for explicit phylogenetic networks, and qpAdm / qpGraph (Patterson 2012).

When should I use Bio Comparative Genomics Introgression Detection?

Bio Comparative Genomics Introgression Detection fits situations like: testing inter-species gene flow; dating admixture events; identifying introgressed segments; building phylogenetic networks for reticulate evolution.

How do I install Bio Comparative Genomics Introgression Detection in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-introgression-detection -a claude-code`. Or copy the skill folder (comparative-genomics/introgression-detection in GPTomics/bioSkills) into .claude/skills/bio-comparative-genomics-introgression-detection in your project. Claude Code loads it when a task matches its description.

How do I install Bio Comparative Genomics Introgression Detection in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-introgression-detection -a codex`. Or copy the skill folder (comparative-genomics/introgression-detection in GPTomics/bioSkills) into .agents/skills/bio-comparative-genomics-introgression-detection in your project. Codex loads it when a task matches its description.

Can I use Bio Comparative Genomics Introgression Detection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-introgression-detection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-comparative-genomics-introgression-detection, .gemini/skills/bio-comparative-genomics-introgression-detection, .github/skills/bio-comparative-genomics-introgression-detection and .opencode/skills/bio-comparative-genomics-introgression-detection in your project.

What does Bio Comparative Genomics Introgression Detection need to run?

Going by SKILL.md and its folder, Bio Comparative Genomics Introgression Detection needs a shell for the scripts in its folder and the command-line tools its instructions call (git, python, pip, make, conda and mvn). Our summary lists: Python 3; A Bash shell.

Does Bio Comparative Genomics Introgression Detection access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Comparative Genomics Introgression Detection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Comparative Genomics Introgression Detection use?

Bio Comparative Genomics Introgression Detection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Comparative Genomics Introgression Detection use?

About 8k tokens (SKILL.md is roughly 32k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Comparative Genomics Introgression Detection?

Skills that share tags, products or a category with Bio Comparative Genomics Introgression Detection: PyDESeq2 Differential Expression (davila7/claude-code-templates, 32k stars), Ukb Ppp Region Fetch (ClawBio/ClawBio, 1.2k stars), Tiledbvcf (K-Dense-AI/scientific-agent-skills, 48k stars) and Volcano Plot Script (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Comparative Genomics Introgression Detection?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.