Agent skill

Bio Comparative Genomics Positive Selection

by GPTomics in GPTomics/bioSkills

Detect positive (diversifying / episodic / pervasive) selection using codon dN/dS frameworks.

MITAuto-check passedResearch & Science

Install Bio Comparative Genomics Positive Selection

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-positive-selection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-comparative-genomics-positive-selection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/comparative-genomics/positive-selection .claude/skills/bio-comparative-genomics-positive-selection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-comparative-genomics-positive-selection
GitHub stars
1.2k
Used in
2 other repos
Token cost
~9.7k tokens
SKILL.md length
4,024 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Detect positive (diversifying / episodic / pervasive) selection using codon dN/dS frameworks.

  • Testing adaptive evolution at codons
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Decision Tree by Experimental… and Per-Method Failure Modes, plus 12 more sections
  • Runs Python scripts from its folder; calls python, pip and conda; reaches benhaller.com and web.cbio.uct.ac.za
  • Running GARD recombination pre-screen

What it does

Bio Comparative Genomics Positive Selection is an agent skill from GPTomics/bioSkills. Detect positive (diversifying / episodic / pervasive) selection using codon dN/dS frameworks. Implements PAML codeml site models (M0/M1a/M2a/M7/M8/M8a), branch models, branch-site model A (Zhang 2005), and HyPhy methods (BUSTED, BUSTED-S, BUSTED-MH, BUSTED-PH, MEME, FEL, FUBAR, aBSREL, SLAC, RELAX, GARD, FUBAR-MH). Includes McDonald-Kreitman framework (asymptotic alpha, impMKT, polyDFE, DFE-alpha, GRAPES) for within-species + divergence inference, RERconverge for trait-correlated rate shifts, CSUBST for…

Its SKILL.md is about 9.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/selection_analysis.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Testing adaptive evolution at codons
  • Running GARD recombination pre-screen
  • Controlling alignment-error and gBGC false positives
  • Reconciling PAML vs HyPhy results

Example prompts

  • “/bio-comparative-genomics-positive-selection”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • benhaller.com
    • web.cbio.uct.ac.za
    • bioweb.supagro.inra.fr

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Comparative Genomics Positive Selection loads about 9.7k tokens when it runs. Until then it costs about 218 tokens; SKILL.md has 4,024 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~218
When it runs · the whole SKILL.md, loaded when a task matches
~9.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 4,024 words, ~9,654 tokens.

Download SKILL.mdSave it as .claude/skills/bio-comparative-genomics-positive-selection/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-comparative-genomics-positive-selection
description
Detect positive (diversifying / episodic / pervasive) selection using codon dN/dS frameworks. Implements PAML codeml site models (M0/M1a/M2a/M7/M8/M8a), branch models, branch-site model A (Zhang 2005), and HyPhy methods (BUSTED, BUSTED-S, BUSTED-MH, BUSTED-PH, MEME, FEL, FUBAR, aBSREL, SLAC, RELAX, GARD, FUBAR-MH). Includes McDonald-Kreitman framework (asymptotic alpha, impMKT, polyDFE, DFE-alpha, GRAPES) for within-species + divergence inference, RERconverge for trait-correlated rate shifts, CSUBST for convergent substitution, and PhyloAcc for accelerated noncoding evolution. Use when testing adaptive evolution at codons, branches, or full gene; running GARD recombination pre-screen; controlling alignment-error and gBGC false positives; reconciling PAML vs HyPhy results; or performing genome-scale selection scans.
tool_type
mixed
primary_tool
PAML

Version Compatibility

Reference examples tested with: PAML 4.10.7+, HyPhy 2.5.62+ (BUSTED-MH from Lucaci 2023 MBE 40:msad150; FUBAR-MH from same), datamonkey.org 2024+ for web jobs, IQ-TREE 2.3.6+, MACSE V2.07+, PRANK 170427+, MAFFT 7.526+, PREQUAL 1.02+, HmmCleaner 0.243+, GARD (HyPhy bundled), RDP5 5.59+, ete4 4.1.0+, BioPython 1.84+, scipy 1.13+, polyDFE 2.0+, DFE-alpha 2.16+, GRAPES 1.1.1+, RERconverge 0.3.0+, CSUBST 1.6.0+, PhyloAcc 2.4.0+. Quest-for-Selection benchmark refreshed annually.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: codeml (PAML; check by codeml /dev/null -- prints version banner), hyphy --version, gard --help
  • Python: pip show pyhyphy; introspect ete4 API for tree-labeling
  • R: packageVersion('RERconverge'); ?correlateWithBinaryPhenotype

If code throws branch-site test LRT non-positive, omega2 hit upper bound 999, MEME ML mixed gradient, the most common cause is alignment error or saturated dS -- inspect alignment with TCS / Guidance2 and dS-vs-divergence-time. PAML 4.10 changed several control-file keywords from 4.9 (getSE = 1 syntax tightened).

Positive Selection Analysis

"Is this gene / branch / site under positive selection?" -> dN/dS (omega = nonsynonymous-to-synonymous substitution rate ratio) framework with explicit choice of WHICH question is being asked (gene-wide / branch-specific / site-specific / episodic) and WHICH null is being rejected. The "test failed because of selection" claim has more known confounders than any other comparative-genomics inference; mandatory pre-screens are: recombination (GARD), alignment errors (PREQUAL or HmmCleaner), saturation (dS distribution), and gBGC (W->S substitution bias). Skipping any one inflates Type-I error to ~20-50% (Anisimova & Yang 2007 MBE 24:1219; Pond 2006 Mol Biol Evol 23:1891).

  • CLI: codeml PAML site, branch, branch-site models
  • CLI: hyphy busted hyphy meme hyphy fel hyphy fubar hyphy absrel hyphy relax hyphy gard
  • Web: datamonkey.org for HyPhy jobs without local install
  • R: RERconverge::correlateWithBinaryPhenotype() for trait-rate associations
  • CLI: csubst analyze for convergent substitution
  • R/CLI: phyloacc for noncoding accelerated evolution

Algorithmic Taxonomy

MethodQuestionNull modelStrengthFails when
PAML codeml M0 (Yang 1997 CABIOS 13:555)Gene-wide single-omega estimate-- (point estimate)Standard reference omega; baseline testSite heterogeneity (use M3+)
codeml M1a vs M2a (Yang 2000 Genetics 155:431)Any site under selection?Nearly neutral, 2-categoryConservative; LRT df=2Low power for episodic selection
codeml M7 vs M8More-sensitive site testBeta(0,1)Higher power than M1a/M2aHigher false-positive rate; relaxed-constraint mimics selection
codeml M8 vs M8a (Swanson 2003 MBE 20:18)Conservative site test (omega2 = 1 null)Beta + omega2=1Cleanest LRT df=1; preferred site testLower power than M7 vs M8
codeml branch-site mod A (Zhang 2005 MBE 22:2472)Selection on pre-specified foreground branchA1 (omega2=1 fixed)Most powerful for episodic per-branch selectionForeground specified post hoc -> Type-I inflation
codeml clade model (Bielawski & Yang 2004 J Mol Evol 59:121)Different omega between named cladesM3 with shared categoriesTests for shifted selection regimeRequires clade pre-specification
codeml free-ratioPer-branch omega estimates (exploratory)M0Visualizes branch-wise variationUnidentifiable for short branches; no formal LRT
HyPhy BUSTED (Murrell 2015 MBE 32:1365)Any episodic selection on any branch site?No omega+ classSite + branch joint; foreground assignableSensitive to alignment errors
HyPhy BUSTED-S (Wisotsky 2020 MBE 37:2430)BUSTED with synonymous-rate variation--Corrects for SRV; reduces false positivesSlightly less power than BUSTED
HyPhy BUSTED-MH (Lucaci 2023 MBE 40:msad150)BUSTED with multi-nucleotide substitutions--Captures complex (multi-hit) substitutions; reduces false positives from MNMsNewer; limited benchmarking
HyPhy BUSTED-PHTwo phenotypes; selection on one not other--Tests phenotype-specific selectionRequires phenotype branch label
HyPhy MEME (Murrell 2012 PLoS Genet 8:e1002764)Per-site episodic selectionFELDetects sites under episodic positive selectionHigher false-positive rate at p threshold
HyPhy FEL (Kosakovsky Pond 2005 MBE 22:1208)Per-site pervasive selection--Fast; counts substitutions per siteNo episodic detection
HyPhy FUBAR (Murrell 2013 MBE 30:1196)Bayesian per-site pervasive selection--Scales to 1000s of sequences; posterior probabilityNo episodic detection
HyPhy SLACCounting-based fast estimator--Very fast; rough estimateLower power; no statistical model
HyPhy aBSREL (Smith 2015 MBE 32:1342)Branch-specific selection without pre-specification--Adaptive per-branch omega categories; corrects multiple testingMultiple-testing burden across many branches
HyPhy RELAX (Wertheim 2015 MBE 32:820)Selection relaxation (k<1) or intensification (k>1)--Detects RELAXED selection; cannot be done by other testsNot designed for adaptive evolution per se
HyPhy GARD (Pond 2006 MBE 23:1891)Recombination breakpoint detectionNo recombinationMANDATORY pre-screen for any selection testComputationally heavy; > 50 sequences slow
McDonald-Kreitman (McDonald & Kreitman 1991 Nature 351:652)Adaptive substitution rate alpha from poly + div dataNeutral mutation accumulationPer-gene alpha; population genetics nativeSlightly deleterious bias (downward); fixed by asymptotic alpha
Asymptotic alpha (Messer & Petrov 2013 PNAS 110:8615)MK with slightly deleterious correction--Unbiased alpha; works at low MAF SFSRequires SFS data
impMKT (Murga-Moreno 2022 G3 12:jkac206)MK with conservative imputation--Gene-level evidence; faster than alpha asymptoticLess unbiased than asymptotic alpha
polyDFE (Tataru & Bataillon 2019 Bioinformatics 35:2868)Full DFE + alpha jointly--Quantifies the distribution of fitness effectsComputational cost; requires polymorphism data
DFE-alpha (Eyre-Walker & Keightley 2009 MBE 26:2097)Faster DFE method--Standard DFE inference; many simulated DFEsRequires demographic correction
GRAPES (Galtier 2016 PLoS Genet 12:e1005774)DFE on neutral + selected sites--Joint demography + alpha; robustGenome-scale dataset required
RERconverge (Kowalczyk 2019 Bioinformatics 35:4815; Redlich 2024 MBE 41:msae210)Relative-rate shifts correlated with categorical phenotype--Phylogenome-wide trait associationsInherits all dN/dS confounders
CSUBST (Fukushima & Pollock 2023 Nat Eco Evo 7:155)Convergent substitutions across independent lineages--Combinatorial-substitution omegaC ratio; null-correctedRequires multi-clade dataset
PhyloAcc (Hu 2019 MBE 36:1086; Thomas 2024)Bayesian convergent accelerated noncoding rate--For noncoding elements (CNEs); convergent rate shiftsCDS analyses prefer codon-based methods
phyloP (Pollard 2010 GR 20:110)Per-site noncoding rate test--Simple; widely used for noncodingNo convergence; site-by-site
PRANK + codeml pipelineCodon-aware MSA + codeml--Standard publication-grade workflowSlow for large datasets

Methodology evolves; verify the latest HyPhy / PAML manuals and the Álvarez-Carretero "Beginner's Guide" (Álvarez-Carretero et al 2023 MBE 40:msad041) before locking on a single method. The BUSTED-MH and FUBAR-MH (multi-hit) extensions specifically address known Type-I inflation from multi-nucleotide substitutions and are now recommended over basic BUSTED / FUBAR.

Decision Tree by Experimental Scenario

ScenarioRecommended approachWhy
Single gene, mammalian (~60 Myr), pre-specified foreground branchcodeml branch-site mod A AND HyPhy aBSREL on foregroundMutual validation; mod A LRT df=1 + aBSREL adaptive site classes
Single gene, deep eukaryote (~500+ Myr), no foreground hypothesisGARD pre-screen -> BUSTED-MH gene-wide -> MEME for sitesEpisodic-selection-only methods; saturation-aware (HyPhy under MG94 codon model)
Genome-wide scan, vertebratescodeml M7 vs M8 OR HyPhy FUBAR-MH per gene; FDR-correctPervasive-selection sites; multi-hit correction critical at scale
Episodic selection scanHyPhy MEME genome-wide (per gene); FDR-correctSite-level episodic detection
Branch-specific selection on unspecified branchesHyPhy aBSRELAdaptive per-branch test with built-in multiple-testing
Comparing selection regimes between two phenotypesHyPhy BUSTED-PH or RELAXPhenotype-specific or relaxation-detection
Recently diverged species (low divergence)MK / asymptotic alpha (population genetics)Codon dN/dS unreliable at low divergence; SFS-based instead
Within-species, dense polymorphism + divergencepolyDFE / GRAPES / asymptotic-MKFull DFE + alpha jointly; preferred for adaptive-substitution rate
Coding selection genome-wide, with SFS availablegrapes -m AUTO_ALLDemography-aware alpha; standard population-genetics-aware adaptive-substitution scan
Noncoding accelerated evolution (CNEs / ECRs)PhyloAcc, phyloP-accCodon-based unsuitable; PhyloAcc Bayesian convergence
Convergent substitutions across independent lineagesCSUBSTCombinatorial-substitution omegaC; null-corrected
Trait-correlated rate shifts genome-wideRERconvergeCategorical / binary phenotype; correlates RERs across thousands of genes
Suspected positive selection but dS > 2Use protein-level method or reduce taxon samplingCodon-based methods unreliable at saturation; protein-only ASR can still work
Recombination expected (immune genes, viral genomes)GARD pre-screen mandatoryRecombination + tree-based selection -> false positives (Anisimova 2003)
Convergent codon substitution at specific sitesTDG09 (Tamuri 2009) or PCOC (Rey 2018)TDG09 detects site-specific shifts in selective constraint between trait-defined lineage groups; PCOC detects convergent amino-acid substitution
Drug-target evolution screenaBSREL on candidate genes; cross-validate with MEMERecent positive selection at drug-target loci
Pathogen / immune-evasion gene with high dS variationBUSTED-S (synonymous rate variation aware)dS variation across sites violates basic BUSTED assumptions
Plasmodium / Trypanosoma / Plasmid analysisBUSTED-MH (multi-hit aware)Multi-nucleotide substitutions common in these; basic BUSTED inflates false positives

Per-Method Failure Modes

Recombination producing false positive selection

Trigger: Running codeml or BUSTED on a gene with recombination breakpoints (viral genes, immune genes, paralog families).

Mechanism: All single-tree codon models assume one phylogeny across all sites. Recombination produces different trees for different segments; treating them as one tree forces the model to invent rate variation that mimics positive selection (Anisimova et al 2003 Genetics 164:1229).

Symptom: PAML M8 strongly rejects M7 (LRT > 50), with omega2 = 999 (PAML upper bound) at several "selected sites"; HyPhy BUSTED highly significant; sites clustered in specific gene regions.

Fix: MANDATORY: run GARD before any positive selection test. If GARD detects breakpoints (p < 0.05), partition the alignment at breakpoints and analyze each segment separately, or use the recombination-aware MEME with the partitioned tree set. RDP5 (Martin 2021 Virus Evol 7:veaa087) is an alternative for viral genomes. GARD output .json lists breakpoint positions and posterior support.

Alignment errors producing false positives

Trigger: Using default MAFFT or MUSCLE alignment on divergent CDS sequences; skipping codon-aware aligner.

Mechanism: Frame-shifted or misaligned codons introduce apparent non-synonymous substitutions at every position; codon-aware tools see these as positive selection (Schneider 2009 GBE 1:114; Markova-Raina & Petrov 2011 GR 21:863).

Symptom: "Selected sites" cluster in alignment regions with > 30% gaps; per-site posteriors in BEB / FUBAR concentrate in ambiguous columns; PREQUAL or Guidance2 marks these regions as poorly aligned; protein alignment shows obvious mismatches.

Fix: Use codon-aware aligner: PRANK (Loytynoja 2014 Methods Mol Biol 1079:155) is the standard for selection analysis (correctly models insertions); MACSE V2 (Ranwez 2018 MBE 35:2582) handles frameshifts and pseudogenes natively; OMM_MACSE wrapper combines them. After alignment, filter with PREQUAL (segment-level) or HmmCleaner (Di Franco 2019 BMC Evol Biol 19:21); do NOT use block-filtering (Gblocks, trimAl) which removes informative sites. Segment-level filtering preferred for selection (Di Franco 2019).

Saturated synonymous sites

Trigger: Comparing distantly related taxa (deep eukaryotic divergence, > 100 Myr); dS > 3 across most pairs.

Mechanism: Synonymous sites have undergone multiple substitutions; the observed dS underestimates true dS. The model can't recover the true rate; omega = dN/dS becomes unstable at the upper bound or low (depending on which direction the bias goes).

Symptom: PAML M0 omega = 999 or near-zero; per-branch dS variance huge; sites with omega > 1 in M8 BEB are at conserved residues (paradox).

Fix: Reduce taxon sampling to species with dS < 2 on internal branches. For deep selection inference on conserved residues, use protein-level methods (BUSTED with --model GTR AA codon translation; aBSREL with protein model option) or restrict to subclade with reasonable saturation. Yang 2007 PAML manual recommends dS < 1.5 per branch.

gBGC inflating apparent positive selection

Trigger: Mammalian / vertebrate gene with W->S substitution bias on a fast-evolving lineage.

Mechanism: GC-biased gene conversion fixes A/T -> G/C alleles preferentially in regions of high recombination, independent of selection (Galtier & Duret 2007 Trends Genet 23:273; Capra 2013 PLoS Genet 9:e1003684). Standard codon models attribute this to positive selection because nonsynonymous substitutions are unequally distributed across codon positions.

Symptom: Branch with apparent positive selection sits in high-recombination region; W->S / S->W substitution ratio > 1.5; selected sites concentrate at non-degenerate codon positions; HyPhy MEME-MH and BUSTED-MH attribute signal to multi-hit rather than positive selection.

Fix: Test for gBGC: W->S substitution rates on selected branch / S->W rates; report ratio. Re-run selection analysis with HyPhy BUSTED-MH (multi-hit aware); if signal vanishes, the original "selection" was gBGC + multi-hit substitutions. For genome-wide scans, mask sub-telomeric / high-recombination regions.

Branch-site test foreground specification

Trigger: Running codeml branch-site mod A after looking at the data to choose foreground branch.

Mechanism: The branch-site test is designed for a single a priori foreground; post hoc specification inflates Type-I by ~5x because the choice was informed by the data.

Symptom: Branch-site test highly significant for the "interesting" branch; aBSREL on same data shows no significant branch (aBSREL has built-in multiple-testing correction).

Fix: Specify foreground branches in registered protocol before looking at data. For exploratory branch-wise analysis, use aBSREL (Smith 2015 MBE 32:1342) which adaptively assigns branch-specific omega classes with multiple-testing built in. If branch-site test was post hoc, apply Bonferroni correction across all branches tested + report explicitly.

LRT critical value confusion

Trigger: Computing branch-site test p-value using standard chi-square df=2.

Mechanism: The branch-site test compares mod A (4 omega classes) against mod A1 (omega2 fixed at 1). The LRT statistic distribution is a 50:50 mixture of point-mass-at-0 and chi-square(df=1), not chi-square(df=2) (Self & Liang 1987 JASA 82:605; Zhang 2005 MBE 22:2472; Wong 2004 Genetics 168:1041). Using df=2 makes the test conservative; using df=1 standard makes it anticonservative.

Symptom: Branch-site p-values incorrectly inflated or deflated; users report finding selection at very stringent thresholds.

Fix: Use the 50:50 mixture critical value: 2.71 at p=0.05 (NOT 3.84). PAML's chi2 1 LRT command applies the mixture. Many published applications use chi-square df=2 conservatively, which loses power but doesn't inflate; chi-square df=1 directly is wrong and inflates Type-I.

omega2 hitting upper bound (999)

Trigger: PAML codeml output shows omega2 = 999 for an "under selection" site class.

Mechanism: PAML codeml uses an internal upper bound of 999 (= "infinity" in single precision). Hitting it indicates numerical issue: extremely few synonymous sites in the selected class, dS underestimation, or numerical optimization failure.

Symptom: Sites flagged as positive selection have omega2 = 999; BEB posteriors for those sites are weirdly distributed.

Fix: Re-run with multiple starting values of omega (fix_omega=0, vary omega = 0.1, 0.5, 1.0, 2.0, 5.0 across runs); check that all converge to same omega. Inspect alignment at flagged sites for unusual residue conservation. If omega = 999 persists, the gene may have rare-substitution patterns; switch to BUSTED-MH which accounts for multi-hit substitutions.

Multiple-testing burden in genome scans

Trigger: Running selection tests across thousands of genes without correction.

Mechanism: With ~5000 protein-coding genes in a typical analysis, 250 will be significant at p=0.05 under H0. The false-discovery rate without correction is 50%.

Symptom: Implausibly large gene lists "under selection"; functional categories enriched are non-specific (e.g. all immune genes by FDR).

Fix: Apply FDR correction (Benjamini-Hochberg). Genes in syntenic regions are non-independent; use Benjamini-Yekutieli for stronger control under dependence. For HyPhy site-level methods, the per-site p < 0.1 default is a starting point; multiple-test correction within a gene is typically not applied (sites within a gene are dependent), but cross-gene correction is necessary. Holm-Bonferroni for strict Type-I.

Convergent substitution misinterpreted as positive selection

Trigger: Lineage-specific selection found at a residue that has independently changed in multiple unrelated lineages.

Mechanism: Convergent substitutions at the same site in independent lineages produce signals in branch-site and other tests; this is convergence, not adaptive evolution per se (though convergent residues often ARE adaptive).

Symptom: Same residue flagged in multiple unrelated lineages by branch-site test; alignment shows convergent substitutions.

Fix: Switch from selection test to convergence test: CSUBST (Fukushima & Pollock 2023 Nat Eco Evo 7:155) for combinatorial substitution analysis; RERconverge (Redlich 2024 MBE 41:msae210) for relative-rate-vs-phenotype across categorical traits; PCOC (Rey 2018) for biophysical convergence. Report both convergence test and selection test results.

Show full SKILL.md (1,604 more words)Show less

Quantitative Thresholds

QuantityThresholdSource / Rationale
dN/dS interpretationomega < 1 purifying; omega = 1 neutral; omega > 1 positive (per site, branch, or gene depending on model)Yang & Bielawski 2000 TREE 15:496; foundational
Branch-site test LRT critical value2.71 at p=0.05 (50:50 mixture chi^2)Self-Liang 1987 JASA 82:605; Zhang 2005 MBE 22:2472
Site-level p-value defaultp <= 0.1 (FEL, MEME, FUBAR); FUBAR posterior >= 0.9Murrell 2012/2013; Datamonkey conventions
BEB posterior probability>= 0.95 significant; >= 0.99 highly significantYang & Bielawski 2000
dS upper limit for reliabilitydS < 1.5 per branch; dS < 3 overallYang 2007 PAML manual
Minimum sequences for codeml>= 8 with sufficient divergenceAnisimova et al 2001 MBE 18:1585
Branch-site test minimum lineages>= 20 in tree; >= 4 background branchesYang 2007
GARD breakpoint significancep < 0.05 to partition alignmentPond 2006; mandatory pre-screen
MK alpha thresholdalpha > 0 indicates adaptive substitutions; report 95% CISmith & Eyre-Walker 2002
Asymptotic alpha minimum SFS density>= 50 sites per frequency binMesser & Petrov 2013
FDR genome-wide selection scanq < 0.05 Benjamini-HochbergStandard
MEME minimum site-level supportp < 0.1; +/-3 sequences with substitutionsMurrell 2012
aBSREL p-valuep < 0.05 (corrected by Holm-Bonferroni internally)Smith 2015
RELAX k interpretationk < 1 relaxed; k > 1 intensifiedWertheim 2015
Codon usage bias ENCENC < 35 high bias; consider effect on dSWright 1990 Gene 87:23
W->S substitution ratio for gBGC> 1.5 suggests gBGCOperational convention
BUSTED-MH multi-hit thresholdomega_DH > 1 indicates multi-hit patternLucaci 2023
HyPhy SRV (Synonymous Rate Variation)Use BUSTED-S when dS varies across sites > 2xWisotsky 2020

Selection Scan Standard Pipeline

Goal: Test all coding genes in a clade for evidence of positive selection, with full quality control.

Approach: Align with PRANK -> filter with PREQUAL -> pre-screen with GARD -> run BUSTED-MH (gene-wide) + MEME (sites) + aBSREL (branches); FDR-correct across genes; verify top candidates pass alignment / saturation / gBGC checks.

bash
# Per-gene pipeline (parallelizable)
for og in orthogroups/*.fa; do
    base=$(basename $og .fa)

    # 1. Codon-aware MSA
    prank -d=$og -o=msa/$base.prank -codon -F

    # 2. Filter alignment errors (segment-level)
    PREQUAL -i msa/$base.prank.best.fas -o msa_filt/$base

    # 3. Recombination pre-screen
    hyphy gard --alignment msa_filt/$base.filtered --output gard/$base.json
    # If breakpoints found: partition and treat per-segment

    # 4. Gene-wide test (multi-hit aware)
    hyphy busted --alignment msa_filt/$base.filtered \
        --tree species_tree.nwk --output busted_mh/$base.json \
        --srv Yes --multiple-hits Double+Triple

    # 5. Site-level
    hyphy meme --alignment msa_filt/$base.filtered \
        --tree species_tree.nwk --output meme/$base.json

    # 6. Branch-level
    hyphy absrel --alignment msa_filt/$base.filtered \
        --tree species_tree.nwk --output absrel/$base.json
done

# 7. Aggregate and FDR
python aggregate_selection_scan.py busted_mh/ meme/ absrel/ > selection_results.tsv
python
'''Aggregate genome-wide selection scan results; FDR-correct.'''
import json, glob, pandas as pd
from scipy.stats import false_discovery_control

def parse_busted(p):
    d = json.load(open(p))
    return {'p_value': d.get('test results', {}).get('p-value'),
            'LRT': d.get('test results', {}).get('LRT'),
            'omega_DH': d.get('fits', {}).get('Unconstrained model', {}).get('omega3')}

def count_meme_sig(p, alpha=0.1):
    d = json.load(open(p))
    mle = d.get('MLE', {}).get('content', {}).get('0', {})
    headers = [h[0] for h in d.get('MLE', {}).get('headers', [[]])]
    pi = headers.index('p-value') if 'p-value' in headers else -1
    return sum(1 for v in mle.values() if pi >= 0 and v[pi] < alpha)

rows = []
for path in glob.glob('busted_mh/*.json'):
    gene = path.split('/')[-1].replace('.json', '')
    rows.append({'gene': gene, **parse_busted(path),
                 'meme_sig_sites': count_meme_sig(f'meme/{gene}.json')})
df = pd.DataFrame(rows)
df['busted_fdr'] = false_discovery_control(df['p_value'].fillna(1.0), method='bh')
df['adaptive'] = (df['busted_fdr'] < 0.05) & (df['meme_sig_sites'] > 0)
df.sort_values('busted_fdr').to_csv('selection_results.tsv', sep='\t', index=False)

PAML Branch-Site Test (Operational)

Goal: Test for episodic positive selection on a pre-specified foreground branch.

Approach: Mark foreground in newick (#1) -> codeml branch-site mod A vs A1 -> LRT against 50:50 mixture chi^2(0):chi^2(1).

bash
# Mark foreground branch: use ete4 or manually
python -c "
from ete4 import Tree
t = Tree('species_tree.nwk', format=1)
target = t.search_nodes(name='target_species')[0]
target.name = target.name + ' #1'
print(t.write(format=1))
" > foreground.nwk

# Branch-site mod A (alternative)
cat > codeml_modA.ctl << 'EOF'
seqfile = alignment.phy
treefile = foreground.nwk
outfile = mod_A.mlc
runmode = 0
seqtype = 1
CodonFreq = 2
model = 2
NSsites = 2
fix_kappa = 0
kappa = 2
fix_omega = 0
omega = 0.4
RateAncestor = 1
cleandata = 0
EOF
codeml codeml_modA.ctl

# Null model A1 (omega_2 = 1)
cp codeml_modA.ctl codeml_modA1.ctl
sed -i 's/^omega = 0.4/omega = 1/' codeml_modA1.ctl
sed -i 's/^fix_omega = 0/fix_omega = 1/' codeml_modA1.ctl
sed -i 's/outfile = mod_A.mlc/outfile = mod_A1.mlc/' codeml_modA1.ctl
codeml codeml_modA1.ctl
python
'''Branch-site test LRT with 50:50 mixture critical value.'''
from scipy.stats import chi2

def branch_site_lrt(lnL_alt, lnL_null):
    lrt = 2 * (lnL_alt - lnL_null)
    if lrt <= 0:
        return {'LRT': lrt, 'p_value': 0.5}
    # 50:50 mixture of chi^2(0) and chi^2(1)
    p = 0.5 * (1 - chi2.cdf(lrt, df=1))
    return {'LRT': lrt, 'p_value': p}

Foreground branch must be specified before viewing data; for genome-wide screens with no a priori branch, use aBSREL instead. Bayes Empirical Bayes (BEB) sites with posterior > 0.95 on positive-selection class are the per-site call.

McDonald-Kreitman with Asymptotic Alpha

Goal: Estimate adaptive substitution rate alpha = 1 - (Ds Pn) / (Dn Ps), corrected for slightly deleterious bias.

Approach: Compute counts of synonymous and nonsynonymous polymorphisms (P) and divergences (D); fit asymptotic alpha by binning by minor-allele frequency and extrapolating.

r
# Standard MK
mk_alpha <- function(Dn, Ds, Pn, Ps) {
    1 - (Ds * Pn) / (Dn * Ps)
}

# Asymptotic alpha via the Messer-Petrov 2013 web tool
# (https://benhaller.com/messerlab/asymptoticMK.html) or the impMKT R package
# (Murga-Moreno 2022 G3 12:jkac206) which wraps the asymptotic computation.
library(impMKT)
# Inputs: per-frequency-bin (Pn, Ps) plus genome-wide (Dn, Ds)
freq_bins <- seq(0.01, 0.5, 0.01)
pn_by_freq <- c(...)  # nonsyn polymorphism count per bin
ps_by_freq <- c(...)  # syn polymorphism count per bin
fit <- asymptoticMK(
    Dn = total_dn, Ds = total_ds,
    Pn = pn_by_freq, Ps = ps_by_freq,
    x = freq_bins
)
fit$alpha_asymptotic       # adaptive substitution rate, corrected
fit$alpha_original         # original MK (biased)

For full DFE inference (alpha + distribution of fitness effects), use polyDFE (Tataru-Bataillon 2019):

bash
polyDFE -d data.txt -m C -i estimates.init -o output_basename

DFE-alpha (Eyre-Walker 2009) and GRAPES (Galtier 2016) are alternatives; GRAPES is most robust for genome-wide adaptive-substitution scans.

RERconverge for Trait-Correlated Rate Shifts

Goal: Identify genes whose evolutionary rate correlates with a binary or categorical phenotype across the species tree.

Approach: Compute per-gene relative evolutionary rates -> correlate against phenotype -> Bonferroni or FDR-correct across genes.

r
library(RERconverge)

# Read alignments and tree
trees <- readTrees('orthogroup_trees.txt', minSpecies = 10)
rer <- getAllResiduals(trees, useSpecies = species_names, transform = 'sqrt',
                       weighted = TRUE, scale = TRUE)

# Define binary phenotype (e.g., echolocation in mammals)
phen_paths <- foreground2Paths(c('Bat1', 'Bat2', 'Dolphin'), trees, clade = 'terminal')
phen_vec <- foreground2Tree(c('Bat1', 'Bat2', 'Dolphin'), trees, clade = 'terminal')

# Correlate
cors <- correlateWithBinaryPhenotype(rer, phen_paths, min.sp = 10, min.pos = 2,
                                      weighted = 'auto')
top_genes <- cors[order(cors$P), ][1:50, ]

For categorical traits (more than binary), Redlich 2024 MBE 41:msae210 extends RERconverge.

Reconciliation: When Methods Disagree

PatternLikely causeAction
codeml M8 significant, BUSTED nullM8 vs M7 inflated by relaxed constraint mimicking selectionTrust BUSTED; check M8a vs M8 instead (stricter null)
codeml branch-site significant, aBSREL nullBranch-site test foreground post hocaBSREL with built-in multiple-testing is correct; downgrade claim
BUSTED significant, MEME no sitesEpisodic at sites BUSTED can't pinpoint; or basic BUSTED detected SRV not selectionRun BUSTED-S; if signal vanishes, was SRV; if persists, gene-wide episodic
MEME positive, FEL nullEpisodic selection (MEME-specific)Trust MEME for episodic; FEL only detects pervasive
Multiple tests positive at same siteHigh-confidence site under selectionReport; consider experimental validation
Test positive but PREQUAL flagged 20% of alignmentAlignment artifactRe-filter (HmmCleaner); re-test; downgrade if positive site is in filtered region
Test positive but in high-recombination regiongBGCW->S substitution test; if gBGC-attributable, downgrade
BUSTED-MH null where BUSTED significantMulti-hit substitutions misattributedTrust BUSTED-MH; original positive was multi-hit pattern
RELAX k > 1 with branch-site test nullSelection regime intensification (more purifying)RELAX captures regime shift; branch-site missed because foreground different
Branch-site significant on Drosophila branch but no signal in mammalsLineage-specific adaptation; or dS saturation in mammalsInspect dS distribution; if mammals dS < 0.5 across branch, signal is real; if dS > 2, saturation explanation
asymptotic alpha < 0DFE has high deleterious load; or demographic violationCheck polyDFE / GRAPES with demographic correction

Operational rule for publication: GARD pre-screen documented as negative + PREQUAL/HmmCleaner filtering applied + dS < 1.5 per branch + W->S ratio not elevated + BUSTED-MH significant (gene-wide) + MEME flags sites + aBSREL flags branches with consistent direction = publication-ready evidence. Single-method significance (especially M8 vs M7 alone) should be downgraded.

Cohort Gotchas

  • Immune / MHC loci: intra-genic recombination is high; GARD pre-screen mandatory; high apparent positive selection often reflects gene conversion between alleles, not adaptive change
  • Viral genomes: rapid evolution + recombination + multi-hit substitutions common; BUSTED-MH and FUBAR-MH essential; use RDP5 for recombination detection
  • Plasmodium / Trypanosoma: high codon-usage bias and multi-hit substitutions; use BUSTED-MH and BUSTED-S
  • Mammalian X-chromosome: higher dS than autosomes (male-driven evolution); gBGC asymmetry by chromosome; reduce dS threshold for X-linked genes
  • Recent human / population genetics: dS dramatically underestimated at recent divergence; use SFS-based methods (asymptotic alpha, polyDFE)
  • Convergent evolution traits (echolocation, marine): RERconverge / CSUBST / PhyloAcc-noncoding designed for these; codon-based methods alone miss the convergent signal

Anticipated Reviewer Pushback

PushbackStandard response
"GARD pre-screen?"Yes; no breakpoints (or partitioned at p < 0.05 breakpoints); per-segment results consistent
"Alignment filtering?"PRANK codon-aware MSA; PREQUAL segment filter applied; Guidance2 scores reported
"Saturation?"dS distribution shown; max per-branch dS < 1.5; analysis restricted to subclades meeting this
"Branch-site test foreground post hoc?"Foreground pre-registered OR exploratory analysis acknowledged + aBSREL with built-in multiple-testing used
"Multiple-testing correction?"FDR (Benjamini-Hochberg) across genes; per-site within gene not corrected (dependence)
"gBGC?"W->S substitution ratio not elevated; non-sub-telomeric; BUSTED-MH null in candidates rules out
"Multi-hit?"BUSTED-MH used; if signal persists, robust to multi-hit confounder
"Why this LRT df?"Branch-site test uses 50:50 mixture (Self-Liang 1987; Zhang 2005); critical value 2.71 at p=0.05
"Sensitivity to model choice?"Cross-validated PAML vs HyPhy; consistent across both; reported both p-values

Common Errors

Error / symptomCauseSolution
codeml runs but rst file emptyRateAncestor = 0 or path not writableSet RateAncestor = 1; check output directory
codeml omega2 = 999Numerical pathology / saturated dSVary starting omega; reduce taxon sampling to dS < 2
codeml LRT negative (-0.001)Numerical noise at convergenceRound; treat as no signal; rerun with different starting values
HyPhy "tree branches don't match alignment"Mismatched taxa namesUse exact same labels in tree and alignment
HyPhy MEME returns "no sites significant"High alignment uncertainty; or no episodic selectionRe-filter alignment; try BUSTED-S for gene-wide signal
GARD takes forever> 50 sequencesReduce to representative subset; or use RDP5 for viral data
MK alpha negativeDemographic issue or DFE has many slightly deleteriousUse polyDFE / GRAPES with demography correction
RERconverge "too few species per gene"Stringent defaultReduce min.sp = 5; document
CSUBST omega_C unstableFew combinations; small cladeNeed >= 5 clades for stable convergence estimate
PhyloAcc convergence failureInsufficient lineagesRe-run with relaxed prior; check input MAF distribution

Tool Installation Notes

bash
conda install -c bioconda paml hyphy gard prank prequal hmmcleaner
# RDP5: http://web.cbio.uct.ac.za/~darren/rdp.html
# MACSE V2: wget https://bioweb.supagro.inra.fr/macse/releases/macse_v2.07.jar
pip install ete4 pyhyphy csubst
Rscript -e "install.packages(c('asymptoticMK', 'polyDFE'))"
Rscript -e "remotes::install_github('nclark-lab/RERconverge')"
# polyDFE / GRAPES / DFE-alpha source binaries at respective github / bioconda channels

For genome-wide scans (> 5000 genes), parallelize per-gene analyses with Snakemake / Nextflow.

References

  • Yang Z 1997 CABIOS 13:555 (PAML codeml)
  • Yang Z et al 2000 Genetics 155:431 (codon models M0-M8)
  • Yang Z & Bielawski JP 2000 TREE 15:496 (codon model framework)
  • Zhang J et al 2005 MBE 22:2472 (branch-site mod A); Wong WSW et al 2004 Genetics 168:1041 (LRT mixture); Self SG & Liang K-Y 1987 JASA 82:605 (LRT boundary)
  • Swanson WJ et al 2003 MBE 20:18 (M8a null); Bielawski JP & Yang Z 2004 J Mol Evol 59:121 (clade models)
  • Anisimova M & Yang Z 2007 MBE 24:1219 (multiple-testing / branch-site power); Anisimova M et al 2003 Genetics 164:1229 (recombination FP); Anisimova M, Bielawski JP & Yang Z 2001 MBE 18:1585 (LRT power)
  • Pond SLK et al 2006 MBE 23:1891 (GARD); Martin DP et al 2021 Virus Evol 7:veaa087 (RDP5)
  • Kosakovsky Pond SL & Frost SDW 2005 MBE 22:1208 (FEL); Murrell B et al 2012 PLoS Genet 8:e1002764 (MEME); Murrell B et al 2013 MBE 30:1196 (FUBAR)
  • Murrell B et al 2015 MBE 32:1365 (BUSTED); Wisotsky SR et al 2020 MBE 37:2430 (BUSTED-S); Lucaci AG et al 2023 MBE 40:msad150 (BUSTED-MH)
  • Smith MD et al 2015 MBE 32:1342 (aBSREL); Wertheim JO et al 2015 MBE 32:820 (RELAX)
  • McDonald JH & Kreitman M 1991 Nature 351:652 (MK); Smith NGC & Eyre-Walker A 2002 Nature 415:1022 (alpha); Messer PW & Petrov DA 2013 PNAS 110:8615 (asymptotic alpha)
  • Murga-Moreno J et al 2022 G3 12:jkac206 (impMKT); Tataru P & Bataillon T 2019 Bioinformatics 35:2868 (polyDFE); Eyre-Walker A & Keightley PD 2009 MBE 26:2097 (DFE-alpha); Galtier N 2016 PLoS Genet 12:e1005774 (GRAPES)
  • Galtier N & Duret L 2007 Trends Genet 23:273 (gBGC); Capra JA et al 2013 PLoS Genet 9:e1003684 (gBGC genome-scale)
  • Schneider A et al 2009 GBE 1:114 + Markova-Raina P & Petrov D 2011 GR 21:863 (alignment-error FP)
  • Loytynoja A 2014 Methods Mol Biol 1079:155 (PRANK); Ranwez V et al 2018 MBE 35:2582 (MACSE V2); Whelan S et al 2018 Bioinformatics 34:3929 (PREQUAL); Di Franco A et al 2019 BMC Evol Biol 19:21 (HmmCleaner)
  • Yang Z 2007 PAML manual; Álvarez-Carretero S et al 2023 MBE 40:msad041 (Beginner's Guide PAML)
  • Kowalczyk A et al 2019 Bioinformatics 35:4815 + Redlich R et al 2024 MBE 41:msae210 (RERconverge)
  • Fukushima K & Pollock DD 2023 Nat Eco Evo 7:155 (CSUBST); Hu Z et al 2019 MBE 36:1086 (PhyloAcc); Pollard KS et al 2010 GR 20:110 (phyloP); Rey C et al 2018 MBE 35:2296 (PCOC)
  • comparative-genomics/ortholog-inference - Single-copy ortholog alignments as input
  • comparative-genomics/ancestral-reconstruction - Branch-specific ancestral sequence inference
  • comparative-genomics/gene-tree-species-tree-reconciliation - Reconciled gene trees as PAML input
  • alignment/multiple-alignment - PRANK / MACSE codon-aware MSA
  • alignment/alignment-trimming - PREQUAL / HmmCleaner segment filtering
  • phylogenetics/modern-tree-inference - Tree inference required for codeml
  • population-genetics/selection-statistics - SFS-based alpha + DFE methods
  • causal-genomics/heritability-partitioning - LDSC partition includes positive-selection annotations
  • variant-calling/variant-annotation - Functional annotation of selected sites

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in comparative-genomics/positive-selection of GPTomics/bioSkills.

  • SKILL.md
  • examples/selection_analysis.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Comparative Genomics Positive Selection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Comparative Genomics Positive Selection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Comparative Genomics Positive Selection this skillGPTomics/bioSkills1.2k2 repos~9.7kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Comparative Genomics Positive Selection

What does Bio Comparative Genomics Positive Selection do?

Detect positive (diversifying / episodic / pervasive) selection using codon dN/dS frameworks. Bio Comparative Genomics Positive Selection is an agent skill from GPTomics/bioSkills. Detect positive (diversifying / episodic / pervasive) selection using codon dN/dS frameworks.

When should I use Bio Comparative Genomics Positive Selection?

Bio Comparative Genomics Positive Selection fits situations like: testing adaptive evolution at codons; running GARD recombination pre-screen; controlling alignment-error and gBGC false positives; reconciling PAML vs HyPhy results.

How do I install Bio Comparative Genomics Positive Selection in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-positive-selection -a claude-code`. Or copy the skill folder (comparative-genomics/positive-selection in GPTomics/bioSkills) into .claude/skills/bio-comparative-genomics-positive-selection in your project. Claude Code loads it when a task matches its description.

How do I install Bio Comparative Genomics Positive Selection in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-positive-selection -a codex`. Or copy the skill folder (comparative-genomics/positive-selection in GPTomics/bioSkills) into .agents/skills/bio-comparative-genomics-positive-selection in your project. Codex loads it when a task matches its description.

Can I use Bio Comparative Genomics Positive Selection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-positive-selection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-comparative-genomics-positive-selection, .gemini/skills/bio-comparative-genomics-positive-selection, .github/skills/bio-comparative-genomics-positive-selection and .opencode/skills/bio-comparative-genomics-positive-selection in your project.

What does Bio Comparative Genomics Positive Selection need to run?

Going by SKILL.md and its folder, Bio Comparative Genomics Positive Selection needs Python for the scripts in its folder and the command-line tools its instructions call (python, pip and conda). Our summary lists: Python 3.

Does Bio Comparative Genomics Positive Selection access the network?

SKILL.md names 3 domains. In commands or code: benhaller.com, web.cbio.uct.ac.za and bioweb.supagro.inra.fr; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Comparative Genomics Positive Selection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Comparative Genomics Positive Selection use?

Bio Comparative Genomics Positive Selection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Comparative Genomics Positive Selection use?

About 9.7k tokens (SKILL.md is roughly 39k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Comparative Genomics Positive Selection?

Skills that share tags, products or a category with Bio Comparative Genomics Positive Selection: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Comparative Genomics Positive Selection?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.