Agent skill

Bio Phylo Modern Tree Inference

by GPTomics in GPTomics/bioSkills

Infers maximum-likelihood phylogenetic trees with IQ-TREE2 and RAxML-NG -- model selection (ModelFinder), branch support (UFBoot2, SH-aLRT), concordance factors (gCF/sCF), partitioning, topology…

MITAuto-check passedResearch & Science

Install Bio Phylo Modern Tree Inference

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-phylo-modern-tree-inference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-phylo-modern-tree-inference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/phylogenetics/modern-tree-inference .claude/skills/bio-phylo-modern-tree-inference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-phylo-modern-tree-inference
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.3k tokens
SKILL.md length
2,600 words
Files
5
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Infers maximum-likelihood phylogenetic trees with IQ-TREE2 and RAxML-NG -- model selection (ModelFinder), branch support (UFBoot2, SH-aLRT), concordance factors (gCF/sCF), partitioning, topology…

  • Works in 3 steps: High support is consistent with being… → More data fixes variance, not bias.… → Concordance factors are the honest…
  • Inferring an ML tree
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Model Selection, plus 9 more sections
  • Runs Shell scripts from its folder

What it does

Bio Phylo Modern Tree Inference is an agent skill from GPTomics/bioSkills. Infers maximum-likelihood phylogenetic trees with IQ-TREE2 and RAxML-NG -- model selection (ModelFinder), branch support (UFBoot2, SH-aLRT), concordance factors (gCF/sCF), partitioning, topology tests, and long-branch-attraction control. Covers why an ML tree inherits every flaw of the assumed model and the fixed alignment, why reported support measures repeatability under resampling and not correctness, why UFBoot uses a =95 cutoff and not the bootstrap-70 rule, and why a node with UFBoot 100 but gCF ~35 is…

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `examples/iqtree_basic.sh`, `examples/partitioned_analysis.sh` and `examples/raxml_analysis.sh`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Inferring an ML tree
  • Selecting a substitution
  • Partition model
  • Interpreting support measures

Example prompts

  • “Use the bio-phylo-modern-tree-inference skill to infer maximum-likelihood phylogenetic trees with IQ-TREE2 and RAxML-NG -- model selection…”
  • “/bio-phylo-modern-tree-inference”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. High support is consistent with being wrong. Bootstrap, UFBoot, SH-aLRT, and aBayes all ask whether a branch reappears when the data or…
  2. More data fixes variance, not bias. Adding sites shrinks sampling error and makes the estimate more confident, but the systematic error…
  3. Concordance factors are the honest measure at genome scale. gCF/sCF ask what FRACTION of genes or sites actually contains a branch. A node…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Phylo Modern Tree Inference loads about 5.3k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 2,600 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~225
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,600 words, ~5,339 tokens.

Download SKILL.mdSave it as .claude/skills/bio-phylo-modern-tree-inference/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
bio-phylo-modern-tree-inference
description
Infers maximum-likelihood phylogenetic trees with IQ-TREE2 and RAxML-NG -- model selection (ModelFinder), branch support (UFBoot2, SH-aLRT), concordance factors (gCF/sCF), partitioning, topology tests, and long-branch-attraction control. Covers why an ML tree inherits every flaw of the assumed model and the fixed alignment, why reported support measures repeatability under resampling and not correctness, why UFBoot uses a >=95 cutoff and not the bootstrap-70 rule, and why a node with UFBoot 100 but gCF ~35 is essentially unresolved ILS rather than a clade. Use when inferring an ML tree, selecting a substitution or partition model, choosing or interpreting support measures, testing an a-priori topology, or diagnosing LBA. Routes model-free distance trees to distance-calculations, posteriors to bayesian-inference, and species trees under ILS to species-trees.
tool_type
cli
primary_tool
IQ-TREE2

Version Compatibility

Reference examples tested with: IQ-TREE 2.2+ / 2.3+, RAxML-NG 1.2+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: iqtree2 --version then iqtree2 --help to confirm flags
  • CLI: raxml-ng --version then raxml-ng --help to confirm flags

If code throws an unrecognized-argument or model-parse error, introspect the installed tool and adapt the example to match the actual API rather than retrying.

IQ-TREE2 uses single-dash documented forms (-alrt, -bnni, -B); -B/-T are v2.x (v1.x used -bb/-nt). Do NOT write --alrt. The likelihood site-concordance flag --scfl requires IQ-TREE 2.2.2+ (older builds have only the parsimony --scf).

Modern ML Tree Inference -- ML Support Measures Repeatability, Not Correctness

"Build a maximum-likelihood tree with support from my alignment" -> Select a substitution model, search topology and branch lengths that maximize the likelihood, then attach support that quantifies repeatability and concordance that quantifies genealogical agreement.

  • CLI: iqtree2 -s aln.fasta -m MFP -B 1000 -bnni -alrt 1000 (model selection + dual support, all built in)
  • CLI: raxml-ng --all --msa aln.fasta --model GTR+G --bs-metric fbp,tbe (very large trees, transfer bootstrap, precise branch lengths)

Scope: ML estimation of topology, branch lengths, model selection, branch support, concordance factors, partitioning, topology tests, and LBA control. Model-free distance/NJ trees and distance correction -> distance-calculations. Posterior distributions, MCMC, and CAT-GTR -> bayesian-inference. Per-locus gene trees summarized into a species tree under ILS -> species-trees. Time-scaled trees -> divergence-dating.

The Single Most Important Modern Insight

An ML tree is the topology and branch lengths that maximize the likelihood under an ASSUMED substitution model, conditioned on a FIXED multiple-sequence alignment. It inherits every flaw of both: a misaligned column is a fabricated character the model dutifully fits, and a misspecified model biases the point estimate toward a wrong topology that more data only sharpens. The reported support measures REPEATABILITY under resampling, not correctness. Three load-bearing facts:

  1. High support is consistent with being wrong. Bootstrap, UFBoot, SH-aLRT, and aBayes all ask whether a branch reappears when the data or tree is perturbed. Under model misspecification every replicate reproduces the same bias, so support climbs toward 100% precisely as the inference becomes more wrong. At genome scale, sampling error vanishes and UFBoot ~100 on most branches is the default, not a signal.
  2. More data fixes variance, not bias. Adding sites shrinks sampling error and makes the estimate more confident, but the systematic error from saturation, compositional heterogeneity, and ILS is bias that CONCENTRATES with scale. The cure for confident-wrong trees is better models, better alignments, and better diagnostics -- never more bootstrap replicates.
  3. Concordance factors are the honest measure at genome scale. gCF/sCF ask what FRACTION of genes or sites actually contains a branch. A node with UFBoot 100 and gCF 35 means the concatenated likelihood is certain but only ~35% of loci endorse that branch -- biologically unresolved (ILS or introgression), and the bootstrap answered a question nobody should have asked. Report CFs on every phylogenomic tree; treat bootstrap as necessary-not-sufficient.

Tool Taxonomy

ToolCitationRoleWhen
IQ-TREE2Minh 2020ML search + ModelFinder + UFBoot2 + SH-aLRT + gCF/sCF + AU test + C60/PMSF, all built inthe default for almost all work
RAxML-NGKozlov 2019ML search, transfer bootstrap, terrace-aware, MPI/checkpointingvery large trees, TBE, precise branch lengths, long HPC runs
PhyMLGuindon 2010ML search, origin of SH-aLRTlegacy/teaching; SH-aLRT is now in IQ-TREE2
FastTreePrice 2010approximate ML, single-pass NNI/SPRa fast first-pass tree on thousands of sequences; not for final support

IQ-TREE2 vs RAxML-NG (both hill-climb the same likelihood; they differ in built-in features and scaling, not correctness):

NeedUse
Model selection, UFBoot2, SH-aLRT, concordance factors, AU test, mixture/PMSFIQ-TREE2 (built in; nothing else bundles all of this)
Very large trees (thousands of taxa), low memory, MPIRAxML-NG
Transfer bootstrap (TBE) for rogue-taxon mega-treesRAxML-NG (--bs-metric tbe)
Most precise branch lengths for downstream datingRAxML-NG
Robust checkpoint/restart on long runsRAxML-NG

Common production pattern: model-select, compute CFs, and topology-test in IQ-TREE2; do heavy bootstrap on a mega-tree in RAxML-NG with --bs-metric fbp,tbe.

Model Selection

ModelFinder (Kalyaanamoorthy 2017) scores the substitution matrix and the rate-heterogeneity model jointly, including FreeRate categories the old jModelTest/ProtTest generation never tested, and ranks by BIC (the default; k*ln(n) penalty favors simpler models that generalize on large n).

  • -m MFP -- ModelFinder Plus: test all models by BIC, then search with the winner. The standard. (-m MF selects only; avoid -m TEST/-m TESTONLY, the legacy jModelTest-style limited set.)
  • +G vs +R. +G (discrete Gamma, one alpha) is the workhorse for single short genes. +R (FreeRate) freely estimates each category's rate AND weight, capturing the non-Gamma, often multimodal rate distributions of concatenated phylogenomic data; expect +R3 to +R6 to win on BIC at scale.
  • The +I+G trap. The invariant-sites proportion (+I) and the Gamma shape (+G) describe overlapping rate distributions, so the likelihood surface has a flat ridge: the two estimates are individually near-meaningless and start-dependent. Prefer +R (its slowest category absorbs near-invariant sites); use +G alone unless BIC genuinely demands +I+G.
  • Partition model selection. -m MFP+MERGE fits per-partition models AND greedily merges partitions that fit the same model (BIC-chosen scheme, the successor to PartitionFinder greedy). Pair with -rcluster 10 (relaxed clustering: only test the top 10% most-similar pairs) for many partitions.
  • Site-heterogeneous mixtures for deep data. Empirical matrices (LG, WAG) assume one residue-frequency vector for the whole alignment; real proteins do not, and that across-site compositional heterogeneity is the chief driver of deep LBA. The ML answer is the C10..C60 profile-mixture series (LG+C60+F+G), made tractable by PMSF (Wang 2018): a guide-tree pass computes one posterior-mean profile per site, and the real search uses those frozen profiles. PMSF is the standard recommendation for deep / LBA-prone protein phylogenomics.

Branch Support -- the Heart

Five measures, three different questions; the cardinal sin is cross-comparing their cutoffs.

MetricWhat it perturbs / measuresStrong cutoffTool / flagFailure mode
Standard bootstrap (FBP)resample sites; clade frequency (binary)>=70 (folklore)RAxML-NG --bs-metric fbp; IQ-TREE -bslow; crushed by rogue taxa in big trees
Ultrafast bootstrap 2 (UFBoot)RELL-resampled log-Ls; ~unbiased clade prob>=95 (NOT 70)IQ-TREE2 -B 1000 (+-bnni)inflates under model violation -> use -bnni
SH-aLRTlocal NNI likelihood ratio (no resampling)>=80IQ-TREE2 -alrt 1000conservative on very short branches
aBayesposterior from 3 NNI Ls, flat prior>=0.95IQ-TREE2 -abayesanti-conservative; never the sole criterion
Transfer bootstrap (TBE)gradual transfer distance under resamplingno fixed cutoff; > FBPRAxML-NG --bs-metric tbepermissive; "fuzzy" branch identity

UFBoot2 (Hoang 2018) uses the RELL trick (resample site log-likelihoods, reuse a candidate tree set) to run hundreds of times faster than the Felsenstein bootstrap (1985), and its values are CLOSER to unbiased clade probabilities -- which is exactly why the strong-support cutoff is 95, not the conservative-bootstrap 70. Treating UFBoot 70 as "good" is a category error. -bnni re-optimizes each replicate tree by NNI to rein in the inflation that model violation causes; use -B 1000 -bnni routinely. UFBoot is on a DIFFERENT scale from the standard bootstrap -- never read it with the BP-70 rule.

SH-aLRT (Guindon 2010) does not resample data; for each branch it tests whether the ML likelihood beats its two best NNI rearrangements. UFBoot (data perturbation) and SH-aLRT (tree perturbation) have different failure modes, so the community-standard joint criterion requires both:

A branch is strongly supported iff SH-aLRT >= 80% AND UFBoot >= 95%.

For large rogue-taxon-prone trees, the binary Felsenstein bootstrap lets a single wandering tip crush an otherwise-recovered deep branch; transfer bootstrap (TBE, Lemoine 2018, RAxML-NG --bs-metric tbe) replaces the in/out indicator with a gradual transfer distance and rescues those branches, at the cost of being more permissive.

Concordance Factors

Bootstrap quantifies statistical confidence given the concatenated data; concordance factors (Minh 2020) quantify how much of the actual data carries a branch.

  • gCF (gene concordance factor): the percentage of decisive single-locus gene trees that contain the exact branch. Needs per-locus gene trees.
  • sCF (site concordance factor): the percentage of decisive sites supporting the branch, from sampled quartets; works on a single concatenated alignment with no gene trees.
  • sCFL (likelihood sCF, Mo 2023): uses ancestral-state likelihoods rather than parsimony quartet counting, substantially reducing (not abolishing) homoplasy and taxon-sampling bias. Prefer --scfl over the old --scf on IQ-TREE 2.2.2+.
bash
# one gene tree per locus from a directory of locus alignments (-S = separate, no concatenation)
iqtree2 -S loci_dir -m MFP -B 1000 -T AUTO --prefix loci   # -B 1000 = UFBoot per gene tree, needed to contract weak branches before ASTRAL

# gene + likelihood site concordance against a fixed concatenated tree (-te fixes the tree)
iqtree2 -te concat.treefile -s concat.fasta --gcf loci.treefile --scfl 100 -T 4 --prefix concord
#  --gcf loci.treefile   per-locus gene trees for gCF
#  --scfl 100            100 sampled quartets per branch for likelihood sCF (higher = more stable)

Outputs concord.cf.tree (Newick with gCF/sCF labels) and concord.cf.stat (per-branch gCF, gDF1, gDF2, gDFP, sCF). A node with UFBoot 100 but gCF ~35 (gDF1 ~33, gDF2 ~30) is genes split three ways: the concatenated point estimate barely edges the alternatives and the node is biologically unresolved -- the signature of ILS or introgression, not a clade. Report CFs alongside support on every phylogenomic tree.

Topology Tests

For testing an a-priori hypothesis ("can I reject that X and Y are monophyletic?") against the ML tree, by comparing a set of fixed trees. Build the constrained tree with -g constraint.tree, then evaluate both trees:

bash
# trees.nex holds the unconstrained ML tree + the constrained/alternative trees
iqtree2 -s aln.fasta -m <model> -z trees.nex -n 0 -zb 10000 -au --prefix autest
#  -z trees.nex   trees to compare        -n 0   no fresh search, just evaluate
#  -zb 10000      RELL replicates (>=1000) -au    add the AU test (must accompany -zb)

The AU test (Shimodaira 2002) uses multiscale bootstrap resampling to correct both the selection bias of the KH test (invalid on the data-selected ML tree) and the over-conservatism of the SH test (which rejects less as the candidate set is padded). Use AU by default. p-AU < 0.05 means that tree is REJECTED; p-AU >= 0.05 means it is in the 95% confidence set (failure to reject is not acceptance -- weak data fails to reject many trees). Report SH/KH only for completeness.

Show full SKILL.md (1,077 more words)Show less

Partitioning

When splitting an alignment into partitions, the branch-length linkage choice is what people get wrong:

ModeFlagBranch lengthsUse when
Edge-equal-q part.nexidentical across partitionspartitions share rate (rare, restrictive)
Edge-linked proportional-p part.nexshared topology, per-partition rate multiplierDEFAULT -- genes evolve at different speeds, share history
Edge-unlinked-Q part.nexfully independent per partitiongenuine heterotachy; parameter-hungry, overfits

-p (edge-linked proportional) is the standard: one rate multiplier per partition over a shared topology. Over-partitioning spends degrees of freedom without bias reduction and inflates variance; the antidote is to start fine (gene x codon position) and let -m MFP+MERGE -rcluster 10 find the coarsest BIC-justified scheme. Prefer a merged scheme over a hand-picked maximal one; prefer -p over -Q.

Per-Method Failure Modes

Long-Branch Attraction

Trigger: Two or more independently fast-evolving lineages on long branches separated by a short internode. Mechanism: Convergent/homoplastic substitutions on the long branches look like shared ancestry; a site-homogeneous model cannot separate convergence from homology and groups them -- bias that GROWS with more sites. Symptom: Fast taxa group with the outgroup or each other at 100% bootstrap; the grouping collapses under a better model or when a long-branch taxon is removed. Fix: Site-heterogeneous model (C60/PMSF) first; remove the fastest sites and watch the node; drop the long-branch taxon or use a closer outgroup; cross-check with SR4/Dayhoff recoding. Believe a deep node only when it survives all of these, not when it merely has UFBoot 100 under LG+G.

Model Underspecification

Trigger: A single inadequate model on heterogeneous data; the best site-homogeneous model by BIC still inadequate at depth. Mechanism: Wrong matrix, missing rate heterogeneity, or no partitioning biases the topology while support stays high. Symptom: Biologically implausible nodes with full support that move under a richer model. Fix: -m MFP (+MERGE for multi-locus), FreeRate +R, and at amino-acid depth a C60/PMSF mixture. Best-by-BIC among site-homogeneous models is not sufficient deep in the tree.

Over-Partitioning

Trigger: Hundreds of hand-defined partitions, especially under -Q. Mechanism: Each partition's model is estimated from too little data; parameter and branch-length estimates get noisy without reducing bias. Symptom: Slow runs, noisy estimates, degraded support. Fix: -m MFP+MERGE -rcluster 10 to the coarsest BIC-justified scheme; use -p, not -Q.

The Support-Accuracy Gap

Trigger: UFBoot/bootstrap ~100 everywhere, including implausible or conflicting nodes. Mechanism: Support measures repeatability of a possibly-biased estimate; concatenation pools conflicting gene signals so the likelihood is certain while the loci disagree. Symptom: Full support but gCF ~33 / sCF near its ~33% floor on contested nodes. Fix: Compute gCF/sCFL; treat UFBoot 100 + gCF ~33 as UNRESOLVED; require SH-aLRT >=80 AND UFBoot >=95; use -bnni. If most genes reject the ML resolution, route to a coalescent species tree -> species-trees.

Quantitative Thresholds

QuantityThresholdSource
UFBoot strong support>=95 (NOT the bootstrap 70)Hoang 2018
SH-aLRT strong support>=80Guindon 2010
Joint ruleSH-aLRT >=80 AND UFBoot >=95community standard
Standard bootstrap "good">=70 (different metric, folklore)Felsenstein 1985 / Hillis 1993
aBayes strong>=0.95 (anti-conservative; never alone)Anisimova 2011
Bootstrap replicatesUFBoot -B >=1000; SH-aLRT -alrt >=1000; AU -zb 10000IQ-TREE docs
gCF reading>~75 agree; ~50 conflicted; ~33 effective polytomy; <33 with higher alternative = possible wrong resolutionMinh 2020
sCF floor~33% (three quartet resolutions); ~33 = no site signalMinh 2020
AU testp-AU < 0.05 => tree REJECTEDShimodaira 2002
Model selectionrank by BIC (-m MFP default), k*ln(n) penaltyKalyaanamoorthy 2017

Common Errors

Error / symptomCauseSolution
Unknown argument --alrtwrote the GNU double-dash formIQ-TREE2 uses single-dash -alrt, -bnni, -B
-bb / -nt not recognizedv1.x flags on a v2.x binaryuse -B (bootstrap) and -T (threads) in 2.x
Reading UFBoot 80 as "supported"applied the bootstrap-70 rule to a different scaleuse UFBoot >=95 AND SH-aLRT >=80
Fully-supported deep node distrusted by reviewerno concordance factors reportedcompute gCF/sCFL; treat high-support/low-CF as unresolved
--scfl unrecognizedIQ-TREE older than 2.2.2upgrade, or fall back to parsimony --scf
Concatenated tree confidently wrong on a rapid radiationILS; concatenation is inconsistent in the anomaly zoneinfer per-locus gene trees and a coalescent species tree -> species-trees
AU test "fails to reject" the alternativeweak data, or candidate set paddeddo not pad the set; failure to reject is not acceptance

References

Felsenstein J. 1978. Cases in which parsimony or compatibility methods will be positively misleading. Systematic Zoology 27(4):401-410. Felsenstein J. 1985. Confidence limits on phylogenies: an approach using the bootstrap. Evolution 39(4):783-791. Guindon S, Dufayard J-F, Lefort V, Anisimova M, Hordijk W, Gascuel O. 2010. New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of PhyML 3.0. Systematic Biology 59(3):307-321. Price MN, Dehal PS, Arkin AP. 2010. FastTree 2: approximately maximum-likelihood trees for large alignments. PLoS ONE 5(3):e9490. Anisimova M, Gil M, Dufayard J-F, Dessimoz C, Gascuel O. 2011. Survey of branch support methods demonstrates accuracy, power, and robustness of fast likelihood-based approximation schemes. Systematic Biology 60(5):685-699. Shimodaira H. 2002. An approximately unbiased test of phylogenetic tree selection. Systematic Biology 51(3):492-508. Kalyaanamoorthy S, Minh BQ, Wong TKF, von Haeseler A, Jermiin LS. 2017. ModelFinder: fast model selection for accurate phylogenetic estimates. Nature Methods 14(6):587-589. Wang H-C, Minh BQ, Susko E, Roger AJ. 2018. Modeling site heterogeneity with posterior mean site frequency profiles accelerates accurate phylogenomic estimation. Systematic Biology 67(2):216-235. Hoang DT, Chernomor O, von Haeseler A, Minh BQ, Vinh LS. 2018. UFBoot2: improving the ultrafast bootstrap approximation. Molecular Biology and Evolution 35(2):518-522. Lemoine F, Domelevo Entfellner J-B, Wilkinson E, Correia D, Davila Felipe M, De Oliveira T, Gascuel O. 2018. Renewing Felsenstein's phylogenetic bootstrap in the era of big data. Nature 556(7702):452-456. Kozlov AM, Darriba D, Flouri T, Morel B, Stamatakis A. 2019. RAxML-NG: a fast, scalable and user-friendly tool for maximum likelihood phylogenetic inference. Bioinformatics 35(21):4453-4455. Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, Lanfear R. 2020. IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic era. Molecular Biology and Evolution 37(5):1530-1534. Minh BQ, Hahn MW, Lanfear R. 2020. New methods to calculate concordance factors for phylogenomic datasets. Molecular Biology and Evolution 37(9):2727-2733. Mo YK, Lanfear R, Hahn MW, Minh BQ. 2023. Updated site concordance factors minimize effects of homoplasy and taxon sampling. Bioinformatics 39(1):btac741.

  • distance-calculations - model-corrected distances and fast NJ trees as a model-free alternative
  • bayesian-inference - posteriors, MCMC convergence, and CAT-GTR site-heterogeneous models
  • species-trees - coalescent species-tree estimation when concordance factors reveal ILS
  • divergence-dating - time-scaled trees from the ML topology
  • tree-manipulation - rooting, pruning, and collapsing low-support nodes
  • alignment/alignment-io - the alignment whose homology assumption the ML tree trusts as fixed

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in phylogenetics/modern-tree-inference of GPTomics/bioSkills.

  • SKILL.md
  • examples/iqtree_basic.sh
  • examples/partitioned_analysis.sh
  • examples/raxml_analysis.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Phylo Modern Tree Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Phylo Modern Tree Inference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Phylo Modern Tree Inference this skillGPTomics/bioSkills1.2k1 repos~5.3kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Phylo Modern Tree Inference

What does Bio Phylo Modern Tree Inference do?

Infers maximum-likelihood phylogenetic trees with IQ-TREE2 and RAxML-NG -- model selection (ModelFinder), branch support (UFBoot2, SH-aLRT), concordance factors (gCF/sCF), partitioning, topology…. Bio Phylo Modern Tree Inference is an agent skill from GPTomics/bioSkills. Infers maximum-likelihood phylogenetic trees with IQ-TREE2 and RAxML-NG -- model selection (ModelFinder), branch support (UFBoot2, SH-aLRT), concordance factors (gCF/sCF), partitioning, topology tests, and long-branch-attraction control.

When should I use Bio Phylo Modern Tree Inference?

Bio Phylo Modern Tree Inference fits situations like: inferring an ML tree; selecting a substitution; partition model; interpreting support measures.

How do I install Bio Phylo Modern Tree Inference in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-phylo-modern-tree-inference -a claude-code`. Or copy the skill folder (phylogenetics/modern-tree-inference in GPTomics/bioSkills) into .claude/skills/bio-phylo-modern-tree-inference in your project. Claude Code loads it when a task matches its description.

How do I install Bio Phylo Modern Tree Inference in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-phylo-modern-tree-inference -a codex`. Or copy the skill folder (phylogenetics/modern-tree-inference in GPTomics/bioSkills) into .agents/skills/bio-phylo-modern-tree-inference in your project. Codex loads it when a task matches its description.

Can I use Bio Phylo Modern Tree Inference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-phylo-modern-tree-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-phylo-modern-tree-inference, .gemini/skills/bio-phylo-modern-tree-inference, .github/skills/bio-phylo-modern-tree-inference and .opencode/skills/bio-phylo-modern-tree-inference in your project.

What does Bio Phylo Modern Tree Inference need to run?

Going by SKILL.md and its folder, Bio Phylo Modern Tree Inference needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Bio Phylo Modern Tree Inference access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Phylo Modern Tree Inference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Phylo Modern Tree Inference use?

Bio Phylo Modern Tree Inference is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Phylo Modern Tree Inference use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Phylo Modern Tree Inference?

Skills that share tags, products or a category with Bio Phylo Modern Tree Inference: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Phylo Modern Tree Inference?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.