Agent skill

Bio Microbiome Differential Abundance

by GPTomics in GPTomics/bioSkills

Tests which individual taxa differ between groups on an amplicon ASV/feature table (phyloseq) using compositionally-aware methods - ALDEx2 (Dirichlet-MC CLR, conservative), ANCOM-BC2/ANCOMBC…

MITAuto-check passedResearch & Science

Install Bio Microbiome Differential Abundance

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-microbiome-differential-abundance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-microbiome-differential-abundance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/microbiome/differential-abundance .claude/skills/bio-microbiome-differential-abundance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-microbiome-differential-abundance
GitHub stars
1.2k
Used in
1 other repo
Token cost
~6.1k tokens
SKILL.md length
2,524 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Tests which individual taxa differ between groups on an amplicon ASV/feature table (phyloseq) using compositionally-aware methods - ALDEx2 (Dirichlet-MC CLR, conservative), ANCOM-BC2/ANCOMBC…

  • Works in 3 steps: A relative-abundance increase is not an… → Uncorrected Wilcoxon/t-test on raw… → There is no settled best tool. The…
  • Finding differentially abundant taxa
  • SKILL.md covers Version Compatibility, The Single Most Important…, The Benchmark Landscape (no… and Tool Taxonomy, plus 12 more sections
  • Runs R scripts from its folder

What it does

Bio Microbiome Differential Abundance is an agent skill from GPTomics/bioSkills. Tests which individual taxa differ between groups on an amplicon ASV/feature table (phyloseq) using compositionally-aware methods - ALDEx2 (Dirichlet-MC CLR, conservative), ANCOM-BC2/ANCOMBC (sampling-fraction bias correction, structural zeros, passedss, default padjmethod=holm), MaAsLin2/MaAsLin3 (multivariable GLM, random effects, prevalence/abundance split), LinDA (CLR mixed-model regression), ZicoSeq (permutation FDR), LEfSe, and q2-composition ancombc. Covers why the hit list depends more on the DA tool than…

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Finding differentially abundant taxa
  • Handling covariates
  • Longitudinal designs
  • Choosing a method

Example prompts

  • “Use the bio-microbiome-differential-abundance skill to test which individual taxa differ between groups on an amplicon ASV/feature table (phyloseq)…”
  • “/bio-microbiome-differential-abundance”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A relative-abundance increase is not an absolute increase. Microbiome counts are compositional - the sequencer fixes the total, so one…
  2. Uncorrected Wilcoxon/t-test on raw relative abundances is wrong twice in one line - closure (reference-frame) AND multiple testing. But a…
  3. There is no settled best tool. The benchmarks optimize different criteria, so they rank tools differently. Consensus-of-tools is the only…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Microbiome Differential Abundance loads about 6.1k tokens when it runs. Until then it costs about 266 tokens; SKILL.md has 2,524 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~266
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,524 words, ~6,126 tokens.

Download SKILL.mdSave it as .claude/skills/bio-microbiome-differential-abundance/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-microbiome-differential-abundance
description
Tests which individual taxa differ between groups on an amplicon ASV/feature table (phyloseq) using compositionally-aware methods - ALDEx2 (Dirichlet-MC CLR, conservative), ANCOM-BC2/ANCOMBC (sampling-fraction bias correction, structural zeros, passed_ss, default p_adj_method=holm), MaAsLin2/MaAsLin3 (multivariable GLM, random effects, prevalence/abundance split), LinDA (CLR mixed-model regression), ZicoSeq (permutation FDR), LEfSe, and q2-composition ancombc. Covers why the hit list depends more on the DA tool than the biology (Nearing benchmark) so the deliverable is a CONSENSUS of >=2 tools, why a relative change is not absolute without a load anchor, the prevalence-filter knob, BH/FDR plus an effect-size floor, and why DESeq2/edgeR misfire here. Use when finding differentially abundant taxa, handling covariates or longitudinal designs, or choosing a method. Whole-community diversity -> diversity-analysis; shotgun DA -> metagenomics/metagenome-visualization; CoDA theory -> metagenomics/abundance-estimation
tool_type
r
primary_tool
ALDEx2

Version Compatibility

Reference examples tested with: ALDEx2 1.34+, ANCOMBC 2.4+, Maaslin2 1.16+, MicrobiomeStat 1.2+ (LinDA), GUniFrac 1.8+ (ZicoSeq), phyloseq 1.46+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

ANCOM-BC2 changed argument names between ancombc() and ancombc2(), and its default p_adj_method is holm, not BH - confirm both against the installed version. The MaAsLin3 maaslin3() API differs from MaAsLin2's Maaslin2().

Differential Abundance Testing

"Find which taxa differ between my groups" -> Run two or more compositionally-aware DA tools and report their consensus - because the significant-taxa list is a property of the tool as much as of the sample, and a relative-abundance change is not an absolute change.

  • R: ALDEx2::aldex(counts, conds, test='t', effect=TRUE, denom='all') then a second tool (ANCOMBC::ancombc2() or MicrobiomeStat::linda())

Scope: per-taxon DA on an amplicon feature table. Whole-community alpha/beta/PERMANOVA -> diversity-analysis. Shotgun profiler-table DA -> metagenomics/metagenome-visualization. Shared compositional/closure/CLR/zero theory -> metagenomics/abundance-estimation. Collapse ASVs to genus/species first -> taxonomy-assignment. QIIME2 CLI route -> qiime2-workflow.

The Single Most Important Modern Insight -- Which Taxa Are "Significant" Depends More on the Tool Than on the Biology

Run ALDEx2, ANCOM-BC2, MaAsLin2, and LinDA on the same ASV table and the four significant-taxa lists overlap but disagree (Nearing 2022 Nat Commun 13:342, across 38 datasets). So the deliverable is NOT "the differential taxa" - it is the CONSENSUS of >=2 compositionally-aware tools, every tool NAMED: the intersection is high-confidence, the union is exploratory, and a single-tool hit is tentative. Picking the tool with the prettiest volcano is p-hacking by software (uncorrected multiplicity hidden in the method menu). Three corollaries:

  1. A relative-abundance increase is not an absolute increase. Microbiome counts are compositional - the sequencer fixes the total, so one taxon blooming forces every other taxon's proportion down (the blooming-taxon illusion). "Taxon X increased" is a statement about its SHARE unless an external load anchor (spike-in / flow cytometry / qPCR, see metagenomics/abundance-estimation) or MaAsLin3's absolute-abundance mode licenses an absolute claim.
  2. Uncorrected Wilcoxon/t-test on raw relative abundances is wrong twice in one line - closure (reference-frame) AND multiple testing. But a BH-corrected simple test, honestly labelled as relative, can replicate BETTER than a sophisticated model (Pelto 2025): the forbidden thing is the uncorrected, closure-blind form, not simple tests per se.
  3. There is no settled best tool. The benchmarks optimize different criteria, so they rank tools differently. Consensus-of-tools is the only stance that survives all of them.

The Benchmark Landscape (no settled winner)

BenchmarkOptimized forVerdict
Nearing 2022 Nat Commun 13:342cross-method consistencyALDEx2 + ANCOM-II most consistent and most conservative; LEfSe/edgeR flag far more, agree less
Yang & Chen 2022 Microbiome 10:130FDR-power balanceZicoSeq / LinDA / ANCOM-BC-family best
Yang & Chen 2023 Brief Bioinform 24:bbac607correlated (repeated-measures) designsuse a mixed-model-capable tool (LinDA, MaAsLin2, ANCOM-BC2)
Pelto 2025 Brief Bioinform 26(2):bbaf130cross-study replicabilityelementary BH-corrected methods most replicable; ANCOM-BC2 worst

Report the disagreement AS the result; verify current best practice against the latest tool docs rather than hard-coding one method.

Tool Taxonomy

ToolCitationMechanism / roleWhen
ALDEx2Fernandes 2014 Microbiome 2:15Dirichlet Monte-Carlo posterior + CLR; tests each draw; reports expected effect + BH-adjusted pconservative two-group anchor; small-to-moderate n
ANCOM-BC2Lin & Peddada 2024 Nat Methods 21:83estimates per-sample sampling fraction and bias-corrects; structural zeros; pseudo-count sensitivity (passed_ss)interpretable LFC + CI; covariates; multi-group
MaAsLin2Mallick 2021 PLoS Comput Biol 17:e1009442general (mixed) linear model on transformed abundancemultivariable / longitudinal / metadata-rich
MaAsLin3Nickols 2026 Nat Methods 23:554splits abundance (level when present) from prevalence (present/absent); absolute-abundance modeprevalence-vs-abundance separation; load data available
LinDAZhou 2022 Genome Biol 23:95CLR regression with mode-based bias correction; asymptotic FDRlarge cohorts; fast; native mixed model
ZicoSeqYang & Chen 2022 Microbiome 10:130reference-taxa normalization + permutation FDR; winsorizationcovariates; non-parametric permutation p; strong FP control
LEfSeSegata 2011 Genome Biol 12:R60Kruskal-Wallis + LDA effect sizeexploratory biomarker discovery; NOT a formal FDR-controlled test
DESeq2Love 2014 Genome Biol 15:550RNA-seq median-of-ratios size factorcaveat only; geometric-mean reference dies on sparse zero-heavy tables

Decision Tree by Scenario

ScenarioRecommendedWhy
Two groups, want a trustworthy conservative anchorALDEx2Dirichlet-MC + CLR; most reproducible/conservative (Nearing); gate on effect size
Need interpretable LFC + CI, structural zeros, multi-groupANCOM-BC2models and corrects per-sample sampling fraction; global/pairwise/Dunnett/trend; passed_ss
Large cohort, covariates, speed, mixed modelLinDACLR regression + bias mode; asymptotic FDR; fast; random effects in the formula
Covariates + permutation-grounded non-parametric pZicoSeqreference-taxa frame + permutation FDR
Longitudinal / many covariates / flexible GLMMaAsLin2fixed_effects + random_effects; normalization/transform menu
Prevalence-vs-abundance separation or absolute abundanceMaAsLin3logistic prevalence model + abundance model; load-data hook
Inside a QIIME2 CLI pipelineqiime composition ancombc + tabulate/da-barplotnative artifact flow (v1 ANCOM-BC; go to R for v2 passed_ss/multi-group) -> qiime2-workflow
Repeated / paired samplesany tool above WITH a random effectignoring subject structure is pseudo-replication
ALWAYSrun >=2 of the above, report the consensustool choice drives the hit list more than biology (Nearing 2022)
Shotgun species table, not amplicon-> metagenomics/metagenome-visualizationsame CoDA theory; different upstream pipeline
Uncorrected t-test/Wilcoxon on TSS proportionsDO NOTclosure biases the test and there is no FDR control

Filter Before Testing (a modeling knob, not housekeeping)

Goal: Drop rare features before testing so the BH denominator is not crushed and log/CLR transforms are well-behaved.

Approach: Keep features present in at least 10-25% of samples (and optionally a mean-abundance floor); declare the threshold and confirm the headline result is not knife-edge-sensitive to it. Every tool exposes this (prv_cut, min_prevalence, prev.filter).

r
library(phyloseq)
ps <- readRDS('phyloseq_object.rds')
# prv_cut 0.10: a feature must appear in >= 10% of samples; raising to 0.25 removes more tests
# (smaller BH correction, more power on survivors) but discards rare-but-real taxa - a declared choice
keep <- filter_taxa(ps, function(x) sum(x > 0) >= 0.10 * nsamples(ps), TRUE)

ALDEx2: The Conservative Floor of the Consensus

Goal: Identify taxa that differ between two groups while propagating the sampling uncertainty of low-count features.

Approach: Draw mc.samples Monte-Carlo instances from a Dirichlet posterior of the counts (this IS the zero handling - no explicit pseudocount), CLR-transform each instance against the geometric mean of all features (denom='all'), run the test on every draw, and report the EXPECTED effect size and BH-adjusted p over the draws.

r
library(ALDEx2)
counts <- as.matrix(otu_table(ps))            # integer counts, taxa in ROWS
if (!taxa_are_rows(ps)) counts <- t(counts)
groups <- as.character(sample_data(ps)$Group)

# mc.samples 128: standard Monte-Carlo draws; 256+ for publication (more stable expected p)
res <- aldex(counts, groups, mc.samples = 128, test = 't', effect = TRUE, denom = 'all')
# we.eBH = Welch expected BH-adjusted p (report this, NOT we.ep); wi.eBH = Wilcoxon equivalent
# effect = median standardized effect = median(diff.btw / max(diff.win)); the primary decision variable
hits <- res[res$we.eBH < 0.05 & abs(res$effect) > 1, ]   # q AND effect floor (Gloor: gate on effect, not p alone)

Gate on effect size AND q, not p alone: with large n trivially small CLR differences become "significant," and Gloor's own guidance is that |effect| > 1 is a strong ~2-SD signal. For >2 groups use aldex.kw(); for covariates the aldex.glm() + model.matrix route works but ALDEx2 is weakest here - prefer ANCOM-BC2/LinDA/MaAsLin2 for serious covariate or random-effect modeling.

ANCOM-BC2: Bias-Corrected LFC With a Sensitivity Safeguard

Goal: Estimate an interpretable bias-corrected log-fold-change per taxon, with covariate adjustment, structural-zero handling, and a flag for hits that are hostage to the pseudo-count.

Approach: Model log(observed count) as a function of covariates, estimate each sample's log sampling fraction as an offset and subtract it, then refit across a range of pseudo-counts and record how often each q-value flips (passed_ss).

r
library(ANCOMBC)
out <- ancombc2(data = ps, fix_formula = 'Group + Age + Sex',
                rand_formula = NULL,        # '(1 | SubjectID)' for repeated measures - see Failure Modes
                p_adj_method = 'BH',        # DEFAULT is 'holm'; set 'BH' deliberately for FDR
                prv_cut = 0.10, lib_cut = 1000,
                group = 'Group', struc_zero = TRUE, pseudo_sens = TRUE,
                global = FALSE, pairwise = FALSE, n_cl = 2)
res <- out$res
# a confident hit is BOTH significant AND robust to the pseudo-count. ANCOM-BC2 suffixes the
# diff_/passed_ss_ columns with the literal model-matrix coefficient (variable + factor level,
# verbatim case, e.g. 'Grouptreated') - match it by pattern rather than hard-coding the case.
dcol <- grep('^diff_Group', names(res), value = TRUE)[1]
robust <- res[res[[dcol]] & res[[sub('^diff_', 'passed_ss_', dcol)]], ]

passed_ss is the most valuable ANCOM-BC2-specific feature: a CLR/log model on sparse data is hostage to the zero-replacement constant, and passed_ss quantifies that per taxon. A hit with passed_ss == FALSE depends on the arbitrary pseudo-count - do not report it as confident. For >2 groups set global=TRUE (omnibus), pairwise=TRUE (mdFDR-controlled pairs), dunnet=TRUE, or trend=TRUE; results land in out$res_global/res_pair/res_dunn/res_trend.

LinDA: Fast CLR Regression With Native Mixed Models

Goal: Get FDR-controlled log2-fold-changes on a large cohort, including repeated-measures designs, without Monte-Carlo or EM cost.

Approach: Fit ordinary linear regression on the CLR-transformed table covariate by covariate, estimate the compositional bias as the mode of the per-feature coefficients and subtract it; a random effect in the formula makes it a linear mixed model.

r
library(MicrobiomeStat)
otu <- as.data.frame(otu_table(ps)); if (!taxa_are_rows(ps)) otu <- t(otu)
meta <- as.data.frame(sample_data(ps))
fit <- linda(feature.dat = otu, meta.dat = meta,
             formula = '~ Group + Age + (1 | SubjectID)',   # random effect -> mixed model
             feature.dat.type = 'count', prev.filter = 0.10, alpha = 0.05)
fit$output[[1]]   # names(fit$output) are the model-matrix coefficient columns (e.g. 'Grouptreated' - the factor level keeps its case); per-feature: log2FoldChange, lfcSE, stat, pvalue, padj, reject

LinDA is the natural fast modern entry in a consensus panel and the cleanest route to mixed models. Yang & Chen rate it among the best FDR-power trade-offs.

MaAsLin2 / MaAsLin3 and ZicoSeq (the rest of the panel)

Goal: Fit covariate-rich or longitudinal differential-abundance models, or add a permutation-based panel member, when ALDEx2/ANCOM-BC2/LinDA do not cover the design.

Approach: Use MaAsLin2/3 for multivariable GLMs with random effects, or ZicoSeq for a non-parametric permutation-FDR test against empirically selected reference taxa.

MaAsLin2 fits a flexible per-feature GLM; its package DEFAULT is TSS + LOG + LM (not CLR), and random_effects is the canonical route for longitudinal designs. NOTE the orientation gotcha: it expects features in COLUMNS, samples in rows.

r
library(Maaslin2)
fit <- Maaslin2(input_data = as.data.frame(t(otu)), input_metadata = meta,
                output = 'maaslin2_out', fixed_effects = c('Group', 'Age'),
                random_effects = c('SubjectID'),
                normalization = 'TSS', transform = 'LOG', analysis_method = 'LM',
                min_prevalence = 0.10, max_significance = 0.05)
# writes all_results.tsv / significant_results.tsv with columns feature, metadata, coef, pval, qval

MaAsLin3 (maaslin3()) splits each feature into an abundance model (level when present) and a logistic prevalence model (present/absent) tested jointly, and can ingest total-load measurements for absolute-abundance inference. ZicoSeq (GUniFrac::ZicoSeq()) winsorizes, posterior-samples, normalizes against empirically selected reference taxa, and returns permutation FDR (zc$p.adj.fdr) - a non-parametric panel member that accepts covariates via adj.name.

Consensus: Intersect the Tools

Goal: Convert two or more per-tool hit sets into a confidence-graded result instead of one tool's answer.

Approach: Collect the significant feature SETS (BH within each tool), then report the intersection as high-confidence, the union as exploratory, and tabulate, per taxon, how many of N tools agree and which ones. Never pool p-values across tools.

r
sig_aldex <- rownames(res)[res$we.eBH < 0.05 & abs(res$effect) > 1]
sig_linda <- rownames(fit$output[[1]])[fit$output[[1]]$reject]   # [[1]] = the group coefficient (named 'Grouptreated')
confident  <- intersect(sig_aldex, sig_linda)   # high-confidence
exploratory <- union(sig_aldex, sig_linda)       # report with the tool that found each

Per-Method Failure Modes

Cherry-picking the tool with the prettiest result

Trigger: running several tools and reporting only the one(s) that flag the favored taxon. Mechanism: that is uncorrected multiplicity hidden in the method menu (p-hacking by software). Symptom: "the recommended method found X" with no mention of the tools that disagreed. Fix: decide the panel a priori, report ALL tools, intersect for confident hits, disclose disagreement.

Show full SKILL.md (1,014 more words)Show less
Uncorrected Wilcoxon/t-test on relative abundances

Trigger: a per-taxon Wilcoxon/t-test on TSS proportions with no FDR correction. Mechanism: closure makes a naive test call every taxon "decreased" when one blooms, and hundreds of uncorrected tests inflate false positives. Symptom: dozens of "significant" taxa, all in the same direction, no q-values. Fix: use a CoDA/reference-frame tool; if a simple test is used, BH-correct it and label the comparison as relative (Pelto 2025).

Pseudo-replication of repeated measures

Trigger: longitudinal/paired samples treated as independent rows. Mechanism: fewer independent units than rows inflates significance. Symptom: implausibly small p-values on a small subject count. Fix: a random effect - ANCOM-BC2 rand_formula='(1|SubjectID)', MaAsLin2 random_effects='SubjectID', LinDA (1|SubjectID) in the formula. Cross-check ANCOM-BC2 mixed-model output against LinDA/MaAsLin2 (GitHub issue #111 reported rand_formula correctness problems in some versions).

Prevalence filter set blindly

Trigger: an undeclared prevalence cut, or none at all. Mechanism: the cut decides which taxa are even tested and thus the BH landscape - it is a modeling choice. Symptom: the hit list changes materially between prv_cut=0.1 and 0.25. Fix: declare and justify the threshold; confirm the headline result survives moving it.

Relative change reported as absolute

Trigger: "taxon X doubled" from a closed table with no load data. Mechanism: one taxon blooming compresses every other proportion. Symptom: whole-community "depletion" that is really one taxon rising. Fix: anchor to load (spike-in/flow/qPCR) or MaAsLin3 absolute mode; otherwise state the claim is relative.

DESeq2/edgeR on a sparse 16S table

Trigger: RNA-seq median-of-ratios / TMM on a zero-heavy ASV table. Mechanism: the geometric-mean size-factor reference collapses on zeros and the "most features unchanged" assumption is violated. Symptom: degenerate size factors, errors, or inflated hit counts that disagree with CoDA tools (Nearing). Fix: use a compositional tool; if DESeq2 is unavoidable, the poscounts estimator is the minimum mitigation - present as a caveat, not a recipe.

ANCOM-BC2 hit held hostage by the pseudo-count

Trigger: reporting diff_* == TRUE without checking passed_ss_*. Mechanism: significance depends on the arbitrary zero-replacement constant. Symptom: a hit that vanishes when the pseudo-count changes. Fix: require diff_* & passed_ss_* for a confident call.

Quantitative Thresholds

ThresholdSourceRationale
Prevalence cut 10-25% (prv_cut/min_prevalence/prev.filter)Nearing 2022; tool defaults (0.10)rare features carry little information and crush the BH denominator; declare the value and test sensitivity
BH q <= 0.05 across taxa, within each toolBenjamini-Hochberg 1995 JRSS B 57:289hundreds-thousands of features make uncorrected p meaningless; do not pool p across tools
ALDEx2 `effect> 1` (with q <= 0.05)
ALDEx2 mc.samples = 128 (256+ for publication)Fernandes 2014 Microbiome 2:15Monte-Carlo draws; more draws stabilize the expected p
ANCOM-BC2 passed_ss == TRUE requiredLin & Peddada 2024 Nat Methods 21:83flags hits whose significance is hostage to the pseudo-count
Consensus of >=2 compositionally-aware toolsNearing 2022 Nat Commun 13:342tool choice drives the hit list more than biology; intersection = confident
ZicoSeq permutations perm.no >= 99Yang & Chen 2022 Microbiome 10:130permutation FDR resolution; raise for finer tail p

Common Errors

Error / symptomCauseSolution
ALDEx2 returns NA effects / errorsproportions or non-integer matrix passedfeed integer COUNTS with taxa in rows
passed_ss column missingpseudo_sens = FALSEset pseudo_sens = TRUE (the default)
Far fewer hits than expectedANCOM-BC2 p_adj_method left at holmset p_adj_method = 'BH' deliberately if FDR is wanted
MaAsLin2 finds nothing / orientation errorfeatures in rows, not columnstranspose so samples are rows, features columns
Mixed-model hits disagree across toolsrand_formula correctness varies by versioncross-check ANCOM-BC2 against LinDA/MaAsLin2
Tools disagree on the hit listnormal - tool choice drives resultsreport the consensus and the disagreement, do not cherry-pick
Many "depleted" taxa in a host/plant samplehost mitochondria/chloroplast 16S inflates the tablefilter Mitochondria/Chloroplast features (see taxonomy-assignment) before DA
Contaminant ASVs among the hits (low-biomass)reagent kitome not removed before DArun decontam upstream with negative controls (amplicon-processing; metagenomics/contamination-controls)

References

  • Fernandes AD, Reid JNS, Macklaim JM, McMurrough TA, Edgell DR, Gloor GB. 2014. Unifying the analysis of high-throughput sequencing datasets: characterizing RNA-seq, 16S rRNA gene sequencing and selective growth experiments by compositional data analysis. Microbiome 2:15.
  • Gloor GB, Macklaim JM, Fernandes AD. 2016. Displaying variation in large datasets: plotting a visual summary of effect sizes. J Comput Graph Stat 25:971-979.
  • Lin H, Peddada SD. 2020. Analysis of compositions of microbiomes with bias correction (ANCOM-BC). Nat Commun 11:3514.
  • Lin H, Peddada SD. 2024. Multigroup analysis of compositions of microbiomes with covariate adjustments and repeated measures (ANCOM-BC2). Nat Methods 21:83-91.
  • Mallick H, Rahnavard A, McIver LJ, et al. 2021. Multivariable association discovery in population-scale meta-omics studies. PLoS Comput Biol 17:e1009442.
  • Nickols WA, Kuntz T, Shen J, et al. 2026. MaAsLin 3: refining and extending generalized multivariable linear models for meta-omic association discovery. Nat Methods 23:554-564.
  • Zhou H, He K, Chen J, Zhang X. 2022. LinDA: linear models for differential abundance analysis of microbiome compositional data. Genome Biol 23:95.
  • Yang L, Chen J. 2022. A comprehensive evaluation of microbial differential abundance analysis methods: current status and potential solutions. Microbiome 10:130.
  • Yang L, Chen J. 2023. Benchmarking differential abundance analysis methods for correlated microbiome sequencing data. Brief Bioinform 24:bbac607.
  • Pelto J, Auranen K, Kujala JV, Lahti L. 2025. Elementary methods provide more replicable results in microbial differential abundance analysis. Brief Bioinform 26(2):bbaf130.
  • Nearing JT, Douglas GM, Hayes MG, et al. 2022. Microbiome differential abundance methods produce different results across 38 datasets. Nat Commun 13:342.
  • Segata N, Izard J, Waldron L, et al. 2011. Metagenomic biomarker discovery and explanation. Genome Biol 12:R60.
  • Love MI, Huber W, Anders S. 2014. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15:550.
  • Benjamini Y, Hochberg Y. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Series B 57:289-300.
  • diversity-analysis - Whole-community alpha/beta/PERMANOVA; answer "do the communities differ" before "which taxa differ"
  • taxonomy-assignment - Collapse ASVs to genus/species before per-taxon testing
  • amplicon-processing - Produces the ASV feature table tested here
  • qiime2-workflow - The qiime composition ancombc CLI route
  • metagenomics/abundance-estimation - Shared compositional/closure/CLR/zero/load-anchor theory
  • metagenomics/metagenome-visualization - The same DA mechanics on shotgun profiler tables
  • experimental-design/multiple-testing - FDR control and multiplicity across taxa

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in microbiome/differential-abundance of GPTomics/bioSkills.

  • SKILL.md
  • examples/aldex2_analysis.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Microbiome Differential Abundance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Microbiome Differential Abundance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Microbiome Differential Abundance this skillGPTomics/bioSkills1.2k1 repos~6.1kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Microbiome Differential Abundance

What does Bio Microbiome Differential Abundance do?

Tests which individual taxa differ between groups on an amplicon ASV/feature table (phyloseq) using compositionally-aware methods - ALDEx2 (Dirichlet-MC CLR, conservative), ANCOM-BC2/ANCOMBC…. Bio Microbiome Differential Abundance is an agent skill from GPTomics/bioSkills. Tests which individual taxa differ between groups on an amplicon ASV/feature table (phyloseq) using compositionally-aware methods - ALDEx2 (Dirichlet-MC CLR, conservative), ANCOM-BC2/ANCOMBC (sampling-fraction bias correction, structural zeros, passedss, default padjmethod=holm), MaAsLin2/MaAsLin3 (multivariable GLM, random effects, prevalence/abundance split), LinDA (CLR mixed-model regression), ZicoSeq (permutation FDR), LEfSe, and q2-composition ancombc.

When should I use Bio Microbiome Differential Abundance?

Bio Microbiome Differential Abundance fits situations like: finding differentially abundant taxa; handling covariates; longitudinal designs; choosing a method.

How do I install Bio Microbiome Differential Abundance in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-microbiome-differential-abundance -a claude-code`. Or copy the skill folder (microbiome/differential-abundance in GPTomics/bioSkills) into .claude/skills/bio-microbiome-differential-abundance in your project. Claude Code loads it when a task matches its description.

How do I install Bio Microbiome Differential Abundance in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-microbiome-differential-abundance -a codex`. Or copy the skill folder (microbiome/differential-abundance in GPTomics/bioSkills) into .agents/skills/bio-microbiome-differential-abundance in your project. Codex loads it when a task matches its description.

Can I use Bio Microbiome Differential Abundance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-microbiome-differential-abundance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-microbiome-differential-abundance, .gemini/skills/bio-microbiome-differential-abundance, .github/skills/bio-microbiome-differential-abundance and .opencode/skills/bio-microbiome-differential-abundance in your project.

What does Bio Microbiome Differential Abundance need to run?

Going by SKILL.md and its folder, Bio Microbiome Differential Abundance needs R for the scripts in its folder.

Does Bio Microbiome Differential Abundance access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Microbiome Differential Abundance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Microbiome Differential Abundance use?

Bio Microbiome Differential Abundance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Microbiome Differential Abundance use?

About 6.1k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Microbiome Differential Abundance?

Skills that share tags, products or a category with Bio Microbiome Differential Abundance: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Microbiome Differential Abundance?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.