Agent skill

Bio Microbiome Diversity Analysis

by GPTomics in GPTomics/bioSkills

Alpha and beta diversity of an amplicon (16S/ITS) ASV/OTU community table - observed features, Shannon, Pielou evenness, Faith PD, Bray-Curtis, Jaccard, weighted/unweighted/generalized UniFrac…

MITAuto-check passedResearch & Science

Install Bio Microbiome Diversity Analysis

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-microbiome-diversity-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-microbiome-diversity-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/microbiome/diversity-analysis .claude/skills/bio-microbiome-diversity-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-microbiome-diversity-analysis
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.3k tokens
SKILL.md length
2,246 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Alpha and beta diversity of an amplicon (16S/ITS) ASV/OTU community table - observed features, Shannon, Pielou evenness, Faith PD, Bray-Curtis, Jaccard, weighted/unweighted/generalized UniFrac…

  • Works in 3 steps: p-sampling-depth is a sample-deletion… → UniFrac and Faith PD are only as real as… → Rarefy for diversity, never for…
  • Summarizing whole-community richness/evenness
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool / Metric Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Shell and R scripts from its folder; calls pip

What it does

Bio Microbiome Diversity Analysis is an agent skill from GPTomics/bioSkills. Alpha and beta diversity of an amplicon (16S/ITS) ASV/OTU community table - observed features, Shannon, Pielou evenness, Faith PD, Bray-Curtis, Jaccard, weighted/unweighted/generalized UniFrac, Aitchison/RPCA - via QIIME2 core-metrics-phylogenetic, phyloseq/vegan, and scikit-bio. Covers the three knobs that set the answer before it is seen (rarefaction sampling depth, the tree, the metric), why core-metrics silently deletes samples below the sampling depth, why de novo trees lose to SEPP fragment-insertion and…

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/core_metrics_qiime2.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Summarizing whole-community richness/evenness
  • Testing group differences in community structure

Example prompts

  • “/bio-microbiome-diversity-analysis”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. p-sampling-depth is a sample-deletion knob in a normalization costume. core-metrics rarefies every sample to the depth and, per the QIIME2…
  2. UniFrac and Faith PD are only as real as the tree, and the tree is a model not a property of the data. A de novo MAFFT+FastTree tree from…
  3. Rarefy for diversity, never for differential abundance. Rarefaction-to-even-depth is defensible for alpha/beta (Schloss 2024); for DA it…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Microbiome Diversity Analysis loads about 5.3k tokens when it runs. Until then it costs about 262 tokens; SKILL.md has 2,246 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~262
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,246 words, ~5,348 tokens.

Download SKILL.mdSave it as .claude/skills/bio-microbiome-diversity-analysis/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-microbiome-diversity-analysis
description
Alpha and beta diversity of an amplicon (16S/ITS) ASV/OTU community table - observed features, Shannon, Pielou evenness, Faith PD, Bray-Curtis, Jaccard, weighted/unweighted/generalized UniFrac, Aitchison/RPCA - via QIIME2 core-metrics-phylogenetic, phyloseq/vegan, and scikit-bio. Covers the three knobs that set the answer before it is seen (rarefaction sampling depth, the tree, the metric), why core-metrics silently deletes samples below the sampling depth, why de novo trees lose to SEPP fragment-insertion and Greengenes2, why unweighted and weighted UniFrac can flip the story, why observed features is an ASV count not a species count, the QIIME2-log2 vs R-ln Shannon mismatch, and pairing PERMANOVA (adonis2) with betadisper. Use when summarizing whole-community richness/evenness or testing group differences in community structure. Per-taxon testing -> differential-abundance. Shotgun tables -> metagenomics/metagenome-visualization. Shared CoDA/rarefaction theory -> metagenomics/abundance-estimation.
tool_type
mixed
primary_tool
phyloseq

Version Compatibility

Reference examples tested with: phyloseq 1.46+, vegan 2.6+, picante 1.8+, GUniFrac 1.8+, scikit-bio 0.6+, QIIME2 2024.2+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: qiime <plugin> <action> --help to confirm flags
  • Python: pip show scikit-bio then help(skbio.diversity.beta_diversity) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

scikit-bio 0.6.0 renamed OTU to taxon across the API and drifted metric kwargs (otu_ids= vs newer forms) - discover names with skbio.diversity.get_beta_diversity_metrics() before hard-coding. UniFrac/Faith PD results inherit the tree (de novo vs SEPP vs Greengenes2 reference build) AND the chosen sampling depth - record both alongside the QIIME2 release that produced the .qza artifacts.

Diversity Analysis

"Compare microbial diversity across my samples" -> Summarize within-sample richness/evenness (alpha) and between-sample dissimilarity (beta) - but only after declaring the rarefaction depth, the tree, and the metric, because each is a knob that sets the answer before it is seen.

  • CLI: qiime diversity core-metrics-phylogenetic --i-phylogeny rooted-tree.qza --i-table table.qza --p-sampling-depth N --m-metadata-file md.tsv --output-dir cm/
  • R: estimate_richness(ps_rare) for alpha; UniFrac(ps_rare, weighted=) / vegdist() then adonis2() + betadisper() for beta

Scope: whole-community summary (a number or ordination per sample) of an amplicon ASV/OTU table plus a tree. Per-taxon between-group testing -> differential-abundance. Shotgun profiler tables (MetaPhlAn/Bracken) -> metagenomics/metagenome-visualization. The shared CoDA and rarefaction-debate theory lives in metagenomics/abundance-estimation; the Hill-number and PERMANOVA-dispersion theory in metagenomics/metagenome-visualization - cross-referenced here, not re-derived. Tree handling -> phylogenetics/tree-io.

The Single Most Important Modern Insight -- A Diversity Number Is the Output of Three Knobs Turned Before the Answer Appears

An alpha or beta diversity value is not a measurement of the community; it is the output of three choices made before the number appears - the rarefaction DEPTH, the TREE, and the METRIC. Turn them differently and the conclusion can change. The job is to declare all three and show the result survives a second reasonable choice, not to run core-metrics-phylogenetic and read the p-value. The quietest and most dangerous knob is the depth:

  1. --p-sampling-depth is a sample-deletion knob in a normalization costume. core-metrics rarefies every sample to the depth and, per the QIIME2 docs, silently drops every sample whose total count is below it - no warning, just fewer points in the PCoA. The dropped samples are the lowest-yield ones (the lowest-biomass swab, the sickest patient, the failed extraction), so the loss is almost never random. Pick the depth from the feature-table summary plus the alpha-rarefaction plateau, report the depth AND the dropped samples, and confirm the conclusion at a nearby depth. Rarefying to min(sample_sums) is the worst of both worlds - one tiny library drags everyone to noise.
  2. UniFrac and Faith PD are only as real as the tree, and the tree is a model not a property of the data. A de novo MAFFT+FastTree tree from ~250 bp reads is poorly resolved and arbitrarily midpoint-rooted (Janssen 2018); SEPP fragment-insertion into a full-length reference, or Greengenes2 placement, gives stable topology and correct associations - and SEPP/GG2 align 16S with shotgun (McDonald 2024). SEPP also drops fragments that fail to insert, a second silent table-shrink.
  3. Rarefy for diversity, never for differential abundance. Rarefaction-to-even-depth is defensible for alpha/beta (Schloss 2024); for DA it discards count information a compositional model needs (McMurdie 2014). Keep the raw counts; rarefy only into the diversity branch; route DA to differential-abundance on the unrarefied table.

Tool / Metric Taxonomy

Metric / toolCitationWhat it measures / doesWhen
Observed features-ASV richness (Hill q=0); most depth-sensitive; an ASV count, not speciesrichness, but report denoising params; prefer Hill q1/q2
Shannon-entropy = richness+evenness (Hill q=1 = exp(H')); QIIME2 log2/bits, R ln/natsbalanced diversity; report exp(H') to dodge the base
Pielou evennessPielou 1966 J Theor Biol 13:131H'/ln(S); 0-1; isolates evenness from richnesswhen evenness is the question
Faith PDFaith 1992 Biol Conserv 61:1sum of branch lengths spanning observed taxa; phylogenetic q=0amplicon-native richness; needs a tree
Jaccard-presence/absence dissimilarity; no treemembership turnover; depth/rare-ASV sensitive
Bray-Curtis-abundance dissimilarity; no tree; compositionally incoherentabundance default; intuitive, label the caveat
Unweighted UniFracLozupone 2005 Appl Environ Microbiol 71:8228branch length unique to one community (presence/absence)rare/divergent lineages + topology; needs a tree
Weighted UniFracLozupone 2007 Appl Environ Microbiol 73:1576branch length weighted by abundance differenceabundant-lineage shifts; needs a tree
Generalized UniFracChen 2012 Bioinformatics 28:2106alpha in [0,1] interpolating unweighted-weightedalpha=0.5 compromise; powerful for moderately abundant lineages
Aitchison / RPCAMartino 2019 mSystems 4:e00016-19CLR + matrix completion; ordination with feature loadingscompositionally coherent; sparse data; no pseudocount
SEPP insertionJanssen 2018 mSystems 3:e00021-18places ASVs into a full-length reference treethe preferred tree for short reads
Greengenes2McDonald 2024 Nat Biotechnol 42:715unified genome+16S reference treemakes 16S UniFrac comparable to shotgun

Decision Tree by Scenario

ScenarioRecommendedWhy
Need a phylogenetic metric (UniFrac, Faith PD)SEPP-into-reference or Greengenes2 treede novo from short reads is unstable (Janssen 2018)
De novo tree is the only optiontreat unweighted UniFrac with suspiciontopology noise on ~250 bp reads dominates it
Change is in rare/low-abundance lineagesunweighted UniFrac, observed featurespresence/absence + topology see rare taxa
Change is a bloom of dominant taxaweighted UniFrac, Bray-Curtisabundance-weighted metrics see dominant shifts
Do not want to metric-shopgeneralized UniFrac alpha=0.5 + report both un/weightedChen 2012 compromise; single-metric hit is tentative
Richness vs evenness questionobserved/Faith (q0) AND Shannon-exp (q1) / InvSimpson (q2)span the richness-evenness spectrum
Compositional, want axis-driving taxaRPCA (DEICODE/gemelli)CLR ordination with interpretable loadings
Picking a rarefaction depthfeature-table summarize + alpha-rarefaction plateaudepth must retain samples AND saturate richness
Per-taxon "which bug changed"-> differential-abundancediversity is whole-community; DA is per-feature
Shotgun profiler table, not amplicon-> metagenomics/metagenome-visualizationno per-feature tree; different idiom

Choosing the Sampling Depth (the biggest lever)

Goal: Pick a rarefaction depth that saturates richness while retaining an acceptable fraction of samples, and know exactly which samples were dropped.

Approach: Read the per-sample frequency distribution from the feature-table summary, find where the alpha-rarefaction curve plateaus, set the depth there, then declare the depth and the dropped-sample list.

bash
qiime feature-table summarize --i-table table.qza --o-visualization table.qzv   # per-sample frequencies; the depth lives here

qiime diversity alpha-rarefaction \
    --i-table table.qza --i-phylogeny rooted-tree.qza \
    --p-max-depth 20000 \   # set near the median sample depth; the curve panel shows survivors per depth
    --m-metadata-file metadata.tsv --o-visualization alpha-rarefaction.qzv

qiime diversity core-metrics-phylogenetic \
    --i-phylogeny rooted-tree.qza --i-table table.qza \
    --p-sampling-depth 10000 \   # on the observed-features plateau; SILENTLY DROPS samples below this
    --m-metadata-file metadata.tsv --output-dir core-metrics-results

core-metrics-phylogenetic rarefies the table, computes the four alpha vectors (faith_pd_vector, observed_features_vector, shannon_vector, evenness_vector) and four beta matrices (unweighted_unifrac_, weighted_unifrac_, jaccard_, bray_curtis_distance_matrix), and produces a PCoA + Emperor plot for each beta metric. The non-phylogenetic twin qiime diversity core-metrics drops Faith PD and both UniFracs and needs no tree.

Building the Tree (a modeling choice, not a fixed step)

Goal: Obtain a phylogeny over the ASVs that does not inject topology noise into UniFrac/Faith PD.

Approach: Prefer SEPP fragment-insertion into a full-length reference (or Greengenes2 placement) over a de novo build from short reads; for de novo, mask the alignment and accept that unweighted UniFrac will be shaky.

bash
qiime fragment-insertion sepp \
    --i-representative-sequences rep-seqs.qza \
    --i-reference-database sepp-refs-gg-13-8.qza \
    --p-threads 4 \
    --o-tree insertion-tree.qza --o-placements insertion-placements.qza

qiime fragment-insertion filter-features \
    --i-table table.qza --i-tree insertion-tree.qza \
    --o-filtered-table table-sepp.qza --o-removed-table removed-table.qza   # fragments that failed to insert are DROPPED

De novo is qiime phylogeny align-to-tree-mafft-fasttree (MAFFT align -> mask -> FastTree2 -> midpoint root) - acceptable only when no reference package fits the marker/region, and unweighted UniFrac on it must be treated as suspect.

Alpha Diversity in R (counts on the rarefied table)

Goal: Compute richness and evenness per sample and test for a group difference without confounding by sequencing depth.

Approach: Rarefy to a chosen depth, estimate Hill-spanning metrics, test with a non-parametric test (escalate to a linear/mixed model for covariates), and report effective species exp(H').

r
library(phyloseq); library(vegan)

ps_rare <- rarefy_even_depth(ps, sample.size = chosen_depth, rngseed = 42, replace = FALSE)
alpha <- estimate_richness(ps_rare, measures = c('Observed', 'Shannon', 'InvSimpson'))   # q0, exp gives q1, q2
alpha$Group <- sample_data(ps_rare)$Group
alpha$Shannon_eff <- exp(alpha$Shannon)   # effective species; base-invariant in interpretation (Hill q=1)

kruskal.test(Shannon ~ Group, data = alpha)   # non-parametric; escalate to lme4/nlme for covariates or repeated measures

Faith PD in R uses picante::pd(otu_matrix, tree, include.root = TRUE). The Shannon from estimate_richness is in natural log (nats); QIIME2 reports log2 (bits) - report exp(Shannon) to compare across the two.

Beta Diversity in R (report weighted AND unweighted)

Goal: Quantify between-sample dissimilarity with phylogenetic and abundance-weighted views, then test the group effect while ruling out a dispersion artifact.

Approach: Compute both UniFrac variants (and generalized UniFrac alpha=0.5), ordinate by PCoA, run adonis2 for location, and ALWAYS pair it with betadisper for spread.

r
wu  <- UniFrac(ps_rare, weighted = TRUE)    # abundant-lineage view
uwu <- UniFrac(ps_rare, weighted = FALSE)   # rare-lineage + topology view
# generalized UniFrac alpha=0.5 (Chen 2012 compromise):
gu  <- as.dist(GUniFrac::GUniFrac(t(as(otu_table(ps_rare), 'matrix')), phy_tree(ps_rare), alpha = 0.5)$unifracs[, , 'd_0.5'])

meta <- data.frame(sample_data(ps_rare))
adonis2(wu ~ Group, data = meta, permutations = 999)   # >=999 permutations; significance = LOCATION
permutest(betadisper(wu, meta$Group))                  # MANDATORY: is it dispersion, not location?

If betadisper is significant the adonis2 result is ambiguous (location vs spread) - state it. The PERMANOVA-dispersion theory is shared; see metagenomics/metagenome-visualization. For a compositionally coherent ordination with feature loadings use RPCA (DEICODE qiime deicode rpca / gemelli). The Python engine is scikit-bio (skbio.diversity.beta_diversity, skbio.stats.ordination.pcoa, skbio.stats.distance.permanova).

Per-Method Failure Modes

Show full SKILL.md (925 more words)Show less
Sampling-depth sample-massacre

Trigger: a --p-sampling-depth higher than some samples' totals. Mechanism: core-metrics drops every sample below the depth with no warning. Symptom: fewer points in the PCoA than samples in the metadata; the lost ones skew low-biomass. Fix: pick the depth from the rarefaction plateau, report the dropped-sample list, confirm at a nearby depth.

De novo tree noise

Trigger: UniFrac/Faith PD on a MAFFT+FastTree tree from short reads. Mechanism: ~250 bp reads give an unstable topology and arbitrary midpoint root. Symptom: unweighted-UniFrac separation that vanishes under SEPP insertion or weighted UniFrac. Fix: use SEPP-into-reference or Greengenes2; treat de novo unweighted UniFrac as suspect.

Unweighted-vs-weighted flip

Trigger: reporting only the UniFrac variant that gives p<0.05. Mechanism: unweighted listens to rare/short branches, weighted to abundant lineages. Symptom: the two disagree and the chosen one is the significant one. Fix: report both plus generalized alpha=0.5; state which lineage axis each implicates.

Rarefy-then-reuse-for-DA

Trigger: feeding the rarefied table to a differential-abundance tool. Mechanism: rarefaction discards count information the DA model needs. Symptom: underpowered or distorted DA. Fix: keep raw counts; rarefy only into the diversity branch; route DA to differential-abundance.

Observed-features-as-species

Trigger: comparing raw ASV counts across runs/studies as "richness". Mechanism: ASV count tracks DADA2 truncation/maxEE/pooling and intragenomic 16S copy variants, not just biology. Symptom: richness shifts with denoising settings. Fix: prefer Hill q1/q2; report observed features with the denoising parameters stated.

Shannon base mismatch

Trigger: comparing a QIIME2 Shannon to an R Shannon. Mechanism: QIIME2 uses log2 (bits), R diversity/estimate_richness natural log (nats). Symptom: numbers differ by a constant factor and look like a real effect. Fix: state the base, convert, or report exp(H').

PERMANOVA dispersion

Trigger: a significant adonis2 read as a composition shift. Mechanism: pseudo-F responds to within-group spread, not only centroid location (shared theory; metagenomics/metagenome-visualization). Symptom: significant adonis2 with significant betadisper. Fix: always run betadisper/permutest alongside; report both.

Quantitative Thresholds

ThresholdSourceRationale
Sampling depth on the observed-features plateauJanssen 2018; QIIME2 docsdepth must saturate richness while retaining samples; report dropped list
Do NOT use min(sample_sums) as the depthMcMurdie 2014one tiny library drags every sample to under-saturated noise
Generalized UniFrac alpha = 0.5Chen 2012 Bioinformatics 28:2106most powerful for moderately abundant lineages; beats running un/weighted jointly
Report Hill q = 0, 1, 2 together(shared; metagenomics/metagenome-visualization)spans richness (q0) -> evenness-weighted (q2)
PERMANOVA permutations >= 999vegan docsresolution floor for p ~ 0.001; use 9999 for publication
Pair adonis2 with betadisperAnderson & Walsh 2013 (shared)distinguishes a location shift from a dispersion difference
Rarefy for diversity, not for DAMcMurdie 2014; Schloss 2024per-analysis decision, not a global switch

Common Errors

Error / symptomCauseSolution
PCoA has fewer points than samples--p-sampling-depth dropped low-count sampleslower the depth or report the loss; never assume zero drops
UniFrac errors / Faith PD missingno phy_tree slot in the phyloseq objectattach a SEPP/GG2 (preferred) or de novo tree
Unweighted UniFrac significant, weighted notchange is in rare lineages, or de novo tree noisereport both; verify the tree; treat single-metric hit as tentative
R and QIIME2 Shannon disagreelog base differs (nats vs bits)report exp(H'); convert by log2(e)
adonis2 p<0.001 but groups visually overlapdispersion difference, not locationrun betadisper; report it
scikit-bio otu_ids= deprecation warning0.6 renamed OTU to taxon; otu_ids= kept as a deprecated aliasget_beta_diversity_metrics() and help() to find current kwargs
Diversity tracks host/plant contenthost mitochondria/chloroplast 16S not removedfilter Mitochondria/Chloroplast features (see taxonomy-assignment) before computing diversity
"Community" in a near-sterile/low-biomass samplereagent kitome not removedsequence controls + run decontam upstream (amplicon-processing; metagenomics/contamination-controls)

References

  • Faith DP. 1992. Conservation evaluation and phylogenetic diversity. Biol Conserv 61:1-10.
  • Pielou EC. 1966. The measurement of diversity in different types of biological collections. J Theor Biol 13:131-144.
  • Lozupone C, Knight R. 2005. UniFrac: a new phylogenetic method for comparing microbial communities. Appl Environ Microbiol 71:8228-8235.
  • Lozupone CA, Hamady M, Kelley ST, Knight R. 2007. Quantitative and qualitative beta diversity measures lead to different insights into factors that structure microbial communities. Appl Environ Microbiol 73:1576-1585.
  • Chen J, Bittinger K, Charlson ES, Hoffmann C, Lewis J, Wu GD, Collman RG, Bushman FD, Li H. 2012. Associating microbiome composition with environmental covariates using generalized UniFrac distances. Bioinformatics 28:2106-2113.
  • Janssen S, McDonald D, Gonzalez A, et al. 2018. Phylogenetic placement of exact amplicon sequences improves associations with clinical information. mSystems 3:e00021-18.
  • Mirarab S, Nguyen N, Warnow T. 2012. SEPP: SATe-enabled phylogenetic placement. Pac Symp Biocomput 2012:247-258.
  • McDonald D, Jiang Y, Balaban M, et al. 2024. Greengenes2 unifies microbial data in a single reference tree. Nat Biotechnol 42:715-718.
  • Martino C, Morton JT, Marotz CA, Thompson LR, Tripathi A, Knight R, Zengler K. 2019. A novel sparse compositional technique reveals microbial perturbations. mSystems 4:e00016-19.
  • McDonald D, Vazquez-Baeza Y, Koslicki D, et al. 2018. Striped UniFrac: enabling microbiome analysis at unprecedented scale. Nat Methods 15:847-848.
  • McMurdie PJ, Holmes S. 2014. Waste not, want not: why rarefying microbiome data is inadmissible. PLoS Comput Biol 10:e1003531.
  • Schloss PD. 2024. Rarefaction is currently the best approach to control for uneven sequencing effort in amplicon sequence analyses. mSphere 9:e00354-23.
  • McMurdie PJ, Holmes S. 2013. phyloseq: an R package for reproducible interactive analysis and graphics of microbiome census data. PLoS One 8:e61217.
  • amplicon-processing - Generate the ASV table and representative sequences upstream
  • taxonomy-assignment - Label the ASVs summarized here
  • differential-abundance - Per-taxon between-group testing on the unrarefied counts
  • qiime2-workflow - The QIIME2 CLI home for core-metrics and tree building
  • phylogenetics/tree-io - Read, write, and root the UniFrac/Faith PD tree
  • metagenomics/abundance-estimation - Shared CoDA and rarefaction-debate theory
  • metagenomics/metagenome-visualization - Shared Hill-number and PERMANOVA-dispersion theory; diversity/ordination on shotgun profiler tables
  • data-visualization/ggplot2-fundamentals - Custom ordination and diversity plots

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in microbiome/diversity-analysis of GPTomics/bioSkills.

  • SKILL.md
  • examples/core_metrics_qiime2.sh
  • examples/diversity_analysis.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Microbiome Diversity Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Microbiome Diversity Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Microbiome Diversity Analysis this skillGPTomics/bioSkills1.2k1 repos~5.3kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Microbiome Diversity Analysis

What does Bio Microbiome Diversity Analysis do?

Alpha and beta diversity of an amplicon (16S/ITS) ASV/OTU community table - observed features, Shannon, Pielou evenness, Faith PD, Bray-Curtis, Jaccard, weighted/unweighted/generalized UniFrac…. Bio Microbiome Diversity Analysis is an agent skill from GPTomics/bioSkills. Alpha and beta diversity of an amplicon (16S/ITS) ASV/OTU community table - observed features, Shannon, Pielou evenness, Faith PD, Bray-Curtis, Jaccard, weighted/unweighted/generalized UniFrac, Aitchison/RPCA - via QIIME2 core-metrics-phylogenetic, phyloseq/vegan, and scikit-bio.

When should I use Bio Microbiome Diversity Analysis?

Bio Microbiome Diversity Analysis fits situations like: summarizing whole-community richness/evenness; testing group differences in community structure.

How do I install Bio Microbiome Diversity Analysis in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-microbiome-diversity-analysis -a claude-code`. Or copy the skill folder (microbiome/diversity-analysis in GPTomics/bioSkills) into .claude/skills/bio-microbiome-diversity-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Bio Microbiome Diversity Analysis in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-microbiome-diversity-analysis -a codex`. Or copy the skill folder (microbiome/diversity-analysis in GPTomics/bioSkills) into .agents/skills/bio-microbiome-diversity-analysis in your project. Codex loads it when a task matches its description.

Can I use Bio Microbiome Diversity Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-microbiome-diversity-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-microbiome-diversity-analysis, .gemini/skills/bio-microbiome-diversity-analysis, .github/skills/bio-microbiome-diversity-analysis and .opencode/skills/bio-microbiome-diversity-analysis in your project.

What does Bio Microbiome Diversity Analysis need to run?

Going by SKILL.md and its folder, Bio Microbiome Diversity Analysis needs a shell and R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.

Does Bio Microbiome Diversity Analysis access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Microbiome Diversity Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Microbiome Diversity Analysis use?

Bio Microbiome Diversity Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Microbiome Diversity Analysis use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Microbiome Diversity Analysis?

Skills that share tags, products or a category with Bio Microbiome Diversity Analysis: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Microbiome Diversity Analysis?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.