Agent skill

Bio Ecological Genomics Biodiversity Metrics

by GPTomics in GPTomics/bioSkills

Quantifies biodiversity from species abundance/incidence tables using Hill numbers (iNEXT) with coverage-based rarefaction-extrapolation (Chao & Jost 2012), asymptotic richness via…

MITAuto-check passedResearch & Science

Install Bio Ecological Genomics Biodiversity Metrics

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-ecological-genomics-biodiversity-metrics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-ecological-genomics-biodiversity-metrics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ecological-genomics/biodiversity-metrics .claude/skills/bio-ecological-genomics-biodiversity-metrics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-ecological-genomics-biodiversity-metrics
GitHub stars
1.2k
Used in
1 other repo
Token cost
~6.8k tokens
SKILL.md length
2,425 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Quantifies biodiversity from species abundance/incidence tables using Hill numbers (iNEXT) with coverage-based rarefaction-extrapolation (Chao & Jost 2012), asymptotic richness via…

  • Comparing diversity across sites with unequal sampling effort
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithmic Taxonomy and Decision Tree by Scenario, plus 12 more sections
  • Runs R scripts from its folder
  • Picking the right richness estimator for singleton-heavy amplicon data

What it does

Bio Ecological Genomics Biodiversity Metrics is an agent skill from GPTomics/bioSkills. Quantifies biodiversity from species abundance/incidence tables using Hill numbers (iNEXT) with coverage-based rarefaction-extrapolation (Chao & Jost 2012), asymptotic richness via Chao1/ACE/jackknife as a lower bound, Baselga turnover/nestedness partition with the Podani alternative as sensitivity check, mandatory Hellinger transformation before ordination (Legendre & Gallagher 2001), Faith PD and SESMPD/SESMNTD with explicit null-model choice, and Maire 2015 functional-diversity dimensionality optimization. Use…

Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Comparing diversity across sites with unequal sampling effort
  • Picking the right richness estimator for singleton-heavy amplicon data
  • Partitioning beta diversity into turnover vs nestedness
  • Reporting Hill-number effective species counts rather than raw entropies

Example prompts

  • “Use the bio-ecological-genomics-biodiversity-metrics skill to quantify biodiversity from species abundance/incidence tables using Hill numbers…”
  • “/bio-ecological-genomics-biodiversity-metrics”

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Ecological Genomics Biodiversity Metrics loads about 6.8k tokens when it runs. Until then it costs about 263 tokens; SKILL.md has 2,425 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~263
When it runs · the whole SKILL.md, loaded when a task matches
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,425 words, ~6,761 tokens.

Download SKILL.mdSave it as .claude/skills/bio-ecological-genomics-biodiversity-metrics/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-ecological-genomics-biodiversity-metrics
description
Quantifies biodiversity from species abundance/incidence tables using Hill numbers (iNEXT) with coverage-based rarefaction-extrapolation (Chao & Jost 2012), asymptotic richness via Chao1/ACE/jackknife as a lower bound, Baselga turnover/nestedness partition with the Podani alternative as sensitivity check, mandatory Hellinger transformation before ordination (Legendre & Gallagher 2001), Faith PD and SES_MPD/SES_MNTD with explicit null-model choice, and Maire 2015 functional-diversity dimensionality optimization. Use when comparing diversity across sites with unequal sampling effort, picking the right richness estimator for singleton-heavy amplicon data, partitioning beta diversity into turnover vs nestedness, reporting Hill-number effective species counts rather than raw entropies, computing SES_MPD with explicit null-model justification, or deciding whether to apply standard metrics to compositional amplicon data. Not for clinical 16S microbiome diversity (see microbiome/diversity-analysis).
tool_type
r
primary_tool
iNEXT

Version Compatibility

Reference examples tested with: iNEXT 3.0+, iNEXT.3D 1.0+, vegan 2.6+, betapart 1.6+, picante 1.8+, ggplot2 3.5+

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Biodiversity Metrics

"Calculate species diversity for my ecological samples" -> Compute Hill-number diversity (numbers-equivalent of richness, Shannon, Simpson) with coverage-based rarefaction/extrapolation, choose the richness estimator appropriate to the singleton/doubleton signature of the data, and partition beta diversity into turnover and nestedness components with a documented partition framework.

  • R: iNEXT::iNEXT() for coverage-based rarefaction/extrapolation
  • R: betapart::beta.multi() for Baselga turnover/nestedness partition
  • R: picante::ses.mpd() for phylogenetic-community SES with explicit null model

The Single Most Important Modern Insight -- Standardize by COVERAGE not by sample size

Chao & Jost 2012 Ecology 93(12):2533-2547 established that comparing diversity across assemblages by rarefying to a common sample size systematically biases comparisons whenever assemblages differ in underlying diversity: a 100-read rarefaction of a 50-species community is essentially saturated (coverage approximately 99%) while the same 100 reads from a 500-species community covers only approximately 40% of the underlying diversity. The two rarefied diversities are not measuring the same thing. Coverage-based rarefaction with iNEXT is the postdoc-grade default; sample-size rarefaction is now considered a methodological anti-pattern for cross-site comparison.

A second insight pairs with this: raw Shannon and Simpson indices are entropies, NOT diversities. Jost 2006 Oikos 113(2):363-375 showed that only their numbers-equivalents (exp(H), 1/D) are comparable as effective species counts and have the intuitive "doubling property" (merging two equally diverse equally abundant assemblages doubles the diversity). Reporting raw Shannon = 3.2 vs 3.0 hides whether the difference is large or trivial; reporting 1D = 24.5 vs 20.1 species-equivalents makes the 22% gap visible.

Algorithmic Taxonomy

MethodEstimandStrengthFails when
Hill numbers q=0,1,2 (iNEXT)Effective species count at order qUnifies richness, Shannon, Simpson; doubling propertyNone; report all three q values
Chao1 = S + f1^2/(2*f2)Lower bound on asymptotic richnessNon-parametric; works at low coverageSingletons dominated by PCR error (amplicon data); f2 near zero
Chao1bc = S + f1(f1-1)/(2(f2+1))Bias-corrected Chao1Stable when f2 = 0Same singleton-bias issue
ACEAsymptotic richness using all rare classes (f1...f10)Uses more rare-class information than Chao1Choice of "rare" cutoff (default 10) is arbitrary
Jackknife1 = S + f1*(n-1)/nAsymptotic richness via resamplingRobust when f2 = 0Sensitive to singleton count alone
Coverage-based rarefaction (iNEXT)Diversity at standardized completenessCorrect comparison across sites with unequal effortExtrapolation beyond 2x reference size is unreliable
Sample-size rarefactionDiversity at common nLegacy familiarSystematically biased when communities differ in true diversity
Faith's PDSum of branch lengths on spanning treeCaptures evolutionary distinctnessSensitive to richness; report alongside SES_PD
Rao's QPairwise functional/phylogenetic distance * abundanceUnifies taxonomic and functional/phylogenetic diversityTrait dimensionality artifacts (see Maire 2015)
FRic / FEve / FDivFunctional richness/evenness/divergenceMultidimensional trait coverageFRic inflates with collinear traits; optimize axis count
Baselga beta partitionSorensen = turnover (Simpson) + nestednessDecomposes beta into ecological processesNot unique; Podani partition gives different components
Podani / Carvalho partitionAlternative turnover + richness-differenceConceptually distinct from BaselgaSame data, different conclusion possible; report both
Hellinger transform + PCA/RDASolves double-zero problemStandard for community ordinationNone; required preprocessing

Decision Tree by Scenario

ScenarioRecommended approachWhy
Cross-site diversity comparison with unequal sampling effortCoverage-based rarefaction in iNEXT to common C (typically 0.95)Sample-size rarefaction is biased when assemblages differ in true diversity
Reporting diversity for publicationHill numbers q=0,1,2 (numbers-equivalents)Comparable, additive, doubling property; raw entropies are not
Amplicon/eDNA data with many singletonsSkip Chao1 OR use after careful denoising; report Good's coverage alongsideSingletons in amplicon data are dominated by PCR error, not undersampling
Singleton-heavy real data (f1 >> f2)Chao1 with wide CI; cross-check with ACEChao1 variance grows as f1^4; ACE uses more rare-class information
f2 = 0 (no doubletons)Chao1bc or jackknife1Original Chao1 is undefined; bias-corrected form is the standard
Diversity at much larger sample size than observedStop extrapolation at 2x reference size (the doubling rule)iNEXT will silently extrapolate further; variance grows superlinearly beyond
Beta diversity for two assemblagesSorensen or Jaccard with Baselga partition; document choiceBray-Curtis is not a true metric and breaks some downstream methods
Beta diversity reported as turnover/nestednessRun BOTH Baselga AND Podani partitions and report bothThe partition is not mathematically unique; one alone hides ambiguity
Before PCA/RDA on community dataHellinger transformation firstSolves double-zero problem; raw Euclidean PCA on counts is malpractice
Phylogenetic community structure (NRI/NTI)SES_MPD with explicit null.model justificationDefault null may not match question; document choice
Functional diversity (FRic)Optimize PCoA axis count via Maire 2015 mAD; report chosen kFRic biased by trait-axis count; defaults (k=2-3) are too few
Amplicon/compositional dataEither presence/absence metrics (Chao2, Sorensen) OR explicit CLR-based diversity per Gloor 2017 Front Microbiol 8:2224Hill-numbers assume absolute abundances; amplicon counts are compositional

Hill Numbers and Why Raw Entropies Are Not Diversities

Goal: Report diversity values that are interpretable as effective species counts and additive under standard partitions.

Approach: Compute Hill numbers qD at q = 0, 1, 2 (Jost 2006 Oikos 113:363-375; iNEXT software paper Hsieh, Ma, Chao 2016 Methods Ecol Evol 7:1451-1456). q controls sensitivity to rare vs common species in a continuous parametric family: q=0 weights all species equally (richness); q=1 is the geometric-mean weighting (Shannon-equivalent exp(H)); q=2 weights down rare species rapidly (Simpson-equivalent 1/D). qD has the doubling property — merging two equally diverse equally abundant assemblages exactly doubles qD. Raw Shannon and Simpson values do NOT.

r
library(iNEXT)
library(vegan)

abundance_data <- list(
    site_A = c(100, 45, 23, 12, 8, 5, 3, 2, 1, 1),
    site_B = c(80, 60, 40, 30, 20, 15, 10, 8, 5, 3, 2, 1),
    site_C = c(200, 10, 5, 2, 1, 1, 1)
)

# Hill numbers q=0,1,2 with coverage-based rarefaction/extrapolation
# nboot=200 is the publication-quality floor; default nboot=50 is too few for CIs
result <- iNEXT(abundance_data, q = c(0, 1, 2), datatype = 'abundance', nboot = 200)

# Report effective species counts at standardized coverage (postdoc-grade default)
# coverage=0.95: 95% of individuals in the underlying community belong to detected species
est <- estimateD(abundance_data, datatype = 'abundance', base = 'coverage', level = 0.95)
est

# For vegan equivalence: numbers-equivalent of Shannon
shannon_eff <- exp(diversity(community_matrix, index = 'shannon'))
# Inverse Simpson is already in numbers-equivalent units
invsimp <- diversity(community_matrix, index = 'invsimpson')

Asymptotic Richness — Chao1 is a Lower Bound, NOT a Point Estimate

Goal: Estimate the lower bound on true richness when sampling is incomplete and report it correctly.

Approach: Compute Chao1 from the singleton/doubleton ratio (Chao 1984), check that singletons are real biology rather than PCR/sequencing error, and report as a LOWER BOUND with Good's coverage to indicate sampling adequacy. Switch to ACE or jackknife1 when f2 is near zero or singletons are unreliable.

The original Chao1 derivation (Chao 1984) under a Gamma-Poisson mixture model produces a NON-PARAMETRIC LOWER BOUND on richness, not a point estimate. The bound is tight only under specific homogeneity assumptions. Reporting "Chao1 estimated richness = 320" without "at least" is widespread in published literature but statistically wrong.

r
library(iNEXT)

# AsyEst returns Chao1 (q=0), Chao-Shannon (q=1), Chao-Simpson (q=2) with CIs
asymp <- iNEXT(abundance_data, q = c(0, 1, 2), datatype = 'abundance')$AsyEst

# CRITICAL: also report Good's coverage to indicate whether the bound is informative
# Coverage < 0.85 means heavily under-sampled; Chao1 CI will be huge and bound is uninformative
coverage <- estimateD(abundance_data, datatype = 'abundance',
                      base = 'coverage', level = 0.95)
# Inspect estimateD output's Coverage column at the OBSERVED sample size

# Singleton check: if f1 >> f2 and singletons are likely PCR artifacts (amplicon data),
# do NOT report Chao1 — it will measure "how much PCR error" not "how much undiscovered diversity"
f1_check <- sapply(abundance_data, function(x) sum(x == 1))
f2_check <- sapply(abundance_data, function(x) sum(x == 2))
cat('Singletons f1:', f1_check, '\n')
cat('Doubletons f2:', f2_check, '\n')
cat('f1^2/(2*f2) Chao1 bound term:', f1_check^2 / (2 * pmax(f2_check, 0.5)), '\n')

Coverage-Based Rarefaction with the Doubling Rule

Goal: Compare diversity across sites at a common sampling completeness, with extrapolation bounded by statistical reliability.

Approach: Use iNEXT's coverage-based rarefaction-extrapolation interpolating to the minimum coverage across sites; bound extrapolation at 2x the reference sample size (Chao et al. 2014 Ecol Monogr 84:45-67 derive variance bounds that grow superlinearly beyond 2x).

r
# Type 3 (diversity vs coverage) is the cross-site comparison plot
# nboot=200 minimum for publication-quality CIs (default 50 is too few)
ggiNEXT(result, type = 3) + theme_bw() +
    labs(title = 'Coverage-Based Hill-Number Diversity')

# The doubling rule: endpoint should be at most 2 * max(observed sample size)
# iNEXT default endpoint = 2 * max(sample size); do NOT override above this
# Beyond 2x, the variance estimate grows superlinearly and CIs become unreliable

# To extract numerical results at standardized coverage:
est_95 <- estimateD(abundance_data, datatype = 'abundance',
                    base = 'coverage', level = 0.95)

The Double-Zero Problem and Hellinger Transformation

Goal: Compute community dissimilarity in a way that does not treat shared absences as evidence of similarity.

Approach: Apply the Hellinger transformation (Legendre & Gallagher 2001 Oecologia 129:271-280) before computing Euclidean distance, PCA, or RDA. This is the single most important preprocessing decision for community ordination.

Standard Euclidean distance and Pearson correlation treat "both samples have 0 abundance of species X" as evidence of similarity. In community ecology this is wrong — two deserts both lacking a rainforest species are not thereby similar. Hellinger transformation removes this artifact while preserving total-abundance information.

r
library(vegan)

# Hellinger transformation: y_ij' = sqrt(y_ij / row_sum_i)
# Euclidean distance on Hellinger-transformed data = Hellinger distance
species_hell <- decostand(community_matrix, method = 'hellinger')

# Now PCA/RDA on the transformed matrix is biologically meaningful
pca_result <- rda(species_hell)

# Alternative: chord transformation (similar effect, slightly different scaling)
species_chord <- decostand(community_matrix, method = 'normalize')

Beta Diversity Partition — Run BOTH Baselga AND Podani

For the broader "multiple meanings of beta diversity" framework, see Anderson et al. 2011 Ecol Lett 14:19-28.

Goal: Decompose total beta diversity into ecologically interpretable components, acknowledging that the partition is not mathematically unique.

Approach: Compute the Baselga partition (Sorensen = turnover + nestedness) with betapart, and the alternative Podani/Carvalho partition (richness-difference framework), and report both. The two frameworks give different ecological interpretations of the same data; presenting only one hides ambiguity.

r
library(betapart)

pa_matrix <- ifelse(community_matrix > 0, 1, 0)

# --- Baselga partition (Baselga 2010 Glob Ecol Biogeogr 19:134-143) ---
# beta.SIM (turnover/Simpson) + beta.SNE (nestedness) = beta.SOR (total Sorensen)
pair_sor <- beta.pair(pa_matrix, index.family = 'sorensen')
multi_sor <- beta.multi(pa_matrix, index.family = 'sorensen')

# --- Abundance-based: Bray-Curtis balanced + gradient ---
# beta.bray.bal (balanced variation; analogous to turnover)
# beta.bray.gra (abundance gradient; analogous to nestedness)
pair_abund <- beta.pair.abund(community_matrix, index.family = 'bray')

# --- Podani/Carvalho framework (different decomposition of same data) ---
# Use betapart::beta.pair with the .fam family option, or carvalho package
# These produce richness-difference components instead of nestedness
# Reporting only Baselga without acknowledging Podani exists is incomplete

Phylogenetic Diversity — Faith's PD, MPD, MNTD with Null-Model Choice

Goal: Quantify evolutionary distinctness and phylogenetic community structure with an explicit, justified null model.

Approach: Compute Faith's PD on a community matrix and ultrametric phylogeny (Faith 1992 Biol Conserv 61:1-10), then SES_MPD and SES_MNTD (Webb 2002 Annu Rev Ecol Syst 33:475-505) standardizing against a documented null model. The choice of null model is the dominant scientific decision — different nulls answer different ecological questions.

r
library(picante)

# Faith's PD: sum of branch lengths spanning the focal species
# include.root=TRUE counts root branch; FALSE excludes (matters for within-clade comparisons)
faith_pd <- pd(community_matrix, phylo_tree, include.root = TRUE)

# SES_MPD with EXPLICIT null model (do not accept default silently)
# null.model='taxa.labels': shuffles species across tree, holds sample richness constant
# null.model='richness': randomizes within sample, preserves species occurrence frequencies
# null.model='independentswap': preserves both row and column sums of community matrix
# Each null answers a different question; document the choice in methods
ses_mpd_taxa <- ses.mpd(community_matrix, cophenetic(phylo_tree),
                        null.model = 'taxa.labels', runs = 999, iterations = 1000)
ses_mpd_indep <- ses.mpd(community_matrix, cophenetic(phylo_tree),
                         null.model = 'independentswap', runs = 999, iterations = 1000)

# SES_MPD < 0 (NRI > 0): phylogenetic clustering
# SES_MPD > 0 (NRI < 0): phylogenetic overdispersion
# DO NOT infer "clustering = environmental filtering" from sign alone;
# Mayfield & Levine 2010 Ecol Lett 13:1085-1093 showed competition can cluster
# when traits track phylogeny and similar species coexist via R*-rule dynamics

Functional Diversity with Trait-Axis Dimensionality Optimization

Goal: Compute multidimensional functional diversity without inflating values via correlated trait axes.

Approach: Run Maire et al. 2015 Glob Ecol Biogeogr 24:728-740 trait-space-quality assessment (mAD metric) to choose the number of PCoA axes that minimizes deviation between original trait distances and axis-reconstructed distances. Use that k for FRic, FEve, FDiv. Default k=2-3 typically biases FRic; optimum is usually 4-6 for typical trait datasets.

r
library(FD)
library(mFD)

# Trait dissimilarity from a trait matrix (Gower for mixed quantitative/categorical)
trait_dist <- gowdis(trait_matrix)

# mFD optimizes the number of trait axes via Maire 2015 mAD criterion
# Smaller mAD = better fidelity between trait distances and PCoA-axis distances
quality_funct_space <- quality.fspaces(trait_dist, maxdim_pcoa = 10,
                                       fdist_scaling = TRUE, fdendro = NULL)
best_k <- which.min(quality_funct_space$quality_fspaces$mad)
cat('Maire-optimal number of trait axes:', best_k, '\n')

# Compute FD with the optimal k
fd_result <- dbFD(trait_matrix, community_matrix, m = best_k, corr = 'cailliez')
fd_result$FRic   # functional richness (convex hull volume)
fd_result$FEve   # functional evenness
fd_result$FDiv   # functional divergence

Per-Method Failure Modes

Show full SKILL.md (1,015 more words)Show less
Singleton-driven Chao1 inflation in amplicon data

Trigger: Computing Chao1 directly on ASV/OTU tables that include singletons of suspected PCR-error origin.

Mechanism: Chao1 assumes singletons are biologically real rare species; under that assumption, more singletons relative to doubletons signals more undiscovered diversity. PCR/sequencing error produces many singletons that look like rare species to Chao1, inflating the bound.

Symptom: Chao1 dramatically exceeds observed richness (e.g., Chao1 = 500 from S_obs = 100) with extremely wide confidence intervals; Good's coverage < 0.85 despite > 10,000 reads per sample.

Fix: Denoise first (DADA2 / UNOISE3 / swarm v2) to remove PCR-error variants, then either skip Chao1 entirely (Callahan 2017 ASV philosophy) or report observed ASV count + Good's coverage instead. Alternative: report Chao2 from incidence (presence across replicates), which is robust to PCR-error singletons.

Rarefaction across mismatched assemblage sizes

Trigger: Sample-size rarefaction to the smallest assemblage when sites differ in true diversity.

Mechanism: A 100-read rarefaction of a 50-species community is nearly saturated (coverage approximately 99%); the same 100 reads from a 500-species community covers approximately 40% of true diversity. The two rarefied values measure different completeness levels.

Symptom: Rarefied richness ordering disagrees with intuitive site-by-site comparison; small-sample-size sites appear artificially equal in rarefied diversity to high-diversity sites.

Fix: Use coverage-based rarefaction via estimateD(..., base = 'coverage', level = 0.95). The 0.95 coverage target is the modern default; 0.99 for high-precision work, 0.85 minimum for accepting a site into the comparison.

Hellinger forgotten before PCA/RDA on community data

Trigger: Running rda(species_matrix) or prcomp(species_matrix) on raw species counts without prior transformation.

Mechanism: The double-zero problem inflates dissimilarity for sample pairs sharing many absent species. Raw-count PCA projects samples primarily by sequencing depth (PC1 = library size proxy), not by community composition.

Symptom: PC1 axis correlates strongly with sample read totals; biological gradients appear only on PC3 or later; ordination is dominated by sites with extreme sample sizes.

Fix: decostand(matrix, method = 'hellinger') before ordination. Always.

FRic inflation from collinear trait axes

Trigger: Computing FRic from 8-10 raw correlated traits without dimensionality optimization.

Mechanism: FRic = convex-hull volume in trait PCoA space. Adding correlated traits increases the apparent dimensionality of the trait space without adding ecological information; convex-hull volume inflates accordingly.

Symptom: FRic increases when adding new correlated traits to the analysis; FRic comparisons across studies that use different trait counts are inconsistent.

Fix: Run mFD::quality.fspaces() and use the mAD-optimal k. Report k in methods.

SES_MPD interpretation by sign alone

Trigger: Reporting "SES_MPD < 0 indicates environmental filtering" without testing alternative explanations.

Mechanism: Mayfield & Levine 2010 Ecol Lett 13(9):1085-1093 showed competition can produce phylogenetic clustering when ecologically similar species (which tend to be closely related) coexist via R*-rule trait differences. The clustering = filtering / overdispersion = competition dichotomy from Webb 2002 is incomplete.

Symptom: Phylogenetic clustering interpretation is challenged at review; trait-similarity vs phylogenetic-similarity correlation has not been examined.

Fix: Report SES_MPD with one explicit null model AND test trait conservatism (Blomberg's K, Pagel's lambda) — if traits track phylogeny strongly, clustering may reflect either filtering OR competition; cite Mayfield & Levine 2010 in interpretation.

Quantitative Thresholds

ThresholdValueSource / rationale
Coverage target for cross-site comparisonC = 0.95Postdoc-grade convention; 0.99 for high-precision, 0.85 minimum to accept site
iNEXT extrapolation limitendpoint <= 2x reference sample sizeChao et al. 2014 doubling rule; variance grows superlinearly beyond
iNEXT bootstrap floor for CIsnboot = 200Default 50 is too few for publication-quality CIs
Chao1 reliabilityf2 > 0 AND singletons biologicalIf f2 = 0, switch to Chao1bc or jackknife1
Hellinger before ordinationAlways for community dataLegendre & Gallagher 2001; non-negotiable
SES_MPD significanceSES
Functional-diversity axis count kMaire 2015 mAD-optimalDefault k=2-3 too few; typically 4-6 optimal
Phylogenetic-tree requirementUltrametricIf not, ape::chronos() or BEAST-based dating

Common errors

ErrorCauseSolution
Chao1 reports infinity or NaNf2 = 0 (no doubletons)Use Chao1bc form or jackknife1
iNEXT extrapolation curve flat then explodesExtrapolated beyond doubling-rule limitSet endpoint <= 2 * max(sample sizes)
Hellinger PCA gives flat resultsDecostand not applied or applied to wrong axisdecostand(matrix, method = 'hellinger') rows = sites
ses.mpd returns all p-values approximately 0.5Wrong null model for question (e.g., taxa.labels on a single-richness dataset)Choose null that varies the quantity tested
FRic dramatically increases with new traitsCorrelated trait axes inflating convex hullOptimize axis count with mFD::quality.fspaces
beta.multi returns NaN for nestednessCommunities entirely disjoint (no shared species)beta.SNE -> 0 trivially; interpret as pure turnover
Bray-Curtis flagged for triangle-inequality violationBray-Curtis is not a true metricSwitch to Sorensen (a metric) for downstream methods requiring metricity

References

  • Chao A, Jost L (2012) Coverage-based rarefaction and extrapolation. Ecology 93(12):2533-2547. doi:10.1890/11-1952.1
  • Chao A, Gotelli NJ, Hsieh TC, Sander EL, Ma KH, Colwell RK, Ellison AM (2014) Rarefaction and extrapolation with Hill numbers. Ecol Monogr 84(1):45-67. doi:10.1890/13-0133.1
  • Hsieh TC, Ma KH, Chao A (2016) iNEXT: an R package for rarefaction and extrapolation of species diversity. Methods Ecol Evol 7(12):1451-1456. doi:10.1111/2041-210X.12613
  • Jost L (2006) Entropy and diversity. Oikos 113(2):363-375. doi:10.1111/j.2006.0030-1299.14714.x
  • Faith DP (1992) Conservation evaluation and phylogenetic diversity. Biol Conserv 61(1):1-10. doi:10.1016/0006-3207(92)91201-3
  • Legendre P, Gallagher ED (2001) Ecologically meaningful transformations for ordination. Oecologia 129(2):271-280. doi:10.1007/s004420100716
  • Baselga A (2010) Partitioning the turnover and nestedness components of beta diversity. Glob Ecol Biogeogr 19(1):134-143. doi:10.1111/j.1466-8238.2009.00490.x
  • Anderson MJ, Crist TO, Chase JM et al. (2011) Navigating the multiple meanings of beta diversity. Ecol Lett 14(1):19-28. doi:10.1111/j.1461-0248.2010.01552.x
  • Webb CO, Ackerly DD, McPeek MA, Donoghue MJ (2002) Phylogenies and community ecology. Annu Rev Ecol Syst 33:475-505. doi:10.1146/annurev.ecolsys.33.010802.150448
  • Mayfield MM, Levine JM (2010) Opposing effects of competitive exclusion on phylogenetic community structure. Ecol Lett 13(9):1085-1093. doi:10.1111/j.1461-0248.2010.01509.x
  • Maire E, Grenouillet G, Brosse S, Villeger S (2015) How many dimensions are needed to accurately assess functional diversity? Glob Ecol Biogeogr 24(6):728-740. doi:10.1111/geb.12299
  • Chao A (1984) Nonparametric estimation of the number of classes in a population. Scand J Stat 11(4):265-270
  • Gloor GB, Macklaim JM, Pawlowsky-Glahn V, Egozcue JJ (2017) Microbiome datasets are compositional: and this is not optional. Front Microbiol 8:2224. doi:10.3389/fmicb.2017.02224
  • ecological-genomics/edna-metabarcoding - Generate ASV/species tables prior to diversity analysis
  • ecological-genomics/community-ecology - Constrained ordination, indicator species, PERMANOVA on transformed data
  • microbiome/diversity-analysis - 16S clinical microbiome diversity metrics with compositional considerations
  • data-visualization/ggplot2-fundamentals - Customize diversity plots and rarefaction curves
  • phylogenetics/tree-io - Ultrametric tree preparation for PD/MPD/MNTD

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in ecological-genomics/biodiversity-metrics of GPTomics/bioSkills.

  • SKILL.md
  • examples/beta_partitioning.R
  • examples/inext_diversity.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Ecological Genomics Biodiversity Metrics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Ecological Genomics Biodiversity Metrics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Ecological Genomics Biodiversity Metrics this skillGPTomics/bioSkills1.2k1 repos~6.8kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Ecological Genomics Biodiversity Metrics

What does Bio Ecological Genomics Biodiversity Metrics do?

Quantifies biodiversity from species abundance/incidence tables using Hill numbers (iNEXT) with coverage-based rarefaction-extrapolation (Chao & Jost 2012), asymptotic richness via…. Bio Ecological Genomics Biodiversity Metrics is an agent skill from GPTomics/bioSkills. Quantifies biodiversity from species abundance/incidence tables using Hill numbers (iNEXT) with coverage-based rarefaction-extrapolation (Chao & Jost 2012), asymptotic richness via Chao1/ACE/jackknife as a lower bound, Baselga turnover/nestedness partition with the Podani alternative as sensitivity check, mandatory Hellinger transformation before ordination (Legendre & Gallagher 2001), Faith PD and SESMPD/SESMNTD with explicit null-model choice, and Maire 2015 functional-diversity dimensionality optimization.

When should I use Bio Ecological Genomics Biodiversity Metrics?

Bio Ecological Genomics Biodiversity Metrics fits situations like: comparing diversity across sites with unequal sampling effort; picking the right richness estimator for singleton-heavy amplicon data; partitioning beta diversity into turnover vs nestedness; reporting Hill-number effective species counts rather than raw entropies.

How do I install Bio Ecological Genomics Biodiversity Metrics in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-ecological-genomics-biodiversity-metrics -a claude-code`. Or copy the skill folder (ecological-genomics/biodiversity-metrics in GPTomics/bioSkills) into .claude/skills/bio-ecological-genomics-biodiversity-metrics in your project. Claude Code loads it when a task matches its description.

How do I install Bio Ecological Genomics Biodiversity Metrics in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-ecological-genomics-biodiversity-metrics -a codex`. Or copy the skill folder (ecological-genomics/biodiversity-metrics in GPTomics/bioSkills) into .agents/skills/bio-ecological-genomics-biodiversity-metrics in your project. Codex loads it when a task matches its description.

Can I use Bio Ecological Genomics Biodiversity Metrics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-ecological-genomics-biodiversity-metrics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-ecological-genomics-biodiversity-metrics, .gemini/skills/bio-ecological-genomics-biodiversity-metrics, .github/skills/bio-ecological-genomics-biodiversity-metrics and .opencode/skills/bio-ecological-genomics-biodiversity-metrics in your project.

What does Bio Ecological Genomics Biodiversity Metrics need to run?

Going by SKILL.md and its folder, Bio Ecological Genomics Biodiversity Metrics needs R for the scripts in its folder.

Does Bio Ecological Genomics Biodiversity Metrics access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Ecological Genomics Biodiversity Metrics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Ecological Genomics Biodiversity Metrics use?

Bio Ecological Genomics Biodiversity Metrics is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Ecological Genomics Biodiversity Metrics use?

About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Ecological Genomics Biodiversity Metrics?

Skills that share tags, products or a category with Bio Ecological Genomics Biodiversity Metrics: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Ecological Genomics Biodiversity Metrics?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.