Agent skill

Bio Metabolomics Statistical Analysis

by GPTomics in GPTomics/bioSkills

Decision-grade statistical analysis for metabolomics intensity tables.

MITAuto-check passedData & Analytics

Install Bio Metabolomics Statistical Analysis

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-metabolomics-statistical-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-metabolomics-statistical-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/metabolomics/statistical-analysis .claude/skills/bio-metabolomics-statistical-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-metabolomics-statistical-analysis
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5k tokens
SKILL.md length
2,128 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Decision-grade statistical analysis for metabolomics intensity tables.

  • Works in 8 steps: Report the transformation + scaling used… → Report R2X, R2Y, Q2 and the number of… → Choose the number of components inside… → …
  • Testing which metabolites differ
  • SKILL.md covers Version Compatibility, The Single Most Important…, Transformation and Scaling --… and Decision Tree by Scenario, plus 10 more sections
  • Runs Python and R scripts from its folder; calls pip

What it does

Bio Metabolomics Statistical Analysis is an agent skill from GPTomics/bioSkills. Decision-grade statistical analysis for metabolomics intensity tables. Covers transformation and scaling (Pareto vs unit-variance as a hidden hypothesis), unsupervised structure (PCA/HCA for QC), permutation-validated PLS-DA/OPLS-DA (R2 vs Q2, double CV, VIP as heuristic), univariate testing (Welch/Mann-Whitney/ANOVA/LMM with covariate adjustment), and dependence-aware multiple testing. Use when testing which metabolites differ, building or validating a discriminant model, choosing a scaling, or correcting many…

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/metabolomics_differential.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Statistics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Testing which metabolites differ
  • Validating a discriminant model
  • Choosing a scaling
  • Correcting many correlated tests

Example prompts

  • “/bio-metabolomics-statistical-analysis”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Report the transformation + scaling used (it changes the loadings, VIPs, and story).
  2. Report R2X, R2Y, Q2 and the number of predictive + orthogonal components.
  3. Choose the number of components inside CV, not by eye on the training fit.
  4. Permutation test (>= 1000) of the full pipeline -> permutation p for Q2 (and R2Y). Permute every step that touched the labels.
  5. For honest generalization error use double (cross-model) CV or an untouched test set; single CV that also tuned the model is optimistic.
  6. Independent validation cohort for any biomarker claim.
  7. Read VIP / S-plot only from a validated model; corroborate with univariate FDR + effect size; report ranking stability across resamples.
  8. Put a PCA score plot beside the PLS-DA one -- separation only under supervision is the artifact signature.

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Metabolomics Statistical Analysis loads about 5k tokens when it runs. Until then it costs about 231 tokens; SKILL.md has 2,128 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~231
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,128 words, ~4,997 tokens.

Download SKILL.mdSave it as .claude/skills/bio-metabolomics-statistical-analysis/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-metabolomics-statistical-analysis
description
Decision-grade statistical analysis for metabolomics intensity tables. Covers transformation and scaling (Pareto vs unit-variance as a hidden hypothesis), unsupervised structure (PCA/HCA for QC), permutation-validated PLS-DA/OPLS-DA (R2 vs Q2, double CV, VIP as heuristic), univariate testing (Welch/Mann-Whitney/ANOVA/LMM with covariate adjustment), and dependence-aware multiple testing. Use when testing which metabolites differ, building or validating a discriminant model, choosing a scaling, or correcting many correlated tests. For sample-wise normalization/drift correction see metabolomics/normalization-qc; for ML classifiers and selection-inside-CV leakage see machine-learning/biomarker-discovery and machine-learning/model-validation; for pathway interpretation see metabolomics/pathway-mapping; for design/power/multiplicity regime see experimental-design/multiple-testing.
tool_type
mixed
primary_tool
ropls

Version Compatibility

Reference examples tested with: ropls 1.34+, scipy 1.12+, statsmodels 0.14+, numpy 1.26+, pandas 2.1+, matplotlib 3.8+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Metabolomics Statistical Analysis

"Tell me which metabolites separate my groups" -> run an honest univariate test with dependence-aware FDR AND a permutation-validated multivariate model, then reconcile the two.

  • R: ropls::opls() (PCA/PLS-DA/OPLS-DA + permutation), wilcox.test()/lm()/lme4::lmer(), p.adjust(method='BH')
  • Python: scipy.stats.ttest_ind(equal_var=False)/mannwhitneyu, statsmodels multipletests(method='fdr_bh'), sklearn.cross_decomposition.PLSRegression + permutation_test_score

The Single Most Important Modern Insight -- A Score Plot Is a Hypothesis, Not a Result

In metabolomics the regime is p >> n (hundreds-to-thousands of features, tens of samples) with strongly correlated features. In that regime any binary labelling of n points in >= n-1 dimensions is linearly separable with probability 1, so PLS-DA and OPLS-DA produce a clean two-cluster score plot even for randomly assigned labels. A beautiful score plot is the generic output of the algorithm and carries essentially zero information. Only a cross-validated Q2 benchmarked against a permutation null distinguishes signal from geometry (Westerhuis 2008; Ruiz-Perez 2020). Three corollaries reorganize the whole skill: (1) R2 is no evidence (it can be driven to 1 by adding components); only permutation-validated Q2 licenses a claim. (2) Scaling is a hidden hypothesis -- variance-driven methods weight a feature by the variance it is allowed to contribute, so Pareto vs unit-variance hands back a different VIP list and a different biological story (van den Berg 2006). (3) Features are not independent (pathways, adducts, isotopologues), so naive BH-independence is violated and one real signal lights up its whole correlated cluster.

Transformation and Scaling -- the Decision That Changes Conclusions

Transformation (nonlinear, per value: corrects heteroscedastic multiplicative MS noise and skew) and scaling (linear, per feature: sets relative weight) are distinct. Mean-centering is the universal first step. Choosing not to scale is the strongest prior of all -- it lets the most abundant metabolite drive PC1.

MethodPer feature jEffectUse when
Centeringsubtract meanoffsets removed, variance unchangedalways (prerequisite for all below)
Auto / unit-variance / "standard"center, / SD_jevery metabolite equal weighta priori all metabolites equally important; classic default -- but inflates near-LOD noise
Paretocenter, / sqrt(SD_j)between raw and UVde facto metabolomics/NMR/OPLS-DA default; curbs dominance with less noise inflation than UV
Rangecenter, / (max-min)abundance dependence removedclean data, few outliers (outlier-sensitive)
VastUV x (mean/SD)up-weights low-CV stable featuresfocus on robust/reproducible features with prior class info
Levelcenter, / mean_jrelative (% change)response as relative change (mean noisy at low abundance)
Log / log10log(x)multiplicative -> additiveconcentrations spanning orders of magnitude (undefined at 0)
gloglinear near 0, log for large xvariance-stabilizingdata with zeros / near-LOD values; preferred over plain log (needs transition param)
Power (sqrt, cube-root)x^(1/2)mild stabilizationmild skew with zeros present

van den Berg 2006: on real data autoscale and range recovered biologically meaningful loadings; Pareto is the pragmatic middle. Decision rule: transform first (if heteroscedastic -- usually yes for MS), then center, then scale; run at least Pareto AND UV, and if the top-VIP list or conclusion flips, the result is scaling-fragile and must be tempered.

Decision Tree by Scenario

Goal / situationDoWhy
First look, QC, batch/outlier checkPCA on scaled data; color scores by batch/injection order; Hotelling T2 ellipseUnsupervised -> cannot overfit the grouping; pooled-QC samples must cluster tightly in the center, else analytical variance dominates
Which single metabolites differ (2 groups)Welch t-test (post-transform) or Mann-Whitney; BH FDR; report fold change + CIInterpretable per-feature effect; FDR-controlled; effect size mandatory in p>>n
2 groups, paired/pre-postPaired t-test or Wilcoxon signed-rankDiscards within-subject pairing if analyzed unpaired -> underpowered
>2 groupsOne-way ANOVA (+Tukey) or Kruskal-Wallis (+Dunn)Match normality assumption
Longitudinal / repeated measuresLinear mixed model (random intercept/slope per subject)Handles unbalanced timepoints, missingness, within-subject correlation
Covariate adjustment (age/sex/BMI/batch)Per-feature linear model y ~ group + covarsHuman metabolome is dominated by age/sex/BMI -- unadjusted, they masquerade as case/control signal (Thevenot 2015)
A discriminant / predictive modelPLS-DA (orthoI=0) or OPLS-DA (predI=1, orthoI=NA) + permutation + double CVSupervised; demands full validation (see checklist)
Built-in feature selectionsparse PLS-DA splsda() (mixOmics) with tune.splsdaSelection must be inside CV -> hand off to machine-learning/biomarker-discovery
Rank discriminant featuresVIP from a permutation-validated model only; corroborate with univariate FDRVIP > 1 is a heuristic, not a test (see failure modes)
Confirm a biomarkerIndependent validation cohortInternal CV does not correct overfitting/forking-paths; discovery performance overestimates external

Scaling + PCA

Goal: Get the honest unsupervised first look that cannot chase the labels, with QC as the primary data-quality readout.

Approach: Transform if heteroscedastic, then PCA with an explicit scaling; inspect QC clustering, batch coloring, and Hotelling T2.

r
library(ropls)
# scaleC default is "standard" (unit-variance/autoscale), NOT Pareto -- set explicitly
pca <- opls(t(feature_matrix), scaleC = 'pareto', fig.pdfC = 'none', info.txtC = 'none')
scores <- getScoreMN(pca)               # samples x components
getSummaryDF(pca)                       # R2X(cum) per component
# Tight pooled-QC clustering in the center = trustworthy run; QC scatter = analytical variance dominates

Permutation-Validated PLS-DA / OPLS-DA

Goal: Decide whether group separation is real, not a geometry artifact, before reading any VIP or S-plot.

Approach: Fit with an explicit scaling, raise permI far above the default of 20, and read pQ2/pR2Y -- a model whose true Q2 sits inside the permutation cloud is indistinguishable from chance.

r
library(ropls)
group <- factor(sample_info$group)
# OPLS-DA: 1 predictive + auto orthogonal; permI default 20 is too few for a reliable pQ2 -> >=1000
oplsda <- opls(t(feature_matrix), group, predI = 1, orthoI = NA,
               scaleC = 'pareto', permI = 1000, crossvalI = 7,
               fig.pdfC = 'none', info.txtC = 'none')
summ <- getSummaryDF(oplsda)            # R2X(cum), R2Y(cum), Q2(cum), pre, ort, pR2Y, pQ2
vip_pred <- getVipVn(oplsda)            # predictive VIP (Galindo-Prieto 2014); orthoL=TRUE for orthogonal
# Claim is licensed only if Q2 high AND pQ2 small. R2Y alone proves nothing.

PLS-DA is orthoI = 0. OPLS-DA has identical predictive power to PLS-DA -- it is a coordinate rotation, not a better model; the orthogonal block often encodes a confounder (inspect what correlates with it). DQ2 (Westerhuis 2008b) is the discriminant-appropriate figure of merit when Q2 penalizes correct-side over-predictions.

PLS-DA / OPLS-DA Validation Checklist

  1. Report the transformation + scaling used (it changes the loadings, VIPs, and story).
  2. Report R2X, R2Y, Q2 and the number of predictive + orthogonal components.
  3. Choose the number of components inside CV, not by eye on the training fit.
  4. Permutation test (>= 1000) of the full pipeline -> permutation p for Q2 (and R2Y). Permute every step that touched the labels.
  5. For honest generalization error use double (cross-model) CV or an untouched test set; single CV that also tuned the model is optimistic.
  6. Independent validation cohort for any biomarker claim.
  7. Read VIP / S-plot only from a validated model; corroborate with univariate FDR + effect size; report ranking stability across resamples.
  8. Put a PCA score plot beside the PLS-DA one -- separation only under supervision is the artifact signature.

Univariate Testing + Correct FDR

Goal: Produce an interpretable, FDR-controlled per-metabolite answer with effect sizes.

Approach: Match the test to the design, compute log2 fold change as a difference of group means on transformed data, then apply BH explicitly (defaults are not BH in either language).

python
import numpy as np
import pandas as pd
from scipy.stats import ttest_ind
from statsmodels.stats.multitest import multipletests

logged = np.log2(intensities.replace(0, np.nan))   # transform before testing
pvals, lfc = [], []
for feat in logged.index:
    a = logged.loc[feat, case].dropna().values
    b = logged.loc[feat, ctrl].dropna().values
    if len(a) >= 3 and len(b) >= 3:
        pvals.append(ttest_ind(a, b, equal_var=False)[1])   # Welch: scipy defaults to Student
        lfc.append(a.mean() - b.mean())                     # geometric-mean ratio on log scale
    else:
        pvals.append(np.nan); lfc.append(np.nan)
res = pd.DataFrame({'feature': logged.index, 'log2fc': lfc, 'pval': pvals}).dropna(subset=['pval'])
# statsmodels default is 'hs' (Holm-Sidak); R p.adjust default is 'holm' -- ALWAYS pass BH explicitly
res['padj'] = multipletests(res['pval'], method='fdr_bh')[1]

BH controls FDR under independence and PRDS; positively-correlated metabolomics features roughly satisfy PRDS, so BH is valid but conservative -- but closure-induced negative correlations (after total-area/PQN normalization) fall outside the clean case, where a permutation FDR sidesteps the dependence assumptions. The effective number of independent tests is far below the feature count (one compound = many adducts/isotopologues/fragments); use an effective-number-of-tests correction (Peluso 2021) rather than Bonferroni-on-features, and collapse features to compounds before counting "how many metabolites changed."

Volcano Plot

Goal: Show significance and magnitude together for all features.

Approach: Plot log2 fold change vs -log10(p), with the FDR cutoff annotated (raw p on the axis is fine only if the FDR line is drawn).

python
import matplotlib.pyplot as plt
hit = (res['padj'] < 0.05) & (res['log2fc'].abs() > 1)   # 2-fold + FDR 5%
plt.scatter(res['log2fc'], -np.log10(res['pval']), c=np.where(hit, 'firebrick', 'gray'), s=12, alpha=0.6)
plt.axhline(-np.log10(0.05), ls='--'); plt.axvline(1, ls='--'); plt.axvline(-1, ls='--')
plt.xlabel('log2 fold change'); plt.ylabel('-log10(p)')

Per-Method Failure Modes

Noise separation (the cardinal sin)
  • Trigger: Reporting a PLS-DA/OPLS-DA score plot as evidence of a group difference.
  • Mechanism: In p>>n any labelling is linearly separable; the algorithm always finds a covariance-maximizing direction, even for random labels.
  • Symptom: Clean two-cluster score plot, high R2Y, but Q2 low/negative or inside the permutation cloud; PCA shows no separation.
  • Fix: Permutation test (>=1000) of the full pipeline; require small pQ2; put the PCA plot beside it.
Show full SKILL.md (871 more words)Show less
VIP misuse
  • Trigger: Selecting biomarkers by VIP > 1 from a single model fit.
  • Mechanism: VIPs are normalized so the mean squared VIP = 1 -- roughly half the features exceed 1 by construction; there is no null, no p-value, and the ranking is unstable under resampling in p>>n.
  • Symptom: Top-20 VIP list reshuffles when the model is re-bootstrapped; VIP-only hits fail to replicate.
  • Fix: Use VIP only from a permutation-validated model; require univariate FDR + effect-size concordance and resampling stability; use the OPLS-specific VIP so a high orthogonal-block VIP (the confounder) is not credited to disease.
Naive FDR under correlation
  • Trigger: BH or Bonferroni applied as if the features were independent.
  • Mechanism: Pathway co-regulation plus adducts/isotopologues/fragments make features strongly correlated; one signal lights up its whole cluster, and closure (after sample-wise normalization) injects negative correlations.
  • Symptom: A "200 significant metabolites" list that encodes a handful of independent signals; over-conservative threshold from Bonferroni-on-features.
  • Fix: Effective-number-of-tests or permutation FDR (Peluso 2021); collapse features to compounds before counting hits; report independent-signal counts.
Log with zeros / detection-rate confound
  • Trigger: Half-min (or zero) imputation followed by log, especially when detection rate differs between groups.
  • Mechanism: "Missing" is left-censored (MNAR); a constant imputed at the LOD then logged spikes the censored region, and a detection-rate difference masquerades as a concentration difference.
  • Symptom: Fake bimodality; a low-abundance "hit" that is really a difference in how often the metabolite was detected.
  • Fix: Report per-group detection rates with any low-abundance hit; prefer glog or a left-censored imputer (QRILC/GSimp) over impute-constant-then-log when detection differs (see metabolomics/normalization-qc).

Quantitative Thresholds

ThresholdSourceRationale
Q2 > 0.5 "good"Triba 2015 (heuristic)Predictive ability rule-of-thumb; not a hard cutoff -- many published models report Q2 < 0.5; report the value, not a verdict
permI >= 1000Szymanska 2012Q2/DQ2 null distributions are skewed; the ropls default of 20 estimates only the granularity of the grid, not a usable pQ2
pQ2 < 0.05Westerhuis 2008Fraction of permuted models with Q2 >= true Q2; the actual evidence the separation is real
crossvalI = 7ropls default7-fold CV; for very small n LOO is common but optimistic
VIP > 1Galindo-Prieto 2014Above-average contributor; a ranking heuristic with no error control -- never a standalone selector
BH FDR < 0.05Benjamini-HochbergExpected false-positive proportion among rejections; the metabolomics discovery default
|log2FC| > 1convention2-fold; effect-size gate orthogonal to the p-value, mandatory in p>>n

Common Errors

Error / symptomCauseSolution
Model "significant" yet noisepermI = 20 (ropls default)Set permI >= 1000; read pQ2/pR2Y from getSummaryDF
Wrong scaling shipped silentlyscaleC default is "standard" (UV), not ParetoSet scaleC = 'pareto' (or the intended scaling) explicitly; report it
PLS-DA vs OPLS-DA "function not found"type is set by orthoI, not a separate functionorthoI = 0 -> PLS; orthoI = NA -> OPLS; predI = 1 for 2-class
FDR is actually HolmR p.adjust default is 'holm' (FWER)Pass method = 'BH'
FDR is actually Holm-Sidakstatsmodels multipletests default is 'hs'Pass method = 'fdr_bh'
Student instead of Welchscipy ttest_ind default equal_var=TrueSet equal_var=False (group variances differ, esp. near LOD)
Reversed/unstable fold changelog2(mean_ratio) uses arithmetic meansDifference of log-means (geometric-mean ratio), consistent with limma/DESeq2
Optimistic CV errorfeature selection done before CVRe-fit selection inside every fold; see machine-learning/model-validation
getVipVn gives orthogonal importanceorthoL = TRUE returns orthogonal VIPUse default (predictive VIP) for discriminant ranking

References

  • van den Berg RA, Hoefsloot HCJ, Westerhuis JA, Smilde AK, van der Werf MJ. 2006. Centering, scaling, and transformations: improving the biological information content of metabolomics data. BMC Genomics 7:142.
  • Westerhuis JA, Hoefsloot HCJ, Smit S, Vis DJ, Smilde AK, et al. 2008. Assessment of PLSDA cross validation. Metabolomics 4:81-89.
  • Westerhuis JA, van Velzen EJJ, Hoefsloot HCJ, Smilde AK. 2008. Discriminant Q2 (DQ2) for improved discrimination in PLSDA models. Metabolomics 4:293-296.
  • Saccenti E, Hoefsloot HCJ, Smilde AK, Westerhuis JA, Hendriks MMWB. 2014. Reflections on univariate and multivariate analysis of metabolomics data. Metabolomics 10:361-374.
  • Broadhurst DI, Kell DB. 2006. Statistical strategies for avoiding false discoveries in metabolomics and related experiments. Metabolomics 2:171-196.
  • Thevenot EA, Roux A, Xu Y, Ezan E, Junot C. 2015. Analysis of the human adult urinary metabolome variations with age, body mass index, and gender by implementing a comprehensive workflow for univariate and OPLS statistical analyses. J Proteome Res 14:3322-3335.
  • Triba MN, Le Moyec L, Amathieu R, Goossens C, Bouchemal N, et al. 2015. PLS/OPLS models in metabolomics: the impact of permutation of dataset rows on the K-fold cross-validation quality parameters. Mol BioSyst 11:13-19.
  • Szymanska E, Saccenti E, Smilde AK, Westerhuis JA. 2012. Double-check: validation of diagnostic statistics for PLS-DA models in metabolomics studies. Metabolomics 8(Suppl 1):3-16.
  • Galindo-Prieto B, Eriksson L, Trygg J. 2014. Variable influence on projection (VIP) for orthogonal projections to latent structures (OPLS). J Chemometr 28:623-632.
  • Ruiz-Perez D, Guan H, Madhivanan P, Mathee K, Narasimhan G. 2020. So you think you can PLS-DA? BMC Bioinformatics 21(Suppl 1):2.
  • Peluso A, Glen R, Ebbels TMD. 2021. Multiple-testing correction in metabolome-wide association studies. BMC Bioinformatics 22:67.
  • Storey JD, Tibshirani R. 2003. Statistical significance for genomewide studies. Proc Natl Acad Sci USA 100:9440-9445.
  • metabolomics/normalization-qc - Sample-wise normalization, drift/batch correction, missing-value imputation upstream of testing
  • metabolomics/pathway-mapping - Functional interpretation of differential metabolites
  • machine-learning/biomarker-discovery - Feature selection inside CV, stability, minimal-optimal vs all-relevant
  • machine-learning/model-validation - Leakage taxonomy, nested CV, calibration vs discrimination
  • experimental-design/multiple-testing - FDR vs FWER regime, discovery vs confirmatory
  • data-visualization/volcano-and-ma-plots - Volcano plot recipes

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in metabolomics/statistical-analysis of GPTomics/bioSkills.

  • SKILL.md
  • examples/metabolomics_differential.py
  • examples/metabolomics_stats.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Metabolomics Statistical Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Metabolomics Statistical Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Metabolomics Statistical Analysis this skillGPTomics/bioSkills1.2k1 repos~5kAutomated safety check: PassMIT
Sandbox Benchvercel/next.js143k—~4.1kAutomated safety check: PassMIT
Statistical Analysisspacering-net/codeg3.9k3 repos~5kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause
AI Daily DigestvigorX777/ai-daily-digest1.6k—~1.3kAutomated safety check: PassNone
Statistical Powerspacering-net/codeg3.9k1 repos~3.6kAutomated safety check: NotesMIT

Similar skills

  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Statistical Analysis

    spacering-net/codeg

    Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting.

    3.9k GitHub starsUsed in 3 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • AI Daily Digest

    vigorX777/ai-daily-digest

    Fetches RSS feeds from 90 top Hacker News blogs (curated by Karpathy), uses AI to score and filter articles, and generates a daily digest in Markdown with Chinese-translated titles, category…

    1.6k GitHub stars~1.3k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Statistical Power

    spacering-net/codeg

    Sample-size and statistical power calculations for planning studies.

    3.9k GitHub starsUsed in 1 repo~3.6k tokens
    Data & AnalyticsAuto-check: notes
  • Agent Session Monitor

    higress-group/higress

    Real-time agent conversation monitoring - monitors Higress access logs, aggregates conversations by session, tracks token usage.

    9.5k GitHub stars~3.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Metabolomics Statistical Analysis

What does Bio Metabolomics Statistical Analysis do?

Decision-grade statistical analysis for metabolomics intensity tables. Bio Metabolomics Statistical Analysis is an agent skill from GPTomics/bioSkills. Decision-grade statistical analysis for metabolomics intensity tables.

When should I use Bio Metabolomics Statistical Analysis?

Bio Metabolomics Statistical Analysis fits situations like: testing which metabolites differ; validating a discriminant model; choosing a scaling; correcting many correlated tests.

How do I install Bio Metabolomics Statistical Analysis in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-metabolomics-statistical-analysis -a claude-code`. Or copy the skill folder (metabolomics/statistical-analysis in GPTomics/bioSkills) into .claude/skills/bio-metabolomics-statistical-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Bio Metabolomics Statistical Analysis in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-metabolomics-statistical-analysis -a codex`. Or copy the skill folder (metabolomics/statistical-analysis in GPTomics/bioSkills) into .agents/skills/bio-metabolomics-statistical-analysis in your project. Codex loads it when a task matches its description.

Can I use Bio Metabolomics Statistical Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-metabolomics-statistical-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-metabolomics-statistical-analysis, .gemini/skills/bio-metabolomics-statistical-analysis, .github/skills/bio-metabolomics-statistical-analysis and .opencode/skills/bio-metabolomics-statistical-analysis in your project.

What does Bio Metabolomics Statistical Analysis need to run?

Going by SKILL.md and its folder, Bio Metabolomics Statistical Analysis needs Python and R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Metabolomics Statistical Analysis access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Metabolomics Statistical Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Metabolomics Statistical Analysis use?

Bio Metabolomics Statistical Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Metabolomics Statistical Analysis use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Metabolomics Statistical Analysis?

Skills that share tags, products or a category with Bio Metabolomics Statistical Analysis: Sandbox Bench (vercel/next.js, 143k stars), Statistical Analysis (spacering-net/codeg, 3.9k stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars) and AI Daily Digest (vigorX777/ai-daily-digest, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Metabolomics Statistical Analysis?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.