Agent skill

Bio Proteomics Differential Abundance

by GPTomics in GPTomics/bioSkills

Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives.

MITAuto-check passedData & Analytics

Install Bio Proteomics Differential Abundance

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-proteomics-differential-abundance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/differential-abundance .claude/skills/bio-proteomics-differential-abundance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-proteomics-differential-abundance
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.7k tokens
SKILL.md length
2,400 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives.

  • Works in 3 steps: Missing values in label-free MS are… → At the n=3-5 replicates proteomics… → Feature/peptide-level modeling beats…
  • Identifying proteins with significant abundance changes between experimental groups
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 10 more sections
  • Runs Python and R scripts from its folder; calls pip

What it does

Bio Proteomics Differential Abundance is an agent skill from GPTomics/bioSkills. Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives. Frames missing values as left-censored MNAR (model, do not impute), makes variance moderation the load-bearing step at n=3-5, and prefers feature/peptide-level testing. Use when identifying proteins with significant abundance changes between experimental groups. Summarization and normalization mechanics are…

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/differential_abundance.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Bioinformatics, Data cleaning and Data visualization. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Identifying proteins with significant abundance changes between experimental groups
  • Tasks that involve Bioinformatics
  • Tasks that involve Data cleaning

Example prompts

  • “Use the bio-proteomics-differential-abundance skill to test for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes…”
  • “/bio-proteomics-differential-abundance”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Missing values in label-free MS are left-censored MNAR -- missing BECAUSE the intensity is low -- and imputing them (especially…
  2. At the n=3-5 replicates proteomics actually uses, per-protein variance has only 2-4 residual df and is unusable raw -- variance moderation…
  3. Feature/peptide-level modeling beats summarize-then-test. Summarizing first (one number per protein per run) discards the within-protein…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Proteomics Differential Abundance loads about 5.7k tokens when it runs. Until then it costs about 174 tokens; SKILL.md has 2,400 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~174
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,400 words, ~5,690 tokens.

Download SKILL.mdSave it as .claude/skills/bio-proteomics-differential-abundance/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-proteomics-differential-abundance
description
Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives. Frames missing values as left-censored MNAR (model, do not impute), makes variance moderation the load-bearing step at n=3-5, and prefers feature/peptide-level testing. Use when identifying proteins with significant abundance changes between experimental groups. Summarization and normalization mechanics are proteomics/quantification; volcano and MA plots are data-visualization/volcano-and-ma-plots; pathway enrichment of the hit list is pathway-analysis/go-enrichment.
tool_type
mixed
primary_tool
limma

Version Compatibility

Reference examples tested with: limma 3.58+, DEqMS 1.20+, proDA 1.20+, ashr 2.2+, pandas 2.2+, scipy 1.12+, statsmodels 0.14+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Differential Protein Abundance -- Moderated Testing on a Log-Intensity Matrix with Honest Missingness

"Find differentially abundant proteins between my conditions" -> Moderated statistical testing on a normalized log-intensity matrix, carrying missingness in the likelihood instead of filling it in -- because the missing values are low BECAUSE the protein is low, and imputing them manufactures false positives.

  • R: limma::eBayes(fit, trend=TRUE, robust=TRUE) for empirical-Bayes moderated t-tests (the protein-level workhorse)
  • R: DEqMS::spectraCounteBayes() when PSM/peptide counts are available (preferred over limma-trend when quant depth varies)
  • R: proDA::test_diff() / msqrob2 / MSstats when missing values are extensive (model the dropout, no imputation)
  • Python: scipy.stats.ttest_ind(equal_var=False) + statsmodels BH (large n only; no moderation)

Scope: this skill owns the statistical TEST -- design/contrast construction, variance moderation, missingness handling, multiple-testing correction, minimum-fold-change testing, and fold-change shrinkage. Peptide-to-protein summarization and normalization mechanics -> proteomics/quantification. Volcano/MA plots -> data-visualization/volcano-and-ma-plots. Enrichment of the hit list -> pathway-analysis/go-enrichment. OUT OF SCOPE: how MaxLFQ/TMP/IRS produce the matrix (quantification); how to draw a volcano (data-visualization).

The Single Most Important Modern Insight -- Model the Missingness, Moderate the Variance, Test at the Feature Level

  1. Missing values in label-free MS are left-censored MNAR -- missing BECAUSE the intensity is low -- and imputing them (especially Perseus/MaxQuant downshift) manufactures SYSTEMATIC false positives. Downshift draws each missing value from a narrow Gaussian (mean = mu - 1.8sigma, SD = 0.3sigma). For an on/off protein (seen in all of group A, missing in all of B) the t-statistic numerator is inflated by construction (mean_B fixed ~1.8 sigma below observed, deterministic) and the denominator is artificially deflated (all imputed B values from one 0.3-sigma Gaussian -> collapsed within-group SD) -> enormous t -> tiny p. Because every on/off protein is treated identically, the false positives are systematic: the volcano-plot "anchor/wing" artifact (rigid near-vertical streaks of pinned points far out on both x-axis sides). The honest statement is "undetected in group B", not "20x lower, p=1e-6". The correct approach is to MODEL the dropout in the likelihood -- proDA (probabilistic dropout), msqrob2, MSstats-AFT -- NOT fill it (Lazar 2016; Ahlmann-Eltze & Anders 2019).
  2. At the n=3-5 replicates proteomics actually uses, per-protein variance has only 2-4 residual df and is unusable raw -- variance moderation is the load-bearing element, not optional. limma borrows a prior d0 across all proteins so a 4-replicate design tests on ~10 df instead of 6; trend=TRUE makes the prior a function of mean intensity (effectively mandatory for label-free, where a single global prior mis-calibrates FDR across the abundance range); robust=TRUE Winsorizes outlier variances (Phipson 2016). DEqMS makes the prior a function of PSM/peptide count and generally outperforms limma-trend when quantification depth varies across proteins (Zhu 2020).
  3. Feature/peptide-level modeling beats summarize-then-test. Summarizing first (one number per protein per run) discards the within-protein between-peptide variance and the correct degrees of freedom: 12 consistent peptides deserve a smaller SE than 12 disagreeing ones, but after summarization both look equally certain, and a protein with 30 observations looks as informative as one with 3. msqrob2/MSstats keep every peptide as a degree of freedom; this is why the same data gives different answers (Goeminne 2016; Sticker 2020; Choi 2014).

Tool Taxonomy

Tool / methodCitationMechanism / roleWhen
limmaRitchie 2015; Phipson 2016EB moderated t; posterior variance blends a prior d0 with the per-protein estimate; trend ties the prior to mean intensity, robust Winsorizes outliersprotein-level summaries, small n, the default workhorse
DEqMSZhu 2020prior variance = loess of log-variance vs log2(count); precision follows quantification DEPTH not just intensityTMT (count=PSM) and label-free DDA (count=peptide); quant depth varies; preferred over limma-trend
proDAAhlmann-Eltze & Anders 2019 (preprint)probabilistic dropout: missing = left-censored, integrated under a per-sample sigmoid dropout curve; EB on location and variance; no imputationlabel-free DDA with many MNAR missing values, small n, proteins absent in one group
msqrob2Sticker 2020; Goeminne 2016peptide-level robust ridge: Huber M-estimation downweights outlier peptides, ridge shrinks effects from few observations, EB variance moderationlabel-free DDA, outlier-peptide / unbalanced-coverage risk; best FDR in hard spike-in regimes
MSstatsChoi 2014feature-level linear mixed model (group fixed + feature + run/subject random); AFT censored handling for missingSRM/PRM/DIA, technical replicates, nested/repeated-measures, labeled designs
Welch t-test + BH--per-protein two-sample t with equal_var=False + Benjamini-Hochberglarge n (>10/group), Python-only; no moderation, unusable at n=3-5
ashrStephens 2017mixture prior with a point mass at zero; posterior means shrink uncertain effects toward zerorecovering "which proteins truly changed and by how much" (not for GSEA ranking)
volcano / MA plot--(route OUT)visualization -> data-visualization/volcano-and-ma-plots
enrichment of hits--(route OUT)functional interpretation -> pathway-analysis/go-enrichment

Decision Tree by Scenario

ScenarioRecommendedWhy
Small n (3-5/group), protein-level summary matrixlimma eBayes(trend=TRUE, robust=TRUE)EB borrows variance across proteins; the trend calibrates FDR across abundance
PSM/peptide counts available (TMT or label-free DDA)DEqMS spectraCounteBayesprior keyed on quant depth removes single-PSM false positives limma admits
Label-free with many MNAR missing values, on/off proteinsproDA test_diffmodels the censored dropout; never imputes; correct verdict for "undetected in one group"
Outlier-peptide risk, unbalanced peptide coveragemsqrob2 (peptide-level robust ridge)keeps feature df; Huber downweights bad peptides; best FDR in spike-in benchmarks
Technical replicates, nested/repeated-measures, labeled (SRM/PRM/DIA)MSstats (feature-level mixed model)random effects capture run/subject structure summarize-then-test discards
Batch presentbatch as a covariate in the design (~ batch + condition)removeBatchEffect is visualization-only; never feed its output to lmFit
Minimum biologically meaningful fold changetreat() + topTreat() (or SAM s0)tests
Large n (>10/group), Python-onlyWelch t-test + BHvariance estimates reliable; no moderation needed at large n

Default when uncertain: protein-level summary matrix at n=3-5 -> limma eBayes(trend=TRUE, robust=TRUE); if PSM/peptide counts exist, escalate to DEqMS; if missingness is extensive and intensity-dependent, escalate to proDA.

limma Workflow (R)

Goal: Identify differentially abundant proteins using moderated statistics that borrow information across all proteins.

Approach: Build the design (batch as a covariate when present), fit the linear model and contrast, apply EB moderation with the intensity trend and robust fitting, then extract BH-corrected results. Never feed removeBatchEffect output to lmFit.

r
library(limma)

design <- model.matrix(~0 + condition + batch, data = sample_info)  # batch in the model, not removed first
colnames(design)[1:2] <- levels(factor(sample_info$condition))

fit <- lmFit(protein_matrix, design)
contrast_matrix <- makeContrasts(Treatment - Control, levels = design)
fit2 <- contrasts.fit(fit, contrast_matrix)
fit2 <- eBayes(fit2, trend = TRUE, robust = TRUE)  # trend mandatory for label-free; robust Winsorizes outliers

results <- topTable(fit2, coef = 1, number = Inf, adjust.method = 'BH')
# columns: logFC, AveExpr, t, P.Value, adj.P.Val, B  (adj.P.Val is the BH p; there is no $FDR)
Minimum-fold-change testing

Goal: Call proteins whose effect exceeds a biologically meaningful threshold, not merely differ from zero.

Approach: Use treat() against the moderated null and read topTreat(). NEVER topTable(lfc=...) nor a post-hoc volcano double filter (abs(logFC) > 1 & adj.P.Val < 0.05); conditioning on both the FC and the p-value selects for high-variance nulls (a collider effect) and inflates realized FDR above 50% (Ebrahimpoor & Goeman 2021).

r
LFC_THRESHOLD <- log2(1.2)  # 1.2-fold floor; treat tests against this null, no double-filter FDR inflation
fit2 <- treat(fit2, lfc = LFC_THRESHOLD)
results <- topTreat(fit2, coef = 1, number = Inf)  # topTreat omits the B column

DEqMS Workflow (R)

Goal: Improve on limma by tying each protein's prior variance to its quantification depth -- proteins measured by more PSMs/peptides are more precise.

Approach: Run limma through eBayes, attach the count vector, then apply DEqMS's count-aware EB. Use PSM count for TMT (quant at MS2) and peptide count for label-free DDA; for multi-batch TMT use the MINIMUM count across batches (the bottleneck batch sets precision).

r
library(DEqMS)

# fit2 is the limma fit through eBayes (above)
fit2$count <- psm_count_per_protein[rownames(fit2$coefficients)]  # PSM for TMT, peptide for LFQ; min across batches
fit3 <- spectraCounteBayes(fit2)

results <- outputResult(fit3, coef_col = 1)
# adds sca.t, sca.P.Value, sca.adj.pval (the count-adjusted statistics; use these, not the limma columns)

proDA Workflow (R)

Goal: Test proteins with extensive MNAR missingness, including on/off proteins, without imputing a single value.

Approach: Fit the probabilistic-dropout model directly on the log-intensity matrix; missing values contribute as left-censored observations under a per-sample dropout curve. Then test the contrast against zero.

r
library(proDA)

fit <- proDA(protein_matrix, design = ~condition, col_data = sample_info,
             reference_level = 'Control')
result_names(fit)  # list testable coefficients first
results <- test_diff(fit, conditionTreatment - conditionControl)
# columns: name, pval, adj_pval, diff (log2FC), t_statistic, se

Python Workflow

Goal: Run the full pipeline in Python when no R is available and n is large enough that moderation is unnecessary.

Approach: Log2-transform, median-normalize, run per-protein Welch t-tests, apply Benjamini-Hochberg. This has NO variance moderation and should not be used at n=3-5 -- escalate to limma/DEqMS for small n.

python
import numpy as np
import pandas as pd
from scipy import stats
from statsmodels.stats.multitest import multipletests

def preprocess(intensities):
    log2_data = np.log2(intensities.replace(0, np.nan))  # zeros -> NaN to avoid -inf
    sample_medians = log2_data.median(axis=0)
    return log2_data - sample_medians + sample_medians.median()

def differential_abundance(normalized, case_cols, ctrl_cols):
    rows = []
    for protein in normalized.index:
        case, ctrl = normalized.loc[protein, case_cols].dropna(), normalized.loc[protein, ctrl_cols].dropna()
        if len(case) >= 2 and len(ctrl) >= 2:
            _, pval = stats.ttest_ind(case, ctrl, equal_var=False)  # Welch; scipy defaults to Student's True
            rows.append({'protein': protein, 'log2fc': case.mean() - ctrl.mean(), 'pvalue': pval})
    df = pd.DataFrame(rows)
    df['padj'] = multipletests(df['pvalue'], method='fdr_bh')[1]  # default is Holm-Sidak; pass fdr_bh explicitly
    return df

Fold-Change Reporting

Goal: Hand the right effect estimate to the right consumer.

Approach: Report the RAW fold change (the best unbiased point estimate) for GSEA/pathway ranking and meta-analysis -- those need the full continuous distribution or FC+SE pairs. Apply shrinkage (ashr) only when recovering "which proteins truly changed and by how much"; it fits a mixture prior with a point mass at zero and shrinks uncertain effects smoothly toward zero. This is preferred over hard-thresholding (zeroing FCs at padj 0.05), which creates an arbitrary step function. No mature Python ashr equivalent exists.

r
library(ashr)

se <- sqrt(fit2$s2.post) * fit2$stdev.unscaled[, 1]
shrunk <- ash(fit2$coefficients[, 1], se, mixcompdist = 'normal')
shrunken_fc <- shrunk$result$PosteriorMean  # report alongside raw logFC, not as a replacement for GSEA
lfsr <- shrunk$result$lfsr

Per-Method Failure Modes

Downshift / any imputation feeding a variance-based test

Trigger: Perseus/MaxQuant downshift (or MinDet/MinProb/QRILC) fills NAs, then limma/t-test runs on the filled matrix. Mechanism: Imputed values come from one narrow Gaussian -> fabricated low within-group variance + deterministic mean offset -> inflated t. Symptom: Volcano "anchor/wing" -- rigid near-vertical streaks of pinned on/off proteins at high significance; realized FDR far above nominal. Fix: Model the missingness instead (proDA / msqrob2 / MSstats-AFT); report on/off proteins as "undetected in group X".

Show full SKILL.md (966 more words)Show less
kNN imputation on left-censored data

Trigger: kNN/mean imputation applied to label-free data with MNAR dropout. Mechanism: Mean-reverting -- pulls a truly-low (missing because low) value UP toward the mean. Symptom: Real down-regulation is compressed; down hits weakened or lost. Fix: Only valid under MCAR/MAR; for MNAR model the dropout. Under uncertainty Lazar 2016 shows the milder MCAR error beats MNAR-imputers slamming random highs to the floor.

removeBatchEffect before testing

Trigger: removeBatchEffect() output fed to lmFit. Mechanism: Subtracts the fitted batch component with no uncertainty propagation -> understated residual variance, inflated EB df; if batch is confounded with biology it deletes real signal. Symptom: Anticonservative p-values; lost true effects when cases/controls split by batch. Fix: Include batch as a covariate in the SAME model (~ batch + condition); use removeBatchEffect only for PCA/visualization.

eBayes(trend=FALSE) on intensity data

Trigger: Plain eBayes (trend off) on a log-intensity matrix. Mechanism: A single global prior over-shrinks high-abundance and under-shrinks low-abundance proteins. Symptom: Mis-calibrated FDR across the abundance range. Fix: eBayes(trend = TRUE, robust = TRUE); escalate to DEqMS when quant depth varies.

Wrong DEqMS count column

Trigger: Razor+unique counts vs MS2-level PSMs, or total-across-batches vs minimum-across-batches. Mechanism: The variance-vs-count prior is fit on the wrong precision proxy. Symptom: Mis-ranked proteins; the count moderation helps the wrong ones. Fix: PSM count for TMT, peptide count for label-free; minimum count across batches for multi-batch TMT.

proDA on MCAR missingness

Trigger: proDA applied where dropout is random (e.g. a TMT channel lost at random), not detection-limited. Mechanism: The left-censored dropout model is mis-specified. Symptom: Biased estimates; the model fits a dropout curve that does not exist. Fix: proDA needs intensity-dependent missingness; for MCAR use limma/DEqMS on the observed values.

FC + significance double filter

Trigger: abs(logFC) > 1 & adj.P.Val < 0.05 applied after the test. Mechanism: |logFC| is large for a true effect OR a large SE; filtering on both the FC and the p (both depend on SE) selects high-variance nulls (collider effect). Symptom: Realized FDR above 50% at nominal 5% (Ebrahimpoor & Goeman 2021). Fix: treat()+topTreat() or SAM s0, which sit inside the statistic before selection.

Quantitative Thresholds

ThresholdSourceRationale
n=3-5 replicates -> 2-4 residual df--raw per-protein variance unusable; moderation is mandatory, not optional
limma adds prior d0 (~4) dfRitchie 2015a 4-replicate design tests on ~10 df vs 6; the borrowed df is the benefit
downshift mean = mu - 1.8sigma, SD = 0.3sigmaPerseus default1.8 places imputed mass ~3.6th percentile; 0.3 gives only 30% of real spread -> manufactured false positives
trend=TRUE effectively mandatory for label-freeRitchie 2015a single global prior mis-calibrates FDR across abundance
min-FC floor log2(1.2) (1.2-fold) via treat()--example floor; common alternatives 1.5-fold (~0.58) or 2-fold (1.0); set by biology, tested against the moderated null
BH adjusted p < 0.05Benjamini-Hochbergcontrols FDR over the WHOLE rejection set, not subsets carved out afterward
DEqMS multi-batch TMT: minimum count across batchesZhu 2020the bottleneck batch sets the realized precision
realized FDR > 50% from FC+significance double filterEbrahimpoor & Goeman 2021top-100 at n=12 exceeded 50% FDR at nominal 5%

Common Errors

Error / symptomCauseSolution
results$FDR is NULLlimma topTable/topTreat have no $FDR columnuse adj.P.Val (BH-adjusted p)
topTreat row has no BtopTreat omits B (a topTable column)read logFC, AveExpr, t, P.Value, adj.P.Val
FDR mis-calibrated across abundanceeBayes with trend=FALSE on intensity dataeBayes(fit, trend = TRUE, robust = TRUE)
min-FC test inflates FDRtopTable(lfc=...) or post-hoc volcano double filtertreat(fit, lfc=log2(1.2)) then topTreat()
anticonservative p after batch correctionremoveBatchEffect output fed to lmFitput batch in the design: ~ batch + condition
DEqMS columns missingforgot fit$count or read limma columnsset fit$count, run spectraCounteBayes, read sca.adj.pval from outputResult
Student's t instead of Welchscipy.stats.ttest_ind defaults equal_var=Truepass equal_var=False
p-values look like Holm-Sidakstatsmodels multipletests defaults to 'hs'pass method='fdr_bh'
volcano "anchor/wing" streaksdownshift/imputation feeding the testmodel dropout (proDA/msqrob2/MSstats-AFT); report on/off proteins as undetected

References

  • Ritchie ME, Phipson B, Wu D, Hu Y, Law CW, Shi W, Smyth GK. 2015. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res 43(7):e47.
  • Phipson B, Lee S, Majewski IJ, Alexander WS, Smyth GK. 2016. Robust hyperparameter estimation protects against hypervariable genes and improves power to detect differential expression. Ann Appl Stat 10(2):946-963.
  • Zhu Y, Orre LM, Zhou Tran Y, et al. 2020. DEqMS: a method for accurate variance estimation in differential protein expression analysis. Mol Cell Proteomics 19(6):1047-1057.
  • Ahlmann-Eltze C, Anders S. 2019. proDA: probabilistic dropout analysis for identifying differentially abundant proteins in label-free mass spectrometry. bioRxiv 661496 (preprint; cite citation("proDA"), never a journal).
  • Choi M, Chang CY, Clough T, Broudy D, Killeen T, MacLean B, Vitek O. 2014. MSstats: an R package for statistical analysis of quantitative mass spectrometry-based proteomic experiments. Bioinformatics 30(17):2524-2526.
  • Goeminne LJE, Gevaert K, Clement L. 2016. Peptide-level robust ridge regression improves estimation, sensitivity, and specificity in data-dependent quantitative label-free shotgun proteomics. Mol Cell Proteomics 15(2):657-668.
  • Sticker A, Goeminne L, Martens L, Clement L. 2020. Robust summarization and inference in proteome-wide label-free quantification. Mol Cell Proteomics 19(7):1209-1219.
  • Lazar C, Gatto L, Ferro M, Bruley C, Burger T. 2016. Accounting for the multiple natures of missing values in label-free quantitative proteomics data sets to compare imputation strategies. J Proteome Res 15(4):1116-1125.
  • Stephens M. 2017. False discovery rates: a new deal. Biostatistics 18(2):275-294.
  • Ebrahimpoor M, Goeman JJ. 2021. Inflated false discovery rate due to volcano plots: problem and solutions. Brief Bioinform 22(5):bbab053.
  • quantification - peptide-to-protein summarization, normalization, and IRS that produce the matrix this skill tests
  • proteomics-qc - quality control and batch-effect assessment before testing
  • protein-inference - razor/shared-peptide ambiguity that drives which protein group gets the quantity
  • ptm-analysis - site-level differential testing for modified peptides
  • differential-expression/de-results - analogous empirical-Bayes interpretation for RNA-seq DE
  • data-visualization/volcano-and-ma-plots - volcano and MA plots of the result table
  • pathway-analysis/go-enrichment - functional enrichment of the significant protein hit list
  • machine-learning/biomarker-discovery - building predictive panels from differential proteins
  • workflows/proteomics-pipeline - end-to-end pipeline that calls this skill as the testing stage

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in proteomics/differential-abundance of GPTomics/bioSkills.

  • SKILL.md
  • examples/differential_abundance.py
  • examples/limma_analysis.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Proteomics Differential Abundance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Proteomics Differential Abundance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Proteomics Differential Abundance this skillGPTomics/bioSkills1.2k1 repos~5.7kAutomated safety check: PassMIT
Bio Splicing QcFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~1.6kAutomated safety check: PassNone
Bio Copy Number Cnv VisualizationFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~2.6kAutomated safety check: PassNone
Math Modeling Data Cleaning and Chartsyushui2022/MathModel-Skill4541 repos~1.7kAutomated safety check: PassMIT
FBA Flux Analyzeraiming-lab/AutoResearchClaw15k—~2.3kAutomated safety check: PassMIT
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent886—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Bio Splicing Qc

    FreedomIntelligence/OpenClaw-Medical-Skills

    Assesses RNA-seq data quality for splicing analysis including junction saturation curves, splice site strength scoring, and junction coverage metrics using RSeQC.

    3.1k GitHub stars~1.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Bio Copy Number Cnv Visualization

    FreedomIntelligence/OpenClaw-Medical-Skills

    Visualize copy number profiles, segments, and compare across samples.

    3.1k GitHub stars~2.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Math Modeling Data Cleaning and Charts

    yushui2022/MathModel-Skill

    Cleans raw or scraped competition data and produces exploratory charts and a figure plan as one stage of a mathematical modeling paper workflow.

    454 GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • FBA Flux Analyzer

    aiming-lab/AutoResearchClaw

    Turns raw flux balance analysis output and a COBRApy model into gene essentiality maps, phenotypic phase planes, flux sampling results, pathway summaries and secretion predictions.

    15k GitHub stars~2.3k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    886 GitHub stars~2.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Bio Metagenomics Visualization

    FreedomIntelligence/OpenClaw-Medical-Skills

    Visualize metagenomic profiles using R (phyloseq, microbiome) and Python (matplotlib, seaborn).

    3.1k GitHub starsUsed in 1 repo~1.8k tokens
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Proteomics Differential Abundance

What does Bio Proteomics Differential Abundance do?

Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives. Bio Proteomics Differential Abundance is an agent skill from GPTomics/bioSkills. Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives.

When should I use Bio Proteomics Differential Abundance?

Bio Proteomics Differential Abundance fits situations like: identifying proteins with significant abundance changes between experimental groups; tasks that involve Bioinformatics; tasks that involve Data cleaning.

How do I install Bio Proteomics Differential Abundance in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a claude-code`. Or copy the skill folder (proteomics/differential-abundance in GPTomics/bioSkills) into .claude/skills/bio-proteomics-differential-abundance in your project. Claude Code loads it when a task matches its description.

How do I install Bio Proteomics Differential Abundance in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a codex`. Or copy the skill folder (proteomics/differential-abundance in GPTomics/bioSkills) into .agents/skills/bio-proteomics-differential-abundance in your project. Codex loads it when a task matches its description.

Can I use Bio Proteomics Differential Abundance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-differential-abundance, .gemini/skills/bio-proteomics-differential-abundance, .github/skills/bio-proteomics-differential-abundance and .opencode/skills/bio-proteomics-differential-abundance in your project.

What does Bio Proteomics Differential Abundance need to run?

Going by SKILL.md and its folder, Bio Proteomics Differential Abundance needs Python and R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Proteomics Differential Abundance access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Proteomics Differential Abundance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Proteomics Differential Abundance use?

Bio Proteomics Differential Abundance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Proteomics Differential Abundance use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Proteomics Differential Abundance?

Skills that share tags, products or a category with Bio Proteomics Differential Abundance: Bio Splicing Qc (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Bio Copy Number Cnv Visualization (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Math Modeling Data Cleaning and Charts (yushui2022/MathModel-Skill, 454 stars) and FBA Flux Analyzer (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Proteomics Differential Abundance?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.