Bio Splicing Qc
FreedomIntelligence/OpenClaw-Medical-Skills
Assesses RNA-seq data quality for splicing analysis including junction saturation curves, splice site strength scoring, and junction coverage metrics using RSeQC.
Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-differential-abundance --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/differential-abundance .claude/skills/bio-proteomics-differential-abundance && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-proteomics-differential-abundance" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/differential-abundance into .claude/skills/bio-proteomics-differential-abundance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-differential-abundance", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/proteomics/differential-abundanceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-differential-abundance --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/proteomics/differential-abundance .agents/skills/bio-proteomics-differential-abundance && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-proteomics-differential-abundance" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/differential-abundance into .agents/skills/bio-proteomics-differential-abundance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-differential-abundance", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-differential-abundance --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/proteomics/differential-abundance .cursor/skills/bio-proteomics-differential-abundance && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-proteomics-differential-abundance" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/differential-abundance into .cursor/skills/bio-proteomics-differential-abundance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-differential-abundance", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path proteomics/differential-abundance--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-differential-abundance --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/proteomics/differential-abundance .gemini/skills/bio-proteomics-differential-abundance && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-proteomics-differential-abundance" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/differential-abundance into .gemini/skills/bio-proteomics-differential-abundance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-differential-abundance", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-proteomics-differential-abundanceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/proteomics/differential-abundance .github/skills/bio-proteomics-differential-abundance && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-proteomics-differential-abundance" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/differential-abundance into .github/skills/bio-proteomics-differential-abundance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-differential-abundance", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-differential-abundance --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/proteomics/differential-abundance .opencode/skills/bio-proteomics-differential-abundance && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-proteomics-differential-abundance" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/differential-abundance into .opencode/skills/bio-proteomics-differential-abundance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-differential-abundance", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-proteomics-differential-abundanceTests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives.
Bio Proteomics Differential Abundance is an agent skill from GPTomics/bioSkills. Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives. Frames missing values as left-censored MNAR (model, do not impute), makes variance moderation the load-bearing step at n=3-5, and prefers feature/peptide-level testing. Use when identifying proteins with significant abundance changes between experimental groups. Summarization and normalization mechanics are…
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/differential_abundance.py` and `usage-guide.md`).
It sits in Data & Analytics, covering Bioinformatics, Data cleaning and Data visualization. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python and R), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Proteomics Differential Abundance loads about 5.7k tokens when it runs. Until then it costs about 174 tokens; SKILL.md has 2,400 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,400 words, ~5,690 tokens.
.claude/skills/bio-proteomics-differential-abundance/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: limma 3.58+, DEqMS 1.20+, proDA 1.20+, ashr 2.2+, pandas 2.2+, scipy 1.12+, statsmodels 0.14+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturespackageVersion('<pkg>') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Find differentially abundant proteins between my conditions" -> Moderated statistical testing on a normalized log-intensity matrix, carrying missingness in the likelihood instead of filling it in -- because the missing values are low BECAUSE the protein is low, and imputing them manufactures false positives.
limma::eBayes(fit, trend=TRUE, robust=TRUE) for empirical-Bayes moderated t-tests (the protein-level workhorse)DEqMS::spectraCounteBayes() when PSM/peptide counts are available (preferred over limma-trend when quant depth varies)proDA::test_diff() / msqrob2 / MSstats when missing values are extensive (model the dropout, no imputation)scipy.stats.ttest_ind(equal_var=False) + statsmodels BH (large n only; no moderation)Scope: this skill owns the statistical TEST -- design/contrast construction, variance moderation, missingness handling, multiple-testing correction, minimum-fold-change testing, and fold-change shrinkage. Peptide-to-protein summarization and normalization mechanics -> proteomics/quantification. Volcano/MA plots -> data-visualization/volcano-and-ma-plots. Enrichment of the hit list -> pathway-analysis/go-enrichment. OUT OF SCOPE: how MaxLFQ/TMP/IRS produce the matrix (quantification); how to draw a volcano (data-visualization).
trend=TRUE makes the prior a function of mean intensity (effectively mandatory for label-free, where a single global prior mis-calibrates FDR across the abundance range); robust=TRUE Winsorizes outlier variances (Phipson 2016). DEqMS makes the prior a function of PSM/peptide count and generally outperforms limma-trend when quantification depth varies across proteins (Zhu 2020).| Tool / method | Citation | Mechanism / role | When |
|---|---|---|---|
| limma | Ritchie 2015; Phipson 2016 | EB moderated t; posterior variance blends a prior d0 with the per-protein estimate; trend ties the prior to mean intensity, robust Winsorizes outliers | protein-level summaries, small n, the default workhorse |
| DEqMS | Zhu 2020 | prior variance = loess of log-variance vs log2(count); precision follows quantification DEPTH not just intensity | TMT (count=PSM) and label-free DDA (count=peptide); quant depth varies; preferred over limma-trend |
| proDA | Ahlmann-Eltze & Anders 2019 (preprint) | probabilistic dropout: missing = left-censored, integrated under a per-sample sigmoid dropout curve; EB on location and variance; no imputation | label-free DDA with many MNAR missing values, small n, proteins absent in one group |
| msqrob2 | Sticker 2020; Goeminne 2016 | peptide-level robust ridge: Huber M-estimation downweights outlier peptides, ridge shrinks effects from few observations, EB variance moderation | label-free DDA, outlier-peptide / unbalanced-coverage risk; best FDR in hard spike-in regimes |
| MSstats | Choi 2014 | feature-level linear mixed model (group fixed + feature + run/subject random); AFT censored handling for missing | SRM/PRM/DIA, technical replicates, nested/repeated-measures, labeled designs |
| Welch t-test + BH | -- | per-protein two-sample t with equal_var=False + Benjamini-Hochberg | large n (>10/group), Python-only; no moderation, unusable at n=3-5 |
| ashr | Stephens 2017 | mixture prior with a point mass at zero; posterior means shrink uncertain effects toward zero | recovering "which proteins truly changed and by how much" (not for GSEA ranking) |
| volcano / MA plot | -- | (route OUT) | visualization -> data-visualization/volcano-and-ma-plots |
| enrichment of hits | -- | (route OUT) | functional interpretation -> pathway-analysis/go-enrichment |
| Scenario | Recommended | Why |
|---|---|---|
| Small n (3-5/group), protein-level summary matrix | limma eBayes(trend=TRUE, robust=TRUE) | EB borrows variance across proteins; the trend calibrates FDR across abundance |
| PSM/peptide counts available (TMT or label-free DDA) | DEqMS spectraCounteBayes | prior keyed on quant depth removes single-PSM false positives limma admits |
| Label-free with many MNAR missing values, on/off proteins | proDA test_diff | models the censored dropout; never imputes; correct verdict for "undetected in one group" |
| Outlier-peptide risk, unbalanced peptide coverage | msqrob2 (peptide-level robust ridge) | keeps feature df; Huber downweights bad peptides; best FDR in spike-in benchmarks |
| Technical replicates, nested/repeated-measures, labeled (SRM/PRM/DIA) | MSstats (feature-level mixed model) | random effects capture run/subject structure summarize-then-test discards |
| Batch present | batch as a covariate in the design (~ batch + condition) | removeBatchEffect is visualization-only; never feed its output to lmFit |
| Minimum biologically meaningful fold change | treat() + topTreat() (or SAM s0) | tests |
| Large n (>10/group), Python-only | Welch t-test + BH | variance estimates reliable; no moderation needed at large n |
Default when uncertain: protein-level summary matrix at n=3-5 -> limma eBayes(trend=TRUE, robust=TRUE); if PSM/peptide counts exist, escalate to DEqMS; if missingness is extensive and intensity-dependent, escalate to proDA.
Goal: Identify differentially abundant proteins using moderated statistics that borrow information across all proteins.
Approach: Build the design (batch as a covariate when present), fit the linear model and contrast, apply EB moderation with the intensity trend and robust fitting, then extract BH-corrected results. Never feed removeBatchEffect output to lmFit.
library(limma)
design <- model.matrix(~0 + condition + batch, data = sample_info) # batch in the model, not removed first
colnames(design)[1:2] <- levels(factor(sample_info$condition))
fit <- lmFit(protein_matrix, design)
contrast_matrix <- makeContrasts(Treatment - Control, levels = design)
fit2 <- contrasts.fit(fit, contrast_matrix)
fit2 <- eBayes(fit2, trend = TRUE, robust = TRUE) # trend mandatory for label-free; robust Winsorizes outliers
results <- topTable(fit2, coef = 1, number = Inf, adjust.method = 'BH')
# columns: logFC, AveExpr, t, P.Value, adj.P.Val, B (adj.P.Val is the BH p; there is no $FDR)Goal: Call proteins whose effect exceeds a biologically meaningful threshold, not merely differ from zero.
Approach: Use treat() against the moderated null and read topTreat(). NEVER topTable(lfc=...) nor a post-hoc volcano double filter (abs(logFC) > 1 & adj.P.Val < 0.05); conditioning on both the FC and the p-value selects for high-variance nulls (a collider effect) and inflates realized FDR above 50% (Ebrahimpoor & Goeman 2021).
LFC_THRESHOLD <- log2(1.2) # 1.2-fold floor; treat tests against this null, no double-filter FDR inflation
fit2 <- treat(fit2, lfc = LFC_THRESHOLD)
results <- topTreat(fit2, coef = 1, number = Inf) # topTreat omits the B columnGoal: Improve on limma by tying each protein's prior variance to its quantification depth -- proteins measured by more PSMs/peptides are more precise.
Approach: Run limma through eBayes, attach the count vector, then apply DEqMS's count-aware EB. Use PSM count for TMT (quant at MS2) and peptide count for label-free DDA; for multi-batch TMT use the MINIMUM count across batches (the bottleneck batch sets precision).
library(DEqMS)
# fit2 is the limma fit through eBayes (above)
fit2$count <- psm_count_per_protein[rownames(fit2$coefficients)] # PSM for TMT, peptide for LFQ; min across batches
fit3 <- spectraCounteBayes(fit2)
results <- outputResult(fit3, coef_col = 1)
# adds sca.t, sca.P.Value, sca.adj.pval (the count-adjusted statistics; use these, not the limma columns)Goal: Test proteins with extensive MNAR missingness, including on/off proteins, without imputing a single value.
Approach: Fit the probabilistic-dropout model directly on the log-intensity matrix; missing values contribute as left-censored observations under a per-sample dropout curve. Then test the contrast against zero.
library(proDA)
fit <- proDA(protein_matrix, design = ~condition, col_data = sample_info,
reference_level = 'Control')
result_names(fit) # list testable coefficients first
results <- test_diff(fit, conditionTreatment - conditionControl)
# columns: name, pval, adj_pval, diff (log2FC), t_statistic, seGoal: Run the full pipeline in Python when no R is available and n is large enough that moderation is unnecessary.
Approach: Log2-transform, median-normalize, run per-protein Welch t-tests, apply Benjamini-Hochberg. This has NO variance moderation and should not be used at n=3-5 -- escalate to limma/DEqMS for small n.
import numpy as np
import pandas as pd
from scipy import stats
from statsmodels.stats.multitest import multipletests
def preprocess(intensities):
log2_data = np.log2(intensities.replace(0, np.nan)) # zeros -> NaN to avoid -inf
sample_medians = log2_data.median(axis=0)
return log2_data - sample_medians + sample_medians.median()
def differential_abundance(normalized, case_cols, ctrl_cols):
rows = []
for protein in normalized.index:
case, ctrl = normalized.loc[protein, case_cols].dropna(), normalized.loc[protein, ctrl_cols].dropna()
if len(case) >= 2 and len(ctrl) >= 2:
_, pval = stats.ttest_ind(case, ctrl, equal_var=False) # Welch; scipy defaults to Student's True
rows.append({'protein': protein, 'log2fc': case.mean() - ctrl.mean(), 'pvalue': pval})
df = pd.DataFrame(rows)
df['padj'] = multipletests(df['pvalue'], method='fdr_bh')[1] # default is Holm-Sidak; pass fdr_bh explicitly
return dfGoal: Hand the right effect estimate to the right consumer.
Approach: Report the RAW fold change (the best unbiased point estimate) for GSEA/pathway ranking and meta-analysis -- those need the full continuous distribution or FC+SE pairs. Apply shrinkage (ashr) only when recovering "which proteins truly changed and by how much"; it fits a mixture prior with a point mass at zero and shrinks uncertain effects smoothly toward zero. This is preferred over hard-thresholding (zeroing FCs at padj 0.05), which creates an arbitrary step function. No mature Python ashr equivalent exists.
library(ashr)
se <- sqrt(fit2$s2.post) * fit2$stdev.unscaled[, 1]
shrunk <- ash(fit2$coefficients[, 1], se, mixcompdist = 'normal')
shrunken_fc <- shrunk$result$PosteriorMean # report alongside raw logFC, not as a replacement for GSEA
lfsr <- shrunk$result$lfsrTrigger: Perseus/MaxQuant downshift (or MinDet/MinProb/QRILC) fills NAs, then limma/t-test runs on the filled matrix. Mechanism: Imputed values come from one narrow Gaussian -> fabricated low within-group variance + deterministic mean offset -> inflated t. Symptom: Volcano "anchor/wing" -- rigid near-vertical streaks of pinned on/off proteins at high significance; realized FDR far above nominal. Fix: Model the missingness instead (proDA / msqrob2 / MSstats-AFT); report on/off proteins as "undetected in group X".
Trigger: kNN/mean imputation applied to label-free data with MNAR dropout. Mechanism: Mean-reverting -- pulls a truly-low (missing because low) value UP toward the mean. Symptom: Real down-regulation is compressed; down hits weakened or lost. Fix: Only valid under MCAR/MAR; for MNAR model the dropout. Under uncertainty Lazar 2016 shows the milder MCAR error beats MNAR-imputers slamming random highs to the floor.
Trigger: removeBatchEffect() output fed to lmFit.
Mechanism: Subtracts the fitted batch component with no uncertainty propagation -> understated residual variance, inflated EB df; if batch is confounded with biology it deletes real signal.
Symptom: Anticonservative p-values; lost true effects when cases/controls split by batch.
Fix: Include batch as a covariate in the SAME model (~ batch + condition); use removeBatchEffect only for PCA/visualization.
Trigger: Plain eBayes (trend off) on a log-intensity matrix.
Mechanism: A single global prior over-shrinks high-abundance and under-shrinks low-abundance proteins.
Symptom: Mis-calibrated FDR across the abundance range.
Fix: eBayes(trend = TRUE, robust = TRUE); escalate to DEqMS when quant depth varies.
Trigger: Razor+unique counts vs MS2-level PSMs, or total-across-batches vs minimum-across-batches. Mechanism: The variance-vs-count prior is fit on the wrong precision proxy. Symptom: Mis-ranked proteins; the count moderation helps the wrong ones. Fix: PSM count for TMT, peptide count for label-free; minimum count across batches for multi-batch TMT.
Trigger: proDA applied where dropout is random (e.g. a TMT channel lost at random), not detection-limited. Mechanism: The left-censored dropout model is mis-specified. Symptom: Biased estimates; the model fits a dropout curve that does not exist. Fix: proDA needs intensity-dependent missingness; for MCAR use limma/DEqMS on the observed values.
Trigger: abs(logFC) > 1 & adj.P.Val < 0.05 applied after the test.
Mechanism: |logFC| is large for a true effect OR a large SE; filtering on both the FC and the p (both depend on SE) selects high-variance nulls (collider effect).
Symptom: Realized FDR above 50% at nominal 5% (Ebrahimpoor & Goeman 2021).
Fix: treat()+topTreat() or SAM s0, which sit inside the statistic before selection.
| Threshold | Source | Rationale |
|---|---|---|
| n=3-5 replicates -> 2-4 residual df | -- | raw per-protein variance unusable; moderation is mandatory, not optional |
| limma adds prior d0 (~4) df | Ritchie 2015 | a 4-replicate design tests on ~10 df vs 6; the borrowed df is the benefit |
| downshift mean = mu - 1.8sigma, SD = 0.3sigma | Perseus default | 1.8 places imputed mass ~3.6th percentile; 0.3 gives only 30% of real spread -> manufactured false positives |
trend=TRUE effectively mandatory for label-free | Ritchie 2015 | a single global prior mis-calibrates FDR across abundance |
| min-FC floor log2(1.2) (1.2-fold) via treat() | -- | example floor; common alternatives 1.5-fold (~0.58) or 2-fold (1.0); set by biology, tested against the moderated null |
| BH adjusted p < 0.05 | Benjamini-Hochberg | controls FDR over the WHOLE rejection set, not subsets carved out afterward |
| DEqMS multi-batch TMT: minimum count across batches | Zhu 2020 | the bottleneck batch sets the realized precision |
| realized FDR > 50% from FC+significance double filter | Ebrahimpoor & Goeman 2021 | top-100 at n=12 exceeded 50% FDR at nominal 5% |
| Error / symptom | Cause | Solution |
|---|---|---|
results$FDR is NULL | limma topTable/topTreat have no $FDR column | use adj.P.Val (BH-adjusted p) |
topTreat row has no B | topTreat omits B (a topTable column) | read logFC, AveExpr, t, P.Value, adj.P.Val |
| FDR mis-calibrated across abundance | eBayes with trend=FALSE on intensity data | eBayes(fit, trend = TRUE, robust = TRUE) |
| min-FC test inflates FDR | topTable(lfc=...) or post-hoc volcano double filter | treat(fit, lfc=log2(1.2)) then topTreat() |
| anticonservative p after batch correction | removeBatchEffect output fed to lmFit | put batch in the design: ~ batch + condition |
| DEqMS columns missing | forgot fit$count or read limma columns | set fit$count, run spectraCounteBayes, read sca.adj.pval from outputResult |
| Student's t instead of Welch | scipy.stats.ttest_ind defaults equal_var=True | pass equal_var=False |
| p-values look like Holm-Sidak | statsmodels multipletests defaults to 'hs' | pass method='fdr_bh' |
| volcano "anchor/wing" streaks | downshift/imputation feeding the test | model dropout (proDA/msqrob2/MSstats-AFT); report on/off proteins as undetected |
citation("proDA"), never a journal).© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in proteomics/differential-abundance of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Proteomics Differential Abundance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Proteomics Differential Abundance this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.7k | Automated safety check: Pass | MIT | |
| Bio Splicing QcFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | — | ~1.6k | Automated safety check: Pass | None | |
| Bio Copy Number Cnv VisualizationFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | — | ~2.6k | Automated safety check: Pass | None | |
| Math Modeling Data Cleaning and Chartsyushui2022/MathModel-Skill | 454 | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| FBA Flux Analyzeraiming-lab/AutoResearchClaw | 15k | — | ~2.3k | Automated safety check: Pass | MIT | |
| Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent | 886 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 |
FreedomIntelligence/OpenClaw-Medical-Skills
Assesses RNA-seq data quality for splicing analysis including junction saturation curves, splice site strength scoring, and junction coverage metrics using RSeQC.
FreedomIntelligence/OpenClaw-Medical-Skills
Visualize copy number profiles, segments, and compare across samples.
yushui2022/MathModel-Skill
Cleans raw or scraped competition data and produces exploratory charts and a figure plan as one stage of a mathematical modeling paper workflow.
aiming-lab/AutoResearchClaw
Turns raw flux balance analysis output and a COBRApy model into gene essentiality maps, phenotypic phase planes, flux sampling results, pathway summaries and secretion predictions.
NVIDIA-AI-Blueprints/deep-researcher-agent
A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.
FreedomIntelligence/OpenClaw-Medical-Skills
Visualize metagenomic profiles using R (phyloseq, microbiome) and Python (matplotlib, seaborn).
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Works with
Categories
Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives. Bio Proteomics Differential Abundance is an agent skill from GPTomics/bioSkills. Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives.
Bio Proteomics Differential Abundance fits situations like: identifying proteins with significant abundance changes between experimental groups; tasks that involve Bioinformatics; tasks that involve Data cleaning.
Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a claude-code`. Or copy the skill folder (proteomics/differential-abundance in GPTomics/bioSkills) into .claude/skills/bio-proteomics-differential-abundance in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a codex`. Or copy the skill folder (proteomics/differential-abundance in GPTomics/bioSkills) into .agents/skills/bio-proteomics-differential-abundance in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-differential-abundance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-differential-abundance, .gemini/skills/bio-proteomics-differential-abundance, .github/skills/bio-proteomics-differential-abundance and .opencode/skills/bio-proteomics-differential-abundance in your project.
Going by SKILL.md and its folder, Bio Proteomics Differential Abundance needs Python and R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Proteomics Differential Abundance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Proteomics Differential Abundance: Bio Splicing Qc (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Bio Copy Number Cnv Visualization (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Math Modeling Data Cleaning and Charts (yushui2022/MathModel-Skill, 454 stars) and FBA Flux Analyzer (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.