Agent skill

Bio Differential Expression De Results

by GPTomics in GPTomics/bioSkills

Extracts, filters, annotates, and exports differential expression results from DESeq2 or edgeR with proper handling of padj=NA (independent filtering, Cook's outliers, all-zero), multiple-testing…

MITAuto-check passedDatabases

Install Bio Differential Expression De Results

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-differential-expression-de-results -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-differential-expression-de-results --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/differential-expression/de-results .claude/skills/bio-differential-expression-de-results && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-differential-expression-de-results
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.7k tokens
SKILL.md length
2,339 words
Files
5
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Extracts, filters, annotates, and exports differential expression results from DESeq2 or edgeR with proper handling of padj=NA (independent filtering, Cook's outliers, all-zero), multiple-testing…

  • Extracting and interpreting DE results
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithmic Taxonomy and Decision Tree by Scenario, plus 13 more sections
  • Runs R scripts from its folder; calls pip
  • Troubleshooting padj=NA

What it does

Bio Differential Expression De Results is an agent skill from GPTomics/bioSkills. Extracts, filters, annotates, and exports differential expression results from DESeq2 or edgeR with proper handling of padj=NA (independent filtering, Cook's outliers, all-zero), multiple-testing correction choice (BH vs Storey q-value vs IHW vs lfsr), TREAT vs post-hoc fold-change filtering, p-value histogram diagnostics, gene annotation via org.db/biomaRt/mygene, GSEA preranked input, ORA background construction, replication reality (Schurch 2016 small-n result), and SABV/sex-stratified reporting. Use when…

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `usage-guide.md`).

It sits in Databases, covering Statistics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Extracting and interpreting DE results
  • Troubleshooting padj=NA
  • Choosing FDR method
  • Preparing ranked lists for pathway analysis

Example prompts

  • “/bio-differential-expression-de-results”

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Differential Expression De Results loads about 5.7k tokens when it runs. Until then it costs about 186 tokens; SKILL.md has 2,339 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~186
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,339 words, ~5,666 tokens.

Download SKILL.mdSave it as .claude/skills/bio-differential-expression-de-results/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
bio-differential-expression-de-results
description
Extracts, filters, annotates, and exports differential expression results from DESeq2 or edgeR with proper handling of padj=NA (independent filtering, Cook's outliers, all-zero), multiple-testing correction choice (BH vs Storey q-value vs IHW vs lfsr), TREAT vs post-hoc fold-change filtering, p-value histogram diagnostics, gene annotation via org.db/biomaRt/mygene, GSEA preranked input, ORA background construction, replication reality (Schurch 2016 small-n result), and SABV/sex-stratified reporting. Use when extracting and interpreting DE results, troubleshooting padj=NA, choosing FDR method, preparing ranked lists for pathway analysis, annotating gene IDs, or comparing DESeq2 vs edgeR outputs.
tool_type
r
primary_tool
DESeq2

Version Compatibility

Reference examples tested with: DESeq2 1.42+, edgeR 4.0+, IHW 1.34+, qvalue 2.34+, ashr 2.2+, AnnotationDbi 1.66+, org.Hs.eg.db 3.18+, biomaRt 2.58+, mygene 1.38+ (Python), dplyr 1.1+, openxlsx 4.2+

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

DE Results

"What are my significant genes?" -> Extract DE estimates and p-values from the fitted model, handle missing padj correctly, apply FDR control appropriate to the design, and produce the table or ranked list the downstream tool actually needs.

The Single Most Important Modern Insight -- padj = NA has three distinct meanings

A NA in the padj column is not a missing value; it is a flag indicating which filter excluded the gene. The three causes -- independent filtering, Cook's distance outlier, and all-zero in a group -- have completely different remediations. Dropping all NA rows blindly silently discards real signal, most often from low-count master regulators (transcription factors expressed at ~10 counts) that pass biology but fail the data-driven baseMean threshold.

padj = NA causeDESeq2 detectionWhat it meansFix if undesired
Independent filteringfinite pvalue, NA padj, baseMean below auto thresholdRemoved before BH adjustment to maximize rejections at alpharesults(dds, independentFiltering = FALSE) OR filterFun = ihw
Cook's distance outlierNA pvalue, NA padj, baseMean > 0, group has >=3 repsOne sample has Cook's > qf(0.99, p, m-p)results(dds, cooksCutoff = FALSE)
All-zero or near-zero in a groupNA pvalue AND baseMean very lowInsufficient information to testFilter at preprocess time; or accept

Independent filtering (Bourgon, Gentleman, Huber 2010 PNAS 107:9546) chooses the baseMean threshold to maximize rejections. The filter MUST be independent of the test statistic under the null -- this is why baseMean (the across-sample mean) is the canonical choice. Using "min count in treatment group" as a filter VIOLATES the independence requirement and inflates type-I error. Most pipelines unknowingly do this; do not.

A second axis: at n>=7 per group, DESeq() also REPLACES outlier counts via replaceOutliers() and refits (default minReplicatesForReplace = 7). Cook's filtering is NOT computed for continuous covariates -- a continuous-covariate analysis has effectively no outlier filtering.

Algorithmic Taxonomy

MethodWhat it computesWhen to useFailure mode
BH (p.adjust(method='BH'), DESeq2 default pAdjustMethod='BH')FDR at fixed alpha; Benjamini-Hochberg 1995Default for most RNA-seq DEAssumes independence or PRDS; many overlapping tests violate
Storey q-value (qvalue::qvalue)q-value using estimated pi_0Genome-scale with many true nullsPi_0 estimation can fail at small test counts
IHW (results(filterFun=ihw), Ignatiadis 2016 Nat Methods 13:577)Weighted BH with covariate-informed weightsModern default for DESeq2; +5-20% discoveries at same FDRCovariate MUST be independent under null
ashr local false sign rate (lfsr / svalue, with svalue=TRUE)P(sign of estimate is wrong)When effect-direction certainty is what mattersConservative lower bound on FDR; not interchangeable with padj
BY (p.adjust(method='BY'))Benjamini-Yekutieli; arbitrary dependenceStrongly correlated testsUniformly conservative; rarely needed for DE
Holm / BonferroniFWERSmall confirmatory test setsFar too conservative for genome-scale
TREAT / lfcThreshold=FDR for "LFC> tau" hypothesis

Decision Tree by Scenario

ScenarioRecommended approachWhy
Standard bulk DE, two groupsDESeq2 results with default BH; report padj < 0.05Default works
Want more power at same FDRresults(dds, filterFun = ihw)IHW typically gains 5-20%
Pre-specified biological fold-change mattersresults(dds, lfcThreshold = log2(1.5), altHypothesis = 'greaterAbs') OR glmTreat(fit, lfc = log2(1.5))Post-hoc padj<0.05 & abs(LFC)>1 does NOT control FDR for the magnitude claim
Ranking for GSEA prerankedstat (Wald Z) for DESeq2 OR shrunken LFCNever use unshrunken LFC -- low-count noise dominates
ORA inputSubset by padj<0.05; background = ALL TESTED genes (post-independent-filtering)Background = "all genes in genome" is wrong; pre-filtering already excluded many
Many NA padj including biologically interesting genesDiagnose: independent filtering vs Cook's vs all-zero; turn off the offending filter only for that gene setBlanket na.omit discards signal
Multi-condition designLRT for "any change" first; pairwise per-level Wald for effect sizesLRT padj is omnibus; LRT LFC is one specific coefficient
Small n (<=3/group)Report as exploratory, top hits onlySchurch 2016: tools miss 20-40% of true positives at n=3
Human / mouse with mixed sexesInclude sex as covariate; run sex-stratified sensitivitySABV mandate; sex effect is real and chromosomal
ProkaryoticUse Prokka/Bakta GFF, KEGG strain codeEnsembl/org.db are eukaryote-only

Extracting Results

Goal: Pull DE estimates and p-values from a fitted DESeq2 or edgeR object into a usable data frame with explicit contrast naming.

Approach: results() (DESeq2) or topTags() (edgeR) with explicit name= / coef=; convert to data.frame; preserve row order if planning to join with annotation.

r
library(DESeq2)
library(dplyr)

resultsNames(dds)
res <- results(dds, name = 'condition_treated_vs_control', alpha = 0.05)
res_shrunk <- lfcShrink(dds, coef = 'condition_treated_vs_control', type = 'apeglm')

res_df <- as.data.frame(res)
res_df$gene <- rownames(res_df)
r
library(edgeR)
tt <- topTags(qlf, n = Inf, sort.by = 'none')$table
tt$gene <- rownames(tt)

sort.by = 'none' in topTags preserves the original gene order -- critical when joining with an annotation table by row index. Default is sort by p-value.

Column name reminder (a recurring cross-tool bug):

ToolLFC columnAdjusted p-value column
DESeq2log2FoldChangepadj
edgeRlogFCFDR
limma topTablelogFCadj.P.Val
limma topTreatlogFCadj.P.Val (post-TREAT)

TREAT vs Post-hoc LFC Filtering

Goal: Make a defensible FDR claim about "biologically meaningful fold change" genes.

Approach: Use TREAT or lfcThreshold= to test a magnitude hypothesis with proper FDR control. Post-hoc filtering of padj<0.05 & abs(LFC)>tau does NOT control FDR for the magnitude claim.

r
res_treat <- results(dds, lfcThreshold = log2(1.5), altHypothesis = 'greaterAbs', alpha = 0.05)

# edgeR equivalent
tr <- glmTreat(fit, coef = 2, lfc = log2(1.5))

What a reviewer is really probing with a "200 genes >2x changed at FDR 5%" claim: is the FDR for the change>2x claim or for the change-non-zero claim? Post-hoc filtering controls FDR only for the latter. TREAT (or lfcThreshold=) controls FDR for the former. McCarthy & Smyth 2009 Bioinformatics 25:765 is the canonical citation.

IHW for Better Power

Goal: Gain 5-20% more discoveries at the same FDR by weighting p-values with a covariate (typically baseMean) that informs power but is independent of the null.

Approach: results(dds, filterFun = ihw) replaces independent filtering with Ignatiadis 2016 hypothesis weighting.

r
library(IHW)
res_ihw <- results(dds, filterFun = ihw, alpha = 0.05)

When IHW does NOT help:

  • Small number of tests (<5000 after filtering) -- not enough data to learn weights
  • Covariate is treatment-correlated (violates the null-independence requirement)
  • Covariate uninformative about test power

Storey q-value as an alternative (different framework -- estimates pi_0 fraction of true nulls):

r
library(qvalue)
qv <- qvalue(res$pvalue[!is.na(res$pvalue)])
res$qvalue <- NA
res$qvalue[!is.na(res$pvalue)] <- qv$qvalues

ashr lfsr (local false sign rate -- probability the estimated direction is wrong):

r
res_ashr <- lfcShrink(dds, coef = 'condition_treated_vs_control',
                       type = 'ashr', svalue = TRUE)
res_ashr$svalue  # FDR-like, based on lfsr; requires svalue=TRUE

svalue=TRUE is required to populate the svalue column; the default returns the standard pvalue/padj columns only. lfsr and padj are NOT interchangeable. lfsr asks "P(sign wrong)"; padj asks "expected fraction of false discoveries". When reporting, state which.

P-value Histogram Diagnostics

Goal: Diagnose model misspecification, hidden batch effects, or over-correction by inspecting the raw p-value distribution.

Approach: Plot raw p-values; under a correctly specified null, the histogram is uniform with an upward spike near zero (the true DE genes).

r
library(ggplot2)
ggplot(res_df, aes(x = pvalue)) +
    geom_histogram(bins = 50, fill = 'steelblue', color = 'white') +
    labs(x = 'P-value', y = 'Frequency', title = 'P-value distribution') +
    theme_bw()
ShapeMeaningAction
Uniform + spike near 0Correct: null genes uniform, true DE near 0Proceed
Anti-conservative (U-shape; both ends spiked)Hidden batch effect, unmodeled confounder, dispersion misspecifiedInspect PCA for batch; add covariate; check plotDispEsts
Conservative (depleted near 0, spike near 1)Over-correction; too many covariates; wrong dispersionSimplify model; check dispersion plot for excess shrinkage
Spike only at p = 1Discrete artifact from very-low-count genesPre-filter more aggressively
Bimodal with spike at 0.5Unusual; suggests a discrete categorical test masqueradingInvestigate

The histogram is one of the cheapest sanity checks in a DE pipeline; always plot it before believing the gene list.

Filtering and Ordering

Goal: Subset to significant genes and rank by p-value, fold change, or expression level for downstream use.

Approach: dplyr-style filter + arrange; handle NA padj explicitly per the three-meanings table at the top.

r
sig <- res_df %>%
    filter(!is.na(padj), padj < 0.05, abs(log2FoldChange) > 1, baseMean > 10) %>%
    arrange(padj)

# Up- vs down-regulated
up   <- sig %>% filter(log2FoldChange > 0)
down <- sig %>% filter(log2FoldChange < 0)

# Summary
n_tested <- sum(!is.na(res$padj))
n_sig    <- sum(res$padj < 0.05, na.rm = TRUE)
cat(sprintf('Tested: %d   Significant (padj<0.05): %d   Up: %d   Down: %d\n',
            n_tested, n_sig, sum(sig$log2FoldChange > 0), sum(sig$log2FoldChange < 0)))

Gene Annotation

Goal: Map gene IDs to symbols, descriptions, and cross-database identifiers for human-readable results.

Approach: Prefer AnnotationDbi::mapIds with org.db (fast, local, version-pinned); fall back to biomaRt or mygene for symbols/aliases not in org.db; for prokaryotes, use Prokka/Bakta GFF.

r
library(org.Hs.eg.db)
library(AnnotationDbi)

res_df$symbol <- mapIds(org.Hs.eg.db, keys = sub('\\..*', '', res_df$gene),
                         keytype = 'ENSEMBL', column = 'SYMBOL', multiVals = 'first')
res_df$entrez <- mapIds(org.Hs.eg.db, keys = sub('\\..*', '', res_df$gene),
                         keytype = 'ENSEMBL', column = 'ENTREZID', multiVals = 'first')

The sub('\\..*', '', ...) strips the Ensembl version. CAUTION: this regex destroys the _PAR_Y suffix in GENCODE 25-43 PAR genes -- use sub('\\.[0-9]+(_PAR_Y)?$', '\\1', ...) to preserve. See expression-matrix/gene-id-mapping for full details.

For HGNC symbols changed since 2020 (SEPT1 -> SEPTIN1, MARCH1 -> MARCHF1, MARC1 -> MTARC1, DEC1 -> DELEC1) old symbol-keyed downstream tools silently drop genes. Always join on stable Ensembl or Entrez IDs; use symbols as display labels only.

For prokaryotes:

r
library(rtracklayer)
gff <- import('annotation.gff3')
gene_info <- as.data.frame(gff[gff$type == 'gene',
                                c('locus_tag', 'Name', 'product')])
res_annotated <- merge(res_df, gene_info, by.x = 'gene',
                        by.y = 'locus_tag', all.x = TRUE)

GSEA Preranked Input

Goal: Produce a ranked list of all genes (no significance filter) for fgsea / clusterProfiler GSEA.

Approach: Rank by Wald statistic (DESeq2 stat) or shrunken LFC. NEVER use a filtered set as GSEA input -- GSEA's permutation null requires the full background.

r
gsea_ranks <- res_df$stat
names(gsea_ranks) <- res_df$gene
gsea_ranks <- sort(gsea_ranks[!is.na(gsea_ranks)], decreasing = TRUE)

# edgeR equivalent
gsea_ranks_edger <- sign(tt$logFC) * -log10(tt$PValue)
names(gsea_ranks_edger) <- rownames(tt)
gsea_ranks_edger <- sort(gsea_ranks_edger[is.finite(gsea_ranks_edger)],
                         decreasing = TRUE)

stat (Wald Z) is preferred over raw LFC for GSEA because it combines effect and precision in one number. Unshrunken LFC is dominated by low-count noise.

Show full SKILL.md (925 more words)Show less

ORA Input

Goal: Run over-representation analysis (enrichGO, enrichKEGG) on a significant gene list with the correct background.

Approach: Subset to padj<0.05; background = ALL TESTED genes (post-independent-filtering), NOT the genome.

r
library(clusterProfiler)

sig_entrez <- na.omit(res_df$entrez[res_df$padj < 0.05])
bg_entrez  <- na.omit(res_df$entrez[!is.na(res_df$padj)])

ora <- enrichGO(gene          = sig_entrez,
                universe      = bg_entrez,
                OrgDb         = org.Hs.eg.db,
                keyType       = 'ENTREZID',
                ont           = 'BP',
                pAdjustMethod = 'BH')

Common mistake: omitting universe= lets clusterProfiler default to "all annotated genes for this organism" -- which includes thousands of genes never tested. The resulting enrichment p-values are wrong (too small). The background MUST be the tested set.

Cross-Tool Concordance Check

r
deseq2_sig <- rownames(subset(deseq2_res, padj < 0.05))
edger_sig  <- rownames(subset(edger_tt,  FDR  < 0.05))

common      <- intersect(deseq2_sig, edger_sig)
deseq2_only <- setdiff(deseq2_sig, edger_sig)
edger_only  <- setdiff(edger_sig, deseq2_sig)

cat(sprintf('DESeq2 sig: %d   edgeR sig: %d   Common: %d (%.1f%%)\n',
            length(deseq2_sig), length(edger_sig), length(common),
            100 * length(common) / min(length(deseq2_sig), length(edger_sig))))

Concordance >70% at the top 500: robust. <60%: suspect filtering, normalization, or design difference -- not a tool difference. Run both pipelines with the same filtering and design to isolate.

Per-Method Failure Modes

Dropped a key gene by removing NAs

Trigger: Pipeline does res_df <- na.omit(res_df); downstream gene of interest is missing from results.

Mechanism: Gene was flagged by independent filtering OR Cook's distance; padj is NA but the biology is real.

Symptom: A gene with clear differential expression in the count matrix is absent from the results table.

Fix: Diagnose which filter fired (independent filtering vs Cook's vs all-zero); rerun results() with the appropriate filter off (independentFiltering = FALSE or cooksCutoff = FALSE).

Reported FDR on a magnitude-filtered gene set

Trigger: Methods section says "genes with padj < 0.05 and abs(LFC) > 1 (FDR < 5%)".

Mechanism: BH controls FDR for the |LFC| > 0 hypothesis, not the |LFC| > 1 hypothesis. The post-hoc filter adds no FDR control.

Symptom: Reviewer challenges the FDR claim; replication studies show many of the filtered genes are not the magnitude expected.

Fix: Use TREAT (glmTreat) or lfcThreshold= to test the magnitude hypothesis with proper FDR control. Re-do the methods sentence to match what was actually computed.

ORA universe wrong

Trigger: ORA p-values look implausibly small for a small significant gene set.

Mechanism: universe= argument omitted; clusterProfiler defaulted to all annotated genes in the organism, including thousands never in the tested set.

Symptom: Many enriched pathways at strict thresholds; results don't replicate; reviewer questions the background.

Fix: Explicitly pass universe = bg_entrez where bg_entrez is the set of tested gene IDs (i.e., those with non-NA padj).

"Significant" gene list in n=3 study doesn't replicate

Trigger: Small RNA-seq study finds 200 DE genes; validation in independent cohort recovers 60.

Mechanism: Schurch 2016 RNA 22:839: at n=3/group, all tools miss 20-40% of true positives compared to n=30. Variability of the gene list itself is high.

Symptom: 30-50% replication of the gene list across independent runs of the SAME data.

Fix: Frame the small-n DE list as hypothesis-generating, not as a stable set of facts. Validate top hits orthogonally before drawing conclusions. Use TREAT for biologically meaningful thresholds to require larger effects.

Sex-confounded design gives spurious chrX/chrY signal

Trigger: Mixed-sex cohort; sex not in the design; many chrY genes call as DE.

Mechanism: Sex distribution differs across the experimental groups; the "treatment effect" partially captures sex.

Symptom: chrY genes (DDX3Y, RPS4Y1, UTY) and XIST dominate the top DE list.

Fix: Include sex in the design (~ sex + condition); rerun. For chrX/chrY-specific analyses, sex MUST be in the model or the analysis is uninterpretable. Mauvais-Jarvis et al. 2020 Lancet 396:565 reviews the SABV requirement.

Common errors

Error / symptomCauseFix
$FDR not found on DESeq2 resultDESeq2 uses padj; edgeR uses FDRCheck tool, use correct column
summary(res) shows different cutoff than results(alpha=)summary(res, alpha=) defaults to 0.1Pass alpha explicitly to summary()
All padj NAAll genes filtered (rare; usually a data problem)Check independentFilteringResults(res); inspect baseMean distribution
Direction of LFC reversedReference level not set; alphabetical defaultrelevel() BEFORE DESeq()
Gene symbol mapping rate <50%Mixed Ensembl versions; recent HGNC renamesVerify Ensembl release, check for SEPT/MARCH/MARC renames
enrichGO reports thousands of pathwaysWrong universe=Pass universe = bg_entrez (tested set, not genome)

References

  • Benjamini Y, Hochberg Y. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B 57(1):289-300. doi:10.1111/j.2517-6161.1995.tb02031.x
  • Storey JD. 2003. The positive false discovery rate: a Bayesian interpretation and the q-value. Ann Stat 31(6):2013-2035. doi:10.1214/aos/1074290335
  • Ignatiadis N, Klaus B, Zaugg JB, Huber W. 2016. Data-driven hypothesis weighting increases detection power in genome-scale multiple testing. Nat Methods 13(7):577-580. doi:10.1038/nmeth.3885
  • Bourgon R, Gentleman R, Huber W. 2010. Independent filtering increases detection power for high-throughput experiments. PNAS 107(21):9546-9551. doi:10.1073/pnas.0914005107
  • Stephens M. 2017. False discovery rates: a new deal. Biostatistics 18(2):275-294. doi:10.1093/biostatistics/kxw041
  • McCarthy DJ, Smyth GK. 2009. Testing significance relative to a fold-change threshold is a TREAT. Bioinformatics 25(6):765-771. doi:10.1093/bioinformatics/btp053
  • Schurch NJ et al. 2016. How many biological replicates are needed in an RNA-seq experiment and which differential expression tool should you use? RNA 22(6):839-851. doi:10.1261/rna.053959.115
  • Love MI, Huber W, Anders S. 2014. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15(12):550. doi:10.1186/s13059-014-0550-8
  • Mauvais-Jarvis F et al. 2020. Sex and gender: modifiers of health, disease, and medicine. Lancet 396(10250):565-582. doi:10.1016/S0140-6736(20)31561-0
  • Bruford EA et al. 2020. Guidelines for human gene nomenclature. Nat Genet 52:754-758. doi:10.1038/s41588-020-0669-3
  • Ziemann M, Eren Y, El-Osta A. 2016. Gene name errors are widespread in the scientific literature. Genome Biol 17:177. doi:10.1186/s13059-016-1044-7
  • Wu T et al. 2021. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. The Innovation 2(3):100141. doi:10.1016/j.xinn.2021.100141
  • deseq2-basics - Generate DESeq2 results; design, contrasts, LRT, shrinkage
  • edger-basics - Generate edgeR results; QL F-test, TREAT, voom
  • de-visualization - P-value histogram, MA plot, volcano with shrunken LFC, heatmap, sample distance
  • batch-correction - Include batch in design (vs Nygaard 2016 cardinal sin)
  • timeseries-de - LRT-with-reduced-model patterns for time
  • expression-matrix/gene-id-mapping - ID conversion, HGNC renames, ortholog mapping
  • expression-matrix/metadata-joins - Sex covariate, paired design, sample swap detection
  • pathway-analysis/go-enrichment - ORA with proper background
  • pathway-analysis/gsea - GSEA preranked input from stat or shrunken LFC
  • pathway-analysis/kegg-pathways - KEGG with strain-specific organism codes
  • data-visualization/volcano-and-ma-plots - Custom volcano with apeglm-shrunken LFC

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in differential-expression/de-results of GPTomics/bioSkills.

  • SKILL.md
  • examples/annotate_results.R
  • examples/export_excel.R
  • examples/filter_results.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Differential Expression De Results next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Differential Expression De Results compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Differential Expression De Results this skillGPTomics/bioSkills1.2k1 repos~5.7kAutomated safety check: PassMIT
Tooluniverse Gwas Study Explorerwu-yc/LabClaw1.1k2 repos~2.9kAutomated safety check: PassNone
Bio Metabolomics Normalization QcFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2.3kAutomated safety check: PassNone
Tg HubOpenMinis/MinisSkills446—~2.3kAutomated safety check: NotesMIT
Autodlopenshift-eng/ai-helpers120—~2.9kAutomated safety check: PassApache-2.0
Databrain Intelligenceinfometa/workbuddyskills348—~8kAutomated safety check: PassNone

Similar skills

  • Compare GWAS studies, perform meta-analyses, and assess replication across cohorts.

    1.1k GitHub starsUsed in 2 repos~2.9k tokens
    DatabasesAuto-check passed
  • Bio Metabolomics Normalization Qc

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control and normalization for metabolomics data. An agent skill from FreedomIntelligence/OpenClaw-Medical-Skills.

    3.1k GitHub starsUsed in 1 repo~2.3k tokens
    DatabasesAuto-check passed
  • Tg Hub

    OpenMinis/MinisSkills

    A skill for reading and writing Telegram data with Python and UV.

    446 GitHub stars~2.3k tokensUpdated 3 days ago
    DatabasesAuto-check: notes
  • Autodl

    openshift-eng/ai-helpers

    A skill your agent uses when querying auto-collected CI data from test runs in BigQuery (cidataautodl dataset) including risk analysis, disruption, CPU metrics, audit logs, operator state, and retry…

    120 GitHub stars~2.9k tokensUpdated 4 days ago
    DatabasesAuto-check passed
  • Databrain Intelligence

    infometa/workbuddyskills

    DataBrain intelligence data query assistant. An agent skill from infometa/workbuddyskills.

    348 GitHub stars~8k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Jmr Rebuttal

    franklee16/academic-research-skills

    Use after a Journal of Marketing Research (JMR) Revise-and-Resubmit — planning the revision and drafting a point-by-point response that satisfies two independent reviewers and the Coeditor…

    223 GitHub starsUsed in 1 repo~971 tokens
    DatabasesAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Differential Expression De Results

What does Bio Differential Expression De Results do?

Extracts, filters, annotates, and exports differential expression results from DESeq2 or edgeR with proper handling of padj=NA (independent filtering, Cook's outliers, all-zero), multiple-testing…. Bio Differential Expression De Results is an agent skill from GPTomics/bioSkills.db/biomaRt/mygene, GSEA preranked input, ORA background construction, replication reality (Schurch 2016 small-n result), and SABV/sex-stratified reporting.

When should I use Bio Differential Expression De Results?

Bio Differential Expression De Results fits situations like: extracting and interpreting DE results; troubleshooting padj=NA; choosing FDR method; preparing ranked lists for pathway analysis.

How do I install Bio Differential Expression De Results in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-de-results -a claude-code`. Or copy the skill folder (differential-expression/de-results in GPTomics/bioSkills) into .claude/skills/bio-differential-expression-de-results in your project. Claude Code loads it when a task matches its description.

How do I install Bio Differential Expression De Results in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-de-results -a codex`. Or copy the skill folder (differential-expression/de-results in GPTomics/bioSkills) into .agents/skills/bio-differential-expression-de-results in your project. Codex loads it when a task matches its description.

Can I use Bio Differential Expression De Results in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-de-results -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-differential-expression-de-results, .gemini/skills/bio-differential-expression-de-results, .github/skills/bio-differential-expression-de-results and .opencode/skills/bio-differential-expression-de-results in your project.

What does Bio Differential Expression De Results need to run?

Going by SKILL.md and its folder, Bio Differential Expression De Results needs R for the scripts in its folder and the command-line tools its instructions call (pip).

Does Bio Differential Expression De Results access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Differential Expression De Results safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Differential Expression De Results use?

Bio Differential Expression De Results is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Differential Expression De Results use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Differential Expression De Results?

Skills that share tags, products or a category with Bio Differential Expression De Results: Tooluniverse Gwas Study Explorer (wu-yc/LabClaw, 1.1k stars), Bio Metabolomics Normalization Qc (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Tg Hub (OpenMinis/MinisSkills, 446 stars) and Autodl (openshift-eng/ai-helpers, 120 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Differential Expression De Results?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.