Agent skill

Bio Differential Expression Batch Correction

by GPTomics in GPTomics/bioSkills

Handles batch effects in bulk RNA-seq via design-matrix inclusion (the correct path for DE), ComBat/ComBat-seq for visualization, SVA for unknown latent factors, RUVSeq for…

MITAuto-check passedResearch & Science

Install Bio Differential Expression Batch Correction

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-differential-expression-batch-correction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-differential-expression-batch-correction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/differential-expression/batch-correction .claude/skills/bio-differential-expression-batch-correction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-differential-expression-batch-correction
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.7k tokens
SKILL.md length
2,378 words
Files
3
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Handles batch effects in bulk RNA-seq via design-matrix inclusion (the correct path for DE), ComBat/ComBat-seq for visualization, SVA for unknown latent factors, RUVSeq for…

  • Designing a DE analysis with batch structure
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithmic Taxonomy and Decision Tree by Scenario, plus 13 more sections
  • Runs R scripts from its folder
  • Troubleshooting batch-dominated PCA

What it does

Bio Differential Expression Batch Correction is an agent skill from GPTomics/bioSkills. Handles batch effects in bulk RNA-seq via design-matrix inclusion (the correct path for DE), ComBat/ComBat-seq for visualization, SVA for unknown latent factors, RUVSeq for negative-control-gene-anchored unwanted variation, and limma::removeBatchEffect for plotting only. Encodes the Nygaard 2016 cardinal sin against testing on a batch-corrected matrix, the choice between SVA/RUVg/RUVs/RUVr, the confounding non-identifiability problem, the single-cell boundary (Harmony/MNN are NOT for bulk), and the Goh 2017…

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Data visualization. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Designing a DE analysis with batch structure
  • Troubleshooting batch-dominated PCA
  • Choosing ComBat vs ComBat-seq
  • Handling unknown batch via SVA

Example prompts

  • “Use the bio-differential-expression-batch-correction skill to handle batch effects in bulk RNA-seq via design-matrix inclusion (the correct path for…”
  • “/bio-differential-expression-batch-correction”

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Differential Expression Batch Correction loads about 5.7k tokens when it runs. Until then it costs about 202 tokens; SKILL.md has 2,378 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~202
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,378 words, ~5,661 tokens.

Download SKILL.mdSave it as .claude/skills/bio-differential-expression-batch-correction/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-differential-expression-batch-correction
description
Handles batch effects in bulk RNA-seq via design-matrix inclusion (the correct path for DE), ComBat/ComBat-seq for visualization, SVA for unknown latent factors, RUVSeq for negative-control-gene-anchored unwanted variation, and limma::removeBatchEffect for plotting only. Encodes the Nygaard 2016 cardinal sin against testing on a batch-corrected matrix, the choice between SVA/RUVg/RUVs/RUVr, the confounding non-identifiability problem, the single-cell boundary (Harmony/MNN are NOT for bulk), and the Goh 2017 harmonization critique. Use when designing a DE analysis with batch structure, troubleshooting batch-dominated PCA, choosing ComBat vs ComBat-seq, handling unknown batch via SVA, integrating across studies, or deciding when (rarely) to subtract batch.
tool_type
r
primary_tool
sva

Version Compatibility

Reference examples tested with: sva 3.50+ (includes ComBat + ComBat_seq), DESeq2 1.42+, edgeR 4.0+, limma 3.58+, RUVSeq 1.36+, ggplot2 3.5+, harmony 1.2+ (single-cell context only)

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Batch Effect Correction

"Remove the batch effect before DE" -> Almost always WRONG. Include batch as a covariate in the design formula (~ batch + condition) so DESeq2/edgeR/limma model it without subtracting. Subtraction is for visualization only.

The Single Most Important Modern Insight -- The Nygaard 2016 cardinal sin

Nygaard, Rødland, Hovig 2016 Biostatistics 17(1):29-39, "Methods that remove batch effects while retaining group differences may lead to exaggerated confidence in downstream analyses." Translation: never run ComBat (or ComBat-seq, or removeBatchEffect, or SVA-subtract-then-test) and then run DE on the corrected matrix.

Mechanism: batch-correction methods fit a model y_ij = alpha + X_ij beta + gamma_i + delta_i epsilon_ij and subtract the batch terms. The downstream DE test then computes p-values as if those degrees of freedom had never been spent. Residual df is lower than what DESeq2/edgeR/limma assume. Type-I error inflates -- the gene list looks more significant than it should.

The right approach: include batch in the design. ~ batch + condition. DESeq2/edgeR/limma will properly partial out the batch effect from the condition estimate and account for the spent degrees of freedom in the inference. The batch is corrected at the inference stage, not by mutating counts.

removeBatchEffect (limma) is for visualization only -- the function's help page says so. ComBat/ComBat-seq output is for visualization, clustering, or downstream tools that cannot take a design matrix (rare; mostly ML).

A second clarification: structural confounding (every "treated" sample in batch 1, every "control" in batch 2) is non-identifiable. No method fixes this -- the "treatment effect" and "batch effect" are mathematically the same vector. The fix is experimental design (randomize batches in advance). ComBat/SVA on a fully-confounded design silently removes the treatment effect along with the batch effect.

Algorithmic Taxonomy

MethodInputMechanismUse for
Design-matrix inclusion (~ batch + condition)Raw counts; known batchPartial out batch in the GLM, spent df accountedDE testing -- the correct path
ComBat (Johnson, Li, Rabinovic 2007)Log-transformed / continuous expressionEmpirical-Bayes location and scale shifts per batchVisualization of microarray / continuous data
ComBat-seq (Zhang, Parmigiani, Johnson 2020)Raw RNA-seq countsNB-GLM equivalent of ComBat; returns integer countsVisualization of RNA-seq counts; cross-study harmonization for ML (with caveat)
limma::removeBatchEffectNormalized expressionLinear regression subtraction with design protectionVisualization only -- explicit help-page warning
SVA (Leek, Storey 2007 + Leek 2012)Normalized expressionEstimate latent surrogate variables explaining residual variance independent of variable-of-interestUnknown batch / hidden technical structure -- add SVs to design
svaseq (Leek 2014)CountsCount-data variant of SVACounts version
RUVg (Risso, Ngai, Speed, Dudoit 2014)Counts; negative control genesFactor analysis of control-gene residual; W as covariateStrong negative controls (ERCC, housekeeping) -- add W to design
RUVsCounts; replicate samplesFactor analysis within replicate groups; assumes replicates differ only by unwanted variationMulti-condition with biological replicates as "controls"
RUVrCounts; design onlyFactor analysis of residuals from initial fitMost data-driven; most likely to absorb biology -- caution
Harmony, MNN, Scanorama, BBKNNSingle-cell embeddingsIterative alignment of clusters across samplesSingle-cell ONLY; not bulk

Decision Tree by Scenario

ScenarioRecommended approachWhy
Known batch, want DE~ batch + condition in design; do NOT subtractCardinal sin avoidance
PCA shows batch separationInclude batch in design; for the figure, removeBatchEffect is OK (visualization only)Two purposes, two tools
Unknown batch structuresva / svaseq; add SVs as covariates in designCaptures latent technical factors
Have ERCC spike-ins or trusted housekeepingRUVSeq::RUVg with control gene indices; add W to designMost principled UV removal
Have replicate samples (technical reps within biological)RUVSeq::RUVsReplicate structure indicates "this differs only by UV"
Cross-study integration (TCGA + ICGC + own data)ComBat-seq for counts, ComBat for log-expression, THEN meta-analysis -- do NOT pool then DEGoh 2017 warning
Fully confounded batch and conditionRe-collect samples; no method fixes thisNon-identifiable
Visualizing batch removal for a figureremoveBatchEffect(expr, batch = batch, design = model.matrix(~condition))Visualization is what it's for
Single-cell dataHarmony / MNN / Scanorama in the single-cell category, not hereDifferent problem

Standard Workflow -- Design-Matrix Inclusion

Goal: Test the condition effect while accounting for known batch structure with proper degrees-of-freedom accounting.

Approach: Include batch as a covariate in the design formula; DESeq2/edgeR/limma will partial it out and compute correct p-values.

r
library(DESeq2)

dds <- DESeqDataSetFromMatrix(countData = counts, colData = coldata,
                               design = ~ batch + condition)
dds <- DESeq(dds)
res <- results(dds, name = 'condition_treated_vs_control')
r
library(edgeR)
y <- DGEList(counts = counts, group = coldata$condition)
keep <- filterByExpr(y, design = model.matrix(~ batch + condition, coldata))
y <- y[keep, , keep.lib.sizes = FALSE]
y <- normLibSizes(y)

design <- model.matrix(~ batch + condition, coldata)
y <- estimateDisp(y, design, robust = TRUE)
fit <- glmQLFit(y, design, robust = TRUE)
qlf <- glmQLFTest(fit, coef = 'conditiontreated')

When known confounders are continuous (RIN, library prep date as days), include them as continuous covariates -- no batch correction needed for those:

r
design = ~ RIN + library_prep_day + condition

ComBat-seq (Visualization or Cross-Study Counts)

Goal: Adjust raw counts to remove known batch effects while preserving biology, for VISUALIZATION or for downstream tools that need a single corrected matrix.

Approach: ComBat_seq(counts, batch, group) returns batch-adjusted integer counts via NB-GLM. Use the output for PCA, clustering, ML -- NOT for DE testing.

r
library(sva)

corrected_counts <- ComBat_seq(counts = as.matrix(counts),
                                batch = coldata$batch,
                                group = coldata$condition,
                                full_mod = TRUE)

vsd_corrected <- vst(DESeqDataSetFromMatrix(corrected_counts, coldata, ~1))
plotPCA(vsd_corrected, intgroup = 'condition')

full_mod = TRUE keeps biological covariates protected (the group argument). With full_mod = FALSE, ComBat-seq removes batch AND group differences -- a serious failure mode.

ComBat (the non-seq version) is for log-transformed or microarray data, NOT raw counts. Running ComBat on raw counts produces fractional values and assumes Gaussian residuals where counts are NB. Wrong tool, silent failure.

SVA -- Unknown Batch

Goal: Discover latent technical factors when batch is not recorded, add them as covariates.

Approach: Estimate the number of surrogate variables; estimate the SVs themselves; add them to the design and re-run DE.

r
library(sva)
library(DESeq2)

dds <- DESeqDataSetFromMatrix(counts, coldata, design = ~ condition)
dds <- estimateSizeFactors(dds)
norm_counts <- counts(dds, normalized = TRUE)

mod  <- model.matrix(~ condition, coldata)
mod0 <- model.matrix(~ 1, coldata)

n_sv <- num.sv(norm_counts, mod, method = 'leek')
svobj <- svaseq(norm_counts, mod, mod0, n.sv = n_sv)

for (i in seq_len(ncol(svobj$sv))) {
    colData(dds)[[paste0('SV', i)]] <- svobj$sv[, i]
}
sv_formula <- as.formula(paste('~', paste(paste0('SV', seq_len(ncol(svobj$sv))),
                                          collapse = ' + '), '+ condition'))
design(dds) <- sv_formula
dds <- DESeq(dds)

CRITICAL CHECK: compute the correlation between each SV and the variable of interest. If any SV correlates with treatment at r > 0.3, do NOT include it -- doing so would partial out biology, deflating the effect of interest.

r
sapply(seq_len(ncol(svobj$sv)),
       function(i) cor(svobj$sv[, i], as.numeric(coldata$condition)))

svaseq is the count-data version (Leek 2014); sva is for log-transformed / microarray data.

RUVSeq -- Negative-Control-Anchored UV

Goal: Estimate unwanted variation using ERCC spike-ins, housekeeping genes, replicate samples, or post-fit residuals; add as covariates in design.

Approach: Pick the RUV variant matching the available controls; add the resulting W matrix to the design.

r
library(RUVSeq)
library(edgeR)

control_idx <- which(rownames(counts) %in% ercc_genes)

set <- newSeqExpressionSet(as.matrix(counts), phenoData = coldata)
set <- betweenLaneNormalization(set, which = 'upper')

ruv <- RUVg(set, control_idx, k = 2)

design <- model.matrix(~ pData(ruv)$W_1 + pData(ruv)$W_2 + condition,
                        data = coldata)
y <- DGEList(counts = counts, group = coldata$condition)
y <- normLibSizes(y)
y <- estimateDisp(y, design, robust = TRUE)
fit <- glmQLFit(y, design, robust = TRUE)
qlf <- glmQLFTest(fit, coef = 'conditiontreated')
VariantControl sourceMost appropriate whenFailure mode
RUVgNegative control genes (ERCC, housekeeping)Strong, validated controls existBad controls -> partials out biology
RUVsReplicate samples / technical repsMulti-condition with reps that should differ only by UVWrong replicate structure absorbs condition effect
RUVrResiduals from an initial GLMNo external controls; most data-drivenMost likely to absorb true biology -- use last

Choice of k (number of unwanted factors): no automatic procedure. Try k = 1, 2, 3 and inspect PCA after RUV correction. Treatment effect should remain visible; if it disappears, k is too high. 5-10 is typical; >10 is suspicious.

removeBatchEffect -- Visualization Only

r
library(limma)

design <- model.matrix(~ condition, coldata)
log_expr_corrected <- removeBatchEffect(log_expr,
                                         batch = coldata$batch,
                                         design = design)

library(ggplot2)
pca <- prcomp(t(log_expr_corrected), scale. = TRUE)
ggplot(data.frame(PC1 = pca$x[,1], PC2 = pca$x[,2],
                  condition = coldata$condition, batch = coldata$batch),
       aes(PC1, PC2, color = condition, shape = batch)) +
    geom_point(size = 3) +
    ggtitle('After removeBatchEffect (visualization only)')

The design argument protects condition while removing batch. Common error: passing batch ALSO in the design ("Coefficients not estimable" warning) -- pass batch via batch= only, NOT also in design=.

DO NOT feed log_expr_corrected back into limma lmFit for DE testing. The function's own help page says so. The right approach: lmFit(log_expr, model.matrix(~ batch + condition, coldata)) -- include batch as a covariate.

Detecting Confounding Early

r
ct <- table(coldata$condition, coldata$batch)
ct
PatternStatusAction
All cells > 0 with roughly equal proportionsBalancedInclude batch as covariate
Some cells low (e.g., 4/4/2/0)Partially confoundedInclude batch; report reduced power
Some cells zero AND one factor entirely in one batchPerfectly confoundedUNFIXABLE -- re-collect or drop the affected comparison

alias() reveals collinear columns in a design matrix:

r
alias(model.matrix(~ batch + condition, coldata))$Complete

The Single-Cell Boundary

Harmony (Korsunsky et al. 2019 Nat Methods 16:1289), MNN (Haghverdi et al. 2018 Nat Biotechnol 36:421), Scanorama, BBKNN are designed for single-cell integration where the goal is aligning cell-type clusters across samples/batches. They modify the embedding (PCA coordinates) for downstream UMAP/clustering -- they assume the structural alignment problem of single-cell (many similar cells; align clusters).

DO NOT apply to bulk RNA-seq. Bulk lacks the cluster structure these methods assume.

Conversely, DO NOT use ComBat / ComBat-seq for single-cell. ComBat assumes much more homogeneity than scRNA exhibits; it does not handle zeros well.

See single-cell/batch-integration for the single-cell methods.

The Goh 2017 Critique on "Harmonization"

Goh, Wang, Wong 2017 Trends Biotechnol 35(6):498-507 warn that batch-correction methods can:

  • Introduce NEW batch-like artifacts (false positives) when the batch structure is unclear
  • Fail when the "most genes not DE" assumption is violated (heat shock, immune activation, viral host shutoff)
  • Inflate downstream confidence (echoing Nygaard 2016)

Their recommendation for cross-study work: per-cohort DE first, then meta-analyze the effect sizes (metafor, limma treat across cohorts). DO NOT pool data, harmonize, then re-test.

"Harmonization" in clinical genomics often means ComBat across cohorts then DE on the harmonized matrix -- the exact failure mode Nygaard 2016 describes. Same problem, different vocabulary.

Show full SKILL.md (960 more words)Show less

Reconciliation: When Methods Disagree

PatternLikely causeAction
Design-inclusion (~ batch + condition) and SVA give different gene listsSVA captured biology (high SV-condition correlation) OR known batch was incompleteCheck SV-vs-condition correlations; if any
Design-inclusion and RUVg disagreeRUVg control genes (ERCC, housekeeping) had real biological variationValidate control set; switch RUVg -> RUVs (replicate-based) or back to design only
ComBat-seq corrected gene list larger than design-included gene listNygaard 2016 cardinal sin: spent df not accounted for in downstream DE on corrected matrixDISCARD the ComBat-seq DE result; report design-included only
Cross-cohort meta-analysis shows different DE per cohortReal biological heterogeneity OR cohort-specific technical driftPer-cohort DE then meta-analyze effect sizes (metafor); do NOT pool then test (Goh 2017)
All methods (design, SVA, RUVg) agree on top 100 genesRobust signalReport the intersection as high-confidence
All methods disagreeConfounded design OR weak signal at high noise floorSuspect non-identifiability; verify alias() and balance

The design-inclusion result (~ batch + condition or ~ confounders + condition) is the reference. SVA/RUV results are sensitivity analyses; when they agree with the design result, confidence rises. When they diverge, investigate before believing either.

Per-Method Failure Modes

Tested on a batch-corrected matrix -- inflated significance

Trigger: Pipeline: ComBat_seq(counts, batch) -> DESeq2 on corrected_counts. Many more "significant" genes than expected.

Mechanism: ComBat-seq removed batch terms; DESeq2 then computed inference as if those df had not been spent. Type-I error inflates.

Symptom: Implausibly long DE gene list; replication in independent cohort recovers <50%; p-value histogram leans anti-conservative.

Fix: Re-run with ~ batch + condition design on RAW counts. Discard the corrected-counts DE result.

SVA captured the biology

Trigger: sva returned 5 SVs; SV2 correlates with condition at r = 0.7; user added all SVs to design.

Mechanism: When biology is correlated with the hidden factor (e.g., disease severity drives both expression AND blood-draw timing), SVA's SVs capture biology along with technical noise. Partialling them out deflates the effect of interest.

Symptom: Condition effect that was clear in raw PCA disappears after SV adjustment; few or no DE genes.

Fix: Compute SV-vs-condition correlations; exclude any SV with |r| > 0.3 from the design. Re-fit.

ComBat applied to raw counts

Trigger: Pipeline uses ComBat(counts, batch) (not ComBat_seq); output is fractional.

Mechanism: ComBat is for log-transformed / Gaussian data. Counts are NB. The Gaussian assumption is wrong and the output is uninterpretable as counts.

Symptom: Corrected matrix has fractional values; DESeq2 errors out ("counts matrix should be integers"); naive round() gives garbage.

Fix: Use ComBat_seq() for counts. Or better, include batch as a design covariate.

Confounded design corrected with SVA -- biology gone

Trigger: All treated samples in batch 1, all control in batch 2; user runs SVA hoping it will rescue the design.

Mechanism: Batch and treatment vectors are identical (or nearly so). SVA estimates "the unwanted factor"; "the unwanted factor" is treatment.

Symptom: SV1 correlates with treatment at r ~ 1; after adjustment, no DE genes.

Fix: Acknowledge the design is non-identifiable. No statistical method fixes structural confounding. Re-collect with randomized batches.

Used Harmony on bulk RNA-seq

Trigger: Bulk RNA-seq with batch effect; user reaches for Harmony because they've used it for single-cell.

Mechanism: Harmony aligns cluster centroids in single-cell embeddings. Bulk has no cluster structure -- typically 6-30 samples, not 10000+ cells.

Symptom: Harmony "succeeds" but the result is meaningless; downstream DE is nonsense.

Fix: Use design-matrix inclusion (~ batch + condition) or ComBat-seq for visualization. Harmony belongs to single-cell/batch-integration.

Common errors

Error / symptomCauseFix
Coefficients not estimable from removeBatchEffectBatch included in both batch= and design=Pass batch via batch= only
ComBat output has fractional valuesWrong tool for countsUse ComBat_seq
SV adjustment kills condition effectSVs captured biologyCheck correlation; exclude high-correlation SVs
Inconsistent dimensions in RUVgcontrol_idx is symbols vs ENSEMBL rownamesMatch index type
full_mod not specified in ComBat-seqDefaults; biology may not be protectedSet full_mod = TRUE and pass group =

References

  • Nygaard V, Rødland EA, Hovig E. 2016. Methods that remove batch effects while retaining group differences may lead to exaggerated confidence in downstream analyses. Biostatistics 17(1):29-39. doi:10.1093/biostatistics/kxv027
  • Johnson WE, Li C, Rabinovic A. 2007. Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics 8(1):118-127. doi:10.1093/biostatistics/kxj037
  • Zhang Y, Parmigiani G, Johnson WE. 2020. ComBat-seq: batch effect adjustment for RNA-seq count data. NAR Genom Bioinform 2(3):lqaa078. doi:10.1093/nargab/lqaa078
  • Leek JT, Storey JD. 2007. Capturing heterogeneity in gene expression studies by surrogate variable analysis. PLoS Genet 3(9):e161. doi:10.1371/journal.pgen.0030161
  • Leek JT, Johnson WE, Parker HS, Jaffe AE, Storey JD. 2012. The sva package for removing batch effects and other unwanted variation in high-throughput experiments. Bioinformatics 28(6):882-883. doi:10.1093/bioinformatics/bts034
  • Leek JT. 2014. svaseq: removing batch effects and other unwanted noise from sequencing data. Nucleic Acids Res 42(21):e161. doi:10.1093/nar/gku864
  • Risso D, Ngai J, Speed TP, Dudoit S. 2014. Normalization of RNA-seq data using factor analysis of control genes or samples. Nat Biotechnol 32(9):896-902. doi:10.1038/nbt.2931
  • Goh WWB, Wang W, Wong L. 2017. Why batch effects matter in omics data, and how to avoid them. Trends Biotechnol 35(6):498-507. doi:10.1016/j.tibtech.2017.02.012
  • Korsunsky I et al. 2019. Fast, sensitive and accurate integration of single-cell data with Harmony. Nat Methods 16(12):1289-1296. doi:10.1038/s41592-019-0619-0
  • Haghverdi L, Lun ATL, Morgan MD, Marioni JC. 2018. Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors. Nat Biotechnol 36(5):421-427. doi:10.1038/nbt.4091
  • Ritchie ME et al. 2015. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res 43(7):e47. doi:10.1093/nar/gkv007
  • deseq2-basics - DE with ~ batch + condition design
  • edger-basics - edgeR pipeline with batch in design
  • de-results - p-value histogram diagnostics for missed batch
  • de-visualization - PCA shows batch effects; sample distance heatmap; figure-time use of removeBatchEffect
  • expression-matrix/metadata-joins - Confounding detection; sample swap detection
  • expression-matrix/normalization - TMM/RLE failure modes overlap with batch failure modes
  • single-cell/batch-integration - Harmony, MNN, Scanorama for single-cell (not bulk)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in differential-expression/batch-correction of GPTomics/bioSkills.

  • SKILL.md
  • examples/batch_correction_workflow.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Differential Expression Batch Correction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Differential Expression Batch Correction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Differential Expression Batch Correction this skillGPTomics/bioSkills1.2k1 repos~5.7kAutomated safety check: PassMIT
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
deepTools NGS Toolkitdavila7/claude-code-templates32k13 repos~4.5kAutomated safety check: PassMIT
FBA Flux Analyzeraiming-lab/AutoResearchClaw15k—~2.3kAutomated safety check: PassMIT
Ukb Ppp Region FetchClawBio/ClawBio1.2k—~4.6kAutomated safety check: PassMIT
Bio Hi C Analysis Hic VisualizationFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2.2kAutomated safety check: PassNone

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • deepTools NGS Toolkit

    davila7/claude-code-templates

    Guides use of deepTools on sequencing data: BAM to bigWig conversion, QC, sample correlation, and heatmaps or profiles around TSS and peaks for ChIP-seq, RNA-seq and ATAC-seq.

    32k GitHub starsUsed in 13 repos~4.5k tokens
    Research & ScienceAuto-check passed
  • FBA Flux Analyzer

    aiming-lab/AutoResearchClaw

    Turns raw flux balance analysis output and a COBRApy model into gene essentiality maps, phenotypic phase planes, flux sampling results, pathway summaries and secretion predictions.

    15k GitHub stars~2.3k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Ukb Ppp Region Fetch

    ClawBio/ClawBio

    Fetch a regional slice of plasma pQTL summary statistics from the UK Biobank Pharma Proteomics Project (UKB-PPP; Sun 2023 Nature) for a specific (protein, ancestry) measurement.

    1.2k GitHub stars~4.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Bio Hi C Analysis Hic Visualization

    FreedomIntelligence/OpenClaw-Medical-Skills

    Visualize Hi-C contact matrices, TADs, loops, and genomic features using matplotlib, cooltools, and HiCExplorer.

    3.1k GitHub starsUsed in 1 repo~2.2k tokens
    Research & ScienceAuto-check passed
  • Experiment Lab

    Citrus-bit/Anaxa

    A skill your agent uses whenever the user wants reproducible CS/AI experiments, model evaluation, regression/classification/clustering analyses, bioinformatics workflows, QC, differential…

    120 GitHub stars~2.1k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Differential Expression Batch Correction

What does Bio Differential Expression Batch Correction do?

Handles batch effects in bulk RNA-seq via design-matrix inclusion (the correct path for DE), ComBat/ComBat-seq for visualization, SVA for unknown latent factors, RUVSeq for…. Bio Differential Expression Batch Correction is an agent skill from GPTomics/bioSkills. Handles batch effects in bulk RNA-seq via design-matrix inclusion (the correct path for DE), ComBat/ComBat-seq for visualization, SVA for unknown latent factors, RUVSeq for negative-control-gene-anchored unwanted variation, and limma::removeBatchEffect for plotting only.

When should I use Bio Differential Expression Batch Correction?

Bio Differential Expression Batch Correction fits situations like: designing a DE analysis with batch structure; troubleshooting batch-dominated PCA; choosing ComBat vs ComBat-seq; handling unknown batch via SVA.

How do I install Bio Differential Expression Batch Correction in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-batch-correction -a claude-code`. Or copy the skill folder (differential-expression/batch-correction in GPTomics/bioSkills) into .claude/skills/bio-differential-expression-batch-correction in your project. Claude Code loads it when a task matches its description.

How do I install Bio Differential Expression Batch Correction in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-batch-correction -a codex`. Or copy the skill folder (differential-expression/batch-correction in GPTomics/bioSkills) into .agents/skills/bio-differential-expression-batch-correction in your project. Codex loads it when a task matches its description.

Can I use Bio Differential Expression Batch Correction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-batch-correction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-differential-expression-batch-correction, .gemini/skills/bio-differential-expression-batch-correction, .github/skills/bio-differential-expression-batch-correction and .opencode/skills/bio-differential-expression-batch-correction in your project.

What does Bio Differential Expression Batch Correction need to run?

Going by SKILL.md and its folder, Bio Differential Expression Batch Correction needs R for the scripts in its folder.

Does Bio Differential Expression Batch Correction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Differential Expression Batch Correction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Differential Expression Batch Correction use?

Bio Differential Expression Batch Correction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Differential Expression Batch Correction use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Differential Expression Batch Correction?

Skills that share tags, products or a category with Bio Differential Expression Batch Correction: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), deepTools NGS Toolkit (davila7/claude-code-templates, 32k stars), FBA Flux Analyzer (aiming-lab/AutoResearchClaw, 15k stars) and Ukb Ppp Region Fetch (ClawBio/ClawBio, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Differential Expression Batch Correction?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.