Agent skill

Batch Effect Correction

by aipoch in aipoch/medical-research-skills

A skill your agent uses when correcting batch effects in merged bulk expression matrices with sample-level batch metadata while preserving biological group structure and generating before-and-after…

MITAuto-check passedResearch & Science

Install Batch Effect Correction

skills CLI
$ npx skills add aipoch/medical-research-skills --skill batch-effect-correction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills batch-effect-correction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'awesome-med-research-skills/Data Analysis/batch-effect-correction' .claude/skills/batch-effect-correction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
batch-effect-correction
GitHub stars
2k
Token cost
~2.8k tokens
SKILL.md length
1,074 words
Files
18 (incl. scripts, references)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when correcting batch effects in merged bulk expression matrices with sample-level batch metadata while preserving biological group structure and generating before-and-after…

  • Works in 4 steps: Validate Input → Align and Prepare Matrix → Run Batch Correction → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Prerequisites, When to Read External Files, Usage and Arguments, plus 9 more sections
  • Runs R scripts from its folder; reaches cloud.r-project.org

What it does

Batch Effect Correction is an agent skill from aipoch/medical-research-skills. Use when correcting batch effects in merged bulk expression matrices with sample-level batch metadata while preserving biological group structure and generating before-and-after QC plots. NOT for: single-cell integration, raw FASTQ processing, differential expression without batch labels, or datasets without biological groups.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 22 other files, including scripts and reference files (for example `eval_report_batch-effect-correction_result.json`, `references/algorithm.md` and `references/cli-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/batch-effect-correction”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Validate Input
  2. Align and Prepare Matrix
  3. Run Batch Correction
  4. Normalize and Export Results

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (R, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • cloud.r-project.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Batch Effect Correction loads about 2.8k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 88 tokens; SKILL.md has 1,074 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 1,074 words, ~2,757 tokens.

Download SKILL.mdSave it as .claude/skills/batch-effect-correction/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
batch-effect-correction
description
Use when correcting batch effects in merged bulk expression matrices with sample-level batch metadata while preserving biological group structure and generating before-and-after QC plots. NOT for: single-cell integration, raw FASTQ processing, differential expression without batch labels, or datasets without biological groups.
license
MIT
skill-author
AIPOCH

Batch Effect Correction

Prerequisites

Run the following before the first analysis to install all required R packages:

bash
Rscript -e "if (!require('BiocManager', quietly=TRUE)) install.packages('BiocManager'); BiocManager::install(c('sva', 'limma')); install.packages('ggplot2', repos='https://cloud.r-project.org')"

Note: sva and limma are Bioconductor packages and require BiocManager for installation. ggplot2 is a standard CRAN package.

The skill cannot run until these packages are installed. In new or bare R environments, always run the prerequisite step first.


When to Read External Files

SituationFile to ReadPurpose
Need algorithm detailsreferences/algorithm.mdComBat workflow, assumptions, and QC logic
Need to run analysisscripts/main.RExecute: Rscript scripts/main.R --input_file ... --group_file ...
Encounter errorsreferences/troubleshooting.mdCommon errors and solutions
Need CLI examplesreferences/cli-guide.mdDetailed CLI usage examples and baseline run record
Need test datatests/data/Sample input files for testing

Usage

bash
Rscript scripts/main.R \
  --input_file ./expression_matrix.csv \
  --group_file ./sample_info.csv \
  --output_dir ./output/ \
  --batch_column batch \
  --group_column group \
  --sample_column sample \
  --log_transform auto \
  --timeout_seconds 600 \
  --seed 42

Arguments

ShortLongTypeDefaultDescription
-i--input_filecharacterrequiredExpression matrix file (genes as rows, samples as columns)
-g--group_filecharacterrequiredSample metadata file (sample ID, group, and batch columns)
-o--output_dircharacter./output/Output directory
-b--batch_columncharacterbatchBatch column name in metadata
-c--group_columncharactergroupBiological group column name in metadata
-n--sample_columncharactersampleSample ID column name in metadata
-l--log_transformcharacterautoLog transform mode: auto, yes, no
-t--timeout_secondsinteger600Elapsed time limit in seconds; use 0 to disable
-s--seedinteger42Random seed for reproducibility

Input Format

Expression Matrix (input_file)

Genes as rows, samples as columns, CSV format with gene ID in the first column.

csv
"","Sample01","Sample02","Sample03"
"GeneA",5.12,4.87,6.03
"GeneB",8.44,8.11,7.95

Requirements:

  • Gene IDs must be unique and non-empty
  • Sample column names must be unique and non-empty
  • Expression values must be numeric and finite
  • Extra expression-matrix sample columns not present in metadata are allowed and will be ignored with a warning
Sample Metadata (group_file)

CSV with sample ID, biological group, and batch columns.

csv
"sample","group","batch"
"Sample01","Control","Batch1"
"Sample02","Case","Batch1"
"Sample03","Case","Batch2"

Requirements:

  • Sample IDs must be unique and non-empty
  • At least 2 biological groups are required
  • At least 2 batches are required
  • Each group and each batch must contain at least 2 samples
  • Metadata may describe a subset of expression-matrix samples; the analysis will keep only metadata-matched samples and warn about ignored expression columns

Output Files

FileDescription
corrected_expression_matrix.csvBatch-corrected expression matrix
matched_sample_info.csvStandardized metadata used in the analysis
batch_before_boxplot.pdfSample distribution boxplot before correction
batch_after_boxplot.pdfSample distribution boxplot after correction
batch_before_pca.pdfPCA scatter plot before correction with batch-colored points
batch_after_pca.pdfPCA scatter plot after correction with batch-colored points
batch_before_clustering.pdfHierarchical clustering before correction
batch_after_clustering.pdfHierarchical clustering after correction
session_info.txtR session and package version info

Workflow

Step 1: Validate Input
  • Check file existence and non-empty input files
  • Validate metadata column presence
  • Verify expression values are numeric and finite
  • Confirm at least 2 groups, 2 batches, and at least 2 samples per group/batch
Step 2: Align and Prepare Matrix
  • Reorder expression columns to match metadata sample order
  • Keep only metadata-matched samples; warn if the expression matrix contains extra samples absent from metadata
  • Decide whether log transformation is needed (auto, yes, or no)
  • Apply log2(x + 1) only when required
Step 3: Run Batch Correction
  • Build the design matrix with biological group information
  • Run sva::ComBat() to remove batch-driven variation
  • Preserve modeled biological group structure during correction
Step 4: Normalize and Export Results
  • Apply limma::normalizeBetweenArrays() after ComBat
  • Write the corrected matrix and matched metadata
  • Save before/after QC plots and session information

Methods

ComBat

Empirical Bayes batch-effect correction using sva::ComBat(). Recommended when merged bulk expression datasets contain known batch labels and at least two biological groups.

Log Transformation

Supports auto, yes, and no. The auto mode applies log2(x + 1) only when the matrix appears to be on a raw-like scale.

normalizeBetweenArrays

Post-correction normalization with limma::normalizeBetweenArrays() to reduce remaining cross-sample distribution differences.

QC Visualization

Generates paired boxplots, PCA scatter plots with conditional batch ellipses, and hierarchical clustering plots before and after correction to assess whether batch-driven structure is reduced.


Show full SKILL.md (456 more words)Show less

Agent Response Contract

After a successful run, report:

  1. Sample count retained after metadata matching and any subset filtering
  2. Batch count and group count used in the ComBat design matrix
  3. Log transformation applied (auto-detected, forced yes, or skipped)
  4. QC assessment: describe whether before/after PCA plots show reduced batch clustering
  5. Artifact paths: corrected_expression_matrix.csv, batch_after_pca.pdf, batch_after_clustering.pdf

Examples

Basic Usage
bash
Rscript scripts/main.R \
  -i expression_matrix.csv \
  -g sample_info.csv \
  -o ./output
With Custom Metadata Columns
bash
Rscript scripts/main.R \
  -i expression_matrix.csv \
  -g metadata.csv \
  -o ./output \
  -n sample_id \
  -c condition \
  -b platform_batch
Disable Log Transform and Timeout
bash
Rscript scripts/main.R \
  -i expression_matrix.csv \
  -g sample_info.csv \
  -o ./output \
  -l no \
  -t 0 \
  -s 42

Error Handling

Common Errors
ErrorCauseSolution
SKILL_FILE_NOT_FOUNDInput file does not existCheck file path
SKILL_EMPTY_FILEInput file exists but contains no dataRecreate or re-export the file
SKILL_MISSING_COLUMNSMetadata file is missing sample, group, or batch columnsCheck header names or pass custom column names
SKILL_SAMPLE_MISMATCHMetadata sample IDs do not match expression matrix columnsVerify sample names between files
SKILL_INVALID_DATADataset fails minimum design checks (< 2 batches, < 2 groups, < 2 samples per batch/group)Review group counts, batch counts, and ID validity
SKILL_INVALID_TYPEExpression values are non-numeric or non-finiteClean matrix values before running
SKILL_TIMEOUTRun exceeded the configured time limitIncrease --timeout_seconds or set it to 0
SKILL_DEPENDENCY_MISSINGRequired R package is not installedInstall with: Rscript -e "BiocManager::install(c('sva','limma')); install.packages('ggplot2')"
SKILL_RUNTIME_ERRORRuntime I/O or filesystem error occurredCheck read/write permissions and environment

IF error persists, READ: references/troubleshooting.md

Troubleshooting note: In environments where packages are not yet installed, SKILL_DEPENDENCY_MISSING will fire before file-validation or --help. Install dependencies first, then re-run to expose file-related errors or access --help.


Input Validation

This skill accepts:

  1. A bulk RNA-seq or microarray expression matrix (CSV, genes as rows, samples as columns)
  2. A sample metadata file (CSV) with sample ID, biological group, and batch columns; at least 2 batches and 2 biological groups are required

If the user's request does not involve batch effect correction on merged bulk expression matrices — for example, asking to integrate single-cell RNA-seq data, process raw FASTQ files, run differential expression without batch labels, or analyze datasets with only one batch — do not proceed with the workflow. Instead respond:

"Batch Effect Correction is designed to remove batch-driven variation from merged bulk expression matrices using ComBat, while preserving biological group structure. Your request appears to be outside this scope. Please provide a multi-batch expression matrix with sample-level batch metadata, or use a more appropriate tool for single-cell integration, differential expression, or raw sequencing processing."


Testing

Test with Sample Data
bash
# Check help (requires packages installed)
Rscript scripts/main.R --help

# Run with bundled test data
Rscript scripts/main.R \
  -i tests/data/expression_matrix_merged.csv \
  -g tests/data/sample_info.csv \
  -o tests/output/
Validation Commands
bash
# Check corrected matrix exists
ls -la tests/output/corrected_expression_matrix.csv

# Check matched metadata exists
ls -la tests/output/matched_sample_info.csv

# Check PCA output exists
ls -la tests/output/batch_after_pca.pdf

Implementation Checklist

  • CLI parsing with optparse
  • set.seed() for reproducibility
  • requireNamespace() dependency checks
  • Session info recording
  • Time-limit support through setTimeLimit()
  • File reading instructions in SKILL.md
  • Modular script structure in scripts/
  • Test data provided
  • Error handling with SKILL_* codes
  • QC plots generated before and after correction
  • References in references/ directory

Last updated: 2026-04-27 | Version: 1.1.0

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts, references) in awesome-med-research-skills/Data Analysis/batch-effect-correction of aipoch/medical-research-skills.

  • SKILL.md
  • eval_report_batch-effect-correction_result.json
  • references/algorithm.md
  • references/cli-guide.md
  • references/troubleshooting.md
  • scripts/functions.R
  • scripts/input_functions.R
  • scripts/main.R
  • scripts/output_utils.R
  • scripts/plotting.R
  • scripts/run_analysis.R
  • scripts/utils.R
  • tests/data/expression_matrix_merged.csv
  • tests/data/sample_info.csv
  • tests/run_tests.R
  • tests/testthat/helper-load-scripts.R
  • … and 2 more

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Batch Effect Correction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Batch Effect Correction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Batch Effect Correction this skillaipoch/medical-research-skills2k—~2.8kAutomated safety check: PassMIT
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
deepTools NGS Toolkitdavila7/claude-code-templates32k13 repos~4.5kAutomated safety check: PassMIT
Bulkrna Cosinor RhythmTianGzlab/OmicsClaw161—~840Automated safety check: PassApache-2.0
PyDESeq2 Differential Expressiondavila7/claude-code-templates32k12 repos~4kAutomated safety check: PassMIT
Gtars Genomic Interval Toolkitdavila7/claude-code-templates32k12 repos~1.9kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • deepTools NGS Toolkit

    davila7/claude-code-templates

    Guides use of deepTools on sequencing data: BAM to bigWig conversion, QC, sample correlation, and heatmaps or profiles around TSS and peaks for ChIP-seq, RNA-seq and ATAC-seq.

    32k GitHub starsUsed in 13 repos~4.5k tokens
    Research & ScienceAuto-check passed
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated today
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    32k GitHub starsUsed in 12 repos~4k tokens
    Research & ScienceAuto-check passed
  • Gtars Genomic Interval Toolkit

    davila7/claude-code-templates

    Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.

    32k GitHub starsUsed in 12 repos~1.9k tokens
    Research & ScienceAuto-check passed
  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    32k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 21 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 21 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 21 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 21 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed

Questions about Batch Effect Correction

What does Batch Effect Correction do?

A skill your agent uses when correcting batch effects in merged bulk expression matrices with sample-level batch metadata while preserving biological group structure and generating before-and-after…. Batch Effect Correction is an agent skill from aipoch/medical-research-skills. Use when correcting batch effects in merged bulk expression matrices with sample-level batch metadata while preserving biological group structure and generating before-and-after QC plots.

When should I use Batch Effect Correction?

Batch Effect Correction fits situations like: tasks that involve Bioinformatics.

How do I install Batch Effect Correction in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill batch-effect-correction -a claude-code`. Or copy the skill folder (awesome-med-research-skills/Data Analysis/batch-effect-correction in aipoch/medical-research-skills) into .claude/skills/batch-effect-correction in your project. Claude Code loads it when a task matches its description.

How do I install Batch Effect Correction in Codex?

Run `npx skills add aipoch/medical-research-skills --skill batch-effect-correction -a codex`. Or copy the skill folder (awesome-med-research-skills/Data Analysis/batch-effect-correction in aipoch/medical-research-skills) into .agents/skills/batch-effect-correction in your project. Codex loads it when a task matches its description.

Can I use Batch Effect Correction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill batch-effect-correction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/batch-effect-correction, .gemini/skills/batch-effect-correction, .github/skills/batch-effect-correction and .opencode/skills/batch-effect-correction in your project.

What does Batch Effect Correction need to run?

Going by SKILL.md and its folder, Batch Effect Correction needs R for the scripts in its folder.

Does Batch Effect Correction access the network?

SKILL.md names 1 domain. In commands or code: cloud.r-project.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Batch Effect Correction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Batch Effect Correction use?

Batch Effect Correction is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Batch Effect Correction use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to Batch Effect Correction?

Skills that share tags, products or a category with Batch Effect Correction: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), deepTools NGS Toolkit (davila7/claude-code-templates, 32k stars), Bulkrna Cosinor Rhythm (TianGzlab/OmicsClaw, 161 stars) and PyDESeq2 Differential Expression (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Batch Effect Correction?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,974 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.