Agent skill

Bio Methylation Array Preprocessing

by GPTomics in GPTomics/bioSkills

Turns raw Illumina Infinium methylation BeadChip IDATs (450K, EPIC, EPICv2) into a defensible beta/M matrix with sesame (openSesame/SigDF) or minfi (RGChannelSet - MethylSet - GenomicRatioSet).

MITAuto-check passedDatabases

Install Bio Methylation Array Preprocessing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-methylation-array-preprocessing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-methylation-array-preprocessing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/methylation-analysis/array-preprocessing .claude/skills/bio-methylation-array-preprocessing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-methylation-array-preprocessing
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.5k tokens
SKILL.md length
1,977 words
Files
3
Skills in repo
553
Repo updated
First seen
Licence
MIT

At a glance

Turns raw Illumina Infinium methylation BeadChip IDATs (450K, EPIC, EPICv2) into a defensible beta/M matrix with sesame (openSesame/SigDF) or minfi (RGChannelSet - MethylSet - GenomicRatioSet).

  • Works in 3 steps: The manifest is the experiment. The… → Type I and Type II betas disagree by… → Raw beta is uninterpretable until…
  • Choosing a normalization for a 450K/EPIC/EPICv2 cohort
  • SKILL.md covers Version Compatibility, The Single Most Important…, Three Modalities of the Same… and Object Models (do not start…, plus 10 more sections
  • Runs R scripts from its folder

What it does

Bio Methylation Array Preprocessing is an agent skill from GPTomics/bioSkills. Turns raw Illumina Infinium methylation BeadChip IDATs (450K, EPIC, EPICv2) into a defensible beta/M matrix with sesame (openSesame/SigDF) or minfi (RGChannelSet - MethylSet - GenomicRatioSet). Covers Type I vs Type II probe chemistry and why raw Type II beta is compressed, the signal-to-beta math (beta = M/(M+U+100)) and M-value logit, detection-p / pOOBAH masking including the out-of-band deletion-artifact catch, dye-bias correction, and the normalization decision (noob, funnorm, quantile, SWAN, BMIQ, dasen…

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Databases, covering Database schema design and Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Choosing a normalization for a 450K/EPIC/EPICv2 cohort
  • Deciding beta vs M
  • Masking failed probes
  • Producing the corrected matrix before testing

Example prompts

  • “Use the bio-methylation-array-preprocessing skill to turn raw Illumina Infinium methylation BeadChip IDATs (450K, EPIC, EPICv2) into a defensible…”
  • “/bio-methylation-array-preprocessing”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The manifest is the experiment. The array interrogates <3% of human CpGs, and a DIFFERENT <3% across 450K (~485K), EPIC (~865K), and…
  2. Type I and Type II betas disagree by design. Type II probes (one bead, two dyes) have a narrower dynamic range and dye-incorporation bias…
  3. Raw beta is uninterpretable until detection-masked. Failed probes - low signal, germline/somatic deletions, cross-reactive, SNP-hit…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Methylation Array Preprocessing loads about 4.5k tokens when it runs. Until then it costs about 235 tokens; SKILL.md has 1,977 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~235
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,977 words, ~4,523 tokens.

Download SKILL.mdSave it as .claude/skills/bio-methylation-array-preprocessing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-methylation-array-preprocessing
description
Turns raw Illumina Infinium methylation BeadChip IDATs (450K, EPIC, EPICv2) into a defensible beta/M matrix with sesame (openSesame/SigDF) or minfi (RGChannelSet -> MethylSet -> GenomicRatioSet). Covers Type I vs Type II probe chemistry and why raw Type II beta is compressed, the signal-to-beta math (beta = M/(M+U+100)) and M-value logit, detection-p / pOOBAH masking including the out-of-band deletion-artifact catch, dye-bias correction, and the normalization decision (noob, funnorm, quantile, SWAN, BMIQ, dasen, sesame QCDPB). Use when reading IDATs, choosing a normalization for a 450K/EPIC/EPICv2 cohort, deciding beta vs M, masking failed probes, or producing the corrected matrix before testing. For probe/sample filtering, EPICv2 replicate collapse, and sample-identity QC see array-qc-filtering; for native long-read 5mC see long-read-sequencing/nanopore-methylation (a different platform).
tool_type
r
primary_tool
sesame

Version Compatibility

Reference examples tested with: sesame 1.20+, minfi 1.48+, ChAMP 2.32+, wateRmelon 2.8+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

The ARRAY VERSION and GENOME BUILD are versions that matter as much as the package. EPICv2 REQUIRES sesame (mainstream minfi does not auto-detect it and returns "Unknown"); the manifest/annotation packages are array-version- and genome-build-specific (450K and EPICv1 are hg19; EPICv2 is hg38-native). sesame pulls platform/address data from its own hub, so sesameDataCache() must run once before processing. Record the array (450K/EPIC/EPICv2) and the genome build in any output, the way a sequencing run records its reference.

Array Preprocessing

"Give me a clean methylation matrix from my IDATs" -> Read the raw two-channel intensities, correct the Type I/II design mismatch, dye/background bias, and failed probes, then emit beta (for reporting) and M (for testing) - because an Infinium beta is a two-chemistry fluorescence ratio, not a methylation value, until those corrections are applied.

  • R: openSesame(idat_dir, prep='QCDPB', func=getBetas) (sesame) or preprocessFunnorm(rgSet) (minfi)

Scope: IDAT -> corrected, masked, normalized beta/M matrix for one array version. Probe filtering (cross-reactive/SNP/sex), EPICv2 replicate collapse, and sample-identity QC -> array-qc-filtering. Per-CpG testing -> differential-cpg-testing. Region calling -> dmr-detection. Native long-read 5mC -> long-read-sequencing/nanopore-methylation. Bisulfite-sequencing (Bismark/WGBS/RRBS) is the other modality in this category, not this skill.

The Single Most Important Modern Insight -- An Infinium Beta Value Is a Two-Chemistry Fluorescence Ratio, Not a Methylation Measurement

An Infinium array does not measure methylation - it measures the relative fluorescence of a methylated vs unmethylated allele at a fixed, manufacturer-chosen set of CpGs, glued together from two incompatible chemistries. A raw beta becomes a comparable methylation estimate only after preprocessing; preprocessing IS the measurement, not optional cleanup. Three corollaries every misuse violates:

  1. The manifest is the experiment. The array interrogates <3% of human CpGs, and a DIFFERENT <3% across 450K (~485K), EPIC (~865K), and EPICv2 (~935K). "Absent" almost always means "not on this array," and a 450K-trained clock or EWAS does not transfer to EPICv2 without intersecting probe sets. Do not start from a supplied beta matrix when IDATs exist - the raw two-channel intensities, control probes, and out-of-band signal that noob/funnorm/pOOBAH need are already gone.
  2. Type I and Type II betas disagree by design. Type II probes (one bead, two dyes) have a narrower dynamic range and dye-incorporation bias, so raw Type II betas are compressed toward 0.5 relative to Type I (two beads, one channel). Mixing the two chemistries without a design correction (BMIQ/SWAN/sesame matchDesign) injects a probe-type artifact that can exceed the biological effect. The diagnostic is a per-type beta-density plot showing two mismatched peaks.
  3. Raw beta is uninterpretable until detection-masked. Failed probes - low signal, germline/somatic deletions, cross-reactive, SNP-hit - return confident-looking betas that are pure noise. pOOBAH / detection-p masking is what separates a number from a measurement of nothing; pOOBAH additionally catches deletion-driven false-intermediate methylation that negative-control detection-p misses.

Organize the work around DELIVERING a defensible matrix (read -> correct -> mask -> normalize), not around listing minfi functions.

Three Modalities of the Same Biology

DNA methylation is measured three ways, each with different tradeoffs - state which one the data is before choosing tools:

ModalityReadoutCoverageCohort-comparabilityThis skill
Infinium array (450K/EPIC/EPICv2)intensity ratio, no depthfixed <3% of CpGs, regulatory-enrichedhigh (shared manifest, no alignment)YES
WGBS / RRBS bisulfitecount ratio, depth-gatedgenome-wide (WGBS) or enriched (RRBS)needs alignment + matched genome-> bismark-alignment, methylkit-analysis
Long-read native (ONT/PacBio)per-molecule modification callsgenome-wide, phasedgrowing-> long-read-sequencing/nanopore-methylation

Arrays dominate human epigenetic epidemiology (essentially every published clock and large EWAS is array-based) because cost is a fraction of WGBS and the fixed manifest makes cohorts directly comparable.

Object Models (do not start from a beta matrix)

The raw output per sample is a pair of binary IDATs (_Grn.idat, _Red.idat); background, dye, and detection-p correction REQUIRE these plus the control probes.

  • minfi: RGChannelSet (raw red/green) -> a preprocess* step -> MethylSet (M/U intensities) -> RatioSet (beta/M) -> GenomicRatioSet (genome-mapped). read.metharray.exp() reads IDATs; getBeta(), getM(), getCN() extract values.
  • sesame: a SigDF (one signal data.frame per sample). readIDATpair() reads one sample; openSesame() drives the whole pipeline across a directory and returns a betas matrix directly.

Tool Taxonomy

ToolCitationMechanism / roleWhen
sesameZhou 2018 Nucleic Acids Res 46:e123SigDF; openSesame QCDPB; pOOBAH OOB masking; EPICv2-nativeEPICv2; best detection masking; the modern default
minfiAryee 2014 Bioinformatics 30:1363RGChannelSet->GenomicRatioSet; noob/funnorm/quantile/SWAN450K/EPICv1; large downstream ecosystem (DMRcate, conumee)
ChAMPTian 2017 Bioinformatics 33:3982end-to-end pipeline; BMIQ defaultone-call newcomer pipeline on 450K/EPICv1
wateRmelonPidsley 2013 BMC Genomics 14:293dasen/nasen + metric-driven normalization evaldasen default; normalization benchmarking

Normalization Decision Tree by Scenario

Separate the two correction layers that get conflated: (a) background + dye bias (within-sample): noob, sesame dyeBias, dasen background step; (b) Type I/II design correction + between-array harmonization: SWAN, BMIQ, quantile, funnorm, dasen quantile step. A complete pipeline does both.

ScenarioRecommendedWhy
EPICv2 (any design)sesame openSesame(prep='QCDPB')EPICv2-native; pOOBAH; minfi mis-handles duplicate IDs
Cancer / cross-tissue (global differences expected)minfi preprocessFunnorm (noob + control-PCs)preserves real global shifts; quantile would erase them
Subtle blood EWAS (no global difference expected)preprocessQuantile or wateRmelon dasenmarginal distributions assumed equal; safe to harmonize
Strong Type I/II design correction wantedBMIQ (Teschendorff 2013) or SWAN (Maksimovic 2012)dilate Type II onto the Type I distribution; pair with a between-array step
Single-sample / clinical / streamingssNoob or per-IDAT openSesamereproducible without re-normalizing the cohort
Probe/sample filtering, EPICv2 collapse, identity-> array-qc-filteringthis skill stops at the corrected matrix
Per-CpG testing on the matrix-> differential-cpg-testingtest on M-values; report delta-beta

There is no universally best normalization (Pidsley 2013 favored dasen; Fortin 2014 favored funnorm for global-difference studies; Welsh 2023 ranked a sesame/pOOBAH pipeline best and quantile worst on EPIC replicate-concordance). Key the choice on array version + whether global differences are expected + single-sample vs cohort, and verify against current benchmarks rather than hard-coding one method.

Signal -> Beta -> M

  • Beta: beta = M / (M + U + alpha), M = methylated-allele intensity, U = unmethylated, alpha = 100 (minfi default) stabilizes the ratio when both intensities are near zero. beta in [0,1] is interpretable but HETEROSCEDASTIC (variance collapses near 0 and 1), violating the constant-variance assumption of linear models.
  • M-value: M = log2((M_int + alpha) / (U_int + alpha)), the logit of beta. Approximately homoscedastic; the correct scale for limma/t-tests (Du 2010 BMC Bioinformatics 11:587). Rule: test on M-values, report delta-beta for effect size - the same rule as bisulfite sequencing.

Process IDATs with sesame (the EPICv2-safe default)

Goal: Produce a corrected, detection-masked betas matrix from a directory of IDAT pairs without manually juggling manifest packages.

Approach: Cache the sesame data hub once, then run openSesame with the default QCDPB prep (qualityMask, inferInfiniumIChannel, dyeBiasNL, pOOBAH, noob, in that order), which auto-detects the platform and returns betas; pOOBAH writes NA into failed probes in place.

r
library(sesame)
sesameDataCache()                          # once per machine; pulls platform/address data
betas <- openSesame('idat_dir', prep = 'QCDPB', func = getBetas)
# prep codes: Q qualityMask  C inferInfiniumIChannel  D dyeBiasNL  P pOOBAH  B noob
# pOOBAH masks (sets NA) probes whose out-of-band signal is indistinguishable from background,
# catching deletion-driven false-intermediate methylation that negative-control detection-p misses
mvals <- log2(betas / (1 - betas))         # M-values for statistical testing (logit of beta)

For EPICv2, openSesame detects the platform automatically; the replicate-probe collapse (betasCollapseToPfx) belongs to the next stage and is documented in array-qc-filtering.

Process IDATs with minfi (450K / EPICv1)

Goal: Build a normalized GenomicRatioSet and extract beta and M, choosing the normalization by whether global methylation differences are expected.

Approach: Read IDATs into an RGChannelSet, compute a detection-p mask before normalizing, then apply funnorm (global differences) or quantile (no global differences); extract beta and M with the offset-100 defaults.

r
library(minfi)
rgSet <- read.metharray.exp(base = 'idat_dir')
detP <- detectionP(rgSet)                   # neg-control-based; pre-normalization probe-failure map

grSet <- preprocessFunnorm(rgSet, nPCs = 2) # noob first, then 2 control-probe PCs; preserves global shifts
# preprocessQuantile(rgSet) instead when NO global difference is expected (subtle blood EWAS)

beta <- getBeta(grSet)                       # GenomicRatioSet holds precomputed betas (offset applied upstream)
mval <- getM(grSet)                          # log2(beta/(1-beta)) on the ratio set
beta[detP[rownames(beta), colnames(beta)] > 0.01] <- NA   # mask probes failing detection-p (0.01)

EPICv2 is NOT handled by mainstream minfi (it returns "Unknown" and duplicates probe IDs); use sesame for EPICv2.

Show full SKILL.md (754 more words)Show less

Per-Method Failure Modes

Starting from a supplied beta matrix

Trigger: processing begins from a .csv/.RData beta matrix instead of IDATs. Mechanism: a beta matrix has discarded the raw two-channel intensities, control probes, and out-of-band signal. Symptom: noob/funnorm/pOOBAH/dye correction cannot run; detection-p cannot be recomputed. Fix: obtain the raw IDAT pairs; treat a beta matrix as a last resort and document that preprocessing could not be applied.

Type I/II mismatch left in the data

Trigger: testing on raw or only background-corrected betas. Mechanism: Type II betas are compressed toward 0.5 relative to Type I. Symptom: "differential" probes that are design artifacts; a two-peak per-type beta density. Fix: apply BMIQ/SWAN or sesame matchDesign (or use openSesame, which corrects channel/dye) before testing; confirm the two per-type peaks align.

minfi on EPICv2

Trigger: read.metharray.exp on EPICv2 IDATs with mainstream minfi. Mechanism: EPICv2 is not auto-detected; 5,483 loci carry duplicate IDs. Symptom: array reads as "Unknown"; getBeta() returns repeated rownames so match()-based merges silently misbehave. Fix: use sesame (EPICv2-native), or install a third-party EPICv2 manifest/anno, tag the annotation manually, and collapse replicates in array-qc-filtering.

Quantile-normalizing a global-difference contrast

Trigger: preprocessQuantile on cancer vs normal or cross-tissue data. Mechanism: between-array quantile assumes equal marginal beta distributions. Symptom: real global hypomethylation flattened away. Fix: use funnorm (control-probe PCs preserve global shifts); reserve quantile/dasen for subtle no-global-difference designs.

Testing on beta instead of M

Trigger: limma/t-tests run directly on beta. Mechanism: beta is heteroscedastic (variance collapses near 0 and 1). Symptom: miscalibrated variance; inflated or deflated p-values at extreme methylation. Fix: test on M-values, report delta-beta for effect size.

Quantitative Thresholds

ThresholdSourceRationale
beta offset alpha = 100Aryee 2014; minfi defaultstabilizes the ratio when M and U are both near zero
detection-p > 0.01 = failedminfi conventionsignal indistinguishable from background; beta is noise
pOOBAH default p ~ 0.05Zhou 2018OOB-based mask; also catches deletion-driven false intermediate methylation
funnorm nPCs = 2Fortin 2014first 2 control-probe PCs absorb technical variation without erasing biology
Test on M-values, report delta-betaDu 2010M is homoscedastic for modeling; beta is interpretable for effect size
sesame prep QCDPB (ordered)Zhou 2018Q quality, C channel, D dye, P pOOBAH, B noob - the validated default order

Common Errors

Error / symptomCauseSolution
Array reads as "Unknown"EPICv2 in mainstream minfiuse sesame; or third-party manifest + manual annotation tag
getBeta() has repeated rownamesEPICv2 duplicate probe IDscollapse replicates (array-qc-filtering); do not merge by ID first
Two-peak beta density per probe typeType I/II design bias uncorrectedBMIQ/SWAN/openSesame before testing
Global signal vanished after normalizationquantile applied to a global-difference studyuse funnorm
sesameDataCache / platform-not-foundhub not cachedrun sesameDataCache() once before processing
Coordinates misalign merging EPICv2 with 450KEPICv2 is hg38, 450K/EPICv1 hg19track build per array; liftover before merging (array-qc-filtering)

References

  • Aryee MJ, Jaffe AE, Corrada-Bravo H, et al. 2014. Minfi: a flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarrays. Bioinformatics 30:1363-1369.
  • Zhou W, Triche TJ Jr, Laird PW, Shen H. 2018. SeSAMe: reducing artifactual detection of DNA methylation by Infinium BeadChips in genomic deletions. Nucleic Acids Res 46:e123.
  • Triche TJ Jr, Weisenberger DJ, Van Den Berg D, Laird PW, Siegmund KD. 2013. Low-level processing of Illumina Infinium DNA methylation BeadArrays. Nucleic Acids Res 41:e90.
  • Fortin JP, Labbe A, Lemire M, et al. 2014. Functional normalization of 450k methylation array data improves replication in large cancer studies. Genome Biol 15:503.
  • Maksimovic J, Gordon L, Oshlack A. 2012. SWAN: subset-quantile within array normalization for Illumina Infinium HumanMethylation450 BeadChips. Genome Biol 13:R44.
  • Teschendorff AE, Marabita F, Lechner M, et al. 2013. A beta-mixture quantile normalization method for correcting probe design bias in Illumina Infinium 450k DNA methylation data. Bioinformatics 29:189-196.
  • Pidsley R, Wong CCY, Volta M, Lunnon K, Mill J, Schalkwyk LC. 2013. A data-driven approach to preprocessing Illumina 450K methylation array data. BMC Genomics 14:293.
  • Du P, Zhang X, Huang CC, et al. 2010. Comparison of Beta-value and M-value methods for quantifying methylation levels by microarray analysis. BMC Bioinformatics 11:587.
  • Tian Y, Morris TJ, Webster AP, et al. 2017. ChAMP: updated methylation analysis pipeline for Illumina BeadChips. Bioinformatics 33:3982-3984.
  • Kaur D, Lee SM, Goldberg D, et al. 2023. Comprehensive evaluation of the Infinium human MethylationEPIC v2 BeadChip. Epigenetics Commun 3:6.
  • array-qc-filtering - Probe and sample QC/filtering downstream of preprocessing
  • differential-cpg-testing - Per-CpG testing on the resulting beta/M matrix
  • dmr-detection - DMRcate array-mode region calling
  • cell-type-deconvolution - Consumes the clean beta matrix
  • epigenetic-clocks - Consumes the clean beta matrix
  • ewas-design - Study design, batch, and inference layer
  • long-read-sequencing/nanopore-methylation - Native long-read methylation (different platform)
  • workflows/methylation-pipeline - End-to-end pipeline

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in methylation-analysis/array-preprocessing of GPTomics/bioSkills.

  • SKILL.md
  • examples/array_preprocessing.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Methylation Array Preprocessing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Methylation Array Preprocessing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Methylation Array Preprocessing this skillGPTomics/bioSkills1.2k1 repos~4.5kAutomated safety check: PassMIT
Tooluniverse Metabolomics Analysiswu-yc/LabClaw1.1k2 repos~5.9kAutomated safety check: PassNone
Gene Protein Expression Matrix Normalizationaipoch/medical-research-skills2k—~1.5kAutomated safety check: PassMIT
Bio Spatial Transcriptomics Spatial Preprocessingmajiayu000/claude-skill-registry6662 repos~1.3kAutomated safety check: PassMIT
Bio Crispr Screens Batch Correctionmajiayu000/claude-skill-registry6661 repos~2.3kAutomated safety check: PassMIT
Tooluniverse Rnaseq Deseq2wu-yc/LabClaw1.1k2 repos~4.5kAutomated safety check: PassNone

Similar skills

  • Analyze metabolomics data including metabolite identification, quantification, pathway analysis, and metabolic flux.

    1.1k GitHub starsUsed in 2 repos~5.9k tokens
    Research & ScienceAuto-check passed
  • Gene Protein Expression Matrix Normalization

    aipoch/medical-research-skills

    A skill your agent uses when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory…

    2k GitHub stars~1.5k tokensUpdated 20 days ago
    Research & ScienceAuto-check passed
  • Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.

    666 GitHub starsUsed in 2 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Bio Crispr Screens Batch Correction

    majiayu000/claude-skill-registry

    Batch effect correction for CRISPR screens. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 1 repo~2.3k tokens
    Research & ScienceAuto-check passed
  • Production-ready RNA-seq differential expression analysis using PyDESeq2.

    1.1k GitHub starsUsed in 2 repos~4.5k tokens
    Research & ScienceAuto-check passed
  • Bio Single Cell Preprocessing

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control, filtering, and normalization for single-cell RNA-seq using Seurat (R) and Scanpy (Python).

    3.1k GitHub starsUsed in 1 repo~2.4k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 553 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Methylation Array Preprocessing

What does Bio Methylation Array Preprocessing do?

Turns raw Illumina Infinium methylation BeadChip IDATs (450K, EPIC, EPICv2) into a defensible beta/M matrix with sesame (openSesame/SigDF) or minfi (RGChannelSet - MethylSet - GenomicRatioSet). Bio Methylation Array Preprocessing is an agent skill from GPTomics/bioSkills. Turns raw Illumina Infinium methylation BeadChip IDATs (450K, EPIC, EPICv2) into a defensible beta/M matrix with sesame (openSesame/SigDF) or minfi (RGChannelSet - MethylSet - GenomicRatioSet).

When should I use Bio Methylation Array Preprocessing?

Bio Methylation Array Preprocessing fits situations like: choosing a normalization for a 450K/EPIC/EPICv2 cohort; deciding beta vs M; masking failed probes; producing the corrected matrix before testing.

How do I install Bio Methylation Array Preprocessing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-methylation-array-preprocessing -a claude-code`. Or copy the skill folder (methylation-analysis/array-preprocessing in GPTomics/bioSkills) into .claude/skills/bio-methylation-array-preprocessing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Methylation Array Preprocessing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-methylation-array-preprocessing -a codex`. Or copy the skill folder (methylation-analysis/array-preprocessing in GPTomics/bioSkills) into .agents/skills/bio-methylation-array-preprocessing in your project. Codex loads it when a task matches its description.

Can I use Bio Methylation Array Preprocessing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-methylation-array-preprocessing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-methylation-array-preprocessing, .gemini/skills/bio-methylation-array-preprocessing, .github/skills/bio-methylation-array-preprocessing and .opencode/skills/bio-methylation-array-preprocessing in your project.

What does Bio Methylation Array Preprocessing need to run?

Going by SKILL.md and its folder, Bio Methylation Array Preprocessing needs R for the scripts in its folder.

Does Bio Methylation Array Preprocessing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Methylation Array Preprocessing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Methylation Array Preprocessing use?

Bio Methylation Array Preprocessing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Methylation Array Preprocessing use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Methylation Array Preprocessing?

Skills that share tags, products or a category with Bio Methylation Array Preprocessing: Tooluniverse Metabolomics Analysis (wu-yc/LabClaw, 1.1k stars), Gene Protein Expression Matrix Normalization (aipoch/medical-research-skills, 2k stars), Bio Spatial Transcriptomics Spatial Preprocessing (majiayu000/claude-skill-registry, 666 stars) and Bio Crispr Screens Batch Correction (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Methylation Array Preprocessing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.