Agent skill

Bio Methylation Ewas Design

by GPTomics in GPTomics/bioSkills

Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible.

MITAuto-check passedResearch & Science

Install Bio Methylation Ewas Design

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-methylation-ewas-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/methylation-analysis/ewas-design .claude/skills/bio-methylation-ewas-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-methylation-ewas-design
GitHub stars
1.2k
Used in
1 other repo
Token cost
~6.3k tokens
SKILL.md length
3,103 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible.

  • Works in 4 steps: Composition before regulation. Whole… → Confounder > signal. Design… → Blood is the lamppost, not the keys.… → …
  • Designing an EWAS
  • SKILL.md covers Version Compatibility, The Single Most Important…, The Confounding Hierarchy… and Design Beats Correction -- The…, plus 12 more sections
  • Runs R and Python scripts from its folder; calls pip

What it does

Bio Methylation Ewas Design is an agent skill from GPTomics/bioSkills. Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible. Covers the confounding hierarchy (cell composition covariates as the dominant confounder, batch/Sentrix chip/array position, age/sex, smoking AHRR cg05575921, ancestry/mQTL, reverse causation), chip randomization (no-rescue theorem), surrogate variable analysis sva/SmartSVA, ComBat, RUVm, over-correction, genomic inflation lambda vs GWAS genomic control, BACON…

Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/randomize_chip_layout.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Designing an EWAS
  • Choosing a covariate set
  • Randomizing a plate layout
  • Interpreting lambda

Example prompts

  • “Use the bio-methylation-ewas-design skill to design and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite…”
  • “/bio-methylation-ewas-design”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Composition before regulation. Whole blood is a leukocyte mixture; a CpG can be 90% methylated in granulocytes and 10% in lymphocytes. Any…
  2. Confounder > signal. Design (randomization, matching) is the primary defense; statistical adjustment is the backstop. No analysis rescues…
  3. Blood is the lamppost, not the keys. Blood is convenient; the disease tissue (brain, adipose, tumor) is usually inaccessible and weakly…
  4. Discovery is a hypothesis; replication is the finding. With tiny effects and pervasive confounding, a single-cohort…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Methylation Ewas Design loads about 6.3k tokens when it runs. Until then it costs about 261 tokens; SKILL.md has 3,103 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~261
When it runs · the whole SKILL.md, loaded when a task matches
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,103 words, ~6,345 tokens.

Download SKILL.mdSave it as .claude/skills/bio-methylation-ewas-design/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-methylation-ewas-design
description
Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible. Covers the confounding hierarchy (cell composition covariates as the dominant confounder, batch/Sentrix chip/array position, age/sex, smoking AHRR cg05575921, ancestry/mQTL, reverse causation), chip randomization (no-rescue theorem), surrogate variable analysis sva/SmartSVA, ComBat, RUVm, over-correction, genomic inflation lambda vs GWAS genomic control, BACON bias/inflation, genome-wide significance threshold 450K/EPIC, FWER vs FDR, pwrEWAS power, meta-analysis, EWAS Catalog/Atlas, methylation risk scores. Use when designing an EWAS, choosing a covariate set, randomizing a plate layout, interpreting lambda, applying BACON, setting a threshold, powering a study, or using an MRS. For the per-site test see differential-cpg-testing; for cell fractions see cell-type-deconvolution; for causal mQTL orientation see causal-genomics/mendelian-randomization.
tool_type
mixed
primary_tool
meffil

Version Compatibility

Reference examples tested with: meffil 1.3+, sva 3.50+, bacon 1.30+, limma 3.58+, missMethyl 1.36+, pwrEWAS 1.16+, pandas 2.2+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Two versions decide everything downstream. The ARRAY (450K vs EPIC v1 ~850K vs EPIC v2 ~935K) sets the CpG universe and therefore the effective number of tests and the genome-wide threshold - confirm against the current manifest before quoting a threshold. The CELL-TYPE REFERENCE panel must match the tissue and age (cord blood, adult blood, and tumor need different references); a mismatched reference silently biases the cell-fraction covariates. BACON slot names, pwrEWAS argument names, and sva method options have shifted across Bioconductor releases - check ?function on the installed build.

EWAS Design

"Run an EWAS for my phenotype" -> Build the covariate model and the chip randomization BEFORE the per-CpG test, because the confounders are larger than the signal - the p-value is the last and least interesting thing.

  • R: meffil.ewas(beta, variable, covariates=data.frame(age, sex, cellprops, chip, plate)) then BACON-correct the test statistics.

Scope: the design-and-inference layer of an EWAS - confounding, randomization, batch/SVA/RUV correction, genomic inflation/BACON, genome-wide thresholds, power, meta-analysis, replication, and methylation risk scores. The per-CpG test mechanics (beta vs M-value, limma) -> differential-cpg-testing. Cell-fraction estimation algorithms (Houseman/IDOL/EpiDISH) -> cell-type-deconvolution. Normalization internals (funnorm/noob) -> array-preprocessing. mQTL Mendelian randomization for causal orientation -> causal-genomics/mendelian-randomization. MRS weight-learning (penalized regression, nested CV) -> the machine-learning category.

The Single Most Important Modern Insight -- An EWAS Hit Is a Cell-Composition Difference Until Proven Otherwise

The defining feature of the field is that the confounders are LARGER than the signal: a true effect is typically a sub-2% absolute methylation difference, while cell mix, batch, age, sex, smoking, and ancestry each move methylation 10-50%. An EWAS is therefore won or lost BEFORE the per-CpG test - in the chip randomization, the covariate set, and the replication cohort. Four corollaries, each of which a common misuse violates:

  1. Composition before regulation. Whole blood is a leukocyte mixture; a CpG can be 90% methylated in granulocytes and 10% in lymphocytes. Any phenotype that shifts the mixture (age, inflammation, infection, stress, smoking, most diseases) produces methylation differences that are composition artifacts, not within-cell regulatory change. Cell-proportion covariates are mandatory, not optional; the default suspicion for any unadjusted hit is "a blood-count difference."
  2. Confounder > signal. Design (randomization, matching) is the primary defense; statistical adjustment is the backstop. No analysis rescues a design that confounded chip/plate with phenotype at the bench.
  3. Blood is the lamppost, not the keys. Blood is convenient; the disease tissue (brain, adipose, tumor) is usually inaccessible and weakly correlated (blood-brain mean r ~0.15, Hannon 2015). A blood EWAS for a brain trait tests a confounded surrogate, and a cross-sectional design cannot orient cause vs consequence (reverse causation).
  4. Discovery is a hypothesis; replication is the finding. With tiny effects and pervasive confounding, a single-cohort genome-wide-significant CpG is a lead, not a result. The EWAS Catalog/Atlas exist so a "novel" hit can be checked against the generic age/smoking/cell-comp CpGs everyone finds.

Organize the analysis around defending these four, not around listing sva/ComBat/bacon functions.

The Confounding Hierarchy (ordered by damage, not convenience)

RankConfounderWhy it dominatesDefenseOwner
1Cell compositionEach leukocyte has a radically different methylome; any mixture shift is a fake signalEstimate cell fractions and ALWAYS adjust-> cell-type-deconvolution (estimation); decide here
2Batch / Sentrix chip / array position / plate / scan dateEPIC runs 8 samples per BeadChip (450K: 12); position and chip imprint methylationRandomize at design time; then SVA/RUVm/ComBat + chip/position covariateshere + array-preprocessing
3Age and sexAge is the most reproducible methylome correlate (the clock); sex drives X/Y and X-inactivationAlways include; analyze sex chromosomes separately; sex is a mislabel QC checkhere
4SmokingLargest reproducible blood exposure signature; confounds half of all health phenotypesAdjust (methylation smoking score > self-report); AHRR cg05575921 is the positive controlhere
5Genetic / ancestry / mQTLMany CpGs are under genetic control; population structure confounds like GWASGenetic PCs; analyze within ancestry; drop SNP-affected/cross-reactive probeshere + array-preprocessing
6Reverse causation / tissue relevanceCross-sectional methylation may be a consequence, not a cause; blood != disease tissueProspective/longitudinal or MZ-discordant design; cross-tissue concordancehere; causal -> causal-genomics/mendelian-randomization

Default covariate set for a blood EWAS: age, sex, cell proportions, chip/position (or SVs), and where relevant genetic PCs and a smoking score. The omission of any one is the most common EWAS error.

Design Beats Correction -- The No-Rescue Theorem

If all cases ran on chip A and all controls on chip B, batch and phenotype are inseparable - ComBat will remove the batch AND the signal, or preserve a spurious one. No statistical adjustment recovers a confounded design (general principle: experimental-design/batch-design).

Goal: Make technical batch orthogonal to phenotype before any sample touches the array.

Approach: Randomize (or block) sample-to-chip, sample-to-position, and sample-to-plate assignment so case/control, age, and sex are balanced across chips; reserve no chip for one group. This single uncorrectable decision matters more than every analysis choice that follows.

r
# Stratified randomization of samples to 8-position EPIC BeadChips (450K: 12), balancing case/control per chip
meta$chip <- NA_integer_
for (grp in split(seq_len(nrow(meta)), meta$case)) {
  meta$chip[sample(grp)] <- rep(seq_len(ceiling(length(grp) / 4)), each = 4, length.out = length(grp))
}
# Then interleave the two case groups across chips so no chip is single-group; verify balance:
with(meta, table(chip, case))   # every chip should hold both groups

Batch / SVA / RUV Correction (and the over-correction trap)

ToolCitationMechanism / roleWhen
svaLeek & Storey 2007 PLoS Genet 3:e161latent surrogate variables built orthogonal to the variable of interestreference-free soak-up of cell mix + unknown batch
SmartSVAChen 2017 BMC Genomics 18:413order-of-magnitude-faster SVA with explicit convergencethe de-facto EWAS SVA for large cohorts
ComBatJohnson 2007 Biostatistics 8:118empirical-Bayes removal of a SPECIFIED batch (chip/plate)known, labelled batch; protect biology via mod=
RUVmMaksimovic 2015 Nucleic Acids Res 43:e106two-stage RUV-inverse using Illumina's ~600 negative control probesarray EWAS; estimates unwanted variation from control probes
funnormFortin 2014 Genome Biol 15:503control-probe PCA at the normalization stagefix unwanted variation BEFORE testing (-> array-preprocessing)

The central tension: cell-composition and technical variation MUST be removed, but aggressive correction removes REAL biology when the unwanted variation overlaps the phenotype. Every tool above can erase a true effect.

The positive-control check (the pragmatic referee). Monitor a known signal through correction. In a blood EWAS containing smokers, smoking->cg05575921 (AHRR, ~18% hypomethylation, Joehanes 2016) MUST survive. If adding surrogate variables makes the QQ plot look clean BUT kills AHRR, the pipeline is over-corrected. A clean QQ with a dead positive control is a broken pipeline, not a good one. Do not keep adding SVs until lambda hits 1.0.

ComBat and limma/EWAS modeling run on M-values (logit of beta), not betas (betas are bounded [0,1] and heteroscedastic); effect sizes are reported back on the beta / delta-beta scale for interpretability. Pass the biological covariate to ComBat's mod= so it is protected. sva: pass the FULL model (mod, including the variable of interest) AND the null model (mod0) so SVs are orthogonal to the phenotype; choose the number with num.sv(method='be').

Genomic Inflation and BACON -- Why GWAS Intuition Fails

Lambda (the genomic inflation factor) = median observed chi-square / expected median. In GWAS, lambda > 1 signals stratification and is corrected by genomic control (divide all statistics by lambda). This reasoning is WRONG for EWAS:

  1. EWAS statistics are routinely BOTH inflated AND deflated for non-GWAS reasons - residual cell-composition variation, the strong correlation among CpGs, un-modeled technical variation, and the fact that a strong exposure (smoking) genuinely associates with a large fraction of the genome (real signal that legitimately inflates lambda). A lambda of 1.2 may be real biology in a smoking EWAS and cell-composition leakage in an under-powered case-control study - the SAME number means different things.
  2. Genomic control assumes a single multiplicative inflation on a mostly-null genome and ignores BIAS (a systematic mean shift off zero). EWAS violates all three assumptions.

BACON (van Iterson 2017 Genome Biol 18:19) fits a Bayesian Gaussian mixture to the observed test statistics, estimates the empirical-null distribution, and reports BOTH a bias (mean shift) AND an inflation (scale) - then standardizes statistics against that empirical null without assuming the genome is mostly null. Report lambda but correct with BACON; show QQ plots before and after. Caveat: BACON can over-deflate a highly polygenic exposure if the alternative component is large - apply it with the QQ plot, not as a reflex.

Genome-Wide Thresholds and the FWER-vs-FDR Decision

ThresholdSourceRationale
450K: P < 2.4e-7Saffari 2018 Genet Epidemiol 42:20empirical effective-test FWER 5%; ~210,000 effective tests (< probe count due to correlation)
EPIC v1: P < 9e-8Mansell 2019 BMC Genomics 20:366null-simulation FWER 5% for the ~850K EPIC array
Pragmatic 1e-7communityround number between the two array-specific values
FDR (BH) q < 0.05Benjamini-Hochbergmore powerful; discovery / exposure scans; report as a sensitivity layer

Naive Bonferroni on the probe count is over-conservative because CpGs are correlated (co-methylation, shared mQTLs); raw p is anti-conservative because there are ~850K tests. Use the array-specific FWER threshold for the cross-study-comparable headline claim and the EWAS Catalog; use BH-FDR for hypothesis-generating discovery. BH's independence assumption is imperfect under probe correlation but robust to positive dependence. Region-level (DMR) multiple-testing is owned by dmr-detection; mechanical p.adjust is owned by differential-cpg-testing. EPIC v2 (~935K probes, renamed/dropped vs v1) may shift the threshold - verify against the manifest.

Power and Meta-Analysis

Goal: Size an EWAS for a realistic (tiny) effect using site-specific methylation variance, not a single assumed sd.

Approach: pwrEWAS (Graw 2019 BMC Bioinformatics 20:218) simulates DNAm using empirical per-CpG variance from a reference dataset, so power reflects the real site-specific variance structure (bimodal sites near 0/1 behave differently from intermediate sites). Specify target delta-beta, sample size, and the genome-wide threshold; expect to need hundreds-to-thousands of samples for a 1-2% effect - which is why meta-analysis dominates.

EWAS meta-analysis is the field standard for power: each cohort runs an identical pre-specified pipeline (meffil enables this), BACON-corrects per cohort, uploads summary statistics, and a central site does fixed-effect inverse-variance pooling with heterogeneity testing (Cochran's Q, I^2). meffil's distinguishing feature is distributed normalization - cohorts normalize locally without sharing individual data, reducing meta-analysis heterogeneity - plus automated selection of the number of normalization PCs.

Study Designs

DesignControls forCost
Case-control (cross-sectional)nothing inherently; adjust on age/sex/ancestry/cell-mix, randomize chipsreverse causation, composition confounding
Longitudinal / prospectivereverse causation (methylation measured before outcome); within-person change via mixed modelneeds follow-up; repeated samples
MZ-discordant twingenetic + shared-environment confounding by design (paired within-pair analysis)discordant pairs are rare -> power-limited
Exposure EWASsmoking is the template/positive-controlexposure misclassification (self-report)
Meta-analysislow power of single cohortsrequires harmonized pipelines + BACON per cohort

Replication and Look-Up

Replication is the real significance bar. After discovery, triage every hit against both databases - a "novel" CpG that is in fact a top smoking/age/blood-cell CpG is almost certainly residual confounding.

  • EWAS Catalog (Battram 2022 Wellcome Open Res 7:41; ewascatalog.org) - published associations at P < 1e-4 (so a Catalog "hit" is a lookup, not a genome-wide claim) plus de-novo EWAS; check whether a CpG was reported for any trait.
  • EWAS Atlas / Open Platform (Li 2019 Nucleic Acids Res 47:D983) - curated associations with a trait-ENRICHMENT tool for interpreting a CpG set.
Show full SKILL.md (1,282 more words)Show less

Methylation Risk Scores (MRS)

An EWAS produces per-CpG associations; an MRS turns many of them into one number - a weighted CpG sum, MRS = sum_j w_j * beta_ij, with weights from a training EWAS or a penalized fit. It is the methylation analogue of a polygenic risk score (PRS), and an epigenetic clock is the age/health special case (-> epigenetic-clocks).

The load-bearing distinction from a PRS. A PRS sums germline variants - fixed at conception, antecedent, plausibly causal. An MRS sums methylation - modifiable, tissue/time-specific, and frequently a CONSEQUENCE of the exposure/trait rather than a cause. The methylation smoking score (Elliott 2014 Clin Epigenetics 6:4; generalized by Sugden 2019) does not predict a propensity to smoke; it MEASURES the footprint smoking left. So an MRS is reverse-causal-by-default: report it as a predictive BIOMARKER / objective exposure proxy, and reserve "risk" and "cause" for prospectively- or MR-supported claims. PRS = germline cause; MRS = state consequence.

The most useful design role is as a better-measured confounder: self-reported smoking is biased and coarse, so adjusting an EWAS for the methylation smoking score controls residual smoking confounding far better whenever the phenotype is smoking-correlated (caveat: if smoking is on the causal path, over-adjusting via the score removes real signal - the same confounder-vs-mediator tension as cell composition). Other DNAm scores exist for BMI/alcohol/education (McCartney 2018) and circulating proteins (EpiScores, Gadd 2022 eLife 11:e71802). Portability fails across array (450K vs EPIC drop CpGs), tissue, and ancestry - report how many score CpGs are present on the array. Defer MRS weight-learning (penalized regression, nested CV, calibration, leakage) to the machine-learning category; an MRS validated in its own training cohort is not validated.

Per-Method Failure Modes

EWAS without cell-composition covariates

Trigger: running the regression on whole-blood betas with no cell-fraction adjustment. Mechanism: any phenotype that shifts the leukocyte mixture moves methylation 10-50%. Symptom: the top hits are known granulocyte/lymphocyte-proportion CpGs. Fix: estimate cell proportions (-> cell-type-deconvolution) and always include them; treat any unadjusted hit as composition until shown otherwise.

Chip/plate confounded with phenotype, then "fixed" by ComBat

Trigger: all cases on chip A, controls on chip B. Mechanism: batch and phenotype are inseparable. Symptom: ComBat removes the signal or fabricates one; no error. Fix: randomize sample-to-chip/position at design time - this is uncorrectable after the bench.

GWAS genomic control applied to EWAS

Trigger: dividing all test statistics by lambda. Mechanism: assumes single multiplicative inflation on a mostly-null genome and ignores bias. Symptom: real polygenic signal (smoking) deflated or residual confounding under-corrected. Fix: use BACON to estimate empirical-null bias AND inflation.

Over-correction with too many surrogate variables

Trigger: adding SVs/PCs until lambda hits 1.0. Mechanism: latent factors correlated with the phenotype deflate true signal. Symptom: clean QQ but the positive control (AHRR) is gone. Fix: monitor cg05575921 through correction; stop when adding SVs starts killing it.

Bonferroni-on-probe-count or raw-p reporting

Trigger: 0.05/850000, or reporting nominal p. Mechanism: probe correlation makes Bonferroni too strict; ignoring 850K tests is too loose. Symptom: missed real hits or a flood of false ones. Fix: array-specific FWER threshold (2.4e-7 / 9e-8) for claims, BH-FDR for discovery.

Blood EWAS interpreted as disease-tissue mechanism

Trigger: a blood EWAS for a brain/adipose/tumor phenotype read as tissue biology. Mechanism: blood-brain methylation r ~0.15. Symptom: hits with no plausible blood mechanism. Fix: justify tissue relevance (BECon/blood-brain tools); frame blood as a biomarker, not mechanism, unless the CpG is cross-tissue concordant.

MRS treated as a germline-like risk score

Trigger: describing a DNAm score for a disease as "risk" like a PRS. Mechanism: methylation is modifiable and frequently downstream. Symptom: a "DNAm risk score" that is actually the footprint of disease/treatment/behavior. Fix: call it a predictive biomarker; orient cause vs consequence only with prospective or MR evidence.

Quantitative Thresholds

ThresholdSourceRationale
450K genome-wide P < 2.4e-7Saffari 2018 Genet Epidemiol 42:20empirical effective-test FWER 5% (~210,000 effective tests)
EPIC genome-wide P < 9e-8Mansell 2019 BMC Genomics 20:366null-simulation FWER 5% for ~850K EPIC
BH-FDR q < 0.05Benjamini-Hochbergdiscovery / exposure scans; report alongside FWER
AHRR cg05575921 ~18% hypomethylation in smokersJoehanes 2016 Circ Cardiovasc Genet 9:436the canonical positive control; must survive correction
EWAS Catalog inclusion P < 1e-4Battram 2022 Wellcome Open Res 7:41a Catalog entry is a lookup, not a genome-wide claim
Detect 1-2% delta-beta -> hundreds-to-thousands of samplespwrEWAS, Graw 2019tiny site-specific effects drive the field to meta-analysis
Run ComBat/limma on M-values, report delta-betaDu 2010 BMC Bioinformatics 11:587betas are bounded/heteroscedastic; M-values stabilize variance

Common Errors

Error / symptomCauseSolution
Top hits are blood-cell CpGsno cell-composition covariatesestimate and adjust cell fractions
ComBat removed the signalbatch confounded with phenotyperandomize at design time; cannot fix after
Lambda misread as pass/failGWAS dogma applied to EWASpair lambda with QQ shape + positive control; correct with BACON
Clean QQ but AHRR goneover-correction with too many SVsstop adding SVs when the positive control dies
Threshold flood or famineBonferroni-on-probe-count or raw puse array-specific FWER / BH-FDR
"Novel" CpG is a known generic hitno Catalog/Atlas triagelook up the CpG before claiming novelty
MRS called a risk scorePRS connotations on modifiable methylationreport as biomarker; orient cause only with prospective/MR design

References

  • van Iterson M, van Zwet EW, Heijmans BT; BIOS Consortium. 2017. Controlling bias and inflation in epigenome- and transcriptome-wide association studies using the empirical null distribution. Genome Biol 18:19.
  • Saffari A, Silver MJ, Zavattari P, et al. 2018. Estimation of a significance threshold for epigenome-wide association studies. Genet Epidemiol 42:20-33.
  • Mansell G, Gorrie-Stone TJ, Bao Y, et al. 2019. Guidance for DNA methylation studies: statistical insights from the Illumina EPIC array. BMC Genomics 20:366.
  • Leek JT, Storey JD. 2007. Capturing heterogeneity in gene expression studies by surrogate variable analysis. PLoS Genet 3:e161.
  • Chen J, Behnam E, Huang J, et al. 2017. Fast and robust adjustment of cell mixtures in epigenome-wide association studies with SmartSVA. BMC Genomics 18:413.
  • Johnson WE, Li C, Rabinovic A. 2007. Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics 8:118-127.
  • Maksimovic J, Gagnon-Bartsch JA, Speed TP, Oshlack A. 2015. Removing unwanted variation in a differential methylation analysis of Illumina HumanMethylation450 array data. Nucleic Acids Res 43:e106.
  • Min JL, Hemani G, Davey Smith G, Relton C, Suderman M. 2018. Meffil: efficient normalization and analysis of very large DNA methylation datasets. Bioinformatics 34:3983-3989.
  • Graw S, Henn R, Thompson JA, Koestler DC. 2019. pwrEWAS: a user-friendly tool for comprehensive power estimation for epigenome wide association studies (EWAS). BMC Bioinformatics 20:218.
  • Joehanes R, Just AC, Marioni RE, et al. 2016. Epigenetic Signatures of Cigarette Smoking. Circ Cardiovasc Genet 9:436-447.
  • Elliott HR, Tillin T, McArdle WL, et al. 2014. Differences in smoking associated DNA methylation patterns in South Asians and Europeans. Clin Epigenetics 6:4.
  • Gadd DA, Hillary RF, McCartney DL, et al. 2022. Epigenetic scores for the circulating proteome as tools for disease prediction. eLife 11:e71802.
  • McCartney DL, Hillary RF, Stevenson AJ, et al. 2018. Epigenetic prediction of complex traits and death. Genome Biol 19:136.
  • Battram T, Yousefi P, Crawford G, et al. 2022. The EWAS Catalog: a database of epigenome-wide association studies. Wellcome Open Res 7:41.
  • Li M, Zou D, Li Z, et al. 2019. EWAS Atlas: a curated knowledgebase of epigenome-wide association studies. Nucleic Acids Res 47:D983-D988.
  • Hannon E, Lunnon K, Schalkwyk L, Mill J. 2015. Interindividual methylomic variation across blood, cortex, and cerebellum. Epigenetics 10:1024-1032.
  • Michels KB, Binder AM, Dedeurwaerder S, et al. 2013. Recommendations for the design and analysis of epigenome-wide association studies. Nat Methods 10:949-955.
  • differential-cpg-testing - The per-site test this design layer feeds
  • cell-type-deconvolution - Cell-fraction covariates (the dominant confounder)
  • array-preprocessing - Normalization choice (funnorm) vs model-level batch correction
  • array-qc-filtering - Probe filtering and chip/position batch diagnosis
  • causal-genomics/mendelian-randomization - mQTL-based causal orientation (reverse causation)
  • experimental-design/batch-design - General randomization and batch-design principles
  • clinical-biostatistics/multiplicity-graphical - FWER for confirmatory trials (contrast with discovery FDR)
  • workflows/methylation-pipeline - End-to-end pipeline

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in methylation-analysis/ewas-design of GPTomics/bioSkills.

  • SKILL.md
  • examples/ewas_sva_bacon.R
  • examples/randomize_chip_layout.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Methylation Ewas Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Methylation Ewas Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Methylation Ewas Design this skillGPTomics/bioSkills1.2k1 repos~6.3kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Methylation Ewas Design

What does Bio Methylation Ewas Design do?

Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible. Bio Methylation Ewas Design is an agent skill from GPTomics/bioSkills. Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible.

When should I use Bio Methylation Ewas Design?

Bio Methylation Ewas Design fits situations like: designing an EWAS; choosing a covariate set; randomizing a plate layout; interpreting lambda.

How do I install Bio Methylation Ewas Design in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a claude-code`. Or copy the skill folder (methylation-analysis/ewas-design in GPTomics/bioSkills) into .claude/skills/bio-methylation-ewas-design in your project. Claude Code loads it when a task matches its description.

How do I install Bio Methylation Ewas Design in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a codex`. Or copy the skill folder (methylation-analysis/ewas-design in GPTomics/bioSkills) into .agents/skills/bio-methylation-ewas-design in your project. Codex loads it when a task matches its description.

Can I use Bio Methylation Ewas Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-methylation-ewas-design, .gemini/skills/bio-methylation-ewas-design, .github/skills/bio-methylation-ewas-design and .opencode/skills/bio-methylation-ewas-design in your project.

What does Bio Methylation Ewas Design need to run?

Going by SKILL.md and its folder, Bio Methylation Ewas Design needs R and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Methylation Ewas Design access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Methylation Ewas Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Methylation Ewas Design use?

Bio Methylation Ewas Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Methylation Ewas Design use?

About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Methylation Ewas Design?

Skills that share tags, products or a category with Bio Methylation Ewas Design: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Methylation Ewas Design?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.