Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-ewas-design --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/methylation-analysis/ewas-design .claude/skills/bio-methylation-ewas-design && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-methylation-ewas-design" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/ewas-design into .claude/skills/bio-methylation-ewas-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-ewas-design", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/ewas-designType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-ewas-design --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/methylation-analysis/ewas-design .agents/skills/bio-methylation-ewas-design && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-methylation-ewas-design" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/ewas-design into .agents/skills/bio-methylation-ewas-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-ewas-design", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-ewas-design --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/methylation-analysis/ewas-design .cursor/skills/bio-methylation-ewas-design && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-methylation-ewas-design" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/ewas-design into .cursor/skills/bio-methylation-ewas-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-ewas-design", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path methylation-analysis/ewas-design--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-ewas-design --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/methylation-analysis/ewas-design .gemini/skills/bio-methylation-ewas-design && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-methylation-ewas-design" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/ewas-design into .gemini/skills/bio-methylation-ewas-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-ewas-design", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-methylation-ewas-designInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/methylation-analysis/ewas-design .github/skills/bio-methylation-ewas-design && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-methylation-ewas-design" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/ewas-design into .github/skills/bio-methylation-ewas-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-ewas-design", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-ewas-design --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/methylation-analysis/ewas-design .opencode/skills/bio-methylation-ewas-design && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-methylation-ewas-design" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/ewas-design into .opencode/skills/bio-methylation-ewas-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-ewas-design", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-methylation-ewas-designDesigns and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible.
Bio Methylation Ewas Design is an agent skill from GPTomics/bioSkills. Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible. Covers the confounding hierarchy (cell composition covariates as the dominant confounder, batch/Sentrix chip/array position, age/sex, smoking AHRR cg05575921, ancestry/mQTL, reverse causation), chip randomization (no-rescue theorem), surrogate variable analysis sva/SmartSVA, ComBat, RUVm, over-correction, genomic inflation lambda vs GWAS genomic control, BACON…
Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/randomize_chip_layout.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (R and Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Methylation Ewas Design loads about 6.3k tokens when it runs. Until then it costs about 261 tokens; SKILL.md has 3,103 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,103 words, ~6,345 tokens.
.claude/skills/bio-methylation-ewas-design/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: meffil 1.3+, sva 3.50+, bacon 1.30+, limma 3.58+, missMethyl 1.36+, pwrEWAS 1.16+, pandas 2.2+.
Before using code patterns, verify installed versions match. If versions differ:
packageVersion('<pkg>') then ?function_name to verify parameterspip show <package> then help(module.function) to check signaturesIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Two versions decide everything downstream. The ARRAY (450K vs EPIC v1 ~850K vs EPIC v2 ~935K) sets the CpG universe and therefore the effective number of tests and the genome-wide threshold - confirm against the current manifest before quoting a threshold. The CELL-TYPE REFERENCE panel must match the tissue and age (cord blood, adult blood, and tumor need different references); a mismatched reference silently biases the cell-fraction covariates. BACON slot names, pwrEWAS argument names, and sva method options have shifted across Bioconductor releases - check ?function on the installed build.
"Run an EWAS for my phenotype" -> Build the covariate model and the chip randomization BEFORE the per-CpG test, because the confounders are larger than the signal - the p-value is the last and least interesting thing.
meffil.ewas(beta, variable, covariates=data.frame(age, sex, cellprops, chip, plate)) then BACON-correct the test statistics.Scope: the design-and-inference layer of an EWAS - confounding, randomization, batch/SVA/RUV correction, genomic inflation/BACON, genome-wide thresholds, power, meta-analysis, replication, and methylation risk scores. The per-CpG test mechanics (beta vs M-value, limma) -> differential-cpg-testing. Cell-fraction estimation algorithms (Houseman/IDOL/EpiDISH) -> cell-type-deconvolution. Normalization internals (funnorm/noob) -> array-preprocessing. mQTL Mendelian randomization for causal orientation -> causal-genomics/mendelian-randomization. MRS weight-learning (penalized regression, nested CV) -> the machine-learning category.
The defining feature of the field is that the confounders are LARGER than the signal: a true effect is typically a sub-2% absolute methylation difference, while cell mix, batch, age, sex, smoking, and ancestry each move methylation 10-50%. An EWAS is therefore won or lost BEFORE the per-CpG test - in the chip randomization, the covariate set, and the replication cohort. Four corollaries, each of which a common misuse violates:
Organize the analysis around defending these four, not around listing sva/ComBat/bacon functions.
| Rank | Confounder | Why it dominates | Defense | Owner |
|---|---|---|---|---|
| 1 | Cell composition | Each leukocyte has a radically different methylome; any mixture shift is a fake signal | Estimate cell fractions and ALWAYS adjust | -> cell-type-deconvolution (estimation); decide here |
| 2 | Batch / Sentrix chip / array position / plate / scan date | EPIC runs 8 samples per BeadChip (450K: 12); position and chip imprint methylation | Randomize at design time; then SVA/RUVm/ComBat + chip/position covariates | here + array-preprocessing |
| 3 | Age and sex | Age is the most reproducible methylome correlate (the clock); sex drives X/Y and X-inactivation | Always include; analyze sex chromosomes separately; sex is a mislabel QC check | here |
| 4 | Smoking | Largest reproducible blood exposure signature; confounds half of all health phenotypes | Adjust (methylation smoking score > self-report); AHRR cg05575921 is the positive control | here |
| 5 | Genetic / ancestry / mQTL | Many CpGs are under genetic control; population structure confounds like GWAS | Genetic PCs; analyze within ancestry; drop SNP-affected/cross-reactive probes | here + array-preprocessing |
| 6 | Reverse causation / tissue relevance | Cross-sectional methylation may be a consequence, not a cause; blood != disease tissue | Prospective/longitudinal or MZ-discordant design; cross-tissue concordance | here; causal -> causal-genomics/mendelian-randomization |
Default covariate set for a blood EWAS: age, sex, cell proportions, chip/position (or SVs), and where relevant genetic PCs and a smoking score. The omission of any one is the most common EWAS error.
If all cases ran on chip A and all controls on chip B, batch and phenotype are inseparable - ComBat will remove the batch AND the signal, or preserve a spurious one. No statistical adjustment recovers a confounded design (general principle: experimental-design/batch-design).
Goal: Make technical batch orthogonal to phenotype before any sample touches the array.
Approach: Randomize (or block) sample-to-chip, sample-to-position, and sample-to-plate assignment so case/control, age, and sex are balanced across chips; reserve no chip for one group. This single uncorrectable decision matters more than every analysis choice that follows.
# Stratified randomization of samples to 8-position EPIC BeadChips (450K: 12), balancing case/control per chip
meta$chip <- NA_integer_
for (grp in split(seq_len(nrow(meta)), meta$case)) {
meta$chip[sample(grp)] <- rep(seq_len(ceiling(length(grp) / 4)), each = 4, length.out = length(grp))
}
# Then interleave the two case groups across chips so no chip is single-group; verify balance:
with(meta, table(chip, case)) # every chip should hold both groups| Tool | Citation | Mechanism / role | When |
|---|---|---|---|
| sva | Leek & Storey 2007 PLoS Genet 3:e161 | latent surrogate variables built orthogonal to the variable of interest | reference-free soak-up of cell mix + unknown batch |
| SmartSVA | Chen 2017 BMC Genomics 18:413 | order-of-magnitude-faster SVA with explicit convergence | the de-facto EWAS SVA for large cohorts |
| ComBat | Johnson 2007 Biostatistics 8:118 | empirical-Bayes removal of a SPECIFIED batch (chip/plate) | known, labelled batch; protect biology via mod= |
| RUVm | Maksimovic 2015 Nucleic Acids Res 43:e106 | two-stage RUV-inverse using Illumina's ~600 negative control probes | array EWAS; estimates unwanted variation from control probes |
| funnorm | Fortin 2014 Genome Biol 15:503 | control-probe PCA at the normalization stage | fix unwanted variation BEFORE testing (-> array-preprocessing) |
The central tension: cell-composition and technical variation MUST be removed, but aggressive correction removes REAL biology when the unwanted variation overlaps the phenotype. Every tool above can erase a true effect.
The positive-control check (the pragmatic referee). Monitor a known signal through correction. In a blood EWAS containing smokers, smoking->cg05575921 (AHRR, ~18% hypomethylation, Joehanes 2016) MUST survive. If adding surrogate variables makes the QQ plot look clean BUT kills AHRR, the pipeline is over-corrected. A clean QQ with a dead positive control is a broken pipeline, not a good one. Do not keep adding SVs until lambda hits 1.0.
ComBat and limma/EWAS modeling run on M-values (logit of beta), not betas (betas are bounded [0,1] and heteroscedastic); effect sizes are reported back on the beta / delta-beta scale for interpretability. Pass the biological covariate to ComBat's mod= so it is protected. sva: pass the FULL model (mod, including the variable of interest) AND the null model (mod0) so SVs are orthogonal to the phenotype; choose the number with num.sv(method='be').
Lambda (the genomic inflation factor) = median observed chi-square / expected median. In GWAS, lambda > 1 signals stratification and is corrected by genomic control (divide all statistics by lambda). This reasoning is WRONG for EWAS:
BACON (van Iterson 2017 Genome Biol 18:19) fits a Bayesian Gaussian mixture to the observed test statistics, estimates the empirical-null distribution, and reports BOTH a bias (mean shift) AND an inflation (scale) - then standardizes statistics against that empirical null without assuming the genome is mostly null. Report lambda but correct with BACON; show QQ plots before and after. Caveat: BACON can over-deflate a highly polygenic exposure if the alternative component is large - apply it with the QQ plot, not as a reflex.
| Threshold | Source | Rationale |
|---|---|---|
| 450K: P < 2.4e-7 | Saffari 2018 Genet Epidemiol 42:20 | empirical effective-test FWER 5%; ~210,000 effective tests (< probe count due to correlation) |
| EPIC v1: P < 9e-8 | Mansell 2019 BMC Genomics 20:366 | null-simulation FWER 5% for the ~850K EPIC array |
| Pragmatic 1e-7 | community | round number between the two array-specific values |
| FDR (BH) q < 0.05 | Benjamini-Hochberg | more powerful; discovery / exposure scans; report as a sensitivity layer |
Naive Bonferroni on the probe count is over-conservative because CpGs are correlated (co-methylation, shared mQTLs); raw p is anti-conservative because there are ~850K tests. Use the array-specific FWER threshold for the cross-study-comparable headline claim and the EWAS Catalog; use BH-FDR for hypothesis-generating discovery. BH's independence assumption is imperfect under probe correlation but robust to positive dependence. Region-level (DMR) multiple-testing is owned by dmr-detection; mechanical p.adjust is owned by differential-cpg-testing. EPIC v2 (~935K probes, renamed/dropped vs v1) may shift the threshold - verify against the manifest.
Goal: Size an EWAS for a realistic (tiny) effect using site-specific methylation variance, not a single assumed sd.
Approach: pwrEWAS (Graw 2019 BMC Bioinformatics 20:218) simulates DNAm using empirical per-CpG variance from a reference dataset, so power reflects the real site-specific variance structure (bimodal sites near 0/1 behave differently from intermediate sites). Specify target delta-beta, sample size, and the genome-wide threshold; expect to need hundreds-to-thousands of samples for a 1-2% effect - which is why meta-analysis dominates.
EWAS meta-analysis is the field standard for power: each cohort runs an identical pre-specified pipeline (meffil enables this), BACON-corrects per cohort, uploads summary statistics, and a central site does fixed-effect inverse-variance pooling with heterogeneity testing (Cochran's Q, I^2). meffil's distinguishing feature is distributed normalization - cohorts normalize locally without sharing individual data, reducing meta-analysis heterogeneity - plus automated selection of the number of normalization PCs.
| Design | Controls for | Cost |
|---|---|---|
| Case-control (cross-sectional) | nothing inherently; adjust on age/sex/ancestry/cell-mix, randomize chips | reverse causation, composition confounding |
| Longitudinal / prospective | reverse causation (methylation measured before outcome); within-person change via mixed model | needs follow-up; repeated samples |
| MZ-discordant twin | genetic + shared-environment confounding by design (paired within-pair analysis) | discordant pairs are rare -> power-limited |
| Exposure EWAS | smoking is the template/positive-control | exposure misclassification (self-report) |
| Meta-analysis | low power of single cohorts | requires harmonized pipelines + BACON per cohort |
Replication is the real significance bar. After discovery, triage every hit against both databases - a "novel" CpG that is in fact a top smoking/age/blood-cell CpG is almost certainly residual confounding.
An EWAS produces per-CpG associations; an MRS turns many of them into one number - a weighted CpG sum, MRS = sum_j w_j * beta_ij, with weights from a training EWAS or a penalized fit. It is the methylation analogue of a polygenic risk score (PRS), and an epigenetic clock is the age/health special case (-> epigenetic-clocks).
The load-bearing distinction from a PRS. A PRS sums germline variants - fixed at conception, antecedent, plausibly causal. An MRS sums methylation - modifiable, tissue/time-specific, and frequently a CONSEQUENCE of the exposure/trait rather than a cause. The methylation smoking score (Elliott 2014 Clin Epigenetics 6:4; generalized by Sugden 2019) does not predict a propensity to smoke; it MEASURES the footprint smoking left. So an MRS is reverse-causal-by-default: report it as a predictive BIOMARKER / objective exposure proxy, and reserve "risk" and "cause" for prospectively- or MR-supported claims. PRS = germline cause; MRS = state consequence.
The most useful design role is as a better-measured confounder: self-reported smoking is biased and coarse, so adjusting an EWAS for the methylation smoking score controls residual smoking confounding far better whenever the phenotype is smoking-correlated (caveat: if smoking is on the causal path, over-adjusting via the score removes real signal - the same confounder-vs-mediator tension as cell composition). Other DNAm scores exist for BMI/alcohol/education (McCartney 2018) and circulating proteins (EpiScores, Gadd 2022 eLife 11:e71802). Portability fails across array (450K vs EPIC drop CpGs), tissue, and ancestry - report how many score CpGs are present on the array. Defer MRS weight-learning (penalized regression, nested CV, calibration, leakage) to the machine-learning category; an MRS validated in its own training cohort is not validated.
Trigger: running the regression on whole-blood betas with no cell-fraction adjustment. Mechanism: any phenotype that shifts the leukocyte mixture moves methylation 10-50%. Symptom: the top hits are known granulocyte/lymphocyte-proportion CpGs. Fix: estimate cell proportions (-> cell-type-deconvolution) and always include them; treat any unadjusted hit as composition until shown otherwise.
Trigger: all cases on chip A, controls on chip B. Mechanism: batch and phenotype are inseparable. Symptom: ComBat removes the signal or fabricates one; no error. Fix: randomize sample-to-chip/position at design time - this is uncorrectable after the bench.
Trigger: dividing all test statistics by lambda. Mechanism: assumes single multiplicative inflation on a mostly-null genome and ignores bias. Symptom: real polygenic signal (smoking) deflated or residual confounding under-corrected. Fix: use BACON to estimate empirical-null bias AND inflation.
Trigger: adding SVs/PCs until lambda hits 1.0. Mechanism: latent factors correlated with the phenotype deflate true signal. Symptom: clean QQ but the positive control (AHRR) is gone. Fix: monitor cg05575921 through correction; stop when adding SVs starts killing it.
Trigger: 0.05/850000, or reporting nominal p. Mechanism: probe correlation makes Bonferroni too strict; ignoring 850K tests is too loose. Symptom: missed real hits or a flood of false ones. Fix: array-specific FWER threshold (2.4e-7 / 9e-8) for claims, BH-FDR for discovery.
Trigger: a blood EWAS for a brain/adipose/tumor phenotype read as tissue biology. Mechanism: blood-brain methylation r ~0.15. Symptom: hits with no plausible blood mechanism. Fix: justify tissue relevance (BECon/blood-brain tools); frame blood as a biomarker, not mechanism, unless the CpG is cross-tissue concordant.
Trigger: describing a DNAm score for a disease as "risk" like a PRS. Mechanism: methylation is modifiable and frequently downstream. Symptom: a "DNAm risk score" that is actually the footprint of disease/treatment/behavior. Fix: call it a predictive biomarker; orient cause vs consequence only with prospective or MR evidence.
| Threshold | Source | Rationale |
|---|---|---|
| 450K genome-wide P < 2.4e-7 | Saffari 2018 Genet Epidemiol 42:20 | empirical effective-test FWER 5% (~210,000 effective tests) |
| EPIC genome-wide P < 9e-8 | Mansell 2019 BMC Genomics 20:366 | null-simulation FWER 5% for ~850K EPIC |
| BH-FDR q < 0.05 | Benjamini-Hochberg | discovery / exposure scans; report alongside FWER |
| AHRR cg05575921 ~18% hypomethylation in smokers | Joehanes 2016 Circ Cardiovasc Genet 9:436 | the canonical positive control; must survive correction |
| EWAS Catalog inclusion P < 1e-4 | Battram 2022 Wellcome Open Res 7:41 | a Catalog entry is a lookup, not a genome-wide claim |
| Detect 1-2% delta-beta -> hundreds-to-thousands of samples | pwrEWAS, Graw 2019 | tiny site-specific effects drive the field to meta-analysis |
| Run ComBat/limma on M-values, report delta-beta | Du 2010 BMC Bioinformatics 11:587 | betas are bounded/heteroscedastic; M-values stabilize variance |
| Error / symptom | Cause | Solution |
|---|---|---|
| Top hits are blood-cell CpGs | no cell-composition covariates | estimate and adjust cell fractions |
| ComBat removed the signal | batch confounded with phenotype | randomize at design time; cannot fix after |
| Lambda misread as pass/fail | GWAS dogma applied to EWAS | pair lambda with QQ shape + positive control; correct with BACON |
| Clean QQ but AHRR gone | over-correction with too many SVs | stop adding SVs when the positive control dies |
| Threshold flood or famine | Bonferroni-on-probe-count or raw p | use array-specific FWER / BH-FDR |
| "Novel" CpG is a known generic hit | no Catalog/Atlas triage | look up the CpG before claiming novelty |
| MRS called a risk score | PRS connotations on modifiable methylation | report as biomarker; orient cause only with prospective/MR design |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in methylation-analysis/ewas-design of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Methylation Ewas Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Methylation Ewas Design this skillGPTomics/bioSkills | 1.2k | 1 repos | ~6.3k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible. Bio Methylation Ewas Design is an agent skill from GPTomics/bioSkills. Designs and defends an epigenome-wide association study (EWAS) on 450K/EPIC array or bisulfite methylation - the layer deciding whether a hit is credible.
Bio Methylation Ewas Design fits situations like: designing an EWAS; choosing a covariate set; randomizing a plate layout; interpreting lambda.
Run `npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a claude-code`. Or copy the skill folder (methylation-analysis/ewas-design in GPTomics/bioSkills) into .claude/skills/bio-methylation-ewas-design in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a codex`. Or copy the skill folder (methylation-analysis/ewas-design in GPTomics/bioSkills) into .agents/skills/bio-methylation-ewas-design in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-methylation-ewas-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-methylation-ewas-design, .gemini/skills/bio-methylation-ewas-design, .github/skills/bio-methylation-ewas-design and .opencode/skills/bio-methylation-ewas-design in your project.
Going by SKILL.md and its folder, Bio Methylation Ewas Design needs R and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Methylation Ewas Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Methylation Ewas Design: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.