Agent skill

Bio Methylation Dmr Detection

by GPTomics in GPTomics/bioSkills

Detects differentially methylated regions (DMRs) from short-read bisulfite (WGBS/RRBS), array, and long-read methylation count tables using dmrseq (permutation region-FDR over the region selection)…

MITAuto-check passedResearch & Science

Install Bio Methylation Dmr Detection

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-methylation-dmr-detection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-methylation-dmr-detection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/methylation-analysis/dmr-detection .claude/skills/bio-methylation-dmr-detection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-methylation-dmr-detection
GitHub stars
1.2k
Used in
1 other repo
Token cost
~6.1k tokens
SKILL.md length
2,509 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Detects differentially methylated regions (DMRs) from short-read bisulfite (WGBS/RRBS), array, and long-read methylation count tables using dmrseq (permutation region-FDR over the region selection)…

  • Works in 3 steps: dmrseq neutralizes the selection; the… → Region q-values are NOT comparable… → Thresholds are conventions, not biology.…
  • Calling region-level methylation differences
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 13 more sections
  • Runs R scripts from its folder

What it does

Bio Methylation Dmr Detection is an agent skill from GPTomics/bioSkills. Detects differentially methylated regions (DMRs) from short-read bisulfite (WGBS/RRBS), array, and long-read methylation count tables using dmrseq (permutation region-FDR over the region selection), DSS callDMR (beta-binomial), methylKit tiles, bsseq BSmooth, DMRcate Gaussian-kernel smoothing, metilene, and comb-p. Covers why a DMR is DEFINED by arbitrary thresholds (min-CpGs, max-gap, delta-beta, q) and a smoothing bandwidth, why selecting extreme runs of CpGs then testing them on the same data is post-selection…

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Calling region-level methylation differences
  • Choosing a DMR caller
  • Controlling region-level FDR
  • Segmenting megabase methylation domains

Example prompts

  • “Use the bio-methylation-dmr-detection skill to detect differentially methylated regions (DMRs) from short-read bisulfite (WGBS/RRBS), array, and…”
  • “/bio-methylation-dmr-detection”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. dmrseq neutralizes the selection; the others combine-and-correct but do not re-select. dmrseq (Korthauer 2019 Biostatistics 20:367) builds…
  2. Region q-values are NOT comparable across tools. "I found 4,000 DMRs at q<0.05" is meaningless without naming the tool and what its q…
  3. Thresholds are conventions, not biology. delta-beta 25%, min-CpGs 3, max-gap 1000bp, q<0.01 are tutorial folklore. The same data yields…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Methylation Dmr Detection loads about 6.1k tokens when it runs. Until then it costs about 254 tokens; SKILL.md has 2,509 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~254
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,509 words, ~6,120 tokens.

Download SKILL.mdSave it as .claude/skills/bio-methylation-dmr-detection/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-methylation-dmr-detection
description
Detects differentially methylated regions (DMRs) from short-read bisulfite (WGBS/RRBS), array, and long-read methylation count tables using dmrseq (permutation region-FDR over the region selection), DSS callDMR (beta-binomial), methylKit tiles, bsseq BSmooth, DMRcate Gaussian-kernel smoothing, metilene, and comb-p. Covers why a DMR is DEFINED by arbitrary thresholds (min-CpGs, max-gap, delta-beta, q) and a smoothing bandwidth, why selecting extreme runs of CpGs then testing them on the same data is post-selection inference, why region q-values are not comparable across tools, and a single-sample domain-segmentation section (PMD, UMR/LMR, MethylSeekR, solo-WCGW) that must run before focal calling on cancer/aging genomes. Use when calling region-level methylation differences, choosing a DMR caller, controlling region-level FDR, or segmenting megabase methylation domains. For per-site testing see differential-cpg-testing; for the methylKit object model see methylkit-analysis.
tool_type
r
primary_tool
dmrseq

Version Compatibility

Reference examples tested with: dmrseq 1.22+, DSS 2.50+, methylKit 1.28+, bsseq 1.38+, DMRcate 2.16+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

The GENOME BUILD is a version that matters. methylKit/bsseq/DMRcate assembly= is metadata, but annotation packages (annotatr::build_annotations(genome='hg38'), TxDb.Hsapiens.UCSC.hg38.knownGene) are build-specific and must match the alignment genome. DMRcate arraytype='EPIC'/'450K' and the IlluminaHumanMethylation annotation package set the array CpG universe. DMRcate defaults shift across Bioconductor releases (the C kernel scaling has no single default - it is platform-dependent); confirm with ?dmrcate on the installed build.

DMR Detection

"Find differentially methylated regions" -> Detect candidate regions, then score them with a null that re-ran the selection - because a DMR is defined by a chain of thresholds, not found, and only a selection-aware q-value is honest.

  • R: dmrseq(bs, testCovariate='condition') (selection-aware) ; DSS::callDMR(), methylKit::tileMethylCounts(), bsseq BSmooth, DMRcate::dmrcate() (combine-and-correct)

Scope: REGION-level differential methylation from any per-CpG methylation+coverage table (WGBS/RRBS counts, array beta/M, long-read modkit bedMethyl), plus single-sample domain segmentation. Per-site DMC/DMP testing -> differential-cpg-testing. The methylKit import/filter/unite object model -> methylkit-analysis. Long-read MM/ML calling that produces the counts -> long-read-sequencing/nanopore-methylation. Functional enrichment of DMR genes -> pathway-analysis/go-enrichment (with the CpG-bias correction noted below).

The Single Most Important Modern Insight -- A DMR Is DEFINED, Not Found, and the Region p-Value Is Only Honest If Its Null Re-Ran the Selection

The naive recipe - compute a per-CpG statistic, select runs of CpGs that look extreme to DEFINE candidate regions, then test those same regions and report a p/q on the same data - reuses the data twice. The regions were CHOSEN because they were extreme; scoring them with the data that selected them inflates significance and produces uncalibrated region FDR. This is post-selection inference, the field's original sin. No error is thrown; the q-values just lie. Three corollaries:

  1. dmrseq neutralizes the selection; the others combine-and-correct but do not re-select. dmrseq (Korthauer 2019 Biostatistics 20:367) builds the null by PERMUTING the condition labels and RE-RUNNING the entire candidate-detection procedure on each permutation, pooling permuted region statistics into one genome-wide null - so its q-value is on REGIONS and accounts for selection. comb-p and DMRcate combine existing per-CpG p-values (Stouffer/Fisher/SLK) and apply a region multiplicity correction, but never re-run selection. methylKit tiles use FIXED windows (boundaries not chosen from the data, so the selection part is sidestepped) but ignore inter-tile correlation. DSS callDMR merges significant CpGs (threshold-then-define).

  2. Region q-values are NOT comparable across tools. "I found 4,000 DMRs at q<0.05" is meaningless without naming the tool and what its q controls. Cross-tool OVERLAP is the real evidence - run two callers and intersect.

  3. Thresholds are conventions, not biology. delta-beta 25%, min-CpGs 3, max-gap 1000bp, q<0.01 are tutorial folklore. The same data yields radically different DMR sets under different settings. Report and justify every knob; never present one config as correct.

Organize the analysis around defending these, not around listing functions.

Tool Taxonomy

ToolCitationMechanism / roleWhen
dmrseqKorthauer 2019 Biostatistics 20:367GLS area statistic + permutation null that re-runs selection; smooths the difference internallyheadline WGBS inference; calibrated region FDR; >=2 reps/group
DSS callDMRFeng 2014 Nucleic Acids Res 42:e69; Park & Wu 2016 Bioinformatics 32:1446per-CpG Bayesian beta-binomial dispersion shrinkage, then merge significant CpGssmall n, complex/multi-factor designs; low-coverage with smoothing
methylKit tilesAkalin 2012 Genome Biol 13:R87fixed windows + logistic/F test per tilefast RRBS/WGBS screening; reuses the methylKit object model
bsseq (BSmooth)Hansen 2012 Genome Biol 13:R83per-sample local-likelihood smoothing -> smoothed t-statisticlow/uneven-coverage WGBS; superseded by dmrseq for calibrated FDR
DMRcatePeters 2015 Epigenetics Chromatin 8:6; Peters 2021 Nucleic Acids Res 49:e109Gaussian-kernel smoothing of the per-CpG statisticarrays (450K/EPIC) and WGBS (different kernel)
metileneJuhling 2016 Genome Res 26:256binary segmentation + 2D Kolmogorov-Smirnov; standalone C CLIfast whole-genome second caller; data-adaptive boundaries; tolerates missingness
comb-pPedersen 2012 Bioinformatics 28:2986ACF + Stouffer-Liptak-Kechris combination + Sidak; Python CLI on any p-valuesarray EWAS regions; tool-agnostic corroboration of any per-CpG p-values

Decision Tree by Scenario

ScenarioRecommendedWhy
WGBS, >=2 reps/group, headline region inferencedmrseqonly caller whose region FDR accounts for selection
WGBS, small n, multi-factor / covariatesDSS (DMLtest.multiFactor -> callDMR)beta-binomial dispersion shrinkage + model formula
Low / uneven coverage WGBSdmrseq (smooths internally) or DSS smoothing=TRUEsmoothing borrows strength across CpGs
RRBS, quick region screenmethylKit tiles (filtered, cov.bases>=3)fixed windows; fast; CpG-island enriched; screening q, not selection-corrected
EPIC/450K arrayDMRcate array mode (cpg.annotate('array')) or comb-p on limma pkernel tuned for array spacing; methylKit/bsseq/dmrseq are count-based
Corroborate any DMR setrun a second caller, intersectregion q is not comparable across tools
Cancer / aging / placenta / cultured-cell WGBSsegment PMDs FIRST (see domain section)a focal caller manufactures fake hypo-DMRs from PMD background
Per-CpG, not region-> differential-cpg-testingsite-level test before aggregating to regions
Long-read modkit bedMethyl input-> long-read-sequencing/nanopore-methylation, then any caller heremodkit calls; feed Nmod + Nvalid_cov here for region statistics

dmrseq: The Selection-Aware Headline Caller

Goal: Call WGBS DMRs with a region-level FDR that survives the region-selection step.

Approach: Build a BSseq object from raw counts (do NOT pre-smooth), filter loci with zero coverage in any sample, then run dmrseq, which detects candidate regions (runs of CpGs whose smoothed methylation-difference coefficient exceeds cutoff), computes one GLS area statistic per region, and generates the null by permuting labels and re-running detection.

r
library(dmrseq)
library(bsseq)

bs <- read.bismark(c('ctrl1.cov.gz', 'ctrl2.cov.gz', 'treat1.cov.gz', 'treat2.cov.gz'),
                   colData = DataFrame(condition = c('ctrl', 'ctrl', 'treat', 'treat')),
                   rmZeroCov = TRUE, strandCollapse = TRUE)

# dmrseq requires every locus to have non-zero coverage in EVERY sample.
bs <- bs[rowSums(getCoverage(bs) == 0) == 0, ]

# cutoff=0.1 only SEEDS candidate detection (10% smoothed difference); significance
# comes from the permutation statistic, so dmrseq does NOT hard-threshold delta-beta.
dmrs <- dmrseq(bs, testCovariate = 'condition', cutoff = 0.1)
sig <- dmrs[dmrs$qval < 0.05]   # qval is a REGION FDR that accounts for selection

dmrseq smooths the DIFFERENCE internally (bpSpan/minInSpan/maxGapSmooth); running BSmooth() first double-smooths and invalidates the model. The permutation null is COARSE at 2-vs-2 (few distinct label permutations) - the package pools across all candidate regions to compensate, but more replicates give a finer null. Use adjustCovariate for nuisance variables and block = TRUE for large-scale differential blocks.

DSS: Beta-Binomial Dispersion Shrinkage

Goal: Call DMRs with per-CpG dispersion shrinkage for small n or a multi-factor design.

Approach: DMLtest does per-CpG Wald tests with a Bayesian beta-binomial dispersion estimate, then callDMR merges significant CpGs into regions.

r
library(DSS)

bs <- makeBSseqData(list(c1, c2, t1, t2), c('C1', 'C2', 'T1', 'T2'))   # each: chr/pos/N/X data.frame
dml <- DMLtest(bs, group1 = c('C1', 'C2'), group2 = c('T1', 'T2'), smoothing = TRUE)   # TRUE for low-cov WGBS

# callDMR defaults (verify on installed build): delta=0, p.threshold=1e-5,
# minlen=50, minCG=3, dis.merge=100, pct.sig=0.5.
# delta=0 means NO effect-size floor - SET it explicitly so tiny shifts are not called.
dmrs <- callDMR(dml, delta = 0.1, p.threshold = 1e-5, minlen = 50, minCG = 3,
                dis.merge = 100, pct.sig = 0.5)   # pct.sig=0.5: >=50% of region CpGs individually significant

callDMR merges significant CpGs (threshold-then-define), so its region p is NOT selection-corrected; its strengths are dispersion shrinkage and multi-factor support (DMLtest.multiFactor with a model formula).

bsseq BSmooth

Goal: Call DMRs on low/uneven-coverage WGBS by smoothing each sample before testing.

Approach: Smooth per sample, compute the smoothed t-statistic, then threshold it into regions.

r
library(bsseq)

bs_smooth <- BSmooth(bs, BPPARAM = MulticoreParam(4), verbose = TRUE)
keep <- rowSums(getCoverage(bs_smooth) >= 2) == ncol(bs_smooth)   # >=2x in every sample
bs_filt <- bs_smooth[keep, ]

tstat <- BSmooth.tstat(bs_filt, group1 = c('C1', 'C2'), group2 = c('T1', 'T2'),
                       estimate.var = 'same', mc.cores = 4)
dmrs <- dmrFinder(tstat, cutoff = c(-4.6, 4.6))   # t-stat cutoff (Hansen 2012 uses quantile-based cutoffs)

dmrFinder needs BSmooth.tstat output, not the smoothed BSseq object directly. BSmooth gives a RANKED DMR list, not a calibrated region FDR - dmrseq (same lab lineage) supersedes it for region inference. Over-smoothing (too-wide bandwidth) washes out focal promoter DMRs.

DMRcate: The Array-vs-WGBS Fork

Goal: Call DMRs by Gaussian-kernel smoothing of the per-CpG statistic, with the correct kernel for the platform.

Approach: Annotate per-CpG statistics through the array OR sequencing entry point, then smooth and extract.

r
library(DMRcate)

# ARRAY (450K/EPIC): beta/M matrix; arraytype sets the CpG universe.
design <- model.matrix(~ condition)
ann_array <- cpg.annotate('array', m_values, what = 'M', arraytype = 'EPIC',
                          analysis.type = 'differential', design = design, coef = 2)
dmrs_array <- extractRanges(dmrcate(ann_array, lambda = 1000, C = 2))   # array kernel

# WGBS: different entry point AND a much smaller kernel - the array C=2 over-smooths
# dense sequencing CpGs ~25x. Annotate from a count/edgeR-DSS path, then C=50.
ann_seq <- sequencing.annotate(bs, design = design, coef = 2)
dmrs_seq <- extractRanges(dmrcate(ann_seq, C = 50))   # WGBS kernel (Peters 2021)

The default lambda=1000, C=2 is ARRAY-only; applying it to WGBS produces massively over-smoothed, merged, inflated DMRs. pcutoff='fdr' returns no DMRs if the upstream limma/DSS yields no significant CpGs.

metilene and comb-p (Second Callers)

metilene is a standalone C CLI taking one tab table (chrom, pos, per-sample methylation rate); binary segmentation + a 2D-KS test find data-adaptive boundaries. comb-p is a Python CLI that takes a BED of per-CpG p-values from ANY upstream test, estimates the p-value autocorrelation, does a Stouffer-Liptak-Kechris combination of neighbors, groups regions, and applies a one-step Sidak correction.

bash
metilene -M 1000 -m 10 -d 0.1 -a g1 -b g2 input.tsv | metilene_output.pl   # -m min CpGs, -d min mean diff
comb-p pipeline -c 4 --seed 0.01 --dist 500 --step 50 -p out methyl_pvals.bed   # --seed = p to start a region

Both combine-and-correct rather than re-select; use them as fast corroboration and intersect with dmrseq.

Thresholds Are Conventions; Region FDR Is Tool-Specific

Every caller exposes the same coupled knobs under different names: min-CpGs (minNumRegion/minCG/-m/min.cpgs/cov.bases), max-gap (maxGap/dis.merge/-M), delta-beta (cutoff/delta/-d/betacutoff/difference), and a significance cutoff. Shrinking max-gap, raising min-CpGs, and raising delta all reduce the DMR count, and the same data yields wildly different DMR sets. The phrase "region-level FDR" means three different objects: a selection-aware permutation FDR (dmrseq), a BH/Sidak correction on combined per-CpG p-values (DMRcate/comb-p), or per-unit q on independent tiles/CpGs (methylKit/DSS). Report all knobs, name the tool, and use cross-tool overlap as the evidence statement.

DMR-to-Gene Mapping and the CpG-Density Enrichment Bias

Goal: Interpret DMRs without inflating enrichment from CpG-rich genes.

Approach: Annotate DMRs to features (annotatr returns one row per DMR-feature overlap; genomation collapses by precedence), then run enrichment with a method that corrects for CpG/probe count.

r
library(annotatr)
annots <- build_annotations(genome = 'hg38', annotations = c('hg38_basicgenes', 'hg38_cpg_islands'))
dmr_ann <- annotate_regions(regions = sig, annotations = annots, ignore.strand = TRUE)   # one row per overlap

# Enrichment: methylation has a CpG-density bias (CpG-rich genes harbor DMRs by chance),
# so a plain hypergeometric GO test is biased. Use missMethyl goregion (probe/CpG-bias-aware).
# missMethyl::goregion(sig_ranges, all.cpg=..., collection='GO', array.type='EPIC')

A single DMR commonly overlaps or sits between several genes; mapping DMR -> gene (nearest TSS vs overlap vs within-X-kb) is a modeling choice that changes the gene list. Hand the corrected enrichment to pathway-analysis/go-enrichment, flagging that methylation input needs a CpG-bias-aware method (missMethyl gometh/goregion), not a generic hypergeometric test.

Show full SKILL.md (1,110 more words)Show less

Single-Sample Domain Structure (NOT Differential)

This is a DIFFERENT problem from the focal between-group callers above. The mammalian methylome partitions at MEGABASE scale into Highly Methylated Domains (HMDs, ~80-90%, ordered) and Partially Methylated Domains (PMDs, ~40-70%, disordered, high-variance), and PMDs coincide with late replication, Lamina-Associated Domains, and the Hi-C B-compartment (Lister 2009 Nature 462:315; Berman 2012 Nat Genet 44:40). Cancer "global hypomethylation" is a DOMAIN phenomenon - focal CpG-island hypermethylation sitting ON a background of megabase PMD hypomethylation - not a focal one. Domain structure is a SINGLE-SAMPLE, structural question answered by SEGMENTERS, not by any between-group DMR caller (methylKit/DSS/dmrseq/DMRcate/metilene/comb-p have no single-sample segmentation mode).

  • MethylSeekR (Burger 2013 Nucleic Acids Res 41:e155) segments one WGBS methylome into UMRs (CpG-rich unmethylated = promoters/CGIs), LMRs (CpG-poor low-methylated ~30% = distal enhancers), and PMDs. Pipeline: readMethylome() -> plotAlphaDistributionOneChr() (diagnostic: does the sample have PMDs?) -> segmentPMDs() (2-state Gaussian HMM, 101-CpG windows) -> calculateFDRs() -> segmentUMRsLMRs(m=0.5, n=..., pmdGRanges=...). PMDs MUST be masked before UMR/LMR calling or PMD disorder spawns spurious LMRs.
  • solo-WCGW (Zhou 2018 Nat Genet 50:591) - an isolated CpG in [A/T]CG[A/T] context - loses methylation fastest and most monotonically with cell division and is the most sensitive PMD/mitotic-clock readout, detecting PMD hypomethylation even in near-normal tissue. Quantify as the mean over the published common-PMD solo-WCGW CpG set, not as a DMR. See epigenetic-clocks for the broader clock taxonomy.

The warning: running a focal DMR caller on a PMD-bearing genome manufactures thousands of fake hypo-DMRs that are really one phenomenon - the PMD background shifting - chopped into pieces by the max-gap/min-CpG knobs. Segment domains FIRST, then EXCLUDE PMD intervals from focal calling or STRATIFY every DMR by in-PMD vs out-of-PMD and report the fraction that is PMD background.

Per-Method Failure Modes

PMD background reported as DMRs

Trigger: focal caller on tumor/aged/placenta/cultured WGBS without domain screening. Mechanism: megabase PMD hypomethylation chopped into pieces by max-gap/min-CpG. Symptom: thousands of large hypo-DMRs in gene-desert, late-replicating, low-CpG-density coordinates. Fix: segment PMDs (MethylSeekR) first; exclude or stratify; report the PMD fraction.

Pre-smoothing before dmrseq

Trigger: BSmooth() then feeding the smoothed object to dmrseq. Mechanism: dmrseq smooths the difference internally; pre-smoothing double-smooths. Symptom: distorted candidate regions and invalid statistics. Fix: feed dmrseq the raw BSseq counts.

DMRcate array defaults on WGBS

Trigger: copying lambda=1000, C=2 onto sequencing data. Mechanism: the array kernel is ~25x too wide for dense WGBS CpGs. Symptom: massively over-smoothed, merged, inflated DMRs. Fix: sequencing.annotate() + C=50 for WGBS (Peters 2021).

Threshold-then-test reported as region FDR

Trigger: greping runs of significant per-CpG calls and reporting the per-CpG q. Mechanism: the regions were selected for extremeness, then tested on the same data. Symptom: anti-conservative, uncalibrated region q. Fix: use dmrseq (selection-aware) for the headline; at minimum state that a tile/merge q is a screening q.

Single-CpG tiles

Trigger: tileMethylCounts at the default cov.bases=0. Mechanism: a window with one covered CpG becomes a "DMR." Symptom: thousands of single-CpG noisy regions. Fix: set cov.bases >= 3.

Cross-tool count comparison

Trigger: comparing "N DMRs at q<0.05" between callers. Mechanism: each tool's q controls a different object. Symptom: apparent disagreement that is really an FDR-definition mismatch. Fix: compare OVERLAP, not counts.

Quantitative Thresholds

ThresholdSourceRationale
dmrseq cutoff 0.1 (candidate seed only)Korthauer 2019seeds detection; significance is the permutation statistic, NOT a delta floor
DSS callDMR(delta=0) default -> SET itPark & Wu 2016; DSS docsdelta=0 calls regions with no effect-size floor; set ~0.1 to require a real shift
min-CpGs per region 3-5conventionsingle-CpG "regions" are DMPs in disguise; trades sensitivity vs specificity
delta-beta 25% ("moderate")methylKit tutorial folkloreNOT derived; biologically meaningful delta is feature- and purity-dependent
methylKit tileMethylCounts(cov.bases>=3)nuance (default is 0)the default 0 lets single-CpG tiles through
coverage floor ~10x per CpGfield standarda single-CpG beta below ~10x is a coin flip
DMRcate WGBS C=50 (array C=2)Peters 2021 Nucleic Acids Res 49:e109dense WGBS CpGs need a far narrower kernel than array probes
MethylSeekR m=0.5, FDR<5%Burger 2013methylation cutoff for hypomethylated regions; FDR target picks the CpG-count threshold n

Common Errors

Error / symptomCauseSolution
dmrseq error about zero-coverage locia locus has 0 coverage in some samplefilter rowSums(getCoverage(bs)==0)==0 first
Thousands of huge hypo-DMRs in cancer WGBSPMD background not segmentedMethylSeekR segmentPMDs first; exclude/stratify
Over-merged WGBS DMRs with DMRcatearray kernel on sequencingsequencing.annotate() + C=50
dmrFinder errors on a smoothed objectneeds BSmooth.tstat outputrun BSmooth.tstat before dmrFinder
DMRcate returns no DMRspcutoff='fdr' and no significant upstream CpGscheck the upstream limma/DSS result first
Biased GO enrichment of DMR genesplain hypergeometric ignores CpG densityuse missMethyl goregion/gometh

References

  • Korthauer K, Chakraborty S, Benjamini Y, Irizarry RA. 2019. Detection and accurate false discovery rate control of differentially methylated regions from whole genome bisulfite sequencing. Biostatistics 20:367-383.
  • Feng H, Conneely KN, Wu H. 2014. A Bayesian hierarchical model to detect differentially methylated loci from single nucleotide resolution sequencing data. Nucleic Acids Res 42:e69.
  • Park Y, Wu H. 2016. Differential methylation analysis for BS-seq data under general experimental design. Bioinformatics 32:1446-1453.
  • Hansen KD, Langmead B, Irizarry RA. 2012. BSmooth: from whole genome bisulfite sequencing reads to differentially methylated regions. Genome Biol 13:R83.
  • Akalin A, Kormaksson M, Li S, et al. 2012. methylKit: a comprehensive R package for the analysis of genome-wide DNA methylation profiles. Genome Biol 13:R87.
  • Peters TJ, Buckley MJ, Statham AL, et al. 2015. De novo identification of differentially methylated regions in the human genome. Epigenetics Chromatin 8:6.
  • Peters TJ, Buckley MJ, Chen Y, et al. 2021. Calling differentially methylated regions from whole genome bisulphite sequencing with DMRcate. Nucleic Acids Res 49:e109.
  • Juhling F, Kretzmer H, Bernhart SH, Otto C, Stadler PF, Hoffmann S. 2016. metilene: fast and sensitive calling of differentially methylated regions from bisulfite sequencing data. Genome Res 26:256-262.
  • Pedersen BS, Schwartz DA, Yang IV, Kechris KJ. 2012. Comb-p: software for combining, analyzing, grouping and correcting spatially correlated P-values. Bioinformatics 28:2986-2988.
  • Lister R, Pelizzola M, Dowen RH, et al. 2009. Human DNA methylomes at base resolution show widespread epigenomic differences. Nature 462:315-322.
  • Berman BP, Weisenberger DJ, Aman JF, et al. 2012. Regions of focal DNA hypermethylation and long-range hypomethylation in colorectal cancer coincide with nuclear lamina-associated domains. Nat Genet 44:40-46.
  • Zhou W, Dinh HQ, Ramjan Z, et al. 2018. DNA methylation loss in late-replicating domains is linked to mitotic cell division. Nat Genet 50:591-602.
  • Burger L, Gaidatzis D, Schubeler D, Stadler MB. 2013. Identification of active regulatory regions from DNA methylation data. Nucleic Acids Res 41:e155.
  • differential-cpg-testing - Per-site testing before region aggregation
  • methylkit-analysis - methylKit object model and tile construction
  • methylation-calling - Produces the input count tables
  • array-preprocessing - Array beta/M-value input for DMRcate array mode
  • epigenetic-clocks - Mitotic-clock / solo-WCGW overlap (domain section)
  • pathway-analysis/go-enrichment - CpG-bias-aware enrichment of DMR genes (missMethyl gometh)
  • long-read-sequencing/nanopore-methylation - Pipe modkit bedMethyl counts here for region statistics
  • workflows/methylation-pipeline - End-to-end bisulfite pipeline

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in methylation-analysis/dmr-detection of GPTomics/bioSkills.

  • SKILL.md
  • examples/dmr_dmrseq.R
  • examples/dmr_methylkit.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Methylation Dmr Detection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Methylation Dmr Detection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Methylation Dmr Detection this skillGPTomics/bioSkills1.2k1 repos~6.1kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Methylation Dmr Detection

What does Bio Methylation Dmr Detection do?

Detects differentially methylated regions (DMRs) from short-read bisulfite (WGBS/RRBS), array, and long-read methylation count tables using dmrseq (permutation region-FDR over the region selection)…. Bio Methylation Dmr Detection is an agent skill from GPTomics/bioSkills. Detects differentially methylated regions (DMRs) from short-read bisulfite (WGBS/RRBS), array, and long-read methylation count tables using dmrseq (permutation region-FDR over the region selection), DSS callDMR (beta-binomial), methylKit tiles, bsseq BSmooth, DMRcate Gaussian-kernel smoothing, metilene, and comb-p.

When should I use Bio Methylation Dmr Detection?

Bio Methylation Dmr Detection fits situations like: calling region-level methylation differences; choosing a DMR caller; controlling region-level FDR; segmenting megabase methylation domains.

How do I install Bio Methylation Dmr Detection in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-methylation-dmr-detection -a claude-code`. Or copy the skill folder (methylation-analysis/dmr-detection in GPTomics/bioSkills) into .claude/skills/bio-methylation-dmr-detection in your project. Claude Code loads it when a task matches its description.

How do I install Bio Methylation Dmr Detection in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-methylation-dmr-detection -a codex`. Or copy the skill folder (methylation-analysis/dmr-detection in GPTomics/bioSkills) into .agents/skills/bio-methylation-dmr-detection in your project. Codex loads it when a task matches its description.

Can I use Bio Methylation Dmr Detection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-methylation-dmr-detection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-methylation-dmr-detection, .gemini/skills/bio-methylation-dmr-detection, .github/skills/bio-methylation-dmr-detection and .opencode/skills/bio-methylation-dmr-detection in your project.

What does Bio Methylation Dmr Detection need to run?

Going by SKILL.md and its folder, Bio Methylation Dmr Detection needs R for the scripts in its folder. Our summary lists: Python 3.

Does Bio Methylation Dmr Detection access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Methylation Dmr Detection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Methylation Dmr Detection use?

Bio Methylation Dmr Detection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Methylation Dmr Detection use?

About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Methylation Dmr Detection?

Skills that share tags, products or a category with Bio Methylation Dmr Detection: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Methylation Dmr Detection?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.