Agent skill

Bio Methylation Cell Type Deconvolution

by GPTomics in GPTomics/bioSkills

Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder.

MITAuto-check passedResearch & Science

Install Bio Methylation Cell Type Deconvolution

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-methylation-cell-type-deconvolution --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/methylation-analysis/cell-type-deconvolution .claude/skills/bio-methylation-cell-type-deconvolution && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-methylation-cell-type-deconvolution
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.3k tokens
SKILL.md length
2,353 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder.

  • Works in 3 steps: A composition difference IS a… → The reference defines the answer. A cell… → Fractions are compositional. They live…
  • Estimating blood/tissue cell fractions
  • SKILL.md covers Version Compatibility, The Single Most Important…, Reference-Based: The Houseman… and Tool Taxonomy, plus 12 more sections
  • Runs R scripts from its folder

What it does

Bio Methylation Cell Type Deconvolution is an agent skill from GPTomics/bioSkills. Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder. Covers reference-based deconvolution (Houseman constrained-projection, minfi estimateCellCounts2 with FlowSorted.Blood.EPIC + IDOL-optimized libraries, EpiDISH RPC/CBS/CP, 12-cell extended, cord-blood nRBC references, EpiSCORE/hepidish for solid tissue), reference-free correction (ReFACTor, RefFreeEWAS, SVA), using fractions as covariates vs the compositionality/collinearity trap, and…

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Estimating blood/tissue cell fractions
  • Adjusting an EWAS for composition
  • Choosing a deconvolution reference
  • Attributing a methylation signal to a cell type

Example prompts

  • “Use the bio-methylation-cell-type-deconvolution skill to estimate cell-type composition from bulk DNA methylation and uses it to defuse the single…”
  • “/bio-methylation-cell-type-deconvolution”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A composition difference IS a methylation difference. Bulk DNAm is the fraction-weighted average of its constituent cell-type methylomes…
  2. The reference defines the answer. A cell type present in the sample but absent from the reference is silently redistributed onto the…
  3. Fractions are compositional. They live on a simplex (sum ~1), so they are not independent: one going up forces others down. Naively…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Methylation Cell Type Deconvolution loads about 5.3k tokens when it runs. Until then it costs about 239 tokens; SKILL.md has 2,353 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~239
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,353 words, ~5,277 tokens.

Download SKILL.mdSave it as .claude/skills/bio-methylation-cell-type-deconvolution/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-methylation-cell-type-deconvolution
description
Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder. Covers reference-based deconvolution (Houseman constrained-projection, minfi estimateCellCounts2 with FlowSorted.Blood.EPIC + IDOL-optimized libraries, EpiDISH RPC/CBS/CP, 12-cell extended, cord-blood nRBC references, EpiSCORE/hepidish for solid tissue), reference-free correction (ReFACTor, RefFreeEWAS, SVA), using fractions as covariates vs the compositionality/collinearity trap, and cell-type-resolved EWAS (CellDMC, TCA, TOAST, omicwas, HIRE). Use when estimating blood/tissue cell fractions, adjusting an EWAS for composition, choosing a deconvolution reference, or attributing a methylation signal to a cell type. For the EWAS confounder-vs-mediator decision see ewas-design; for the IEAA cell-count adjustment of DNAm age see epigenetic-clocks; for clean beta input see array-preprocessing.
tool_type
r
primary_tool
EpiDISH

Version Compatibility

Reference examples tested with: EpiDISH 2.18+, minfi 1.48+, FlowSorted.Blood.EPIC 2.0+, FlowSorted.CordBloodCombined.450k 1.20+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

The REFERENCE package is the version that matters most. A reference package is platform-, tissue-, and age-specific: FlowSorted.Blood.EPIC ships IDOLOptimizedCpGs (EPIC) and IDOLOptimizedCpGs450klegacy (450K) as distinct libraries, the 12-cell library lives in a separate FlowSorted.BloodExtended.EPIC package, and cord blood needs FlowSorted.CordBloodCombined.450k (it carries nucleated red blood cells). estimateCellCounts2 returns a Neu (neutrophil) column where the older minfi::estimateCellCounts returns Gran - the label changes the downstream column names. Record the reference package and version alongside the array build.

Cell-Type Deconvolution

"How much of my methylation signal is just cell composition?" -> Project the bulk beta matrix onto a purified-cell reference to estimate per-sample fractions, then carry those fractions forward as covariates - because a bulk methylome is a composition-weighted average and a composition difference is a methylation difference.

  • R: epidish(beta.m, ref.m = centDHSbloodDMC.m, method = 'RPC')$estF

Scope: estimate cell-type fractions from a clean bulk beta matrix and use them downstream. Clean beta read-in and EPICv2 replicate-probe collapse -> array-preprocessing. The confounder-vs-mediator decision and the EWAS regression itself -> ewas-design. Adjusting a DNAm clock for cell counts (IEAA) -> epigenetic-clocks. Single-cell/sorted atlases for reference building and validation ground truth -> single-cell/preprocessing. Predictive-model training/leakage -> machine-learning/biomarker-discovery.

The Single Most Important Modern Insight -- A Cell-Fraction Estimate Is a Projection, Not a Measurement

A reference-based fraction is not a measurement of a sample's composition; it is a projection of that sample onto cell types someone else purified, on someone else's platform, in someone else's tissue. Three corollaries each common misuse violates:

  1. A composition difference IS a methylation difference. Bulk DNAm is the fraction-weighted average of its constituent cell-type methylomes, and most CpG variance is between cell types, not between conditions. If cases and controls differ in composition - which they almost always do (age, sex, infection, smoking shift the neutrophil-to-lymphocyte ratio) - the EWAS reports cell-count differences as if they were disease methylation. This is the #1 EWAS confounder (Jaffe & Irizarry 2014 Genome Biol 15:R31).
  2. The reference defines the answer. A cell type present in the sample but absent from the reference is silently redistributed onto the nearest reference types - no error, the fractions still sum to ~1. Cord blood without nRBC, a solid tissue against a blood reference, EPICv2 data against a 450K library: all return confident, wrong proportions.
  3. Fractions are compositional. They live on a simplex (sum ~1), so they are not independent: one going up forces others down. Naively co-regressing or correlating all K fractions manufactures spurious negative associations.

Organize the analysis around matching the reference and handling compositionality, not around picking an algorithm. Deconvolution turns an uncontrollable confounder into a measurable covariate - but only as accurately as the reference matches the sample.

Reference-Based: The Houseman Constrained-Projection Foundation

Houseman 2012 (BMC Bioinformatics 13:86) is the origin. From a matrix of FACS/MACS-purified cell-type mean methylation at discriminating CpGs (L-DMRs), solve for each sample a constrained quadratic program: the non-negative fraction vector w (w_i >= 0, sum ~1) minimizing the squared distance between observed beta and reference x w over the L-DMR CpGs. This "CP" (constrained projection) is what every later method is measured against. The L-DMR selection is itself a tuning choice that the IDOL work (Koestler 2016 BMC Bioinformatics 17:120) optimized into a fixed, benchmarked library.

Tool Taxonomy

ToolCitationMechanism / roleWhen
EpiDISH (RPC)Teschendorff 2017 BMC Bioinformatics 18:105robust partial correlation; downweights noisy CpGsgeneral; the robust default across tissues/noise
EpiDISH (CP)Houseman 2012 BMC Bioinformatics 13:86constrained quadratic projectionreproduce the classic Houseman estimate
EpiDISH (CBS)Newman 2015 Nat Methods 12:453CIBERSORT nu-SVRborrowed from expression; an alternative
minfi estimateCellCounts2Salas 2018 Genome Biol 19:64Houseman projection on the IDOL-optimized EPIC/450K libraryfrom an RGChannelSet; modern 6-cell blood (Neu)
FlowSorted.BloodExtended.EPICSalas 2022 Nat Commun 13:76112-cell IDOL librarynaive/memory T, Treg, eosinophil/basophil resolution
hepidishTeschendorff 2017 BMC Bioinformatics 18:105hierarchical Epi/Fib/Immune then immune subtypessolid tissue with immune infiltration
EpiSCORETeschendorff 2020 Genome Biol 21:221scRNA-seq-imputed DNAm referencesolid tissues with no sorted reference
ReFACTorRahmani 2016 Nat Methods 13:443sparse-PCA components as covariatesreference-free; no matched reference exists
RefFreeEWASHouseman 2014 Bioinformatics 30:1431NMF/SVD-style latent cell-mixturereference-free; unlabeled components

Decision Tree by Scenario

ScenarioRecommendedWhy
Adult whole bloodestimateCellCounts2 IDOL (6: Neu/CD4T/CD8T/NK/Bcell/Mono) or EpiDISH RPC (centDHSbloodDMC.m gives 7, adds Eos)benchmarked blood references
Need naive/memory T, Treg, Eos, Bas12-cell FlowSorted.BloodExtended.EPICthe 6-cell library cannot resolve these
Cord blood / newbornFlowSorted.CordBloodCombined.450kadds nRBC; an adult reference is silently wrong
Saliva / buccalepithelial + immune reference (hepidish)saliva is not blood; epithelial fraction dominates
Solid tissue / tumor with infiltrationhepidish or EpiSCOREflat blood reference on solid tissue is meaningless
450K dataEpiDISH cent*450k.m / IDOLOptimizedCpGs450klegacyplatform-matched CpGs; EPIC library drops CpGs
EPICv2 datacollapse replicate probes first -> array-preprocessingsuffixed replicate beads hide the reference CpGs
No matched reference (novel tissue)ReFACTor / RefFreeEWAS + sensitivityreference-free fallback; components are unlabeled
Which cell type drives a signalCellDMC / TCA / TOAST (below)model composition, do not just regress it out
Confounder-vs-mediator decision-> ewas-designupstream: adjust out, or resolve cell-specific?

Estimate Blood Fractions with EpiDISH (RPC)

Goal: Get per-sample fractions of the major immune cell types from a clean beta matrix to use as EWAS covariates.

Approach: Pass the beta matrix and a tissue-matched reference centroid to epidish with method='RPC' (the robust option), then read the sample-by-cell-type matrix from $estF.

r
library(EpiDISH)
data(centDHSbloodDMC.m)    # 7 immune cell types, adult whole blood

out <- epidish(beta.m = beta_matrix, ref.m = centDHSbloodDMC.m, method = 'RPC')
fractions <- out$estF       # samples x cell types; rows sum to ~1

Estimate Blood Fractions from an RGChannelSet (minfi + IDOL)

Goal: Estimate the modern 6-cell IDOL blood composition straight from raw IDAT-derived data.

Approach: Run estimateCellCounts2 on the RGChannelSet with the IDOL probe selection and the platform-matched reference; for 450K data switch the reference library so the same cell types are estimated cross-platform.

r
library(FlowSorted.Blood.EPIC)

counts <- estimateCellCounts2(
  rgSet,
  compositeCellType = 'Blood',
  processMethod = 'preprocessNoob',
  probeSelect = 'IDOL',
  cellTypes = c('CD8T', 'CD4T', 'NK', 'Bcell', 'Mono', 'Neu'),   # Neu, not Gran
  referencePlatform = 'IlluminaHumanMethylationEPIC'
)$counts

Solid Tissue: Hierarchical Deconvolution

Goal: Deconvolve a solid tissue (epithelial + fibroblast + infiltrating immune) rather than forcing a blood reference onto it.

Approach: Use hepidish to first split Epithelial/Fibroblast/total-Immune, then deconvolve the immune fraction into subtypes and multiply through. For tissues with no sorted reference at all, EpiSCORE builds an imputed DNAm reference from a single-cell RNA atlas.

r
library(EpiDISH)
data(centEpiFibIC.m)       # Epithelial / Fibroblast / Immune-Cell
data(centBloodSub.m)       # immune subtypes for the second level

frac <- hepidish(beta.m = beta_matrix, ref1.m = centEpiFibIC.m,
                 ref2.m = centBloodSub.m, h.CT.idx = 3, method = 'RPC')
# h.CT.idx = 3 = the Immune column in ref1 to expand with ref2

Reference-Free Correction

Goal: Capture composition structure when no matched reference exists, accepting unlabeled components.

Approach: ReFACTor selects the most composition-informative CpGs and runs sparse-PCA; use the top components as EWAS covariates. RefFreeEWAS decomposes the matrix into a latent cell-mixture term. Both correct without naming the cell types, so check that genuine top hits survive (they can absorb real signal).

r
library(TCA)
ref <- refactor(beta_matrix, k = 6)   # k = expected number of cell types
covariates <- ref$scores              # top sparse-PC components as EWAS covariates

Using the Fractions: Covariate vs Cell-Type-Resolved

There are two distinct moves once fractions exist, and they answer different questions.

As covariates (the standard EWAS defense). Include the fractions in the per-CpG design matrix so composition is regressed out. Because fractions are compositional (sum ~1), do NOT enter all K - drop one reference cell type (or use a compositional transform) to avoid perfect collinearity. The confounder-vs-mediator decision (regress out, or treat composition as the mechanism) belongs to ewas-design; execution belongs to differential-cpg-testing.

Cell-type-resolved EWAS (which cell type drives the signal). Instead of regressing composition away, model a phenotype x cell-fraction INTERACTION per CpG to ask which cell type carries the differential methylation and in which direction. CellDMC (Zheng 2018 Nat Methods 15:1059) is the simplest member; a family generalizes it:

MethodCitationAdds beyond the interactionOutput
CellDMCZheng 2018 Nat Methods 15:1059per-CpG linear pheno x fraction interactionwhich cell type is DM + direction (a test)
TCARahmani 2019 Nat Commun 10:3417tensor model; per-sample per-cell-type levelscell-type-specific methylation + association test
TOASTLi & Wu 2019 Genome Biol 20:190iterative csDM; improves reference-free compositioncsTest per cell type; runs reference-free
omicwasTakeuchi & Kato 2021 BMC Bioinformatics 22:141nonlinear ridge for the logit scale + fraction collinearitycell-type-specific association statistics
HIRELuo 2019 Nat Commun 10:3113joint multiplicative-composition hierarchical modelrisk-CpG sites per cell type
r
library(EpiDISH)
res <- CellDMC(beta.m = beta_matrix, pheno.v = phenotype, frac.m = fractions)
# res$dmct: per-CpG, which cell type is differentially methylated (-1/0/1)

A cell-type-resolved call is an ill-posed inverse problem regularized by an assumed reference: rare cell types (2-5% of the mixture) are badly underpowered, fraction collinearity destabilizes the interactions, and deconvolution error propagates straight into the attribution (HIRE's argument for estimating composition jointly). Validation is hard without sorted/single-cell ground truth - method papers lean on simulations and reconstructed mixtures, which are circular. Treat an in-silico cell-type-specific hit as a HYPOTHESIS about what to sort next, not a finding; confirm load-bearing attributions in sorted or single-cell DNAm from independent samples (Walker 2025 Brief Bioinform 26:bbaf427).

Show full SKILL.md (951 more words)Show less

Intrinsic epigenetic age acceleration (IEAA) is DNAm age residualized on chronological age AND estimated blood cell counts - so deconvolution is the prerequisite step: estimate fractions here, then hand them to epigenetic-clocks as the cell-count covariates that distinguish cell-intrinsic aging from a composition shift. Do not teach the clock here; compute the fractions and route the IEAA adjustment to epigenetic-clocks.

Per-Method Failure Modes

Missing cell type silently redistributed

Trigger: a sample contains a cell type absent from the reference (cord-blood nRBC, a rare infiltrate, a granulocyte subtype collapsed to Gran). Mechanism: the constrained projection has no column for it, so its signal lands on the nearest present types. Symptom: plausible-looking fractions that sum to ~1 with no warning. Fix: match the reference to tissue+age (FlowSorted.CordBloodCombined.450k for newborns; hepidish/EpiSCORE for solid tissue).

Platform-mismatched reference library

Trigger: 450K data with the EPIC IDOL library, or EPICv2 with either. Mechanism: reference CpGs are partly absent on the other platform, shrinking the L-DMR set used for the projection. Symptom: biased fractions, no error. Fix: IDOLOptimizedCpGs450klegacy / cent*450k.m for 450K; collapse EPICv2 replicate probes first (-> array-preprocessing).

Collinear cell-fraction covariates

Trigger: entering all K fractions (sum ~1) into a design matrix. Mechanism: the simplex constraint makes the K-th fraction a linear function of the others. Symptom: rank-deficient design, dropped coefficient, or spurious negative fraction-fraction correlations. Fix: drop one reference cell type or use a compositional (CLR/ILR) transform.

Reference-free over-correction

Trigger: including too many ReFACTor/RefFreeEWAS components, or using them when a reference exists. Mechanism: unlabeled latent components can absorb true biological signal alongside composition. Symptom: top EWAS hits vanish; false negatives. Fix: prefer reference-based when a reference exists; use reference-free as a fallback/sensitivity check and confirm hits survive.

Cell-type attribution from rare cells

Trigger: reading a CellDMC/TCA call for a 2-5% cell type. Mechanism: a rare cell contributes a fraction-attenuated slice of bulk variance, so its interaction estimate is dominated by deconvolution noise. Symptom: confident-looking csDM in basophils/eosinophils; nulls misread as "no effect." Fix: report each cell type's mean fraction; distrust specific calls for low-abundance types; never infer absence of effect from an underpowered null.

Quantitative Thresholds

ThresholdSourceRationale
IDOL EPIC 6-cell library ~450 CpGsSalas 2018 Genome Biol 19:64benchmarked L-DMR set; R^2 ~0.992 on reconstructed mixtures
method = 'RPC' for EpiDISHTeschendorff 2017 BMC Bioinformatics 18:105robust to outlier/noisy CpGs; more stable than CP across tissues
drop 1 of K fractions as covariatescompositional constraintfractions sum to ~1, so all K are perfectly collinear
cord blood reference must carry nRBCGervin 2019 / CordBloodCombinednRBC abundant in cord blood, absent from adult references
ReFACTor k = expected cell-type countRahmani 2016 Nat Methods 13:443k sets the rank; too high over-corrects, too low under-corrects
csDM credible only for abundant typesWalker 2025 Brief Bioinform 26:bbaf427rare cells are fraction-attenuated and underpowered

Common Errors

Error / symptomCauseSolution
Fractions look fine but EWAS still inflatedunmodeled cell type / wrong referencematch reference to tissue+age+platform
Design matrix rank-deficientall K fractions entered as covariatesdrop one cell type or CLR-transform
estimateCellCounts2 returns Neu, code expects Granminfi vs FlowSorted label differenceuse Neu (estimateCellCounts2) consistently
Reference CpGs not found on EPICv2replicate probes not collapsedcollapse to one value per CpG first
Negative or all-zero fraction for a typeplatform mismatch / absent in samplecheck platform-matched library; inspect mean fraction
csDM hit in a rare cell typeunderpowered interactionreport the fraction; validate by sorting/single-cell

References

  • Houseman EA, Accomando WP, Koestler DC, et al. 2012. DNA methylation arrays as surrogate measures of cell mixture distribution. BMC Bioinformatics 13:86.
  • Jaffe AE, Irizarry RA. 2014. Accounting for cellular heterogeneity is critical in epigenome-wide association studies. Genome Biol 15:R31.
  • Koestler DC, Jones MJ, Usset J, et al. 2016. Improving cell mixture deconvolution by identifying optimal DNA methylation libraries (IDOL). BMC Bioinformatics 17:120.
  • Salas LA, Koestler DC, Butler RA, et al. 2018. An optimized library for reference-based deconvolution of whole-blood biospecimens assayed using the Illumina HumanMethylationEPIC BeadArray. Genome Biol 19:64.
  • Salas LA, Zhang Z, Koestler DC, et al. 2022. Enhanced cell deconvolution of peripheral blood using DNA methylation for high-resolution immune profiling. Nat Commun 13:761.
  • Teschendorff AE, Breeze CE, Zheng SC, Beck S. 2017. A comparison of reference-based algorithms for correcting cell-type heterogeneity in epigenome-wide association studies. BMC Bioinformatics 18:105.
  • Teschendorff AE, Zhu T, Breeze CE, Beck S. 2020. EPISCORE: cell type deconvolution of bulk tissue DNA methylomes from single-cell RNA-Seq data. Genome Biol 21:221.
  • Houseman EA, Molitor J, Marsit CJ. 2014. Reference-free cell mixture adjustments in analysis of DNA methylation data. Bioinformatics 30:1431-1439.
  • Rahmani E, Zaitlen N, Baran Y, et al. 2016. Sparse PCA corrects for cell type heterogeneity in epigenome-wide association studies. Nat Methods 13:443-445.
  • Zheng SC, Breeze CE, Beck S, Teschendorff AE. 2018. Identification of differentially methylated cell types in epigenome-wide association studies. Nat Methods 15:1059-1066.
  • Rahmani E, Schweiger R, Rhead B, et al. 2019. Cell-type-specific resolution epigenetics without the need for cell sorting or single-cell biology. Nat Commun 10:3417.
  • Li Z, Wu H. 2019. TOAST: improving reference-free cell composition estimation by cross-cell type differential analysis. Genome Biol 20:190.
  • Takeuchi F, Kato N. 2021. Nonlinear ridge regression improves cell-type-specific differential expression analysis. BMC Bioinformatics 22:141.
  • Luo X, Yang C, Wei Y. 2019. Detection of cell-type-specific risk-CpG sites in epigenome-wide association studies. Nat Commun 10:3113.
  • Walker EM, Dempster EL, Franklin A, et al. 2025. Guidance for the design and analysis of cell-type-specific DNA methylation epidemiology studies. Brief Bioinform 26:bbaf427.
  • array-preprocessing - Provides the clean beta matrix deconvolution consumes
  • ewas-design - Cell-fraction covariate strategy (confounder vs mediator)
  • epigenetic-clocks - IEAA: adjust the clock for estimated cell composition
  • differential-cpg-testing - Uses cell fractions as design-matrix covariates
  • single-cell/preprocessing - scRNA atlases for reference building (EpiSCORE) and ground truth
  • machine-learning/biomarker-discovery - Predictive-model boundary
  • workflows/methylation-pipeline - End-to-end pipeline

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in methylation-analysis/cell-type-deconvolution of GPTomics/bioSkills.

  • SKILL.md
  • examples/cell_deconvolution.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Methylation Cell Type Deconvolution next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Methylation Cell Type Deconvolution compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Methylation Cell Type Deconvolution this skillGPTomics/bioSkills1.2k1 repos~5.3kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow84k4 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills47k2 repos~2.1kAutomated safety check: PassApache-2.0
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT
Peer Reviewspacering-net/codeg3.9k17 repos~5.9kAutomated safety check: NotesMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    47k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.9k GitHub starsUsed in 17 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Last30days

    mvanhorn/last30days-skill

    Research what people actually say about any topic in the last 30 days.

    64k GitHub stars~7.9k tokensUpdated today
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Methylation Cell Type Deconvolution

What does Bio Methylation Cell Type Deconvolution do?

Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder. Bio Methylation Cell Type Deconvolution is an agent skill from GPTomics/bioSkills. Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder.

When should I use Bio Methylation Cell Type Deconvolution?

Bio Methylation Cell Type Deconvolution fits situations like: estimating blood/tissue cell fractions; adjusting an EWAS for composition; choosing a deconvolution reference; attributing a methylation signal to a cell type.

How do I install Bio Methylation Cell Type Deconvolution in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a claude-code`. Or copy the skill folder (methylation-analysis/cell-type-deconvolution in GPTomics/bioSkills) into .claude/skills/bio-methylation-cell-type-deconvolution in your project. Claude Code loads it when a task matches its description.

How do I install Bio Methylation Cell Type Deconvolution in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a codex`. Or copy the skill folder (methylation-analysis/cell-type-deconvolution in GPTomics/bioSkills) into .agents/skills/bio-methylation-cell-type-deconvolution in your project. Codex loads it when a task matches its description.

Can I use Bio Methylation Cell Type Deconvolution in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-methylation-cell-type-deconvolution, .gemini/skills/bio-methylation-cell-type-deconvolution, .github/skills/bio-methylation-cell-type-deconvolution and .opencode/skills/bio-methylation-cell-type-deconvolution in your project.

What does Bio Methylation Cell Type Deconvolution need to run?

Going by SKILL.md and its folder, Bio Methylation Cell Type Deconvolution needs R for the scripts in its folder.

Does Bio Methylation Cell Type Deconvolution access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Methylation Cell Type Deconvolution safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Methylation Cell Type Deconvolution use?

Bio Methylation Cell Type Deconvolution is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Methylation Cell Type Deconvolution use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Methylation Cell Type Deconvolution?

Skills that share tags, products or a category with Bio Methylation Cell Type Deconvolution: Hypothesis Generation (spacering-net/codeg, 3.9k stars), GitHub Deep Research (bytedance/deer-flow, 84k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars) and Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Methylation Cell Type Deconvolution?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.