Hypothesis Generation
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-cell-type-deconvolution --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/methylation-analysis/cell-type-deconvolution .claude/skills/bio-methylation-cell-type-deconvolution && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-methylation-cell-type-deconvolution" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/cell-type-deconvolution into .claude/skills/bio-methylation-cell-type-deconvolution/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-cell-type-deconvolution", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/cell-type-deconvolutionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-cell-type-deconvolution --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/methylation-analysis/cell-type-deconvolution .agents/skills/bio-methylation-cell-type-deconvolution && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-methylation-cell-type-deconvolution" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/cell-type-deconvolution into .agents/skills/bio-methylation-cell-type-deconvolution/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-cell-type-deconvolution", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-cell-type-deconvolution --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/methylation-analysis/cell-type-deconvolution .cursor/skills/bio-methylation-cell-type-deconvolution && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-methylation-cell-type-deconvolution" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/cell-type-deconvolution into .cursor/skills/bio-methylation-cell-type-deconvolution/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-cell-type-deconvolution", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path methylation-analysis/cell-type-deconvolution--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-cell-type-deconvolution --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/methylation-analysis/cell-type-deconvolution .gemini/skills/bio-methylation-cell-type-deconvolution && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-methylation-cell-type-deconvolution" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/cell-type-deconvolution into .gemini/skills/bio-methylation-cell-type-deconvolution/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-cell-type-deconvolution", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-methylation-cell-type-deconvolutionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/methylation-analysis/cell-type-deconvolution .github/skills/bio-methylation-cell-type-deconvolution && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-methylation-cell-type-deconvolution" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/cell-type-deconvolution into .github/skills/bio-methylation-cell-type-deconvolution/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-cell-type-deconvolution", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-cell-type-deconvolution --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/methylation-analysis/cell-type-deconvolution .opencode/skills/bio-methylation-cell-type-deconvolution && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-methylation-cell-type-deconvolution" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/cell-type-deconvolution into .opencode/skills/bio-methylation-cell-type-deconvolution/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-cell-type-deconvolution", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-methylation-cell-type-deconvolutionEstimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder.
Bio Methylation Cell Type Deconvolution is an agent skill from GPTomics/bioSkills. Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder. Covers reference-based deconvolution (Houseman constrained-projection, minfi estimateCellCounts2 with FlowSorted.Blood.EPIC + IDOL-optimized libraries, EpiDISH RPC/CBS/CP, 12-cell extended, cord-blood nRBC references, EpiSCORE/hepidish for solid tissue), reference-free correction (ReFACTor, RefFreeEWAS, SVA), using fractions as covariates vs the compositionality/collinearity trap, and…
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).
It sits in Research & Science. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (R), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Methylation Cell Type Deconvolution loads about 5.3k tokens when it runs. Until then it costs about 239 tokens; SKILL.md has 2,353 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,353 words, ~5,277 tokens.
.claude/skills/bio-methylation-cell-type-deconvolution/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: EpiDISH 2.18+, minfi 1.48+, FlowSorted.Blood.EPIC 2.0+, FlowSorted.CordBloodCombined.450k 1.20+.
Before using code patterns, verify installed versions match. If versions differ:
packageVersion('<pkg>') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
The REFERENCE package is the version that matters most. A reference package is platform-, tissue-, and age-specific: FlowSorted.Blood.EPIC ships IDOLOptimizedCpGs (EPIC) and IDOLOptimizedCpGs450klegacy (450K) as distinct libraries, the 12-cell library lives in a separate FlowSorted.BloodExtended.EPIC package, and cord blood needs FlowSorted.CordBloodCombined.450k (it carries nucleated red blood cells). estimateCellCounts2 returns a Neu (neutrophil) column where the older minfi::estimateCellCounts returns Gran - the label changes the downstream column names. Record the reference package and version alongside the array build.
"How much of my methylation signal is just cell composition?" -> Project the bulk beta matrix onto a purified-cell reference to estimate per-sample fractions, then carry those fractions forward as covariates - because a bulk methylome is a composition-weighted average and a composition difference is a methylation difference.
epidish(beta.m, ref.m = centDHSbloodDMC.m, method = 'RPC')$estFScope: estimate cell-type fractions from a clean bulk beta matrix and use them downstream. Clean beta read-in and EPICv2 replicate-probe collapse -> array-preprocessing. The confounder-vs-mediator decision and the EWAS regression itself -> ewas-design. Adjusting a DNAm clock for cell counts (IEAA) -> epigenetic-clocks. Single-cell/sorted atlases for reference building and validation ground truth -> single-cell/preprocessing. Predictive-model training/leakage -> machine-learning/biomarker-discovery.
A reference-based fraction is not a measurement of a sample's composition; it is a projection of that sample onto cell types someone else purified, on someone else's platform, in someone else's tissue. Three corollaries each common misuse violates:
Organize the analysis around matching the reference and handling compositionality, not around picking an algorithm. Deconvolution turns an uncontrollable confounder into a measurable covariate - but only as accurately as the reference matches the sample.
Houseman 2012 (BMC Bioinformatics 13:86) is the origin. From a matrix of FACS/MACS-purified cell-type mean methylation at discriminating CpGs (L-DMRs), solve for each sample a constrained quadratic program: the non-negative fraction vector w (w_i >= 0, sum ~1) minimizing the squared distance between observed beta and reference x w over the L-DMR CpGs. This "CP" (constrained projection) is what every later method is measured against. The L-DMR selection is itself a tuning choice that the IDOL work (Koestler 2016 BMC Bioinformatics 17:120) optimized into a fixed, benchmarked library.
| Tool | Citation | Mechanism / role | When |
|---|---|---|---|
| EpiDISH (RPC) | Teschendorff 2017 BMC Bioinformatics 18:105 | robust partial correlation; downweights noisy CpGs | general; the robust default across tissues/noise |
| EpiDISH (CP) | Houseman 2012 BMC Bioinformatics 13:86 | constrained quadratic projection | reproduce the classic Houseman estimate |
| EpiDISH (CBS) | Newman 2015 Nat Methods 12:453 | CIBERSORT nu-SVR | borrowed from expression; an alternative |
| minfi estimateCellCounts2 | Salas 2018 Genome Biol 19:64 | Houseman projection on the IDOL-optimized EPIC/450K library | from an RGChannelSet; modern 6-cell blood (Neu) |
| FlowSorted.BloodExtended.EPIC | Salas 2022 Nat Commun 13:761 | 12-cell IDOL library | naive/memory T, Treg, eosinophil/basophil resolution |
| hepidish | Teschendorff 2017 BMC Bioinformatics 18:105 | hierarchical Epi/Fib/Immune then immune subtypes | solid tissue with immune infiltration |
| EpiSCORE | Teschendorff 2020 Genome Biol 21:221 | scRNA-seq-imputed DNAm reference | solid tissues with no sorted reference |
| ReFACTor | Rahmani 2016 Nat Methods 13:443 | sparse-PCA components as covariates | reference-free; no matched reference exists |
| RefFreeEWAS | Houseman 2014 Bioinformatics 30:1431 | NMF/SVD-style latent cell-mixture | reference-free; unlabeled components |
| Scenario | Recommended | Why |
|---|---|---|
| Adult whole blood | estimateCellCounts2 IDOL (6: Neu/CD4T/CD8T/NK/Bcell/Mono) or EpiDISH RPC (centDHSbloodDMC.m gives 7, adds Eos) | benchmarked blood references |
| Need naive/memory T, Treg, Eos, Bas | 12-cell FlowSorted.BloodExtended.EPIC | the 6-cell library cannot resolve these |
| Cord blood / newborn | FlowSorted.CordBloodCombined.450k | adds nRBC; an adult reference is silently wrong |
| Saliva / buccal | epithelial + immune reference (hepidish) | saliva is not blood; epithelial fraction dominates |
| Solid tissue / tumor with infiltration | hepidish or EpiSCORE | flat blood reference on solid tissue is meaningless |
| 450K data | EpiDISH cent*450k.m / IDOLOptimizedCpGs450klegacy | platform-matched CpGs; EPIC library drops CpGs |
| EPICv2 data | collapse replicate probes first -> array-preprocessing | suffixed replicate beads hide the reference CpGs |
| No matched reference (novel tissue) | ReFACTor / RefFreeEWAS + sensitivity | reference-free fallback; components are unlabeled |
| Which cell type drives a signal | CellDMC / TCA / TOAST (below) | model composition, do not just regress it out |
| Confounder-vs-mediator decision | -> ewas-design | upstream: adjust out, or resolve cell-specific? |
Goal: Get per-sample fractions of the major immune cell types from a clean beta matrix to use as EWAS covariates.
Approach: Pass the beta matrix and a tissue-matched reference centroid to epidish with method='RPC' (the robust option), then read the sample-by-cell-type matrix from $estF.
library(EpiDISH)
data(centDHSbloodDMC.m) # 7 immune cell types, adult whole blood
out <- epidish(beta.m = beta_matrix, ref.m = centDHSbloodDMC.m, method = 'RPC')
fractions <- out$estF # samples x cell types; rows sum to ~1Goal: Estimate the modern 6-cell IDOL blood composition straight from raw IDAT-derived data.
Approach: Run estimateCellCounts2 on the RGChannelSet with the IDOL probe selection and the platform-matched reference; for 450K data switch the reference library so the same cell types are estimated cross-platform.
library(FlowSorted.Blood.EPIC)
counts <- estimateCellCounts2(
rgSet,
compositeCellType = 'Blood',
processMethod = 'preprocessNoob',
probeSelect = 'IDOL',
cellTypes = c('CD8T', 'CD4T', 'NK', 'Bcell', 'Mono', 'Neu'), # Neu, not Gran
referencePlatform = 'IlluminaHumanMethylationEPIC'
)$countsGoal: Deconvolve a solid tissue (epithelial + fibroblast + infiltrating immune) rather than forcing a blood reference onto it.
Approach: Use hepidish to first split Epithelial/Fibroblast/total-Immune, then deconvolve the immune fraction into subtypes and multiply through. For tissues with no sorted reference at all, EpiSCORE builds an imputed DNAm reference from a single-cell RNA atlas.
library(EpiDISH)
data(centEpiFibIC.m) # Epithelial / Fibroblast / Immune-Cell
data(centBloodSub.m) # immune subtypes for the second level
frac <- hepidish(beta.m = beta_matrix, ref1.m = centEpiFibIC.m,
ref2.m = centBloodSub.m, h.CT.idx = 3, method = 'RPC')
# h.CT.idx = 3 = the Immune column in ref1 to expand with ref2Goal: Capture composition structure when no matched reference exists, accepting unlabeled components.
Approach: ReFACTor selects the most composition-informative CpGs and runs sparse-PCA; use the top components as EWAS covariates. RefFreeEWAS decomposes the matrix into a latent cell-mixture term. Both correct without naming the cell types, so check that genuine top hits survive (they can absorb real signal).
library(TCA)
ref <- refactor(beta_matrix, k = 6) # k = expected number of cell types
covariates <- ref$scores # top sparse-PC components as EWAS covariatesThere are two distinct moves once fractions exist, and they answer different questions.
As covariates (the standard EWAS defense). Include the fractions in the per-CpG design matrix so composition is regressed out. Because fractions are compositional (sum ~1), do NOT enter all K - drop one reference cell type (or use a compositional transform) to avoid perfect collinearity. The confounder-vs-mediator decision (regress out, or treat composition as the mechanism) belongs to ewas-design; execution belongs to differential-cpg-testing.
Cell-type-resolved EWAS (which cell type drives the signal). Instead of regressing composition away, model a phenotype x cell-fraction INTERACTION per CpG to ask which cell type carries the differential methylation and in which direction. CellDMC (Zheng 2018 Nat Methods 15:1059) is the simplest member; a family generalizes it:
| Method | Citation | Adds beyond the interaction | Output |
|---|---|---|---|
| CellDMC | Zheng 2018 Nat Methods 15:1059 | per-CpG linear pheno x fraction interaction | which cell type is DM + direction (a test) |
| TCA | Rahmani 2019 Nat Commun 10:3417 | tensor model; per-sample per-cell-type levels | cell-type-specific methylation + association test |
| TOAST | Li & Wu 2019 Genome Biol 20:190 | iterative csDM; improves reference-free composition | csTest per cell type; runs reference-free |
| omicwas | Takeuchi & Kato 2021 BMC Bioinformatics 22:141 | nonlinear ridge for the logit scale + fraction collinearity | cell-type-specific association statistics |
| HIRE | Luo 2019 Nat Commun 10:3113 | joint multiplicative-composition hierarchical model | risk-CpG sites per cell type |
library(EpiDISH)
res <- CellDMC(beta.m = beta_matrix, pheno.v = phenotype, frac.m = fractions)
# res$dmct: per-CpG, which cell type is differentially methylated (-1/0/1)A cell-type-resolved call is an ill-posed inverse problem regularized by an assumed reference: rare cell types (2-5% of the mixture) are badly underpowered, fraction collinearity destabilizes the interactions, and deconvolution error propagates straight into the attribution (HIRE's argument for estimating composition jointly). Validation is hard without sorted/single-cell ground truth - method papers lean on simulations and reconstructed mixtures, which are circular. Treat an in-silico cell-type-specific hit as a HYPOTHESIS about what to sort next, not a finding; confirm load-bearing attributions in sorted or single-cell DNAm from independent samples (Walker 2025 Brief Bioinform 26:bbaf427).
Intrinsic epigenetic age acceleration (IEAA) is DNAm age residualized on chronological age AND estimated blood cell counts - so deconvolution is the prerequisite step: estimate fractions here, then hand them to epigenetic-clocks as the cell-count covariates that distinguish cell-intrinsic aging from a composition shift. Do not teach the clock here; compute the fractions and route the IEAA adjustment to epigenetic-clocks.
Trigger: a sample contains a cell type absent from the reference (cord-blood nRBC, a rare infiltrate, a granulocyte subtype collapsed to Gran). Mechanism: the constrained projection has no column for it, so its signal lands on the nearest present types. Symptom: plausible-looking fractions that sum to ~1 with no warning. Fix: match the reference to tissue+age (FlowSorted.CordBloodCombined.450k for newborns; hepidish/EpiSCORE for solid tissue).
Trigger: 450K data with the EPIC IDOL library, or EPICv2 with either. Mechanism: reference CpGs are partly absent on the other platform, shrinking the L-DMR set used for the projection. Symptom: biased fractions, no error. Fix: IDOLOptimizedCpGs450klegacy / cent*450k.m for 450K; collapse EPICv2 replicate probes first (-> array-preprocessing).
Trigger: entering all K fractions (sum ~1) into a design matrix. Mechanism: the simplex constraint makes the K-th fraction a linear function of the others. Symptom: rank-deficient design, dropped coefficient, or spurious negative fraction-fraction correlations. Fix: drop one reference cell type or use a compositional (CLR/ILR) transform.
Trigger: including too many ReFACTor/RefFreeEWAS components, or using them when a reference exists. Mechanism: unlabeled latent components can absorb true biological signal alongside composition. Symptom: top EWAS hits vanish; false negatives. Fix: prefer reference-based when a reference exists; use reference-free as a fallback/sensitivity check and confirm hits survive.
Trigger: reading a CellDMC/TCA call for a 2-5% cell type. Mechanism: a rare cell contributes a fraction-attenuated slice of bulk variance, so its interaction estimate is dominated by deconvolution noise. Symptom: confident-looking csDM in basophils/eosinophils; nulls misread as "no effect." Fix: report each cell type's mean fraction; distrust specific calls for low-abundance types; never infer absence of effect from an underpowered null.
| Threshold | Source | Rationale |
|---|---|---|
| IDOL EPIC 6-cell library ~450 CpGs | Salas 2018 Genome Biol 19:64 | benchmarked L-DMR set; R^2 ~0.992 on reconstructed mixtures |
| method = 'RPC' for EpiDISH | Teschendorff 2017 BMC Bioinformatics 18:105 | robust to outlier/noisy CpGs; more stable than CP across tissues |
| drop 1 of K fractions as covariates | compositional constraint | fractions sum to ~1, so all K are perfectly collinear |
| cord blood reference must carry nRBC | Gervin 2019 / CordBloodCombined | nRBC abundant in cord blood, absent from adult references |
| ReFACTor k = expected cell-type count | Rahmani 2016 Nat Methods 13:443 | k sets the rank; too high over-corrects, too low under-corrects |
| csDM credible only for abundant types | Walker 2025 Brief Bioinform 26:bbaf427 | rare cells are fraction-attenuated and underpowered |
| Error / symptom | Cause | Solution |
|---|---|---|
| Fractions look fine but EWAS still inflated | unmodeled cell type / wrong reference | match reference to tissue+age+platform |
| Design matrix rank-deficient | all K fractions entered as covariates | drop one cell type or CLR-transform |
| estimateCellCounts2 returns Neu, code expects Gran | minfi vs FlowSorted label difference | use Neu (estimateCellCounts2) consistently |
| Reference CpGs not found on EPICv2 | replicate probes not collapsed | collapse to one value per CpG first |
| Negative or all-zero fraction for a type | platform mismatch / absent in sample | check platform-matched library; inspect mean fraction |
| csDM hit in a rare cell type | underpowered interaction | report the fraction; validate by sorting/single-cell |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in methylation-analysis/cell-type-deconvolution of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Methylation Cell Type Deconvolution next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Methylation Cell Type Deconvolution this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.3k | Automated safety check: Pass | MIT | |
| Hypothesis Generationspacering-net/codeg | 3.9k | 14 repos | ~3.6k | Automated safety check: Notes | MIT | |
| GitHub Deep Researchbytedance/deer-flow | 84k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Nature Paper CardYuan1z0825/nature-skills | 47k | 2 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Content Research Writerweapp-tailwindcss/weapp-tailwindcss | 1.9k | 25 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Peer Reviewspacering-net/codeg | 3.9k | 17 repos | ~5.9k | Automated safety check: Notes | MIT |
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
bytedance/deer-flow
Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.
Yuan1z0825/nature-skills
Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.
weapp-tailwindcss/weapp-tailwindcss
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.
spacering-net/codeg
Structured manuscript/grant review with checklist-based evaluation.
mvanhorn/last30days-skill
Research what people actually say about any topic in the last 30 days.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder. Bio Methylation Cell Type Deconvolution is an agent skill from GPTomics/bioSkills. Estimates cell-type composition from bulk DNA methylation and uses it to defuse the single biggest EWAS confounder.
Bio Methylation Cell Type Deconvolution fits situations like: estimating blood/tissue cell fractions; adjusting an EWAS for composition; choosing a deconvolution reference; attributing a methylation signal to a cell type.
Run `npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a claude-code`. Or copy the skill folder (methylation-analysis/cell-type-deconvolution in GPTomics/bioSkills) into .claude/skills/bio-methylation-cell-type-deconvolution in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a codex`. Or copy the skill folder (methylation-analysis/cell-type-deconvolution in GPTomics/bioSkills) into .agents/skills/bio-methylation-cell-type-deconvolution in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-methylation-cell-type-deconvolution -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-methylation-cell-type-deconvolution, .gemini/skills/bio-methylation-cell-type-deconvolution, .github/skills/bio-methylation-cell-type-deconvolution and .opencode/skills/bio-methylation-cell-type-deconvolution in your project.
Going by SKILL.md and its folder, Bio Methylation Cell Type Deconvolution needs R for the scripts in its folder.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Methylation Cell Type Deconvolution is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Methylation Cell Type Deconvolution: Hypothesis Generation (spacering-net/codeg, 3.9k stars), GitHub Deep Research (bytedance/deer-flow, 84k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars) and Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.