Literature Engineer
WILLOSCAR/research-units-pipeline-skills
Multi-route literature expansion + metadata normalization for evidence-first surveys.
Performs probe filtering and sample-level QC on Illumina Infinium methylation arrays (450K / EPIC / EPICv2) to decide which probes and samples to trust.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-array-qc-filtering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/methylation-analysis/array-qc-filtering .claude/skills/bio-methylation-array-qc-filtering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-methylation-array-qc-filtering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/array-qc-filtering into .claude/skills/bio-methylation-array-qc-filtering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-array-qc-filtering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/array-qc-filteringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-array-qc-filtering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/methylation-analysis/array-qc-filtering .agents/skills/bio-methylation-array-qc-filtering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-methylation-array-qc-filtering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/array-qc-filtering into .agents/skills/bio-methylation-array-qc-filtering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-array-qc-filtering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-array-qc-filtering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/methylation-analysis/array-qc-filtering .cursor/skills/bio-methylation-array-qc-filtering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-methylation-array-qc-filtering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/array-qc-filtering into .cursor/skills/bio-methylation-array-qc-filtering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-array-qc-filtering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path methylation-analysis/array-qc-filtering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-array-qc-filtering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/methylation-analysis/array-qc-filtering .gemini/skills/bio-methylation-array-qc-filtering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-methylation-array-qc-filtering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/array-qc-filtering into .gemini/skills/bio-methylation-array-qc-filtering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-array-qc-filtering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-methylation-array-qc-filteringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/methylation-analysis/array-qc-filtering .github/skills/bio-methylation-array-qc-filtering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-methylation-array-qc-filtering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/array-qc-filtering into .github/skills/bio-methylation-array-qc-filtering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-array-qc-filtering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-methylation-array-qc-filtering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/methylation-analysis/array-qc-filtering .opencode/skills/bio-methylation-array-qc-filtering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-methylation-array-qc-filtering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/methylation-analysis/array-qc-filtering into .opencode/skills/bio-methylation-array-qc-filtering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-methylation-array-qc-filtering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-methylation-array-qc-filteringPerforms probe filtering and sample-level QC on Illumina Infinium methylation arrays (450K / EPIC / EPICv2) to decide which probes and samples to trust.
Bio Methylation Array Qc Filtering is an agent skill from GPTomics/bioSkills. Performs probe filtering and sample-level QC on Illumina Infinium methylation arrays (450K / EPIC / EPICv2) to decide which probes and samples to trust. Drops detection-p-failed and low-bead-count probes, removes cross-reactive/non-specific probes (Chen 2013 / Pidsley 2016 lists via maxprobes), excludes SNP-overlapping probes with dropLociWithSnps, and handles sex-chromosome probes. Collapses EPICv2 replicate probes with betasCollapseToPfx and harmonizes across array versions (EPICv2 hg38 vs 450K/EPIC hg19…
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).
It sits in Research & Science, covering Database schema design and Experimental design. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (R), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Methylation Array Qc Filtering loads about 4.8k tokens when it runs. Until then it costs about 257 tokens; SKILL.md has 1,887 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,887 words, ~4,786 tokens.
.claude/skills/bio-methylation-array-qc-filtering/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: minfi 1.48+, sesame 1.20+, maxprobes 0.0.2+, ChAMP 2.32+.
Before using code patterns, verify installed versions match. If versions differ:
packageVersion('<pkg>') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
The ARRAY VERSION is the version that matters most here. Cross-reactive probe lists, SNP-overlap annotation, the manifest, and the genome build are all array-version-specific. Record whether the data is 450K (hg19), EPIC v1 (hg19), or EPIC v2 (hg38), and which annotation package supplies the probe metadata (IlluminaHumanMethylation450kanno.ilmn12.hg19, ...EPICanno.ilm10b4.hg19, ...EPICv2anno.20a1.hg38). EPICv2 carries ~5,100 replicate probes (2-10 designs per locus) and is hg38-native; both facts break naive cross-version merges if ignored.
"Which probes and samples can I trust on my methylation array?" -> Mask the failed/cross-reactive/SNP-overlapping probes, collapse EPICv2 replicates, and check every sample's predicted sex and rs-SNP fingerprint against the sample sheet - because a raw array beta is uninterpretable until it is detection-masked and probe-filtered, and a sample swap is the failure no downstream model can rescue.
dropLociWithSnps(gset, snps=c('CpG','SBE'), maf=0) then getSex() and getSnpBeta() for identity QCScope: probe filtering + sample-level QC + EPICv2 replicate collapse + cross-version harmonization for Infinium arrays. IDAT reading, background/dye correction, detection-p masking, and normalization -> array-preprocessing (it produces the matrix QC'd here). Explicit chip/position batch CORRECTION and study design -> ewas-design. Per-CpG testing on the filtered matrix -> differential-cpg-testing. Cohort cell-composition QC -> cell-type-deconvolution. Short-read bisulfite or long-read MM/ML calling are different modalities (see the bisulfite skills and long-read-sequencing/nanopore-methylation).
A beta value of 0.5 from a failed probe, a cross-reactive probe, or a SNP-overlapping probe looks exactly like a real intermediate methylation call - confident, reproducible, and wrong. Deciding which probes and samples to TRUST is the measurement's integrity layer, not optional cleanup. Three corollaries every misuse violates:
getSex() prediction that disagrees with the sample sheet, or rs-SNP fingerprints that cluster two "different" samples together, reveals a mislabel that no model corrects after the fact. If Sentrix chip or array-position is confounded with the biological group, the technical and biological signals are mathematically inseparable - randomize at design (-> ewas-design), do not try to rescue it.Organize the work around delivering a trustworthy, merge-safe matrix: filter probes, collapse replicates, verify sample identity. Over-filtering is its own error - dropping every flagged probe discards real signal, so the maf cutoff and the cross-reactive list are calibrated decisions, not a fixed recipe.
| Filter | Tool / function | Citation | What it removes |
|---|---|---|---|
| Failed detection-p | detectionP() (minfi) / pOOBAH (sesame) | Aryee 2014 Bioinformatics 30:1363; Zhou 2018 NAR 46:e123 | probes with signal indistinguishable from background (per applied in array-preprocessing) |
| Low bead count | getNBeads() (minfi) / sesame bead data | Aryee 2014 Bioinformatics 30:1363 | probes built from too few beads (<3), unreliable |
| Cross-reactive / non-specific | dropXreactiveLoci() / xreactive_probes() (maxprobes) | Chen 2013 Epigenetics 8:203; Pidsley 2016 Genome Biol 17:208 | probes co-hybridizing to multiple loci; list is ARRAY-VERSION-specific |
| SNP-overlapping | dropLociWithSnps(), getSnpInfo() (minfi) | Aryee 2014 Bioinformatics 30:1363 | CpG-SNP / SBE-SNP probes that report genotype not methylation |
| Sex-chromosome | annotation chr (minfi) | Aryee 2014 Bioinformatics 30:1363 | chrX/chrY probes (sex-confounded; drop or sex-stratify) |
| EPICv2 replicates | betasCollapseToPfx() (sesame) | Kaur 2023 Epigenetics Commun 3:6 | extra designs per locus; collapse to one value per cg core ID |
| Scenario | Recommended | Why |
|---|---|---|
| 450K cohort, EWAS-grade filter | maxprobes 450K list + dropLociWithSnps + drop chrX/Y | Chen 2013 list is 450K-specific; standard EWAS attrition |
| EPIC v1 cohort | maxprobes EPIC list + dropLociWithSnps + sex-chr decision | Pidsley 2016 list is EPIC-specific; do not reuse the 450K list |
| EPIC v2 cohort | sesame betasCollapseToPfx FIRST, then filter on hg38 anno | collapse replicates before any per-locus filter or FDR |
| Sex is the phenotype | keep chrX/chrY, analyze sex-stratified | dropping sex probes throws away the signal of interest |
| Suspected mislabels / replicates | getSex() vs sheet + getSnpBeta() rs-fingerprint clustering | swaps and duplicates are invisible in methylation alone |
| Merge 450K + EPIC + EPICv2 | intersect probe IDs after collapse; mLiftOver for coordinates | EPICv2 is hg38, others hg19; counts/coords clash otherwise |
| Chip/position confounded with group | -> ewas-design | unrecoverable by filtering; a design problem |
| Need the corrected matrix to filter | -> array-preprocessing | this skill QC's a matrix it does not produce |
Goal: Reduce a corrected GenomicRatioSet to the probes whose beta values reflect methylation rather than noise, genotype, or cross-hybridization.
Approach: Drop SNP-overlapping probes with minfi, remove the array-version-matched cross-reactive list with maxprobes, optionally drop sex-chromosome probes, and record the attrition at each step for the methods section.
library(minfi)
library(maxprobes)
# gset is a corrected GenomicRatioSet from array-preprocessing (IDAT -> noob/funnorm -> ratios)
start_n <- nrow(gset)
# SNP at the CpG interrogation or single-base-extension site reports genotype, not methylation.
# maf=0 drops ANY annotated SNP (conservative EWAS default); raise maf to keep rare variants.
gset <- dropLociWithSnps(gset, snps = c('CpG', 'SBE'), maf = 0)
# Cross-reactive list is ARRAY-VERSION-specific: 'EPIC' (Pidsley 2016) vs '450K' (Chen 2013).
gset <- maxprobes::dropXreactiveLoci(gset)
# Sex-chromosome probes are sex-confounded; drop for autosomal EWAS or analyze sex-stratified.
anno <- getAnnotation(gset)
autosomal <- !(anno$chr %in% c('chrX', 'chrY'))
gset <- gset[autosomal, ]
attrition <- c(start = start_n, after_snp_xreact_sex = nrow(gset))
attritionPer-probe detection-p and low-bead masking depend on the raw two-channel signal and control probes, which only exist at the RGChannelSet/SigDF stage handled in array-preprocessing. That skill applies detectionP() (minfi) or pOOBAH (sesame, which also catches deletion-driven false-intermediate calls) and getNBeads() before producing the corrected matrix. This skill assumes that masking is already done; if a supplied beta matrix has NOT been detection-masked, route back to array-preprocessing rather than trusting the betas. The thresholds (detection-p, fraction-of-samples-failed) live with the masking step, not here.
Goal: Reduce EPICv2's multiple probe designs per locus to one value per legacy CpG before any per-CpG analysis or cross-version merge.
Approach: Collapse replicate betas by probe-ID prefix with sesame, choosing mean (default) or the minimum-detection-p replicate, which also strips the design suffix so IDs revert to the classic cg form.
library(sesame)
# EPICv2 IDs carry a design/replicate suffix (e.g. cg00000029_TC21); ~5,100 loci have 2-10 designs.
# Leaving replicates uncollapsed counts a locus multiple times: inflates its weight, makes
# correlated duplicate "tests" break per-CpG FDR, and corrupts any cross-version merge.
betas_collapsed <- betasCollapseToPfx(betas_epicv2) # averages the replicate designs to one value per cg core ID
# betasCollapseToPfx only AVERAGES (it takes betas and nothing else). To keep the best-detection
# replicate instead, request collapse at the SigDF stage from the IDATs (a beta matrix has already
# discarded the per-probe detection p that minPval needs):
# betas <- openSesame(idat_prefixes, func = getBetas, collapseToPfx = TRUE, collapseMethod = 'minPval')Goal: Merge 450K, EPIC, and EPICv2 cohorts (or apply a 450K-trained clock/EWAS signature to EPICv2) without double-counting loci or clashing genome builds.
Approach: Collapse EPICv2 replicates first, intersect on the shared cg core IDs, then liftover coordinates because EPICv2 is hg38 while 450K/EPIC are hg19.
library(sesame)
# 1. Collapse EPICv2 to cg core IDs (above), then intersect probe sets across versions.
shared <- Reduce(intersect, list(rownames(betas_450k), rownames(betas_epic), rownames(betas_collapsed)))
# 2. Coordinates differ by build: EPICv2 is hg38, 450K/EPIC are hg19. mLiftOver harmonizes
# probe-level data across platforms/builds; intersect IDs first, lift coordinates before merging.
# betas_v2_hg19 <- mLiftOver(betas_collapsed, target_platform = 'HM450')
merged <- cbind(betas_450k[shared, ], betas_epic[shared, ], betas_collapsed[shared, ])
dim(merged) # a 450K-trained clock/EWAS does not transfer to EPICv2 without this intersectionGoal: Catch sample swaps, mislabels, and unintended duplicates before any analysis - the single most common data-integrity failure.
Approach: Predict sex from chrX/chrY intensity and compare to the sample sheet, then cluster samples on the rs-SNP genotyping probes (65 on 450K, ~59 on EPIC) to find duplicates and swaps independent of methylation.
library(minfi)
# Sex from log2(median chrY intensity) - log2(median chrX intensity); two clusters = M/F.
# A predicted sex that disagrees with the sample sheet is the canonical sample-swap flag.
predicted <- getSex(gmset) # gmset = mapped MethylSet/GenomicMethylSet
mismatch <- predicted$predictedSex != sample_sheet$Sex
sample_sheet$Basename[mismatch]
# rs-SNP fingerprint: ~59 explicit rs genotyping probes. Clustering on these betas (each ~0/0.5/1)
# reveals duplicate individuals and swaps regardless of methylation - genotype is identity.
snp_betas <- getSnpBeta(rgset) # rgset = the raw RGChannelSet from array-preprocessing
identity_clusters <- hclust(dist(t(snp_betas)))
plot(identity_clusters) # technical replicates of one person cluster tightlySentrix chip (BeadChip barcode) and array position (Sentrix_Position, the row/column on the chip) are the dominant technical axes in Infinium data. This skill DIAGNOSES whether they associate with top variance components; it does NOT correct them. ChAMP's champ.SVD() regresses the leading singular vectors of the beta matrix against chip, position, plate, and the biological factors, flagging which technical axis loads on real variance. If chip or position is confounded with the biological group, it is mathematically unrecoverable - hand the explicit correction (ComBat/SVA, or chip/position as covariates/random effects) and the design fix to ewas-design.
Trigger: running a per-CpG test without dropLociWithSnps. Mechanism: a SNP at the CpG or SBE site makes the probe report genotype, not methylation. Symptom: reproducible "associations" that are actually genetic (often mQTL-driven, trimodal beta). Fix: dropLociWithSnps(snps=c('CpG','SBE'), maf=0); raise maf only to deliberately keep rare variants.
Trigger: applying the Chen 2013 450K list to EPIC/EPICv2 data (or vice versa). Mechanism: the cross-reactive probe set is array-version-specific. Symptom: wrong probes dropped, real cross-reactive probes retained. Fix: use the array-matched list (xreactive_probes(array_type='EPIC') vs '450K'); maxprobes maps EPICv2 via the collapsed EPIC core IDs.
Trigger: treating EPICv2 betas as if probe IDs were unique. Mechanism: ~5,100 loci have 2-10 designs; the same locus appears multiple times. Symptom: duplicated rownames, inflated locus weight, broken per-CpG FDR, corrupted cross-version merge. Fix: betasCollapseToPfx() first; strip the suffix back to the cg core ID before anything downstream.
Trigger: merging EPICv2 (hg38) coordinates with 450K/EPIC (hg19). Mechanism: EPICv2 annotation is hg38-native. Symptom: loci silently misaligned by the hg19/hg38 offset. Fix: intersect on cg IDs and mLiftOver (or restrict to shared IDs and track the build per version).
Trigger: analyzing without the sex/identity QC. Mechanism: a mislabeled IDAT carries the wrong phenotype. Symptom: weakened or spurious associations; getSex() disagrees with the sheet; rs-fingerprints cluster two "different" samples. Fix: run getSex() vs sample sheet and getSnpBeta() fingerprint clustering as mandatory pre-analysis QC.
Trigger: dropping every flagged probe reflexively. Mechanism: some "cross-reactive" probes are fine for the specific locus of interest; maf=0 removes any-SNP probes including innocuous ones. Symptom: real signal discarded; clock/signature CpGs lost. Fix: treat the maf cutoff and cross-reactive list as calibrated to the question; report attrition and check that target CpGs survive.
| Threshold | Source | Rationale |
|---|---|---|
| detection-p > 0.01 = failed | Aryee 2014 Bioinformatics 30:1363 | signal indistinguishable from background; applied in array-preprocessing |
| bead count < 3 = unreliable | minfi docs | too few beads per probe to trust the intensity |
dropLociWithSnps(maf=0) | minfi docs | maf=0 drops any annotated SNP; raise to keep rare variants (calibrated) |
| ~6% of 450K probes cross-reactive | Chen 2013 Epigenetics 8:203 | ~29-39K loci co-hybridize; array-version-specific list |
| EPICv2 ~5,100 replicate loci (2-10 designs) | Kaur 2023 Epigenetics Commun 3:6 | collapse to one cg core ID before per-locus FDR |
| 65 rs-SNP probes on 450K (~59 on EPIC) | minfi annotation | enough genotype to fingerprint identity and catch swaps |
| getSex on log2 medY - log2 medX | Aryee 2014 Bioinformatics 30:1363 | X/Y intensity clusters by sex; mismatch = swap flag |
| Error / symptom | Cause | Solution |
|---|---|---|
| Duplicated rownames in EPICv2 beta matrix | replicate probes not collapsed | betasCollapseToPfx() before merge/test |
| Reproducible genetic-looking hits | SNP-overlap probes retained | dropLociWithSnps(snps=c('CpG','SBE'), maf=0) |
| Coordinates off when merging cohorts | EPICv2 hg38 vs 450K/EPIC hg19 | intersect cg IDs; mLiftOver before merge |
getSex disagrees with sample sheet | sample swap/mislabel | trace the IDAT; rs-SNP fingerprint to confirm |
| Cross-reactive filter drops too few/many | wrong array_type list | match the list to the array version |
dropXreactiveLoci errors on EPICv2 object | maxprobes keys on EPIC core IDs | collapse EPICv2 to cg core IDs first |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in methylation-analysis/array-qc-filtering of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Methylation Array Qc Filtering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Methylation Array Qc Filtering this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.8k | Automated safety check: Pass | MIT | |
| Literature EngineerWILLOSCAR/research-units-pipeline-skills | 513 | — | ~1.5k | Automated safety check: Pass | None | |
| Tooluniverse Metabolomics Analysiswu-yc/LabClaw | 1.1k | 2 repos | ~5.9k | Automated safety check: Pass | None | |
| Bio Geo Datamajiayu000/claude-skill-registry | 666 | 3 repos | ~4.4k | Automated safety check: Pass | MIT | |
| Gene Protein Expression Matrix Normalizationaipoch/medical-research-skills | 2k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Bio Spatial Transcriptomics Spatial Preprocessingmajiayu000/claude-skill-registry | 666 | 2 repos | ~1.3k | Automated safety check: Pass | MIT |
WILLOSCAR/research-units-pipeline-skills
Multi-route literature expansion + metadata normalization for evidence-first surveys.
wu-yc/LabClaw
Analyze metabolomics data including metabolite identification, quantification, pathway analysis, and metabolic flux.
majiayu000/claude-skill-registry
Query and download from NCBI Gene Expression Omnibus (GEO) and EMBL-EBI's BioStudies/ArrayExpress mirror.
aipoch/medical-research-skills
A skill your agent uses when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory…
majiayu000/claude-skill-registry
Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.
majiayu000/claude-skill-registry
Calculate tumor mutational burden from panel or WES data with proper normalization and clinical thresholds.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Categories
Performs probe filtering and sample-level QC on Illumina Infinium methylation arrays (450K / EPIC / EPICv2) to decide which probes and samples to trust. Bio Methylation Array Qc Filtering is an agent skill from GPTomics/bioSkills. Performs probe filtering and sample-level QC on Illumina Infinium methylation arrays (450K / EPIC / EPICv2) to decide which probes and samples to trust.
Bio Methylation Array Qc Filtering fits situations like: filtering methylation array probes; detecting sample swaps; collapsing EPICv2 replicates; merging 450K/EPIC/EPICv2 cohorts.
Run `npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a claude-code`. Or copy the skill folder (methylation-analysis/array-qc-filtering in GPTomics/bioSkills) into .claude/skills/bio-methylation-array-qc-filtering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a codex`. Or copy the skill folder (methylation-analysis/array-qc-filtering in GPTomics/bioSkills) into .agents/skills/bio-methylation-array-qc-filtering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-methylation-array-qc-filtering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-methylation-array-qc-filtering, .gemini/skills/bio-methylation-array-qc-filtering, .github/skills/bio-methylation-array-qc-filtering and .opencode/skills/bio-methylation-array-qc-filtering in your project.
Going by SKILL.md and its folder, Bio Methylation Array Qc Filtering needs R for the scripts in its folder.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Methylation Array Qc Filtering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Methylation Array Qc Filtering: Literature Engineer (WILLOSCAR/research-units-pipeline-skills, 513 stars), Tooluniverse Metabolomics Analysis (wu-yc/LabClaw, 1.1k stars), Bio Geo Data (majiayu000/claude-skill-registry, 666 stars) and Gene Protein Expression Matrix Normalization (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.