Agent skill

Bio Multi Omics Integration Design

by GPTomics in GPTomics/bioSkills

Chooses a bulk multi-omics integration strategy before any tool runs by mapping the biological question (subtype discovery, shared axis of variation, predictive signature, pairwise correlation) to a…

MITAuto-check passedResearch & Science

Install Bio Multi Omics Integration Design

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-multi-omics-integration-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-multi-omics-integration-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/multi-omics-integration/integration-design .claude/skills/bio-multi-omics-integration-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-multi-omics-integration-design
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.5k tokens
SKILL.md length
2,544 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Chooses a bulk multi-omics integration strategy before any tool runs by mapping the biological question (subtype discovery, shared axis of variation, predictive signature, pairwise correlation) to a…

  • Works in 3 steps: The correspondence gate. Vertical… → The variance gate. Total variance scales… → The validation gate. With n = 40 and…
  • Deciding which integration method fits a question
  • SKILL.md covers Version Compatibility, The Single Most Important…, The Integration Taxonomy --… and Tool Taxonomy, plus 9 more sections
  • Runs R scripts from its folder

What it does

Bio Multi Omics Integration Design is an agent skill from GPTomics/bioSkills. Chooses a bulk multi-omics integration strategy before any tool runs by mapping the biological question (subtype discovery, shared axis of variation, predictive signature, pairwise correlation) to a method class, naming the sample correspondence (paired-vertical, horizontal, mosaic, diagonal), enforcing the n<<p discipline that makes a held-out cohort the endpoint instead of in-cohort cross-validation, and running the per-view variance-imbalance diagnostic. Covers the early/mixed/intermediate/late taxonomy, why…

Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Machine learning. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Deciding which integration method fits a question
  • Whether data is paired
  • How to validate an integrated result

Example prompts

  • “Use the bio-multi-omics-integration-design skill to choose a bulk multi-omics integration strategy before any tool runs by mapping the biological…”
  • “/bio-multi-omics-integration-design”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The correspondence gate. Vertical integration (different omics, SAME samples) and horizontal integration (same features, different…
  2. The variance gate. Total variance scales with feature count and feature scale, so a 850k-CpG block out-votes a 100-metabolite block and…
  3. The validation gate. With n = 40 and 5-fold CV each fold tests 8 samples, so in-cohort cross-validation is optimistically biased to the…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Multi Omics Integration Design loads about 5.5k tokens when it runs. Until then it costs about 259 tokens; SKILL.md has 2,544 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~259
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,544 words, ~5,488 tokens.

Download SKILL.mdSave it as .claude/skills/bio-multi-omics-integration-design/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-multi-omics-integration-design
description
Chooses a bulk multi-omics integration strategy before any tool runs by mapping the biological question (subtype discovery, shared axis of variation, predictive signature, pairwise correlation) to a method class, naming the sample correspondence (paired-vertical, horizontal, mosaic, diagonal), enforcing the n<<p discipline that makes a held-out cohort the endpoint instead of in-cohort cross-validation, and running the per-view variance-imbalance diagnostic. Covers the early/mixed/intermediate/late taxonomy, why vertical and horizontal integration are different problems, and why a shared factor dominated by one omic is not integration. Use when deciding which integration method fits a question, whether data is paired or mosaic, supervised or unsupervised, or how to validate an integrated result. For unsupervised factors see mofa-integration; for supervised signatures see mixomics-analysis; for stratification see similarity-network; for single-cell see single-cell/multimodal-integration.
tool_type
r
primary_tool
MultiAssayExperiment

Version Compatibility

Reference examples tested with: MultiAssayExperiment 1.36+, SummarizedExperiment 1.40+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

The tool versions that matter most are MOFA2 and mixOmics, whose APIs have moved across releases; this skill routes to those tool skills rather than calling them, so the binding version here is MultiAssayExperiment (the container in which the paired-vs-mosaic decision is made).

Multi-Omics Integration Design

"How should I integrate these omics?" -> Map the biological question and the sample correspondence to a method class BEFORE running a tool - because at tens of samples and 10^5-10^6 features a spurious cross-omic signal is the default outcome, not the surprise.

  • R: assemble a MultiAssayExperiment, then choose MOFA2 (shared factors) / mixOmics (signature) / SNF (subtypes) by question

Scope: the integration decision itself - method selection, correspondence (paired/horizontal/mosaic/diagonal), supervised-vs-unsupervised mapping, the n<<p discipline, and the variance-imbalance diagnostic. Running the chosen tool -> mofa-integration, mixomics-analysis, similarity-network. Cross-omic preprocessing -> data-harmonization. Single-cell multimodal -> single-cell/multimodal-integration. Horizontal same-feature meta-analysis -> differential-expression/batch-correction.

The Single Most Important Modern Insight -- Bulk Multi-Omics Is a Small-n, Huge-p Discovery Problem Where a Spurious Cross-Omic Signal Is the Default

A typical bulk cohort has n = 30-300 samples and 10^4-10^6 features per omic, so after stacking blocks n is smaller than p by three to four orders of magnitude. In that regime an integrated signature that has not been validated out-of-sample is overwhelmingly noise that fit the training samples. The deliverable is never "the integrated signature" - it is question-matched structure that survives three gates, each of which a common failure violates:

  1. The correspondence gate. Vertical integration (different omics, SAME samples) and horizontal integration (same features, different cohorts) are different problems. The bulk joint-latent tools are indexed by sample; feed them unpaired same-feature data and they still run but emit factors that are pure batch. Name the correspondence first.
  2. The variance gate. Total variance scales with feature count and feature scale, so a 850k-CpG block out-votes a 100-metabolite block and the shared factors become methylation PCs. Inspect the per-factor, per-view variance-explained table every time; if one view dominates every shared factor, the integration re-discovered the biggest omic.
  3. The validation gate. With n = 40 and 5-fold CV each fold tests 8 samples, so in-cohort cross-validation is optimistically biased to the point of fiction. An integrated subtype or signature is not credible until it reproduces in an INDEPENDENT cohort. The held-out cohort is the finding.

Organize the analysis around defending these three gates, not around picking a favorite tool.

The Integration Taxonomy -- Three Orthogonal Axes, Not One Label

A method is a point in a 3D space, not a single name. Stating where a method sits on each axis prevents the category's two deepest errors (horizontal/vertical confusion and concatenation at n<<p).

AxisValuesWhat it decides
Stage - WHEN blocks combine (Ritchie 2015, Picard 2021)early (concatenate then model), mixed (transform each block then combine), intermediate (jointly model blocks into shared + specific factors), late (model each omic, combine results)early is worst at n<<p and variance imbalance; intermediate joint-latent (MOFA/JIVE/iCluster) is the discovery sweet spot; late is robust to missing blocks but drops feature-level cross-talk
Correspondence - WHAT is tied together (Argelaguet 2021)vertical (diff omics, same samples - THIS category), horizontal (same features, diff cohorts - meta-analysis), mosaic (partial overlap), diagonal (no shared axis - single-cell)conflating horizontal and vertical is the deepest category error; only vertical and mosaic belong here
Supervision - WHETHER an outcome drives itunsupervised (discover subtypes/factors), supervised (predict/discriminate a label)unsupervised plus then-correlate-with-outcome is hypothesis-generating, NOT a validated predictor

MOFA = intermediate, vertical, unsupervised. DIABLO = intermediate, vertical, supervised. SNF = mixed/transformation, vertical, unsupervised. ComBat-across-cohorts = horizontal, unsupervised harmonization (routes OUT to differential-expression/batch-correction).

Tool Taxonomy

Tool / classCitationStage / supervisionWhen
MOFA2Argelaguet 2018 Mol Syst Biol 14:e8124; Argelaguet 2020 Genome Biol 21:111intermediate, unsupervisedshared vs view-specific factors; tolerant of missing omics-per-sample; the default factor model -> mofa-integration
mixOmics DIABLO (block.splsda)Singh 2019 Bioinformatics 35:3055; Rohart 2017 PLoS Comput Biol 13:e1005752intermediate, supervisedsparse cross-omic signature that DISCRIMINATES known groups -> mixomics-analysis
mixOmics sPLS (spls)Rohart 2017 PLoS Comput Biol 13:e1005752intermediate, unsupervisedcovariance-maximizing feature pairs between TWO blocks -> mixomics-analysis
mixOmics MINT (mint.splsda)Rohart 2017 BMC Bioinformatics 18:128horizontalSAME omic across multiple STUDIES (study as a known effect) - not cross-omic
SNF (SNFtool)Wang 2014 Nat Methods 11:333mixed/transformation, unsupervisedpatient stratification; feature count buys no votes; robust as complexity grows -> similarity-network
iCluster / iClusterPlus / moClusterShen 2009 Bioinformatics 25:2906; Meng 2016 J Proteome Res 15:755intermediate, unsupervisedONE joint-latent clustering (vs reconciling K separate clusterings); subtype discovery
JIVELock 2013 Ann Appl Stat 7:523intermediate, unsupervisedexplicit joint + individual + noise decomposition (how much signal is cross-omic)
MFA / mixKernelMariette 2018 Bioinformatics 34:1009mixedblock weighting / kernel fusion to stop one omic dominating

Decision Tree by Scenario

ScenarioRecommendedWhy
Different omics on the SAME samples, no phenotype, find shared axesMOFA2unsupervised factor model; variance decomposition; native missing-block handling -> mofa-integration
Different omics on the same samples, want patient SUBTYPESSNF + spectral clustering (or iCluster)transformation-stage; robust to high p and a noisy omic -> similarity-network
Have a class label, want a cross-omic signature that discriminates itmixOmics DIABLO + held-out cohortsupervised sparse multi-block PLS-DA -> mixomics-analysis
Just two omics, want correlated feature pairsmixOmics sPLSsparse PLS for a block pair -> mixomics-analysis
Quantify how much variation is joint vs omic-specificJIVE (or the MOFA variance table)explicit joint/individual split
SAME omic across multiple studies/cohorts-> differential-expression/batch-correction or mixOmics MINThorizontal integration / meta-analysis, NOT cross-omic
Mosaic cohort (some samples missing an omic)MOFA2 (models the missingness)intersecting to complete cases wastes scarce n -> data-harmonization
Single-cell CITE-seq / 10x Multiome / unpaired diagonal-> single-cell/multimodal-integrationper-cell generative models; n is large; different paradigm
Per-omic DE then overlap the hit lists-> differential-expression, methylation-analysis, proteomicsthat is late integration by intersection, not joint modeling
Validate a discovered subtype against outcome-> clinical-biostatistics/survival-analysissurvival / KM / Cox lives there

Default when uncertain: assemble a MultiAssayExperiment, confirm vertical paired (or mosaic) correspondence, run MOFA2 for an unsupervised map and read its per-view variance-explained table, then escalate to a supervised (DIABLO) or stratification (SNF) tool only if the question demands it.

Name the Correspondence First

Goal: Decide whether the data is a job for this category at all, and whether to model the missingness or intersect to complete cases.

Approach: Assemble the blocks into a MultiAssayExperiment (it coordinates assays, a sample map, and colData), then read off whether samples are fully paired, mosaic, or actually horizontal. Only vertical-paired and mosaic belong here.

r
library(MultiAssayExperiment)

mae <- MultiAssayExperiment(experiments=ExperimentList(rna=rna_mat, prot=prot_mat, methyl=methyl_mat),
                            colData=clinical)
upsetSamples(mae)                 # visualize which samples have which omics (mosaic structure)
table(complete.cases(mae))        # how many samples have EVERY omic
paired <- intersectColumns(mae)   # complete-case fallback - counts the n it would cost

If complete.cases keeps most samples, complete-case methods (mixOmics, SNF) are fine. If a large fraction is mosaic, prefer MOFA2 (it models missing-view samples in its likelihood) over intersecting, because at n<<p discarding incomplete samples is expensive and imputing a whole block fabricates data (data-harmonization owns that decision).

The Variance-Imbalance Diagnostic

Goal: Detect, before trusting any shared factor, whether one omic is set to dominate the integration purely because it has more features or larger scale.

Approach: After per-feature scaling, compare each block's total variance and feature count; a block contributing the overwhelming majority of stacked variance will hijack the shared latent space. The definitive check is the per-view variance-explained table that MOFA2 reports after fitting - if every factor loads on one view, equalize the blocks (MFA weighting, per-block keepX, or move to SNF) and refit.

r
block_var <- sapply(assays_list, function(x) sum(apply(x, 1, var)))   # total variance per block
share     <- block_var / sum(block_var)
share                                                                 # any block >> others = imbalance risk

A block holding most of the stacked variance is a red flag that concatenation-style integration will re-discover it. This is the single best honesty check in the category; never skip the post-fit per-view variance read-out.

The n<<p Discipline

The held-out cohort is the endpoint, not in-cohort cross-validation. Three rules follow from n<<p:

  • Tune with repeated cross-validation, never a single run. At n = 40 a single CV estimate is mostly noise; mixOmics perf/tune.* take nrepeat (10-50) - use it. Generic CV/overfitting theory lives in machine-learning/model-validation.
  • Report out-of-sample performance, not the in-sample fit. A supervised signature's training-set discrimination is guaranteed by construction; only an independent cohort makes it a biomarker.
  • An unsupervised factor that correlates with the outcome is a hypothesis. MOFA found it without the label, which is a strength - but calling it predictive requires held-out validation, not the in-cohort correlation that found it.
  • Prefer regularized/sparse methods over early concatenation plus plain CCA. Classical CCA divides out within-block variance and, when p > n, achieves correlation 1 trivially by overfitting; sparse PLS (covariance plus an L1 penalty) and sparse factor models stay identifiable, so the regularization is what makes the fit real, not a stylistic choice.

Per-Method Failure Modes

Horizontal data fed to a vertical method

Trigger: running MOFA/DIABLO/SNF on same-feature, multi-cohort data ("integrate my three RNA-seq studies"). Mechanism: the shared latent is indexed by sample and has nothing to align across feature-identical cohorts. Symptom: the tool runs and the top factors track cohort/run, not biology. Fix: recognize this as horizontal integration; use MINT, ComBat/sva, or differential-expression/batch-correction.

Unvalidated in-cohort signature reported as a result

Trigger: reporting a DIABLO panel or a MOFA-factor-vs-outcome correlation from one cohort. Mechanism: at n<<p thousands of cross-omic feature pairs clear any threshold under the null; in-cohort CV is optimistically biased. Symptom: a beautiful signature that fails to replicate. Fix: hold out an independent cohort; frame an unvalidated finding as hypothesis-generating, never as a biomarker.

Show full SKILL.md (1,001 more words)Show less
Variance imbalance mistaken for integration

Trigger: concatenating blocks of very different feature counts/scales without equalization. Mechanism: the high-feature/high-variance omic casts the most votes for the shared factors. Symptom: every shared factor loads almost entirely on one view. Fix: read the per-view variance-explained table; equalize via MFA weighting / per-block keepX / per-feature z-scoring, or use SNF where each omic is one n x n network.

Question-method mismatch

Trigger: using a tool whose output does not answer the question (e.g. SNF clusters reported with "driver features"). Mechanism: SNF selects no features, MOFA is unsupervised, DIABLO needs a label. Symptom: claims the method cannot support (SNF drivers without a post-hoc per-omic test; MOFA factors called predictive). Fix: map question -> class first (decision tree); do post-hoc per-omic differential analysis to find SNF subtype drivers.

Cross-omic batch masquerading as shared biology

Trigger: omics generated on different platforms/labs/dates, interpreted without a technical check. Mechanism: the samples that ran together in every assay form a shared technical axis. Symptom: the top shared factor tracks run date / plate / site better than phenotype. Fix: correlate top factors against technical covariates before interpreting; correct per omic or model batch as a covariate (data-harmonization), watching for over-correction.

Mosaic cohort forced to complete cases

Trigger: intersectColumns on a mosaic cohort before integrating. Mechanism: complete-case intersection drops every sample missing any omic. Symptom: n halves and power collapses. Fix: prefer MOFA2's native missing-view handling; reserve intersection for when mosaicism is minor.

Quantitative Thresholds

ThresholdSourceRationale
n<<p by ~3-4 orders of magnitude is the default regimeSubramanian 2020 Bioinform Biol Insights 14tens of samples, 10^4-10^6 features; dictates regularized/sparse methods and held-out validation
Repeated CV nrepeat 10-50 for any tuning at small nmixOmics docs; n<<p variancea single CV run at n~40 is noise; repetition stabilizes the estimate
Per-view variance-explained dominance flag: one view >~80% of every shared factorArgelaguet 2018 Mol Syst Biol 14:e8124 (per-view variance decomposition)a factor dominated by one view is view-specific structure, not integration
Drop MOFA factors below ~1-2% variance explained in every viewArgelaguet 2018 Mol Syst Biol 14:e8124low-variance factors are noise/over-parameterization
Held-out independent cohort for any reported biomarker/subtypeSubramanian 2020 Bioinform Biol Insights 14in-cohort CV at n<<p is optimistically biased; replication is the endpoint
Cross-check the headline with a second method classCantini 2021 Nat Commun 12:124; Tini 2019 Brief Bioinform 20:1269; Pierre-Jean 2020 Brief Bioinform 21:2011no single best method and methods diverge; a real result should survive a second class (e.g. a factor model and a fusion clustering)

Common Errors

Error / symptomCauseSolution
Tool runs on multi-cohort data but factors are all batchhorizontal data in a vertical methoduse MINT / ComBat; this is meta-analysis, not cross-omic integration
n drops sharply after assembling the objectcomplete-case intersection on a mosaic cohortuse MOFA2 missing-view handling; intersect only if mosaicism is minor
Shared factors explain mostly one omicvariance imbalance (feature count / scale)equalize blocks (MFA / per-block keepX / z-score) or use SNF
Signature does not replicate in a new cohortreported from in-cohort CV at n<<phold out an independent cohort before claiming a biomarker
Two methods give different subtypesmethod-dependent result (benchmarks rank methods differently because each optimizes a different criterion - robustness vs clustering recovery vs feature selection)report which method and why; cross-check; no universal best method

References

  • Ritchie MD, Holzinger ER, Li R, Pendergrass SA, Kim D. 2015. Methods of integrating data to uncover genotype-phenotype interactions. Nat Rev Genet 16:85-97.
  • Picard M, Scott-Boyer M-P, Bodein A, Perin O, Droit A. 2021. Integration strategies of multi-omics data for machine learning analysis. Comput Struct Biotechnol J 19:3735-3746.
  • Argelaguet R, Cuomo ASE, Stegle O, Marioni JC. 2021. Computational principles and challenges in single-cell data integration. Nat Biotechnol 39:1202-1215.
  • Argelaguet R, Velten B, Arnol D, et al. 2018. Multi-Omics Factor Analysis - a framework for unsupervised integration of multi-omics data sets. Mol Syst Biol 14:e8124.
  • Argelaguet R, Arnol D, Bredikhin D, et al. 2020. MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biol 21:111.
  • Shen R, Olshen AB, Ladanyi M. 2009. Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysis. Bioinformatics 25:2906-2912.
  • Lock EF, Hoadley KA, Marron JS, Nobel AB. 2013. Joint and individual variation explained (JIVE) for integrated analysis of multiple data types. Ann Appl Stat 7:523-542.
  • Meng C, Helm D, Frejno M, Kuster B. 2016. moCluster: identifying joint patterns across multiple omics data sets. J Proteome Res 15:755-765.
  • Mariette J, Villa-Vialaneix N. 2018. Unsupervised multiple kernel learning for heterogeneous data integration. Bioinformatics 34:1009-1016.
  • Rohart F, Gautier B, Singh A, Le Cao K-A. 2017. mixOmics: an R package for 'omics feature selection and multiple data integration. PLoS Comput Biol 13:e1005752.
  • Singh A, Shannon CP, Gautier B, et al. 2019. DIABLO: an integrative approach for identifying key molecular drivers from multi-omics assays. Bioinformatics 35:3055-3062.
  • Wang B, Mezlini AM, Demir F, et al. 2014. Similarity network fusion for aggregating data types on a genomic scale. Nat Methods 11:333-337.
  • Tini G, Marchetti L, Priami C, Scott-Boyer M-P. 2019. Multi-omics integration - a comparison of unsupervised clustering methodologies. Brief Bioinform 20:1269-1279.
  • Cantini L, Zakeri P, Hernandez C, et al. 2021. Benchmarking joint multi-omics dimensionality reduction approaches for the study of cancer. Nat Commun 12:124.
  • Pierre-Jean M, Deleuze J-F, Le Floch E, Mauger F. 2020. Clustering and variable selection evaluation of 13 unsupervised methods for multi-omics data integration. Brief Bioinform 21:2011-2030.
  • Subramanian I, Verma S, Kumar S, Jere A, Anamika K. 2020. Multi-omics data integration, interpretation, and its application. Bioinform Biol Insights 14:1177932219899051.
  • mofa-integration - Unsupervised shared-factor discovery (the default tool once correspondence is vertical)
  • mixomics-analysis - Supervised DIABLO signatures, sPLS pairs, and MINT multi-study integration
  • similarity-network - Patient stratification via similarity network fusion
  • data-harmonization - Per-block normalization, scaling, batch, and the mosaic missing-omic decision
  • single-cell/multimodal-integration - Single-cell CITE-seq/Multiome integration (different paradigm)
  • differential-expression/batch-correction - Horizontal same-feature meta-analysis and batch correction
  • machine-learning/model-validation - Cross-validation and overfitting theory for supervised integration
  • clinical-biostatistics/survival-analysis - Survival validation of discovered subtypes
  • pathway-analysis/gsea - Enrichment of integrated factor or signature features
  • workflows/multi-omics-pipeline - End-to-end multi-omics integration pipeline

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in multi-omics-integration/integration-design of GPTomics/bioSkills.

  • SKILL.md
  • examples/integration_design_diagnostic.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Multi Omics Integration Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Multi Omics Integration Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Multi Omics Integration Design this skillGPTomics/bioSkills1.2k1 repos~5.5kAutomated safety check: PassMIT
tangermeme Genomic Model Analysisjmschrei/tangermeme316—~1.6kAutomated safety check: PassMIT
Gtars Genomic Interval Toolkitdavila7/claude-code-templates32k11 repos~1.9kAutomated safety check: PassMIT
AlphagenomeK-Dense-AI/scientific-agent-skills48k1 repos~4.4kAutomated safety check: NotesMIT
Bio Spatial Transcriptomics Spatial PreprocessingFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2kAutomated safety check: PassNone
External Model Validationaipoch/medical-research-skills2k—~3.2kAutomated safety check: PassMIT

Similar skills

  • Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.

    316 GitHub stars~1.6k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Gtars Genomic Interval Toolkit

    davila7/claude-code-templates

    Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.

    32k GitHub starsUsed in 11 repos~1.9k tokens
    Research & ScienceAuto-check passed
  • Alphagenome

    K-Dense-AI/scientific-agent-skills

    Looks up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC…

    48k GitHub starsUsed in 1 repo~4.4k tokens
    Research & ScienceAuto-check: notes
  • Bio Spatial Transcriptomics Spatial Preprocessing

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • External Model Validation

    aipoch/medical-research-skills

    A skill your agent uses when validating an existing prognostic risk signature on an external bulk expression cohort with survival outcomes, producing risk scores, Kaplan-Meier curves, risk…

    2k GitHub stars~3.2k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • Popv Cell Annotation

    jaechang-hits/SciAgent-Skills

    Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via…

    371 GitHub starsUsed in 2 repos~6.9k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Multi Omics Integration Design

What does Bio Multi Omics Integration Design do?

Chooses a bulk multi-omics integration strategy before any tool runs by mapping the biological question (subtype discovery, shared axis of variation, predictive signature, pairwise correlation) to a…. Bio Multi Omics Integration Design is an agent skill from GPTomics/bioSkills. Chooses a bulk multi-omics integration strategy before any tool runs by mapping the biological question (subtype discovery, shared axis of variation, predictive signature, pairwise correlation) to a method class, naming the sample correspondence (paired-vertical, horizontal, mosaic, diagonal), enforcing the n<<p discipline that makes a held-out cohort the endpoint instead of in-cohort cross-validation, and running the per-view variance-imbalance diagnostic.

When should I use Bio Multi Omics Integration Design?

Bio Multi Omics Integration Design fits situations like: deciding which integration method fits a question; whether data is paired; how to validate an integrated result.

How do I install Bio Multi Omics Integration Design in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-multi-omics-integration-design -a claude-code`. Or copy the skill folder (multi-omics-integration/integration-design in GPTomics/bioSkills) into .claude/skills/bio-multi-omics-integration-design in your project. Claude Code loads it when a task matches its description.

How do I install Bio Multi Omics Integration Design in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-multi-omics-integration-design -a codex`. Or copy the skill folder (multi-omics-integration/integration-design in GPTomics/bioSkills) into .agents/skills/bio-multi-omics-integration-design in your project. Codex loads it when a task matches its description.

Can I use Bio Multi Omics Integration Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-multi-omics-integration-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-multi-omics-integration-design, .gemini/skills/bio-multi-omics-integration-design, .github/skills/bio-multi-omics-integration-design and .opencode/skills/bio-multi-omics-integration-design in your project.

What does Bio Multi Omics Integration Design need to run?

Going by SKILL.md and its folder, Bio Multi Omics Integration Design needs R for the scripts in its folder.

Does Bio Multi Omics Integration Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Multi Omics Integration Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Multi Omics Integration Design use?

Bio Multi Omics Integration Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Multi Omics Integration Design use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Multi Omics Integration Design?

Skills that share tags, products or a category with Bio Multi Omics Integration Design: tangermeme Genomic Model Analysis (jmschrei/tangermeme, 316 stars), Gtars Genomic Interval Toolkit (davila7/claude-code-templates, 32k stars), Alphagenome (K-Dense-AI/scientific-agent-skills, 48k stars) and Bio Spatial Transcriptomics Spatial Preprocessing (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Multi Omics Integration Design?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.