Agent skill

Bio Single Cell Multimodal Integration

by GPTomics in GPTomics/bioSkills

Integrate multimodal single-cell data (CITE-seq RNA+protein, 10x Multiome RNA+ATAC, unpaired/diagonal RNA+ATAC) and choose the right joint method.

MITAuto-check passedResearch & Science

Install Bio Single Cell Multimodal Integration

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-multimodal-integration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-single-cell-multimodal-integration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/single-cell/multimodal-integration .claude/skills/bio-single-cell-multimodal-integration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-single-cell-multimodal-integration
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.6k tokens
SKILL.md length
1,769 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Integrate multimodal single-cell data (CITE-seq RNA+protein, 10x Multiome RNA+ATAC, unpaired/diagonal RNA+ATAC) and choose the right joint method.

  • Classifying an integration task by anchor structure (paired vs unpaired)
  • SKILL.md covers Version Compatibility, Governing Principle, Classify the Task: Anchor… and Method Decision Table (Paired…, plus 12 more sections
  • Runs R and Python scripts from its folder; calls pip
  • Denoising CITE-seq ADT background before joint embedding

What it does

Bio Single Cell Multimodal Integration is an agent skill from GPTomics/bioSkills. Integrate multimodal single-cell data (CITE-seq RNA+protein, 10x Multiome RNA+ATAC, unpaired/diagonal RNA+ATAC) and choose the right joint method. Use when classifying an integration task by anchor structure (paired vs unpaired), denoising CITE-seq ADT background before joint embedding, picking between WNN, totalVI, MultiVI, MOFA+, GLUE, or Seurat v5 bridge integration, or diagnosing why a modality dominates a joint clustering.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/cite_seq_analysis.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Embeddings. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Classifying an integration task by anchor structure (paired vs unpaired)
  • Denoising CITE-seq ADT background before joint embedding
  • Picking between WNN
  • Seurat v5 bridge integration

Example prompts

  • “/bio-single-cell-multimodal-integration”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Single Cell Multimodal Integration loads about 4.6k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 1,769 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,769 words, ~4,614 tokens.

Download SKILL.mdSave it as .claude/skills/bio-single-cell-multimodal-integration/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-single-cell-multimodal-integration
description
Integrate multimodal single-cell data (CITE-seq RNA+protein, 10x Multiome RNA+ATAC, unpaired/diagonal RNA+ATAC) and choose the right joint method. Use when classifying an integration task by anchor structure (paired vs unpaired), denoising CITE-seq ADT background before joint embedding, picking between WNN, totalVI, MultiVI, MOFA+, GLUE, or Seurat v5 bridge integration, or diagnosing why a modality dominates a joint clustering.
tool_type
mixed
primary_tool
Seurat

Version Compatibility

Reference examples tested with: scanpy 1.10+, Seurat 5.0+, anndata 0.10+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Multimodal Integration

"Jointly analyze my CITE-seq / Multiome / unpaired multi-omic data" -> Classify the task by anchor structure, denoise each modality in its native pipeline, then build one joint representation.

  • R: Seurat::FindMultiModalNeighbors() (WNN), Signac (ATAC LSI), dsb::DSBNormalizeProtein() (ADT denoising), PrepareBridgeReference() (v5 bridge)
  • Python: muon/mudata (MuData container), scvi.model.TOTALVI / MULTIVI, MOFA2/muon.tl.mofa, scglue (diagonal)

Governing Principle

Classify the integration task by its anchor structure FIRST, because the anchor decides which algorithm class is even applicable (Argelaguet 2021).

  • Horizontal: same modality, different cells; anchor = shared features (batch correction, not this skill).
  • Vertical (paired, same cell): multiple modalities measured in the same cells; anchor = shared cells. CITE-seq, 10x Multiome.
  • Diagonal (unpaired): different modalities in different cells, no shared cells and no shared features; correspondence is inferred from prior knowledge. Independent scRNA + scATAC.
  • Mosaic: partially observed grid of (modalities x batches); some blocks present, some missing.

Paired vs unpaired is the master fork: paired correspondence is known a priori (WNN, totalVI, MultiVI, MOFA+), unpaired/diagonal correspondence must be inferred (GLUE, Seurat v5 bridge), mosaic mixes both (MultiVI, StabMap). Two separately-paired datasets that share only one modality (for example a 10x Multiome and a CITE-seq experiment sharing only RNA) are a mosaic problem: anchor on the shared RNA and impute or bridge the modality-specific blocks with StabMap, MultiVI, or Seurat v5 bridge integration rather than forcing a single WNN.

CITE-seq ADT background is a three-part mixture, not one "ambient" term: (1) ambient antibody captured in every droplet including empties, (2) cell-intrinsic non-specific binding (Fc receptors, sticky dying cells) that does NOT appear in empties, (3) spillover/index hopping between barcodes. Denoise ADT (DSB or totalVI's built-in background mixture) BEFORE any joint embedding; raw or CLR-only ADT carries this background into the joint graph.

WNN can be dominated by the noisier modality: weights reward local neighbor predictability, and a handful of high-variance or saturating ADT features can manufacture self-consistent neighborhoods and get up-weighted despite carrying less biology. Report the per-cell weight distribution and check whether clustering survives down-weighting the suspect modality.

Imputed modalities are inferences, not measurements: MultiVI/StabMap/Cobolt impute the missing modality for unpaired cells, and gene-activity scores from ATAC approximate RNA. Differential expression or marker calls on imputed values are model-dependent and must be flagged as such.

Classify the Task: Anchor Structure -> Method Class

Anchor structureWhat is sharedExample assayMethod class
Vertical / pairedSame cellsCITE-seq, 10x MultiomeWNN, totalVI, MultiVI(paired), MOFA+, mojitoo
Diagonal / unpairedNothing (prior graph)Independent scRNA + scATACGLUE, Seurat v5 bridge, LIGER
MosaicSome modalities onlyBatch A RNA+ATAC, batch B RNAMultiVI, StabMap, Cobolt, totalVI(partial)

When methods compete, verify the current best-practice default against the installed tool docs before committing; the field moves and defaults drift across minor versions.

Method Decision Table (Paired CITE-seq / Multiome)

MethodModel / assumptionUse whenFails when
WNN (Seurat)Per-cell, per-modality weights from cross-modality neighbor prediction; one weighted graphFast joint embedding/clustering of one well-normalized paired datasetProtein background not removed upstream; noisy/saturating modality dominates; not for unpaired/mosaic
totalVI (scvi-tools)Conditional VAE; RNA NB/ZINB, each protein a 2-component NB mixture (background+foreground)Need denoised protein, principled DE, batch integration, merging different antibody panelsTiny datasets (VAE overfits); no GPU and very large data; protein-specific background structure not captured by one per-cell factor
MultiVI (scvi-tools)Single joint VAE over RNA+ATAC(+protein); mosaic-capable, imputes missing modalityPaired+unpaired RNA/ATAC mixed (mosaic); want generative DE/DA"batch" key is the modality indicator, not sequencing batch; imputed modalities treated as measured
MOFA+ (MOFA2)Linear Bayesian group factor analysis; sparse factors, per-modality variance explainedInterpreting shared vs modality-specific axes of variation (exploratory/explanatory)Used for clustering/denoising; likelihood mismatched to data; expecting batch correction within a view
mojitooCCA across precomputed per-modality reductions; fast, parameter-freeQuick paired joint reduction from existing PCA/LSI slotsNo knob to down-weight a noisy modality; bounded by input reductions; paired only

Method Decision Table (Unpaired / Diagonal / Mosaic)

MethodModel / assumptionUse whenFails when
GLUE (scglue)Per-modality VAEs + prior feature graph (peak-near-gene); adversarial cell alignmentUnpaired diagonal scRNA + scATAC; want regulatory inference as a byproductGenome-build/coordinate mismatch yields an empty guidance graph and garbage alignment; adversarial over-mixing of distinct states
Seurat v5 bridgeMultiome bridge dataset = dictionary linking query modality to reference modalityMapping a query (scATAC) onto a reference built in another modality (scRNA)Poor/batch-mismatched bridge propagates error; rare query-only populations mislabeled
StabMapMosaic topology from shared features; project all cells via shortest pathsMosaic with informative unshared features that cannot be droppedUnshared-feature chaining compounds error per hop
Cobolt / scMoMaTGenerative shared latent over joint + single-modality datasetsMosaic where a generative latent is preferred over feature chainingDE/marker calls made on imputed values

ADT Normalization: CLR vs DSB

MethodWhat it doesUse whenFails when
CLR (centered log-ratio)Rescales compositionally; Seurat NormalizeData(method="CLR", margin=2)Quick, no empty droplets available; small panelsDoes NOT remove background; geometric-mean denominator distorted by saturating high-abundance ADTs
DSBAmbient correction from empty droplets + per-cell technical denoising via 2-component mixture + isotype controlsRaw/unfiltered matrix available (needs empty droplets); want background removed before embeddingNo empty droplets retained; protein-specific non-specific binding (one per-cell factor under/over-corrects); no clearly bimodal proteins

Seurat's CLR margin is genuinely ambiguous across versions (margin=2 = per-feature is the WNN-tutorial recommendation for large panels); verify with ?NormalizeData on the installed version.

CITE-seq: Denoise ADT, Then Joint Embed (Seurat)

Goal: Remove ADT background with DSB before WNN, because WNN does not denoise protein.

Approach: Estimate ambient from empty droplets and per-cell technical noise from a mixture plus isotype controls, then feed denoised ADT into the standard PCA -> WNN flow.

r
library(dsb)
library(Seurat)

raw <- Read10X('raw_feature_bc_matrix/')           # unfiltered: contains empty droplets
cells <- Read10X('filtered_feature_bc_matrix/')    # called cells

adt_cells <- as.matrix(cells[['Antibody Capture']])
adt_empty <- as.matrix(raw[['Antibody Capture']][, setdiff(colnames(raw[['Antibody Capture']]), colnames(adt_cells))])

# isotype.control.name.vec must name the ACTUAL isotype rows (often IgG1/IgG2a/Mouse-IgG2b-Ctrl); the regex below misses those
# When isotypes are absent or not matched, set use.isotype.control = FALSE (keep denoise.counts = TRUE) and pass real names explicitly
adt_dsb <- DSBNormalizeProtein(
    cell_protein_matrix = adt_cells,
    empty_drop_matrix = adt_empty,
    denoise.counts = TRUE,
    use.isotype.control = TRUE,
    isotype.control.name.vec = grep('[Ii]sotype|IgG', rownames(adt_cells), value = TRUE)
)

CITE-seq: WNN Joint Clustering (Seurat)

Goal: Build one weighted-NN graph from denoised RNA and ADT and cluster on it.

Approach: Reduce each modality independently (PCA on RNA, PCA on the small ADT panel), then learn per-cell modality weights and cluster/embed on the joint graph.

r
obj[['ADT']] <- CreateAssay5Object(data = adt_dsb)        # DSB output is already normalized data
DefaultAssay(obj) <- 'RNA'
obj <- NormalizeData(obj) |> FindVariableFeatures() |> ScaleData() |> RunPCA(reduction.name = 'pca')

DefaultAssay(obj) <- 'ADT'
VariableFeatures(obj) <- rownames(obj[['ADT']])
obj <- ScaleData(obj) |> RunPCA(reduction.name = 'apca', npcs = min(18, nrow(obj[['ADT']]) - 1))

# dims.list matched to informative dims; small ADT panels saturate by ~1:18
obj <- FindMultiModalNeighbors(obj, reduction.list = list('pca', 'apca'), dims.list = list(1:30, 1:18))
obj <- FindClusters(obj, graph.name = 'wsnn', algorithm = 3)   # algorithm 3 = SLM (the tutorial choice), NOT Leiden
obj <- RunUMAP(obj, nn.name = 'weighted.nn', reduction.name = 'wnn.umap')

# Inspect the per-cell weight distribution; a single dominant modality is a red flag
VlnPlot(obj, features = 'RNA.weight', group.by = 'seurat_clusters')

CITE-seq: totalVI (Python, denoise + DE in one model)

Goal: Jointly model RNA + protein with explicit protein background, yielding a denoised latent space and foreground probabilities.

Approach: Register a MuData object, train the conditional VAE, then read the latent representation and per-protein foreground probability.

python
import scvi
import mudata as md

# mdata holds .mod['rna'] (raw counts) and .mod['prot'] (raw ADT counts)
scvi.model.TOTALVI.setup_mudata(
    mdata, rna_layer='counts', protein_layer=None,
    modalities={'rna_layer': 'rna', 'protein_layer': 'prot'}
)
model = scvi.model.TOTALVI(mdata)
model.train()

mdata.obsm['X_totalVI'] = model.get_latent_representation()
fg = model.get_protein_foreground_probability()        # 1 - background mixing weight per protein per cell
denoised_rna, denoised_prot = model.get_normalized_expression()
Show full SKILL.md (710 more words)Show less

Multiome (RNA + ATAC, same cell): Native Pipelines, Then Join

Goal: Process each modality in its own statistics before joining, because RNA and ATAC have incompatible distributions.

Approach: PCA on RNA, TF-IDF + LSI on ATAC (drop depth-correlated components), then WNN. See scatac-analysis for ATAC QC and the binarization/depth-component caveats.

r
library(Signac)
DefaultAssay(obj) <- 'RNA'
obj <- NormalizeData(obj) |> FindVariableFeatures() |> ScaleData() |> RunPCA()

DefaultAssay(obj) <- 'ATAC'
obj <- RunTFIDF(obj) |> FindTopFeatures(min.cutoff = 'q0') |> RunSVD()
DepthCor(obj)                                          # diagnose which LSI components track depth

# dims = 2:30 drops LSI_1 ONLY if DepthCor confirms it tracks depth (usually true, not guaranteed)
obj <- FindMultiModalNeighbors(obj, reduction.list = list('pca', 'lsi'), dims.list = list(1:30, 2:30))
obj <- RunUMAP(obj, nn.name = 'weighted.nn', reduction.name = 'wnn.umap')
obj <- FindClusters(obj, graph.name = 'wsnn', algorithm = 3)

Merging multiome datasets requires a common peak set: re-quantify all cells against unified peaks, or peak-boundary differences manufacture spurious batch structure. The ATAC gene-activity matrix is an approximation, not measured RNA; do not conflate it with the RNA modality.

MOFA+ (interpretable shared/specific factors)

Goal: Decompose modalities into shared latent factors with per-modality variance explained.

Approach: Build a MOFA object from per-modality matrices, set likelihoods to match each data type, run, then interpret factor loadings.

python
import muon as mu

# likelihoods must match data: gaussian for scaled RNA, bernoulli for binarized ATAC, poisson for counts
mu.tl.mofa(mdata, n_factors=15, outfile='mofa_model.hdf5')   # writes mdata.obsm['X_mofa']

Unpaired / Diagonal: GLUE (Python)

Goal: Align independent scRNA and scATAC with no shared cells via a prior feature graph.

Approach: Configure each dataset with a count-appropriate probabilistic model, build a gene-anchored guidance graph, fit GLUE, then read aligned embeddings.

python
import scglue

scglue.models.configure_dataset(rna, 'NB', use_highly_variable=True, use_rep='X_pca')     # NB needs RAW counts
scglue.models.configure_dataset(atac, 'ZINB', use_highly_variable=True, use_rep='X_lsi')
graph = scglue.genomics.rna_anchored_guidance_graph(rna, atac)     # peak-near-gene prior; coords must share genome build
glue = scglue.models.fit_SCGLUE({'rna': rna, 'atac': atac}, graph)
rna.obsm['X_glue'] = glue.encode_data('rna', rna)
atac.obsm['X_glue'] = glue.encode_data('atac', atac)

Verify cell-type structure is preserved (not just modality overlap); adversarial alignment can over-mix distinct populations.

MuData Housekeeping

After per-modality QC, modalities hold different cell sets; muon.pp.intersect_obs(mdata) before any paired analysis. Editing a modality-local mdata.mod['rna'].obs needs mdata.update() to propagate to the global mdata.obs. R round-trips (MuDataSeurat, zellkonverter) are lossy; plan to stay in one ecosystem.

Common Errors

SymptomCauseFix
WNN clustering driven entirely by ADTA few saturating high-variance proteins dominate the neighbor graphReport per-cell weight distribution; down-weight or denoise ADT (DSB); re-check clustering stability
"Background" smear in every ADT clusterRan WNN/CLR without empty-droplet denoisingRun DSB (needs raw/unfiltered matrix) or totalVI before joint embedding
DSB errors / nonsense outputPassed a filtered cell matrix only (no empty droplets)Supply empty_drop_matrix from the raw/unfiltered matrix
Spurious batch structure after merging multiomePer-dataset peak sets, not a unified setRe-quantify all cells against one common peak set
GLUE produces a blob / no alignmentGuidance graph near-empty from genome-build/coordinate mismatchAlign RNA gene coords and ATAC peaks to the same build before building the graph
RNA and protein disagree for a markerOften real post-transcriptional biology (stability, trafficking, lag), not an artifactDo not "correct away"; treat single-gene discordance as informative
MultiVI batch effects persistThe batch_key was set to the modality indicator, not sequencing batchAdd a separate covariate for the real batch
DE on a modality looks too cleanComputed on imputed/gene-activity values, not measurementsFlag imputed-modality DE as model-dependent; validate against a measured modality
  • single-cell/scatac-analysis - ATAC QC, TF-IDF/LSI, gene-activity caveats for the Multiome ATAC half
  • single-cell/preprocessing - per-modality RNA QC and normalization before integration
  • single-cell/clustering - clustering and UMAP on the joint graph
  • single-cell/batch-integration - horizontal (same-modality, cross-sample) correction
  • single-cell/markers-annotation - marker-based interpretation of joint clusters
  • atac-seq/motif-deviation - chromVAR TF activity on the Multiome ATAC modality
  • pathway-analysis/go-enrichment - functional interpretation of modality-specific factors

References

Argelaguet R, Cuomo ASE, Stegle O, Marioni JC. Computational principles and challenges in single-cell data integration. Nat Biotechnol 39(10):1202-1215 (2021). Stoeckius M, Hafemeister C, Stephenson W, et al. Simultaneous epitope and transcriptome measurement in single cells (CITE-seq). Nat Methods 14:865-868 (2017). Mulè MP, Martins AJ, Tsang JS. Normalizing and denoising protein expression data from droplet-based single-cell profiling (DSB). Nat Commun 13:2099 (2022). Hao Y, Hao S, Andersen-Nissen E, et al. Integrated analysis of multimodal single-cell data (WNN). Cell 184(13):3573-3587 (2021). Gayoso A, Steier Z, Lopez R, et al. Joint probabilistic modeling of single-cell multi-omic data with totalVI. Nat Methods 18:272-282 (2021). Ashuach T, Gabitto MI, Koodli RV, et al. MultiVI: deep generative model for the integration of multimodal data. Nat Methods 20(8):1222-1231 (2023). Argelaguet R, Arnol D, Bredikhin D, et al. MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biol 21:111 (2020). Cao Z-J, Gao G. Multi-omics single-cell data integration and regulatory inference with graph-linked unified embedding (GLUE). Nat Biotechnol 40(10):1458-1466 (2022). Hao Y, Stuart T, Kowalski MH, et al. Dictionary learning for integrative, multimodal and scalable single-cell analysis (Seurat v5 bridge). Nat Biotechnol 42:293-304 (2024). Bredikhin D, Kats I, Stegle O. MUON: multimodal omics analysis framework. Genome Biol 23:42 (2022). Ghazanfar S, Guibentif C, Marioni JC. Stabilized mosaic single-cell data integration using unshared features (StabMap). Nat Biotechnol 42(2):284-292 (2024). Yin Y, et al. Characterization and decontamination of background noise in droplet-based single-cell protein expression data with DecontPro. Nucleic Acids Res 52(1):e4 (2024).

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in single-cell/multimodal-integration of GPTomics/bioSkills.

  • SKILL.md
  • examples/cite_seq_analysis.R
  • examples/cite_seq_analysis.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Single Cell Multimodal Integration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Single Cell Multimodal Integration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Single Cell Multimodal Integration this skillGPTomics/bioSkills1.2k1 repos~4.6kAutomated safety check: PassMIT
Evo2JimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
Genimldavila7/claude-code-templates32k11 repos~2.5kAutomated safety check: PassMIT
Celltype Specificity ProfilerClawBio/ClawBio1.2k—~4.3kAutomated safety check: PassMIT
Umap Tsne Analysisaipoch/medical-research-skills2k—~2.7kAutomated safety check: PassMIT

Similar skills

  • Evo2

    JimLiu/science-skills

    Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Geniml

    davila7/claude-code-templates

    This skill should be used when working with genomic interval data (BED files) for machine learning tasks.

    32k GitHub starsUsed in 11 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Given a gene and a single-cell atlas, compute how cell-type-specific its expression is — the tau specificity index, Sarle's expression bimodality coefficient, and the cell types that drive the…

    1.2k GitHub stars~4.3k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Umap Tsne Analysis

    aipoch/medical-research-skills

    A skill your agent uses when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE…

    2k GitHub stars~2.7k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • Single Cell Preprocessing With Omicverse

    FreedomIntelligence/OpenClaw-Medical-Skills

    Walk through omicverse's single-cell preprocessing tutorials to QC PBMC3k data, normalise counts, detect HVGs, and run PCA/embedding pipelines on CPU, CPU–GPU mixed, or GPU stacks.

    3.1k GitHub starsUsed in 1 repo~2.8k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Single Cell Multimodal Integration

What does Bio Single Cell Multimodal Integration do?

Integrate multimodal single-cell data (CITE-seq RNA+protein, 10x Multiome RNA+ATAC, unpaired/diagonal RNA+ATAC) and choose the right joint method. Bio Single Cell Multimodal Integration is an agent skill from GPTomics/bioSkills. Integrate multimodal single-cell data (CITE-seq RNA+protein, 10x Multiome RNA+ATAC, unpaired/diagonal RNA+ATAC) and choose the right joint method.

When should I use Bio Single Cell Multimodal Integration?

Bio Single Cell Multimodal Integration fits situations like: classifying an integration task by anchor structure (paired vs unpaired); denoising CITE-seq ADT background before joint embedding; picking between WNN; seurat v5 bridge integration.

How do I install Bio Single Cell Multimodal Integration in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-multimodal-integration -a claude-code`. Or copy the skill folder (single-cell/multimodal-integration in GPTomics/bioSkills) into .claude/skills/bio-single-cell-multimodal-integration in your project. Claude Code loads it when a task matches its description.

How do I install Bio Single Cell Multimodal Integration in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-multimodal-integration -a codex`. Or copy the skill folder (single-cell/multimodal-integration in GPTomics/bioSkills) into .agents/skills/bio-single-cell-multimodal-integration in your project. Codex loads it when a task matches its description.

Can I use Bio Single Cell Multimodal Integration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-single-cell-multimodal-integration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-single-cell-multimodal-integration, .gemini/skills/bio-single-cell-multimodal-integration, .github/skills/bio-single-cell-multimodal-integration and .opencode/skills/bio-single-cell-multimodal-integration in your project.

What does Bio Single Cell Multimodal Integration need to run?

Going by SKILL.md and its folder, Bio Single Cell Multimodal Integration needs R and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Single Cell Multimodal Integration access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Single Cell Multimodal Integration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Single Cell Multimodal Integration use?

Bio Single Cell Multimodal Integration is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Single Cell Multimodal Integration use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Single Cell Multimodal Integration?

Skills that share tags, products or a category with Bio Single Cell Multimodal Integration: Evo2 (JimLiu/science-skills, 227 stars), Scgpt (JimLiu/science-skills, 227 stars), Geniml (davila7/claude-code-templates, 32k stars) and Celltype Specificity Profiler (ClawBio/ClawBio, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Single Cell Multimodal Integration?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.