Agent skill

Bio Single Cell Batch Integration

by GPTomics in GPTomics/bioSkills

Integrate multiple scRNA-seq samples or batches with Harmony, scVI/scANVI, Seurat (CCA/RPCA), fastMNN, Scanorama, or BBKNN.

MITAuto-check passedResearch & Science

Install Bio Single Cell Batch Integration

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-batch-integration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-single-cell-batch-integration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/single-cell/batch-integration .claude/skills/bio-single-cell-batch-integration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-single-cell-batch-integration
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.2k tokens
SKILL.md length
1,683 words
Files
4
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Integrate multiple scRNA-seq samples or batches with Harmony, scVI/scANVI, Seurat (CCA/RPCA), fastMNN, Scanorama, or BBKNN.

  • The dataset size and design
  • SKILL.md covers Version Compatibility, Governing Principle, When NOT to Integrate and Method Selection, plus 9 more sections
  • Runs R and Python scripts from its folder; calls pip
  • How strongly to correct

What it does

Bio Single Cell Batch Integration is an agent skill from GPTomics/bioSkills. Integrate multiple scRNA-seq samples or batches with Harmony, scVI/scANVI, Seurat (CCA/RPCA), fastMNN, Scanorama, or BBKNN. Resolves which method to use for the dataset size and design, how strongly to correct, when integration is the wrong move (confounded batch/biology), how to score integration with scIB metrics without gaming them, and why corrected expression must not be used for differential expression. Use when integrating batches or datasets, choosing an integration method, diagnosing over-correction, or…

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/harmony_integration.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • The dataset size and design
  • How strongly to correct
  • Integration is the wrong move (confounded batch/biology)
  • How to score integration with scIB metrics without gaming them

Example prompts

  • “/bio-single-cell-batch-integration”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Single Cell Batch Integration loads about 4.2k tokens when it runs. Until then it costs about 145 tokens; SKILL.md has 1,683 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~145
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,683 words, ~4,194 tokens.

Download SKILL.mdSave it as .claude/skills/bio-single-cell-batch-integration/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-single-cell-batch-integration
description
Integrate multiple scRNA-seq samples or batches with Harmony, scVI/scANVI, Seurat (CCA/RPCA), fastMNN, Scanorama, or BBKNN. Resolves which method to use for the dataset size and design, how strongly to correct, when integration is the wrong move (confounded batch/biology), how to score integration with scIB metrics without gaming them, and why corrected expression must not be used for differential expression. Use when integrating batches or datasets, choosing an integration method, diagnosing over-correction, or judging integration quality.
tool_type
mixed
primary_tool
Harmony

Version Compatibility

Reference examples tested with: scanpy 1.10+, Seurat 5.0+, scvi-tools 1.1+, harmonypy 0.0.10+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Batch Integration

"Integrate my batches" -> Learn a shared low-dimensional representation that mixes technical batches while preserving biological cell states, then cluster and visualize on it.

  • Python: sce.pp.harmony_integrate, scvi.model.SCVI, sce.pp.bbknn, scanorama
  • R: RunHarmony, IntegrateLayers (Seurat v5), fastMNN (batchelor)

Governing Principle

Integration trades batch-mixing against biological-signal preservation, and the two cannot be jointly maximized. The algorithm removes variance along directions where batches differ, on the assumption that cell-type composition is shared across batches; it does not "know" which variance is technical. When batch correlates with a real biological axis, the method cannot distinguish them and, by construction, erases biology - this is information-theoretic, not a tuning problem. Over-correction is the silent failure: it absorbs rare cell types into common neighbors, snaps continuous gradients toward shared anchors, and deletes condition-specific populations, all while batch-mixing metrics improve. The cells most worth finding (rare, transitional, novel) are exactly the ones integration most endangers. Batch and biology are unidentifiable under confounding - the only tie-breaker is external information (shared controls, multiplexed designs, known shared types), and the fix for a confounded design is experimental, not computational. Always keep the uncorrected embedding for before/after comparison, and never run differential expression on batch-corrected expression.

When NOT to Integrate

Visualize the uncorrected data first; integration is a bias-variance trade and removing batch variance risks removing biology correlated with batch.

  • Confounded design (each condition is its own batch, e.g. all controls day 1, all treated day 2): no algorithm can separate batch from biology. The diagnostic: cluster the uncorrected data and cross-tabulate clusters x batch x condition; pure-by-batch clusters that are also condition-aligned mean integration is unsafe. The fix is experimental - multiplex conditions across batches (cell hashing; genetic demux via souporcell/vireo; split each condition across capture days).
  • Technical replicates of the same tissue that already mix well: over-correction risk outweighs benefit.
  • Per-sample analyses (CNV/tumor-clone inference): integration would erase the signal of interest.

Over-correction signatures: rare types collapsing into neighbors, lost known gradients, disappearing condition-specific populations, markers no longer separating known cell types.

Method Selection

No method wins universally - the scIB benchmark (Luecken 2022, 68 method/preprocessing combos) scores integration as overall = 0.6 x bio-conservation + 0.4 x batch-removal, deliberately weighting biology higher because erasing it is worse than imperfect mixing. Methodology evolves; verify current best practice and APIs against the installed package docs, and in practice run 2-3 candidates and score them (see Evaluating Integration).

MethodModel / assumptionUse whenFails when
HarmonyIterative soft k-means linear correction in PCA space; outputs an embedding, not countsFew/simple batches, fast, low memory; strong default; best usabilityStrong nonlinear batch effects; high theta over-mixes and collapses distinct types
scVIConditional VAE on raw counts (ZINB), batch as covariate -> batch-invariant latentLarge atlases, many nested batches, strong effects; memory-efficient at scaleSmall data (under-trained); latent dims over-interpreted as "biology minus batch"
scANVISemi-supervised scVI using partial labels to protect biologySome cell labels exist and bio fidelity is paramount (tops bio-conservation)Labels noisy/wrong; training cost; closed-world for the labeled states
Seurat CCAAnchor-based, canonical correlation across datasetsStrong shared structure under large shifts; smaller dataSubstantial non-overlap or many samples -> over-correction (CCA aligns distinct states)
Seurat RPCAReciprocal-PCA anchors; faster, more conservativeLarge/many-sample data, substantial non-overlapUnder-correction when truly shared structure is subtle (raise k.anchor)
fastMNNMutual nearest neighbors in PCA spaceRare-population preservation; moderate dataOrder-sensitive (set merge.order, most-heterogeneous first); legacy mnnCorrect is slow
ScanoramaMutual NN across all dataset pairsPartial cell-type overlap across datasets; balanced bio/batchVery large data (slower than Harmony/BBKNN)
BBKNNModifies only the neighbor graph (batch-balanced kNN)Speed; only clustering/UMAP needed downstreamLeans toward batch removal; no embedding or corrected counts for other uses

scIB headline: top combined performers were scANVI, scVI, Scanorama, scGen; Harmony and Seurat were strong on simpler tasks with the best usability; BBKNN sits at the batch-removal end. "Deep methods are always best" is not supported - Harmony/Seurat win simple/small tasks; deep methods win complex/large/label-rich tasks.

Strength Parameters

Aggressive settings increase mixing and over-correction risk in lockstep - raise correction strength only after confirming under-correction, and re-check rare populations after each change.

ParameterToolEffectRationale
thetaHarmonyHigher -> more aggressive batch mixingDefault is an internal fallback, not the signature default; larger theta over-corrects
k.anchorSeuratHigher -> more anchors, stronger correctionRaise (e.g. 20) only when under-correcting
CCA vs RPCASeuratCCA more sensitive but can over-correct; RPCA conservativePrefer RPCA for large/non-overlapping data
n_latentscVILatent dimensionality of the embedding~10-30; too high refits noise, too low under-fits
merge.orderfastMNNOrder batches are mergedOrder-sensitive; merge most-heterogeneous batch first

Integrate with Harmony

Goal: Correct batch in PCA space and run downstream steps on the corrected embedding. Approach: Joint preprocessing -> PCA -> Harmony -> neighbors/UMAP/clustering on X_pca_harmony.

python
import scanpy as sc
import scanpy.external as sce

adata = sc.read_h5ad('merged.h5ad')
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
sc.pp.highly_variable_genes(adata, n_top_genes=2000, batch_key='batch')
adata.raw = adata
adata = adata[:, adata.var.highly_variable]
sc.pp.scale(adata, max_value=10)
sc.tl.pca(adata, n_comps=50)

sce.pp.harmony_integrate(adata, key='batch')  # writes adata.obsm['X_pca_harmony']
sc.pp.neighbors(adata, use_rep='X_pca_harmony')
sc.tl.umap(adata)
sc.tl.leiden(adata, flavor='igraph', n_iterations=2, directed=False)

In Seurat: RunHarmony(obj, group.by.vars = 'orig.ident', reduction.use = 'pca') writes a harmony reduction; group.by.vars takes a vector to correct multiple covariates.

Integrate with scVI / scANVI

Goal: Learn a batch-invariant latent space from raw counts, optionally protecting known labels. Approach: Put raw counts in a layer, register batch (and labels for scANVI), train, and use the latent embedding downstream.

python
import scvi
import scanpy as sc

adata = sc.read_h5ad('merged.h5ad')
adata.layers['counts'] = adata.X.copy()  # scVI needs raw counts
sc.pp.highly_variable_genes(adata, n_top_genes=2000, flavor='seurat_v3',
                            layer='counts', batch_key='batch')
adata = adata[:, adata.var.highly_variable].copy()

scvi.model.SCVI.setup_anndata(adata, layer='counts', batch_key='batch')
model = scvi.model.SCVI(adata, n_latent=10, gene_likelihood='zinb')
model.train()  # default max_epochs heuristic scales down for large data
adata.obsm['X_scVI'] = model.get_latent_representation()

scanvi = scvi.model.SCANVI.from_scvi_model(model, 'Unknown', labels_key='cell_type')
scanvi.train(max_epochs=20)
adata.obs['scanvi_label'] = scanvi.predict()

The scVI latent space is not "biology with batch removed": it is a learned nonlinear embedding optimized to reconstruct counts while being marginally independent of batch. Its dimensions are entangled, individually uninterpretable, and carry no guaranteed correspondence to any biological quantity - treat it as a coordinate system for neighbors/clustering, not a measurement. Note unlabeled_category ('Unknown') is the second positional argument to from_scvi_model, before labels_key.

Integrate with Seurat v5

Goal: Use Seurat v5's modular layer-based integration with a chosen method. Approach: Split layers by batch, run the standard pipeline, call IntegrateLayers, rejoin.

r
library(Seurat)

merged[['RNA']] <- split(merged[['RNA']], f = merged$batch)
merged <- NormalizeData(merged)
merged <- FindVariableFeatures(merged)
merged <- ScaleData(merged)
merged <- RunPCA(merged)

merged <- IntegrateLayers(merged, method = RPCAIntegration,
                          orig.reduction = 'pca', new.reduction = 'integrated.rpca')
merged <- JoinLayers(merged)
merged <- FindNeighbors(merged, reduction = 'integrated.rpca', dims = 1:30)
merged <- FindClusters(merged, resolution = 0.5)
merged <- RunUMAP(merged, reduction = 'integrated.rpca', dims = 1:30)

Methods are passed as bare symbols: CCAIntegration, RPCAIntegration, HarmonyIntegration, FastMNNIntegration, scVIIntegration. For graph-only correction with BBKNN in Python: sce.pp.bbknn(adata, batch_key='batch') rewrites the neighbor graph in place (very fast, feeds Leiden/UMAP only).

Show full SKILL.md (664 more words)Show less

Evaluating Integration

Goal: Decide whether integration mixed batches without erasing biology. Approach: Score batch-mixing and bio-conservation separately and read them jointly - never optimize a batch metric alone.

python
import scanpy as sc
from sklearn.metrics import silhouette_score

# batch silhouette: lower = batches mixed; cell-type silhouette: higher = biology kept
batch_sil = silhouette_score(adata.obsm['X_scVI'], adata.obs['batch'])
ct_sil = silhouette_score(adata.obsm['X_scVI'], adata.obs['cell_type'])

# scib-metrics Benchmarker scores many methods on a common axis set
# from scib_metrics.benchmark import Benchmarker

Batch-mixing metrics (kBET, graph iLISI) are trivially maximized by over-correction - a method that destroys all structure mixes batches perfectly while annihilating biology. Bio-conservation metrics (ARI, NMI, cell-type ASW, graph cLISI, isolated-label F1) guard against that, which is why the scIB composite down-weights batch-removal to 0.4. Selecting a method on a batch metric alone selects for over-correction; always pair batch metrics with bio metrics and inspect rare populations before/after. Run candidates through scib-metrics (Benchmarker) and pick the most robust for the specific task.

Differential expression: use integration outputs (Harmony/scVI/RPCA embeddings) for clustering and visualization, but run DE on uncorrected, log-normalized counts - never on batch-corrected expression. Harmony and BBKNN produce no corrected counts; Scanorama and fastMNN do, and those must not feed DE. For cross-condition DE, aggregate to pseudobulk per sample x cell type (see differential-expression/deseq2-basics).

Reference Mapping vs De-novo Integration

De-novo integration jointly embeds all datasets symmetrically (everything above). Reference mapping projects a query onto a fixed reference embedding without retraining (scArches architectural surgery; Azimuth FindTransferAnchors + MapQuery) - fast, reproducible, scales to millions, consistent cross-study labels. It is closed-world: a novel state the reference never saw is confidently assigned the nearest reference label, converting a technical or biological surprise into a wrong annotation that looks clean and high-confidence. Use reference mapping when a high-quality annotated atlas exists; use de-novo when no suitable reference exists or the query may hold genuinely novel populations. Always inspect per-cell mapping uncertainty and never trust transferred labels for clusters that map poorly.

Common Errors

SymptomCauseFix
The cell type of interest vanished after integrationOver-correction absorbed a rare/condition-specific populationReduce strength (lower theta / use RPCA / fastMNN/Scanorama); compare to uncorrected embedding
Batches still separate on UMAPUnder-correctionRaise correction strength (k.anchor, switch CCA, more Harmony iterations); confirm batch key is correct
"Integration removed my treatment effect"Confounded batch/condition designStop - batch and biology are unidentifiable; redesign with multiplexing; do not integrate away the contrast
Great iLISI/kBET but biology looks flattenedMetric gaming by over-correctionScore bio-conservation too (ASW-celltype, cLISI, ARI); use the scIB composite, not a batch metric alone
DE between two control samples after integrationDE run on batch-corrected expressionRun DE on uncorrected log-normalized counts; use pseudobulk for cross-condition
scVI latent dimension interpreted as a biological axisLatent space is entangled, not "biology minus batch"Use the embedding only for neighbors/clustering; do not read individual dims
Reference-mapped labels look confident but wrongClosed-world projection of a novel/shifted stateInspect mapping uncertainty; treat poorly-mapping clusters as candidate novelty/batch
Results differ run to runStochastic training / unpinned seeds (scVI, Harmony)Set seeds; for scVI fix max_epochs and report it
  • preprocessing - QC and normalization that must precede integration
  • clustering - Cluster on the integrated embedding, not on raw PCA
  • cell-annotation - Reference mapping and label transfer after integration
  • single-cell/multimodal-integration - Joint analysis across modalities (distinct from batch integration)
  • single-cell/differential-abundance - Test whether composition shifts across conditions after integration
  • single-cell/cnv-inference - Per-patient malignant-cell/CNV inference (do not integrate tumors across patients)
  • differential-expression/deseq2-basics - Pseudobulk DE on uncorrected counts per cell type
  • data-visualization/dimensionality-reduction-plots - Before/after UMAP comparison figures

References

  • Korsunsky et al. (2019). Fast, sensitive and accurate integration of single-cell data with Harmony. Nat Methods 16(12):1289-1296.
  • Lopez et al. (2018). Deep generative modeling for single-cell transcriptomics (scVI). Nat Methods 15(12):1053-1058.
  • Xu et al. (2021). Probabilistic harmonization and annotation of single-cell transcriptomics data with deep generative models (scANVI). Mol Syst Biol 17(1):e9620.
  • Luecken et al. (2022). Benchmarking atlas-level data integration in single-cell genomics (scIB). Nat Methods 19:41-50.
  • Hie, Bryson & Berger (2019). Efficient integration of heterogeneous single-cell transcriptomes using Scanorama. Nat Biotechnol 37:685-691.
  • Polanski et al. (2020). BBKNN: fast batch alignment of single cell transcriptomes. Bioinformatics 36(3):964-965.
  • Haghverdi et al. (2018). Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors (MNN). Nat Biotechnol 36:421-427.
  • Lotfollahi et al. (2022). Mapping single-cell data to reference atlases by transfer learning (scArches). Nat Biotechnol 40(1):121-130.

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in single-cell/batch-integration of GPTomics/bioSkills.

  • SKILL.md
  • examples/harmony_integration.R
  • examples/harmony_integration.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Single Cell Batch Integration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Single Cell Batch Integration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Single Cell Batch Integration this skillGPTomics/bioSkills1.2k1 repos~4.2kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k3 repos~3.4kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
MFA Pipeline Orchestratoraiming-lab/AutoResearchClaw15k—~923Automated safety check: PassMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 3 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Single Cell Batch Integration

What does Bio Single Cell Batch Integration do?

Integrate multiple scRNA-seq samples or batches with Harmony, scVI/scANVI, Seurat (CCA/RPCA), fastMNN, Scanorama, or BBKNN. Bio Single Cell Batch Integration is an agent skill from GPTomics/bioSkills. Integrate multiple scRNA-seq samples or batches with Harmony, scVI/scANVI, Seurat (CCA/RPCA), fastMNN, Scanorama, or BBKNN.

When should I use Bio Single Cell Batch Integration?

Bio Single Cell Batch Integration fits situations like: the dataset size and design; how strongly to correct; integration is the wrong move (confounded batch/biology); how to score integration with scIB metrics without gaming them.

How do I install Bio Single Cell Batch Integration in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-batch-integration -a claude-code`. Or copy the skill folder (single-cell/batch-integration in GPTomics/bioSkills) into .claude/skills/bio-single-cell-batch-integration in your project. Claude Code loads it when a task matches its description.

How do I install Bio Single Cell Batch Integration in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-batch-integration -a codex`. Or copy the skill folder (single-cell/batch-integration in GPTomics/bioSkills) into .agents/skills/bio-single-cell-batch-integration in your project. Codex loads it when a task matches its description.

Can I use Bio Single Cell Batch Integration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-single-cell-batch-integration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-single-cell-batch-integration, .gemini/skills/bio-single-cell-batch-integration, .github/skills/bio-single-cell-batch-integration and .opencode/skills/bio-single-cell-batch-integration in your project.

What does Bio Single Cell Batch Integration need to run?

Going by SKILL.md and its folder, Bio Single Cell Batch Integration needs R and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Single Cell Batch Integration access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Single Cell Batch Integration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Single Cell Batch Integration use?

Bio Single Cell Batch Integration is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Single Cell Batch Integration use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Single Cell Batch Integration?

Skills that share tags, products or a category with Bio Single Cell Batch Integration: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars) and Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Single Cell Batch Integration?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.