Scgpt
JimLiu/science-skills
Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.
Deep generative models for single-cell omics: probabilistic batch correction (scVI), semi-supervised annotation (scANVI), CITE-seq RNA+protein (totalVI), transfer learning (scARCHES), and DE with…
$ npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scvi-tools-single-cell --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell .claude/skills/scvi-tools-single-cell && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scvi-tools-single-cell" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell into .claude/skills/scvi-tools-single-cell/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scvi-tools-single-cell", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cellType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scvi-tools-single-cell --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell .agents/skills/scvi-tools-single-cell && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scvi-tools-single-cell" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell into .agents/skills/scvi-tools-single-cell/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scvi-tools-single-cell", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scvi-tools-single-cell --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell .cursor/skills/scvi-tools-single-cell && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scvi-tools-single-cell" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell into .cursor/skills/scvi-tools-single-cell/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scvi-tools-single-cell", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scvi-tools-single-cell --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell .gemini/skills/scvi-tools-single-cell && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scvi-tools-single-cell" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell into .gemini/skills/scvi-tools-single-cell/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scvi-tools-single-cell", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills scvi-tools-single-cellInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell .github/skills/scvi-tools-single-cell && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scvi-tools-single-cell" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell into .github/skills/scvi-tools-single-cell/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scvi-tools-single-cell", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scvi-tools-single-cell --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell .opencode/skills/scvi-tools-single-cell && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scvi-tools-single-cell" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell into .opencode/skills/scvi-tools-single-cell/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scvi-tools-single-cell", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scvi-tools-single-cellDeep generative models for single-cell omics: probabilistic batch correction (scVI), semi-supervised annotation (scANVI), CITE-seq RNA+protein (totalVI), transfer learning (scARCHES), and DE with…
Scvi Tools Single Cell is an agent skill from jaechang-hits/SciAgent-Skills. Deep generative models for single-cell omics: probabilistic batch correction (scVI), semi-supervised annotation (scANVI), CITE-seq RNA+protein (totalVI), transfer learning (scARCHES), and DE with uncertainty. Unified setup→train→extract API on AnnData. Use harmony-batch-correction for fast linear correction without deep learning; muon for multi-modal MuData workflows.
Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with scvi-tools and AnnData. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
doi.orgdocs.scvi-tools.orggithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scvi Tools Single Cell loads about 7.2k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 1,424 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,424 words, ~7,174 tokens.
.claude/skills/scvi-tools-single-cell/SKILL.md (or your agent's skills folder).scvi-tools is a probabilistic modeling framework for single-cell genomics built on PyTorch. It implements variational autoencoders (VAEs) that learn low-dimensional latent representations of cells while explicitly modeling batch effects, count noise distributions, and multi-modal data. All models share a unified API: setup_anndata() to register data, instantiate the model, train(), then extract latent representations, normalized expression, or differential expression results. Models operate on raw count data in AnnData format and return statistically grounded outputs with uncertainty estimates.
scvi-tools>=1.1, scanpy, anndata.h5ad) with raw counts — not log-normalized. Store counts in adata.layers["counts"] if adata.X has been normalized.pip install scvi-tools scanpy
# GPU acceleration (recommended for >50k cells)
pip install "scvi-tools[cuda12]" # or scvi-tools[cuda11]Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: modelChoice
kind: required
source: user
ask: "What should the model do - integrate and denoise unlabelled cells, transfer labels from a reference, or jointly model protein and RNA?"
default: null
- id: D2
param: batchKey
kind: required
source: data
ask: "Which variable separates samples that were processed apart, so the model can hold it out of the latent space?"
default: null
- id: D3
param: rawCountsLayer
kind: required
source: data
ask: "Which layer holds raw integer counts? These models require counts, not normalized values."
default: null
- id: D4
param: latentDimensions
kind: required
source: user
ask: "How many latent dimensions should the cell state be compressed into?"
default: 10
- id: D5
param: geneLikelihood
kind: optional
source: data
ask: "Which count distribution fits this assay - zero-inflated, plain negative binomial, or Poisson?"
default: "zero-inflated negative binomial"
- id: D6
param: trainingSchedule
kind: optional_conditional
source: data
ask: "How long should training run, and should it stop early when validation loss plateaus?"
default: "400 epochs, no early stopping"
- id: D7
param: deEffectSizeThreshold
kind: optional
source: user
ask: "What fold change should the differential expression test treat as the null it must beat?"
default: 0.25
- id: D8
param: networkArchitecture
kind: never_ask
source: user
reason: "Layer count and hidden width are capacity settings whose defaults are validated across datasets; changing them is a methods decision the user has not raised"
default: "1 layer, 128 hidden units"D3 is the failure that looks like success: handing log-normalized values to a count model trains it on the wrong noise structure, and it still returns a latent space, a UMAP and a cluster set. D2 decides what the latent space is allowed to forget, so a batch variable confounded with the condition removes the effect being studied.
Minimal scVI batch integration on a built-in example dataset:
import scvi
import scanpy as sc
# Load example data with batch labels
adata = scvi.data.heart_cell_atlas_subsampled()
# Preprocessing: filter genes, select HVGs (subset to save training time)
sc.pp.filter_genes(adata, min_counts=3)
adata.layers["counts"] = adata.X.copy() # preserve raw counts
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
sc.pp.highly_variable_genes(adata, n_top_genes=2000, batch_key="cell_source", subset=True)
# Register data, train, extract
scvi.model.SCVI.setup_anndata(adata, layer="counts", batch_key="cell_source")
model = scvi.model.SCVI(adata, n_latent=30)
model.train(max_epochs=200, early_stopping=True)
adata.obsm["X_scVI"] = model.get_latent_representation()
sc.pp.neighbors(adata, use_rep="X_scVI")
sc.tl.umap(adata)
sc.tl.leiden(adata, resolution=0.5)
sc.pl.umap(adata, color=["cell_source", "leiden"])
print(f"Latent shape: {adata.obsm['X_scVI'].shape}")
# Latent shape: (14000, 30)All models share the same registration pattern. Call setup_anndata() on the model class before instantiation to tell scvi-tools where to find counts, batch labels, and covariates.
import scvi
# Minimal: counts in a layer, batch column in obs
scvi.model.SCVI.setup_anndata(
adata,
layer="counts", # Key in adata.layers with raw counts; None = adata.X
batch_key="batch", # Column in adata.obs for technical batch
)
# Extended: additional categorical and continuous covariates
scvi.model.SCVI.setup_anndata(
adata,
layer="counts",
batch_key="batch",
categorical_covariate_keys=["donor", "protocol"], # Discrete biological/technical vars
continuous_covariate_keys=["percent_mito", "log_n_counts"], # Continuous covariates
)
# Inspect registered summary
print(adata.uns["_scvi"]["summary_stats"])
# {'n_vars': 2000, 'n_cells': 45000, 'n_batch': 6, 'n_extra_categorical_covs': 2, ...}The core unsupervised model. Learns a batch-corrected latent space and a denoised expression layer. Starting point for any multi-batch scRNA-seq analysis.
import scvi
model = scvi.model.SCVI(
adata,
n_latent=30, # Latent space dimensions; 10–50 typical
n_layers=2, # Hidden layers in encoder/decoder; 1–3
n_hidden=128, # Neurons per hidden layer; 64–256
gene_likelihood="zinb", # "zinb" (zero-inflated NB), "nb", or "poisson"
dispersion="gene", # "gene" or "gene-batch" for batch-specific dispersion
)
model.train(max_epochs=200, early_stopping=True)
# Extract results
latent = model.get_latent_representation() # ndarray (n_cells, n_latent)
normalized = model.get_normalized_expression(
library_size=1e4, n_samples=25, return_mean=True
)
print(f"Latent: {latent.shape}")
print(f"Denoised expression: {normalized.shape}")
# Latent: (45000, 30)
# Denoised expression: (45000, 2000)
adata.obsm["X_scVI"] = latent
adata.layers["scvi_normalized"] = normalizedExtends scVI with label supervision for cell type transfer learning. Accepts partially labeled data (unannotated cells labeled "Unknown") and predicts cell types for unlabeled cells.
import scvi
# Register with cell type labels; mark unannotated cells as "Unknown"
scvi.model.SCANVI.setup_anndata(
adata,
layer="counts",
batch_key="batch",
labels_key="cell_type", # Column in adata.obs with known labels
unlabeled_category="Unknown", # Sentinel value for unannotated cells
)
# Recommended: initialize from a pretrained scVI model
scvi_model = scvi.model.SCVI(adata, n_latent=30)
scvi_model.train(max_epochs=200, early_stopping=True)
model = scvi.model.SCANVI.from_scvi_model(scvi_model, unlabeled_category="Unknown")
model.train(max_epochs=20) # Short fine-tuning on top of scVI
# Predict cell types
labels_pred = model.predict() # Hard labels (str per cell)
probs = model.predict(soft=True) # DataFrame: probability per cell type
confidence = probs.max(axis=1)
adata.obs["scANVI_pred"] = labels_pred
adata.obs["scANVI_confidence"] = confidence
print(f"Median confidence: {confidence.median():.3f}")
print(f"High-confidence cells (>0.7): {(confidence > 0.7).sum()}")Joint probabilistic model for CITE-seq data. Denoises both RNA and protein counts and estimates protein foreground probability (signal vs background).
import scvi
# Protein counts must be in adata.obsm (not adata.X or layers)
# adata.obsm["protein_expression"] shape: (n_cells, n_proteins)
scvi.model.TOTALVI.setup_anndata(
adata,
layer="counts",
batch_key="batch",
protein_expression_obsm_key="protein_expression",
)
model = scvi.model.TOTALVI(adata, latent_distribution="normal")
model.train(max_epochs=200, early_stopping=True)
# Extract joint latent space (RNA + protein)
latent = model.get_latent_representation()
# Denoised RNA and protein separately
rna_norm, protein_norm = model.get_normalized_expression(
n_samples=25, return_mean=True
)
# Protein foreground probability (probability that signal > background)
foreground_prob = model.get_protein_foreground_probability(n_samples=25, return_mean=True)
print(f"RNA normalized: {rna_norm.shape}")
print(f"Protein normalized: {protein_norm.shape}")
print(f"Foreground prob: {foreground_prob.shape}") # (n_cells, n_proteins)
adata.obsm["X_totalVI"] = latentProbabilistic DE using the generative model rather than raw counts. Supports composite hypotheses testing whether effect size exceeds a biologically meaningful threshold (delta).
# DE between two cell type groups
de_df = model.differential_expression(
groupby="cell_type",
group1="CD4 T", # Numerator group
group2="CD8 T", # Denominator group; None = rest of cells
mode="change", # "change": composite |LFC| > delta; "vanilla": any LFC
delta=0.25, # Minimum |LFC| to be considered DE
fdr_target=0.05, # FDR level for calling significance
all_stats=True, # Include mean expression per group
)
# Filter significant DE genes
sig = de_df[de_df["is_de_fdr_0.05"] & (de_df["lfc_mean"].abs() > 0.5)]
sig_sorted = sig.sort_values("lfc_mean", ascending=False)
print(f"Total DE genes (FDR<5%, |LFC|>0.5): {len(sig)}")
print(sig_sorted[["lfc_mean", "bayes_factor", "proba_de"]].head(10))# One-vs-rest DE across all clusters (generates a dict of DataFrames)
de_all = {}
for ct in adata.obs["cell_type"].unique():
de_all[ct] = model.differential_expression(
idx1=[adata.obs["cell_type"] == ct], # Boolean index
mode="change",
delta=0.25,
)
sig_n = de_all[ct]["is_de_fdr_0.05"].sum()
print(f" {ct}: {sig_n} DE genes vs rest")Adapts a pretrained reference model to a new query dataset without re-training from scratch. Preserves the reference embedding structure and maps query cells into the same latent space.
import scvi
# --- Reference training (done once, save the model) ---
scvi.model.SCVI.setup_anndata(ref_adata, layer="counts", batch_key="batch")
ref_model = scvi.model.SCVI(ref_adata, n_latent=30)
ref_model.train(max_epochs=200, early_stopping=True)
ref_model.save("./reference_model/", overwrite=True)
# Reference annotation (use scANVI for label transfer)
scvi.model.SCANVI.setup_anndata(
ref_adata, layer="counts", batch_key="batch",
labels_key="cell_type", unlabeled_category="Unknown",
)
ref_scanvi = scvi.model.SCANVI.from_scvi_model(ref_model, unlabeled_category="Unknown")
ref_scanvi.train(max_epochs=20)
ref_scanvi.save("./reference_scanvi/", overwrite=True)
# --- Query mapping (new dataset, minimal training) ---
ref_scanvi = scvi.model.SCANVI.load("./reference_scanvi/", adata=ref_adata)
# Prepare query: must share gene set with reference
query_adata.obs["cell_type"] = "Unknown" # All query cells are unlabeled
query_model = scvi.model.SCANVI.load_query_data(query_adata, ref_scanvi)
query_model.train(
max_epochs=100,
plan_kwargs={"weight_decay": 0.0}, # Critical: prevents catastrophic forgetting
)
# Map query into reference latent space
query_latent = query_model.get_latent_representation(query_adata)
query_labels = query_model.predict(query_adata)
print(f"Query latent: {query_latent.shape}")
print(f"Query label predictions: {query_labels[:5]}")After extracting the latent representation, use scanpy for UMAP, clustering, and marker genes — with the batch-corrected embedding as input.
import scanpy as sc
# UMAP and clustering on scVI latent representation
adata.obsm["X_scVI"] = model.get_latent_representation()
sc.pp.neighbors(
adata,
use_rep="X_scVI", # Use scVI latent (not PCA)
n_neighbors=30,
n_pcs=None, # n_pcs ignored when use_rep is set
)
sc.tl.umap(adata)
sc.tl.leiden(adata, resolution=0.5)
# Marker genes using raw counts or normalized expression
# Option A: scanpy Wilcoxon on log-normalized expression
sc.tl.rank_genes_groups(adata, groupby="leiden", method="wilcoxon", n_genes=25)
# Option B: scVI probabilistic DE per cluster (more accurate)
markers = {}
for cluster in adata.obs["leiden"].unique():
de = model.differential_expression(
idx1=[adata.obs["leiden"] == cluster],
mode="change", delta=0.25,
)
markers[cluster] = de[de["is_de_fdr_0.05"]].sort_values("lfc_mean", ascending=False)
# Visualization
sc.pl.umap(adata, color=["leiden", "batch", "cell_type"], ncols=3)
sc.pl.rank_genes_groups_dotplot(adata, n_genes=5, groupby="leiden")All scvi-tools models follow the same 4-step workflow:
1. ModelClass.setup_anndata(adata, ...) → Register layers, batch keys, covariates
2. model = ModelClass(adata, ...) → Set architecture hyperparameters
3. model.train(max_epochs=..., ...) → Fit model (GPU auto-detected)
4. model.get_*() → Extract results: latent, normalized, DE| Data | Model | Core Feature | When to Use |
|---|---|---|---|
| scRNA-seq | scVI | Batch correction, denoising | Default for any multi-batch scRNA-seq |
| scRNA-seq (partial labels) | scANVI | Cell type transfer | Have reference labels; want to annotate query |
| CITE-seq (RNA+protein) | totalVI | Joint RNA+protein | 10x CITE-seq, REAP-seq |
| Reference → query | scARCHES | Transfer learning | Map new data to existing atlas |
| Spatial | DestVI | Spot deconvolution | 10x Visium, Slide-seq |
| scRNA-seq QC | Solo | Doublet detection | Pre-analysis QC step |
scvi-tools models learn count distributions directly from raw data. Log-normalized input produces incorrect results. Always preserve raw counts before normalization:
# Correct workflow: save counts, then normalize for scanpy steps
adata.layers["counts"] = adata.X.copy() # Save raw before any transformation
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
# Then: scvi.model.SCVI.setup_anndata(adata, layer="counts", ...)Goal: Integrate multiple scRNA-seq datasets, remove batch effects, identify cell clusters, and find marker genes.
import scvi
import scanpy as sc
# Load and concatenate datasets
adata1 = sc.read_h5ad("dataset1.h5ad")
adata2 = sc.read_h5ad("dataset2.h5ad")
adata3 = sc.read_h5ad("dataset3.h5ad")
adata = sc.concat(
[adata1, adata2, adata3],
label="batch",
keys=["dataset1", "dataset2", "dataset3"],
)
# Preprocessing: preserve raw counts, select HVGs per batch
adata.layers["counts"] = adata.X.copy()
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
sc.pp.highly_variable_genes(
adata,
n_top_genes=4000,
batch_key="batch", # Select HVGs consistently across batches
subset=True,
)
print(f"HVGs selected: {adata.n_vars}")
# scVI integration
scvi.model.SCVI.setup_anndata(adata, layer="counts", batch_key="batch")
model = scvi.model.SCVI(adata, n_latent=30, n_layers=2)
model.train(max_epochs=300, early_stopping=True, early_stopping_patience=15)
print(f"Training history: {model.history['elbo_train'].tail()}")
# Batch-corrected analysis
adata.obsm["X_scVI"] = model.get_latent_representation()
adata.layers["scvi_norm"] = model.get_normalized_expression(n_samples=25, return_mean=True)
sc.pp.neighbors(adata, use_rep="X_scVI", n_neighbors=30)
sc.tl.umap(adata)
sc.tl.leiden(adata, resolution=0.5)
# Verify batch mixing (batches should overlap on UMAP)
sc.pl.umap(adata, color=["batch", "leiden"], save="_batch_integration.png")
model.save("./scvi_model/", overwrite=True)
print(f"Clusters: {adata.obs['leiden'].nunique()}")Goal: Model CITE-seq RNA + protein data jointly, obtain denoised protein estimates, and generate a joint embedding for annotation.
import scvi
import scanpy as sc
import pandas as pd
# Load CITE-seq AnnData (RNA in X, protein counts in obsm)
adata = sc.read_h5ad("citeseq_data.h5ad")
# Expected structure:
# adata.X or adata.layers["counts"] : RNA raw counts (n_cells, n_genes)
# adata.obsm["protein_expression"] : Protein raw counts (n_cells, n_proteins)
# Preprocessing
adata.layers["counts"] = adata.X.copy()
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
sc.pp.highly_variable_genes(adata, n_top_genes=4000, batch_key="batch", subset=True)
# totalVI setup and training
scvi.model.TOTALVI.setup_anndata(
adata,
layer="counts",
batch_key="batch",
protein_expression_obsm_key="protein_expression",
)
model = scvi.model.TOTALVI(adata, latent_distribution="normal")
model.train(max_epochs=200, early_stopping=True)
# Extract joint embedding and denoised values
adata.obsm["X_totalVI"] = model.get_latent_representation()
rna_norm, protein_norm = model.get_normalized_expression(n_samples=25, return_mean=True)
foreground = model.get_protein_foreground_probability(n_samples=25, return_mean=True)
adata.layers["rna_denoised"] = rna_norm
protein_df = pd.DataFrame(
protein_norm,
index=adata.obs_names,
columns=adata.uns["protein_names"],
)
protein_df.to_csv("denoised_protein_expression.csv")
print(f"Protein foreground: min={foreground.min():.3f}, max={foreground.max():.3f}")
# Joint clustering on RNA + protein latent
sc.pp.neighbors(adata, use_rep="X_totalVI", n_neighbors=30)
sc.tl.umap(adata)
sc.tl.leiden(adata, resolution=0.8)
# Protein-guided annotation (use denoised protein for clean signal)
protein_markers = {"B cell": "CD19", "T cell": "CD3E", "Monocyte": "CD14"}
for cell_type, marker in protein_markers.items():
if marker in protein_df.columns:
adata.obs[f"denoised_{marker}"] = protein_df[marker].values
sc.pl.umap(adata, color=["leiden"] + [f"denoised_{m}" for m in protein_markers.values()])| Parameter | Model / Function | Default | Range / Options | Effect |
|---|---|---|---|---|
n_latent | All models | 10 | 10–50 | Latent space dimensionality; 20–30 typical for most datasets |
n_layers | All models | 1 | 1–3 | Depth of encoder/decoder networks |
n_hidden | All models | 128 | 64–256 | Width of each hidden layer |
gene_likelihood | scVI, scANVI | "zinb" | "zinb", "nb", "poisson" | Count noise distribution; "nb" faster if sparsity is low |
max_epochs | train() | 400 | 50–1000 | Training iterations (ELBO steps) |
early_stopping | train() | False | True/False | Halt when validation ELBO plateaus |
early_stopping_patience | train() | 45 | 5–100 | Epochs to wait before stopping |
batch_size | train() | 128 | 64–512 | Mini-batch size; reduce if OOM |
lr | train() | 1e-3 | 1e-4–1e-2 | Initial learning rate |
delta | differential_expression() | 0.25 | 0.1–1.0 | Min |
n_samples | get_normalized_expression() | 1 | 1–100 | Posterior samples; ≥25 for stable estimates |
Always store raw counts before normalization: Run adata.layers["counts"] = adata.X.copy() immediately after loading data, before any normalize_total or log1p call. Passing log-normalized data silently produces incorrect latent spaces.
Select highly variable genes before training: Use sc.pp.highly_variable_genes(n_top_genes=2000–4000, batch_key="batch") to reduce training time and model noise. Including all genes rarely improves results and significantly slows training.
Register all known batch variables: If data has multiple technical sources (sequencing plate, donor, protocol), include them as batch_key or categorical_covariate_keys. Unregistered batch effects contaminate the latent space and downstream DE.
Use the scVI → scANVI two-step pipeline: Initializing scANVI with from_scvi_model() after training scVI is faster, more stable, and produces better embeddings than training scANVI from scratch. Always train scVI first when using scANVI.
Save trained models immediately: model.save("./model_dir/") persists the full model. Reloading with SCVI.load() is seconds vs. minutes of retraining. Models are typically 10–50 MB.
Use n_latent=20–30 as starting point: Values below 10 lose resolution; above 50 tend to overfit without added biological insight. Increase to 50 only for very heterogeneous atlases (>500k cells, many cell types).
Filter low-confidence scANVI predictions: Cells where predict(soft=True).max(axis=1) < 0.7 are ambiguous; treat them as "Unknown" rather than accepting the argmax label.
When to use: Remove doublets before integration or DE to avoid spurious clusters.
import scvi
scvi.model.SCVI.setup_anndata(adata, layer="counts")
vae = scvi.model.SCVI(adata, n_latent=20)
vae.train(max_epochs=100)
solo = scvi.external.SOLO.from_scvi_model(vae)
solo.train(max_epochs=200)
doublet_preds = solo.predict() # DataFrame: {"singlet": prob, "doublet": prob}
adata.obs["doublet_score"] = doublet_preds["doublet"].values
adata.obs["is_doublet"] = doublet_preds["prediction"].values
n_doublets = (adata.obs["is_doublet"] == "doublet").sum()
print(f"Detected {n_doublets} doublets ({n_doublets / adata.n_obs:.1%})")
adata = adata[adata.obs["is_doublet"] == "singlet"].copy()
print(f"Cells after doublet removal: {adata.n_obs}")When to use: Estimate cell type proportions in spatial transcriptomics spots using a matched scRNA-seq reference.
import scvi
# Step 1: Train CondSCVI on reference single-cell data
scvi.model.CondSCVI.setup_anndata(sc_adata, layer="counts", labels_key="cell_type")
sc_model = scvi.model.CondSCVI(sc_adata, weight_obs=False)
sc_model.train(max_epochs=200)
# Step 2: Deconvolve spatial spots
scvi.model.DestVI.setup_anndata(st_adata, layer="counts")
st_model = scvi.model.DestVI.from_rna_model(st_adata, sc_model)
st_model.train(max_epochs=2500)
proportions = st_model.get_proportions() # DataFrame (n_spots, n_cell_types)
st_adata.obsm["cell_type_proportions"] = proportions.values
print(f"Proportions per spot: {proportions.shape}")
print(proportions.head())When to use: Obtain smooth, denoised expression values for visualization or downstream machine learning when raw sparse counts are too noisy.
import scvi
scvi.model.SCVI.setup_anndata(adata, layer="counts", batch_key="batch")
model = scvi.model.SCVI(adata, n_latent=30)
model.train(max_epochs=200, early_stopping=True)
# n_samples >= 25 gives stable posterior mean; increase to 100 for publication
denoised = model.get_normalized_expression(
n_samples=50,
library_size="latent", # "latent" uses learned library size; int for fixed normalization
return_mean=True,
)
adata.layers["denoised"] = denoised
print(f"Denoised layer added: {denoised.shape}")
# Use adata.layers["denoised"] for heatmaps, pseudotime, or ML features| Problem | Cause | Solution |
|---|---|---|
ValueError: adata must contain raw counts | Log-normalized data passed instead of raw counts | Save raw before normalizing: adata.layers["counts"] = adata.X.copy() then setup_anndata(layer="counts") |
| Training loss oscillates without decreasing | Learning rate too high or very heterogeneous data | Try lr=1e-4; check that adata.X and the registered layer are not sparse with negatives |
| CUDA out of memory | Dataset too large for GPU VRAM | Reduce batch_size=64, n_hidden=64; subset to fewer HVGs; use CPU for small datasets |
| Poor batch integration (batches separate on UMAP) | Batch key not registered, or too few epochs | Verify batch_key column exists in adata.obs; increase max_epochs; add early_stopping=True |
| scANVI predicts same label for all cells | Too few labeled cells per type or learning rate issue | Need ≥50 labeled cells per type; use from_scvi_model() workflow; check unlabeled_category matches exactly |
load() fails with KeyError or shape mismatch | AnnData var_names differ from training data | Ensure query adata.var_names exactly matches training data; do not subset genes after training |
| Slow training on CPU for large datasets | Large dataset without GPU acceleration | Install scvi-tools[cuda12]; pass accelerator="gpu" to train(); or subsample to 50k cells for prototyping |
scvi.model.SCANVI.load_query_data shape error | Query and reference have different gene sets | Align genes: query_adata = query_adata[:, ref_adata.var_names] before calling load_query_data |
sc.pp.neighbors(use_rep="X_scVI")© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Scvi Tools Single Cell next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scvi Tools Single Cell this skilljaechang-hits/SciAgent-Skills | 374 | 2 repos | ~7.2k | Automated safety check: Pass | BSD-3-Clause | |
| ScgptJimLiu/science-skills | 228 | 4 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| ScanpyK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~5.1k | Automated safety check: Pass | BSD-3-Clause | |
| Cellxgene CensusK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.4k | Automated safety check: Notes | MIT | |
| AnndataK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.9k | Automated safety check: Notes | BSD-3-Clause | |
| Scanpyaipoch/medical-research-skills | 1.9k | — | ~3.9k | Automated safety check: Pass | MIT |
JimLiu/science-skills
Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.
K-Dense-AI/scientific-agent-skills
Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…
K-Dense-AI/scientific-agent-skills
Queries the CZ CELLxGENE Census programmatically for versioned public single-cell and spatial transcriptomics data.
K-Dense-AI/scientific-agent-skills
Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.
aipoch/medical-research-skills
Standard single-cell RNA-seq analysis pipeline. An agent skill from aipoch/medical-research-skills.
aipoch/medical-research-skills
Deep generative models for single-cell omics; use when you need probabilistic batch correction (scVI), transfer learning, uncertainty-aware differential expression, or multimodal integration…
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Deep generative models for single-cell omics: probabilistic batch correction (scVI), semi-supervised annotation (scANVI), CITE-seq RNA+protein (totalVI), transfer learning (scARCHES), and DE with…. Scvi Tools Single Cell is an agent skill from jaechang-hits/SciAgent-Skills. Deep generative models for single-cell omics: probabilistic batch correction (scVI), semi-supervised annotation (scANVI), CITE-seq RNA+protein (totalVI), transfer learning (scARCHES), and DE with uncertainty.
Scvi Tools Single Cell fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell in jaechang-hits/SciAgent-Skills) into .claude/skills/scvi-tools-single-cell in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/scvi-tools-single-cell in jaechang-hits/SciAgent-Skills) into .agents/skills/scvi-tools-single-cell in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill scvi-tools-single-cell -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scvi-tools-single-cell, .gemini/skills/scvi-tools-single-cell, .github/skills/scvi-tools-single-cell and .opencode/skills/scvi-tools-single-cell in your project.
Going by SKILL.md and its folder, Scvi Tools Single Cell needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 3 domains. As links in the text: doi.org, docs.scvi-tools.org and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Scvi Tools Single Cell is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Scvi Tools Single Cell: Scgpt (JimLiu/science-skills, 228 stars), Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars), Cellxgene Census (K-Dense-AI/scientific-agent-skills, 48k stars) and Anndata (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.