Agent skill

Scvi Tools

by JimLiu in JimLiu/science-skills

Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression.

Apache-2.0Auto-check passedResearch & Science

Install Scvi Tools

skills CLI
$ npx skills add JimLiu/science-skills --skill scvi-tools -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JimLiu/science-skills scvi-tools --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scvi-tools .claude/skills/scvi-tools && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scvi-tools
GitHub stars
227
Used in
4 other repos
Token cost
~2.1k tokens
SKILL.md length
604 words
Files
2
Skills in repo
27
Repo updated
First seen
Licence
Apache-2.0

At a glance

Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression.

  • Tasks that involve Bioinformatics
  • SKILL.md covers How to run, Differential expression, Output format and Remote compute, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Scvi Tools is an agent skill from JimLiu/science-skills. Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression. Reach for this skill to integrate scRNA-seq batches, embed cells for clustering, transfer annotations from a reference onto a query, or score differentially expressed genes per cluster. For spatial deconvolution / mapping use the cell2location, DestVI, or Tangram methods instead.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `kernel.py`).

It sits in Research & Science, covering Bioinformatics. It works with scvi-tools and Scanpy. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/scvi-tools”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit fb309c3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scvi Tools loads about 2.1k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 604 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JimLiu/science-skills at commit fb309c3, republished under its Apache-2.0 licence (© JimLiu). 604 words, ~2,132 tokens.

Download SKILL.mdSave it as .claude/skills/scvi-tools/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
scvi-tools
description
Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression. Reach for this skill to integrate scRNA-seq batches, embed cells for clustering, transfer annotations from a reference onto a query, or score differentially expressed genes per cluster. For spatial deconvolution / mapping use the cell2location, DestVI, or Tangram methods instead.
license
Apache-2.0
requirements
gpu
metadata.display-name
scvi-tools

scvi-tools — scVI / scANVI

scvi-tools (Gayoso et al. 2022, github.com/scverse/scvi-tools, BSD-3-Clause) wraps a family of deep generative models for single-cell omics. The scRNA-seq core is scVI (unsupervised batch-corrected latent embedding) and scANVI (scVI + a classifier head for semi-supervised cell-type label transfer). Both expect raw integer UMI counts and emit a low-dimensional X_scVI / X_scANVI that drops into the scanpy neighbors → leiden → umap pipeline.

How to run

scVI — batch-corrected latent space
python
import scanpy as sc
import scvi

adata = sc.read_h5ad("dataset.h5ad")
adata.layers["counts"] = adata.X.copy()          # preserve raw BEFORE any normalize/log1p
sc.pp.normalize_total(adata); sc.pp.log1p(adata) # optional, for HVG / plotting only
sc.pp.highly_variable_genes(adata, n_top_genes=2000, batch_key="batch", subset=True)

scvi.model.SCVI.setup_anndata(adata, layer="counts", batch_key="batch")
model = scvi.model.SCVI(adata, n_latent=30)
model.train(max_epochs=200, early_stopping=True, accelerator="gpu", devices=1)

adata.obsm["X_scVI"] = model.get_latent_representation()
adata.layers["scvi_normalized"] = model.get_normalized_expression(library_size=1e4)
scANVI — label transfer from a partially-annotated reference
python
lvae = scvi.model.SCANVI.from_scvi_model(
    model, labels_key="cell_type", unlabeled_category="Unknown",
)
lvae.train(max_epochs=20, n_samples_per_label=100, accelerator="gpu", devices=1)

adata.obsm["X_scANVI"] = lvae.get_latent_representation()
adata.obs["pred_cell_type"] = lvae.predict()

accelerator="gpu", devices=1 is the PyTorch-Lightning spelling; the legacy use_gpu= kwarg was removed in scvi-tools 1.x and now raises TypeError.

Differential expression

python
de = model.differential_expression(
    groupby="leiden", group1="3",   # group2=None → vs. all other cells
    mode="change", delta=0.25,
)
top = de.sort_values("proba_de", ascending=False).head(50)

For one-vs-rest leave group2 out — "rest" is scanpy's rank_genes_groups convention, not scvi-tools'; here group2 is a literal category name and "rest" would match zero cells.

scvi-tools ≥1.4 defaults to mode="vanilla", whose result columns are exactly:

['proba_m1', 'proba_m2', 'bayes_factor', 'scale1', 'scale2', 'raw_mean1',
 'raw_mean2', 'non_zeros_proportion1', 'non_zeros_proportion2',
 'raw_normalized_mean1', 'raw_normalized_mean2', 'comparison', 'group1',
 'group2']

— no lfc_*, no proba_de, no is_de_fdr_*. Pass mode="change" to get lfc_mean / lfc_median / proba_de / is_de_fdr_0.05. Sort on proba_de (or on bayes_factor if you deliberately stayed in vanilla mode).

Output format

KeyWhat
adata.obsm["X_scVI"]n_cells × n_latent batch-corrected embedding
adata.obsm["X_scANVI"]label-aware embedding (better separates known classes)
adata.obs["pred_cell_type"]scANVI predicted label per cell
adata.layers["scvi_normalized"]decoded expression, library-size normalized
DE dataframeper-gene lfc_* / proba_de (with mode="change")

Remote compute

A100-class GPU recommended for >50k cells. The prebuilt singlecell_gpu Modal env ships scvi-tools 1.4.2 + scanpy 1.11.5 + anndata 0.11.4 — read compute_details({provider: 'byoc:modal', mode: 'read'}) for the current image ref, then:

python
c = host.compute.create('byoc:modal', provider_params={'modal': {
    'image':  '<image ref from compute_details>',   # e.g. im-...
    'gpu':    'A100',
    'cpu':    8,
    'memory': 32768,
    'timeout': 3600,
}})
job = c.submit_job(
    intent="scVI+scANVI on 80k cells — 1×A100, ~15 min",
    inputs=[
        {"src": "dataset.h5ad", "dst_filename": "dataset.h5ad"},
        {"src": "pipeline.py",  "dst_filename": "pipeline.py"},
    ],
    command="python pipeline.py",
    outputs=["out/**"],
    timeout_seconds=2400,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

h5ad_safe_obs is auto-loaded into the local analysis kernel only — in pipeline.py running on Modal, paste the helper at the top of the script (or inline the pd.Index(np.asarray(..., dtype=object)) coercion) before .write_h5ad().

Then call the wait_for_notification brain-tool. When compute_done arrives, save_artifacts(payload["featured_files"]). For the full result dict, re-enter the kernel and bind the compute handle (not the job) separately — .close() lives on the handle, not on the job:

python
h = host.compute.create('byoc:modal')
res = h.attach_job(job_id).result()   # output_files, remote_workdir, ...
h.close()

See the remote-compute-modal skill for orchestration details.

Gotchas

GotchaWhat happens / fix
differential_expression() defaults to mode="vanilla" (scvi-tools ≥1.4)KeyError: 'lfc_mean' / 'proba_de' when sorting — pass mode="change" to get lfc_*/proba_de/is_de_fdr_*; in vanilla mode sort on bayes_factor.
adata.obs index/columns are string[pyarrow] (ArrowStringArray).write_h5ad() dies with IORegistryError: No method registered for writing <class 'pandas.arrays.ArrowStringArray'> (anndata #2377). Coerce before writing: adata.obs = h5ad_safe_obs(adata.obs) (kernel helper — local kernel only; inline the coercion in remote pipeline.py). .astype(str) alone is not enough — on a pyarrow-backed Index/Series it returns another Arrow-backed array; round-trip through np.asarray(..., dtype=object). anndata.settings.allow_write_nullable_strings = True does not cover Arrow-backed strings.
use_gpu= kwargRemoved in 1.x → TypeError: train() got an unexpected keyword argument 'use_gpu'. Use accelerator="gpu", devices=1.
Log-normalized data fed to setup_anndataSilent garbage — scVI's NB likelihood needs raw integer counts. Stash counts in adata.layers["counts"] before normalize/log1p and pass layer="counts".
Show full SKILL.md (183 more words)Show less

Troubleshooting

SymptomFix
KeyError: 'lfc_mean' (or 'proba_de', 'is_de_fdr_0.05') on DE resultAdd mode="change" to differential_expression(); the default vanilla mode has no LFC columns.
IORegistryError: No method registered for writing <class 'pandas.arrays.ArrowStringArray'> on .write_h5ad()adata.obs = h5ad_safe_obs(adata.obs) (and adata.var if needed) before writing. The allow_write_nullable_strings flag does not help here.
TypeError: ... unexpected keyword argument 'use_gpu'Replace with accelerator="gpu", devices=1.
ValueError: ... non-negative integers / NB loss explodeslayer="counts" points at log/float data — restore raw counts.
MisconfigurationException: No supported gpu backend foundNo CUDA visible — drop accelerator/devices to fall back to CPU, or dispatch via Remote compute.
UnicodeEncodeError: 'ascii' codec can't encode character ... writing a summary / printingContainer has no LANG so Python defaults to ASCII. Open files with encoding="utf-8" and/or sys.stdout.reconfigure(encoding="utf-8") at script top. The prebuilt singlecell_gpu env sets PYTHONIOENCODING=utf-8, so this only bites user-built images.
AttributeError: ... object has no attribute 'close' on a job handleYou chained host.compute.create(...).attach_job(...) and called .close() on the job. Bind the compute handle separately and close that — see Remote compute above.

Next: cluster on X_scVI with scanpy (sc.pp.neighbors(use_rep="X_scVI") → sc.tl.leiden → sc.tl.umap); for spatial deconvolution train cell2location / DestVI / Tangram on the scRNA-seq reference.

© JimLiu, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/scvi-tools of JimLiu/science-skills.

  • SKILL.md
  • kernel.py

Open the folder on GitHubat commit fb309c3

Used in 4 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in JimLiu/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Scvi Tools next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scvi Tools compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scvi Tools this skillJimLiu/science-skills2274 repos~2.1kAutomated safety check: PassApache-2.0
ScanpyK-Dense-AI/scientific-agent-skills48k1 repos~5.1kAutomated safety check: PassBSD-3-Clause
Cellxgene CensusK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: NotesMIT
Scvi ToolsK-Dense-AI/scientific-agent-skills48k1 repos~2.6kAutomated safety check: PassBSD-3-Clause
AnndataK-Dense-AI/scientific-agent-skills48k1 repos~3.9kAutomated safety check: NotesBSD-3-Clause
Anndata Data Structurejaechang-hits/SciAgent-Skills3712 repos~5.8kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Scanpy

    K-Dense-AI/scientific-agent-skills

    Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…

    48k GitHub starsUsed in 1 repo~5.1k tokens
    Research & ScienceAuto-check passed
  • Cellxgene Census

    K-Dense-AI/scientific-agent-skills

    Queries the CZ CELLxGENE Census programmatically for versioned public single-cell and spatial transcriptomics data.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Research & ScienceAuto-check: notes
  • Scvi Tools

    K-Dense-AI/scientific-agent-skills

    Fits probabilistic models for single-cell omics, including scVI batch integration, scANVI annotation, totalVI CITE-seq, MultiVI RNA/ATAC integration, and posterior differential expression.

    48k GitHub starsUsed in 1 repo~2.6k tokens
    Research & ScienceAuto-check passed
  • Anndata

    K-Dense-AI/scientific-agent-skills

    Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.

    48k GitHub starsUsed in 1 repo~3.9k tokens
    Research & ScienceAuto-check: notes
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 2 repos~5.8k tokens
    Research & ScienceAuto-check passed
  • Scvelo

    lamm-mit/scienceclaw

    RNA velocity analysis with scVelo. An agent skill from lamm-mit/scienceclaw.

    244 GitHub starsUsed in 4 repos~524 tokens
    Research & ScienceAuto-check passed

More from JimLiu/science-skills

All 27 skills in this repo
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    Auto-check passed
  • Compute Env Setup

    JimLiu/science-skills

    Set up a compute environment on a remote provider so Claude Science jobs can run there.

    227 GitHub starsUsed in 2 repos~4.4k tokens
    Auto-check passed
  • Borzoi

    JimLiu/science-skills

    Predict genome-wide functional tracks (RNA-seq, CAGE, DNase, ChIP) from DNA sequence with Borzoi.

    227 GitHub starsUsed in 4 repos~973 tokens
    Auto-check passed
  • Evo2

    JimLiu/science-skills

    Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check passed
  • Fair Esm2

    JimLiu/science-skills

    Embed proteins with Meta AI's ESM-2 (fair-esm package). An agent skill from JimLiu/science-skills.

    227 GitHub starsUsed in 4 repos~1.2k tokens
    Auto-check passed
  • Openfold3

    JimLiu/science-skills

    Structure prediction using OpenFold3, an open-weights PyTorch reproduction of AlphaFold3 from the AlQuraishi Lab.

    227 GitHub starsUsed in 4 repos~1.8k tokens
    Auto-check passed

Questions about Scvi Tools

What does Scvi Tools do?

Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression. Scvi Tools is an agent skill from JimLiu/science-skills. Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression.

When should I use Scvi Tools?

Scvi Tools fits situations like: tasks that involve Bioinformatics.

How do I install Scvi Tools in Claude Code?

Run `npx skills add JimLiu/science-skills --skill scvi-tools -a claude-code`. Or copy the skill folder (skills/scvi-tools in JimLiu/science-skills) into .claude/skills/scvi-tools in your project. Claude Code loads it when a task matches its description.

How do I install Scvi Tools in Codex?

Run `npx skills add JimLiu/science-skills --skill scvi-tools -a codex`. Or copy the skill folder (skills/scvi-tools in JimLiu/science-skills) into .agents/skills/scvi-tools in your project. Codex loads it when a task matches its description.

Can I use Scvi Tools in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JimLiu/science-skills --skill scvi-tools -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scvi-tools, .gemini/skills/scvi-tools, .github/skills/scvi-tools and .opencode/skills/scvi-tools in your project.

What does Scvi Tools need to run?

Going by SKILL.md and its folder, Scvi Tools needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Scvi Tools access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scvi Tools safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scvi Tools use?

Scvi Tools is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scvi Tools use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scvi Tools?

Skills that share tags, products or a category with Scvi Tools: Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars), Cellxgene Census (K-Dense-AI/scientific-agent-skills, 48k stars), Scvi Tools (K-Dense-AI/scientific-agent-skills, 48k stars) and Anndata (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scvi Tools?

JimLiu (a GitHub user) maintains it in JimLiu/science-skills, which has 227 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on July 1, 2026.

Source: JimLiu/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.