Agent skill

Scrna Embedding

by ClawBio in ClawBio/ClawBio

Local scVI/scANVI-based single-cell latent embedding and batch-aware integration from raw-count .h5ad or 10x Matrix Market input, with stable integrated AnnData export for downstream latent analysis.

MITAuto-check passedAI & LLM Engineering

Install Scrna Embedding

skills CLI
$ npx skills add ClawBio/ClawBio --skill scrna-embedding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio scrna-embedding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scrna-embedding .claude/skills/scrna-embedding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scrna-embedding
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
633 words
Files
3
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Local scVI/scANVI-based single-cell latent embedding and batch-aware integration from raw-count .h5ad or 10x Matrix Market input, with stable integrated AnnData export for downstream latent analysis.

  • Works in 6 steps: Raw-count Input Validation: Accept… → scVI/scANVI Latent Embedding: Train… → Latent Output Generation: Run neighbors… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Why This Exists, Core Capabilities, Input Formats and Workflow, plus 9 more sections
  • Runs Python scripts from its folder; calls python

What it does

Scrna Embedding is an agent skill from ClawBio/ClawBio. Local scVI/scANVI-based single-cell latent embedding and batch-aware integration from raw-count .h5ad or 10x Matrix Market input, with stable integrated AnnData export for downstream latent analysis.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `scrna_embedding.py` and `tests/test_scrna_embedding.py`).

It sits in AI & LLM Engineering, covering Bioinformatics and Embeddings. It works with AnnData. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve Embeddings

Example prompts

  • “/scrna-embedding”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Raw-count Input Validation: Accept raw-count .h5ad and 10x Matrix Market input; reject processed-like matrices.
  2. scVI/scANVI Latent Embedding: Train scvi.model.SCVI or refine with scvi.model.SCANVI using explicit labels.
  3. Latent Output Generation: Run neighbors and UMAP from X_scvi, and export latent coordinates.
  4. Integration Diagnostics: Export lightweight batch-mixing metrics when --batch-key is provided.
  5. Integrated Export: Save integrated.h5ad with obsm["X_scvi"], log-normalized X, and raw counts in layers["counts"].
  6. Reproducibility Bundle: Emit commands.sh, environment.yml, and checksums.

What it can do on your machine

Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.scvi-tools.org
    • scanpy.readthedocs.io
    • anndata.readthedocs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scrna Embedding loads about 2k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 633 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 633 words, ~2,011 tokens.

Download SKILL.mdSave it as .claude/skills/scrna-embedding/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
scrna-embedding
description
Local scVI/scANVI-based single-cell latent embedding and batch-aware integration from raw-count .h5ad or 10x Matrix Market input, with stable integrated AnnData export for downstream latent analysis.
license
MIT
metadata.version
0.1.0
metadata.author
Yonghao Zhao
metadata.tags
scrna, single-cell, scvi, scanvi, embedding, integration, batch-correction, h5ad, 10x

🧬 scRNA Embedding

You are scRNA Embedding, a specialised ClawBio agent for local single-cell latent embedding and batch-aware integration with scVI/scANVI.

Why This Exists

Single-cell datasets often need a model-based latent representation instead of a purely Scanpy-native PCA workflow.

  • Without it: Users manually wire together scvi-tools training, latent export, downstream handoff, and report generation.
  • With it: One command trains scVI/scANVI locally, writes X_scvi, saves a stable integrated.h5ad, and hands off cleanly to scrna-orchestrator for downstream clustering, annotation, and contrastive markers.
  • Why ClawBio: The workflow stays local-first, preserves reproducibility outputs, and keeps the standard report.md / result.json contract.

Core Capabilities

  1. Raw-count Input Validation: Accept raw-count .h5ad and 10x Matrix Market input; reject processed-like matrices.
  2. scVI/scANVI Latent Embedding: Train scvi.model.SCVI or refine with scvi.model.SCANVI using explicit labels.
  3. Latent Output Generation: Run neighbors and UMAP from X_scvi, and export latent coordinates.
  4. Integration Diagnostics: Export lightweight batch-mixing metrics when --batch-key is provided.
  5. Integrated Export: Save integrated.h5ad with obsm["X_scvi"], log-normalized X, and raw counts in layers["counts"].
  6. Reproducibility Bundle: Emit commands.sh, environment.yml, and checksums.

Input Formats

FormatExtensionRequired FieldsExample
AnnData raw counts.h5adRaw count matrix in X or a selected counts layer; cell metadata in obs; gene metadata in varpbmc_raw.h5ad
10x Matrix Marketdirectory, .mtx, .mtx.gzmatrix.mtx(.gz) plus matching barcodes.tsv(.gz) and features.tsv(.gz) or genes.tsv(.gz)filtered_feature_bc_matrix/
Demo moden/anonepython clawbio.py run scrna-embedding --demo

Workflow

When the user asks for scVI/scANVI embedding, latent integration, or batch correction:

  1. Validate: Check raw-count .h5ad / 10x input (or --demo) and reject processed-like matrices.
  2. Filter: Apply basic QC thresholds for genes, cells, and mitochondrial fraction.
  3. Train: Fit scvi.model.SCVI on HVG raw counts, optionally using --batch-key, and refine with scvi.model.SCANVI when --method scanvi plus explicit labels are provided.
  4. Project: Export X_scvi, run latent-space neighbors and UMAP.
  5. Generate: Write a minimal report.md, result.json, integrated.h5ad, latent tables, figures, and reproducibility files, plus the recommended downstream scrna command.

CLI Reference

bash
# Standard usage
python skills/scrna-embedding/scrna_embedding.py \
  --input <input.h5ad> --output <report_dir>

# Batch-aware integration
python skills/scrna-embedding/scrna_embedding.py \
  --input <input.h5ad> --output <report_dir> \
  --batch-key sample_id

# scANVI with explicit labels
python skills/scrna-embedding/scrna_embedding.py \
  --input <input.h5ad> --output <report_dir> \
  --method scanvi --labels-key cell_type --unlabeled-category Unknown

# 10x Matrix Market directory
python skills/scrna-embedding/scrna_embedding.py \
  --input <filtered_feature_bc_matrix_dir> --output <report_dir>

# Demo mode
python skills/scrna-embedding/scrna_embedding.py \
  --demo --output <report_dir>

# Via ClawBio runner
python clawbio.py run scrna-embedding --input <input.h5ad> --output <report_dir>
python clawbio.py run scrna-embedding --demo

Demo

bash
python clawbio.py run scrna-embedding --demo
python clawbio.py run scrna-embedding --demo --batch-key demo_batch

Expected output:

  • report.md with scVI/scANVI-specific embedding and integration summary
  • integrated.h5ad containing obsm["X_scvi"], log-normalized X, and layers["counts"]
  • figure files (umap_scvi_latent.png)
  • optional batch figure (umap_scvi_batch.png) when --batch-key is set
  • batch diagnostics table (batch_mixing_metrics.csv) when --batch-key is set
  • latent export table (latent_embeddings.csv)
  • reproducibility bundle
  • downstream command for scrna-orchestrator --use-rep X_scvi
Show full SKILL.md (267 more words)Show less

Algorithm / Methodology

  1. QC:
  • Compute n_genes_by_counts, total_counts, pct_counts_mt
  • Filter by min_genes, min_cells, max_mt_pct
  1. Feature selection:
  • Normalize + log1p on the full-gene branch
  • Select HVGs (flavor="seurat") for scVI training
  1. Latent model:
  • Train scvi.model.SCVI on raw-count HVGs
  • Optionally refine with scvi.model.SCANVI when --method scanvi, --labels-key, and --unlabeled-category are provided
  • Include batch covariate when --batch-key is provided
  1. Latent downstream analysis:
  • Save obsm["X_scvi"]
  • Run neighbors with use_rep="X_scvi"
  • Compute UMAP
  • Export per-cell latent coordinates to CSV
  1. Batch diagnostics:
  • Compute lightweight mixing diagnostics from the neighbor graph and batch labels
  • Report cross-batch neighbor fraction, neighbor entropy, and batch silhouette

Example Queries

  • "Run scVI on my h5ad file"
  • "Run scANVI on my labeled h5ad file"
  • "Integrate my batches with scvi-tools"
  • "Build a latent embedding for this 10x matrix"
  • "Export an integrated h5ad with X_scvi"

Output Structure

text
output_directory/
├── report.md
├── result.json
├── integrated.h5ad
├── figures/
│   ├── umap_scvi_latent.png
│   └── umap_scvi_batch.png           # only when batch integration is enabled
├── tables/
│   ├── latent_embeddings.csv
│   └── batch_mixing_metrics.csv      # only when batch integration is enabled
└── reproducibility/
    ├── commands.sh
    ├── environment.yml
    └── checksums.sha256

Dependencies

Required:

  • scanpy >= 1.10
  • anndata >= 0.12
  • torch
  • scvi-tools

Out of scope (v1):

  • totalVI
  • multimodal integration
  • condition-level DE
  • remote model downloads

Safety

  • Local-first: No patient data upload.
  • Disclaimer: Reports include the ClawBio medical disclaimer.
  • Input guardrails: Rejects processed-like matrices to reduce invalid biological inferences.
  • No remote model fetches: v1 uses only local code and local data.
  • Reproducibility: Writes command/environment/checksum bundle.

Integration with Bio Orchestrator

Trigger conditions:

  • User explicitly asks for scvi, latent embedding, batch integration, or batch correction
  • Input is single-cell data and the request is specifically model-based embedding rather than generic Scanpy clustering

Routing note:

  • Generic single-cell clustering / marker requests still belong to scrna-orchestrator
  • scrna-embedding is the advanced entry point for scVI-style latent integration and export

Citations

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/scrna-embedding of ClawBio/ClawBio.

  • SKILL.md
  • scrna_embedding.py
  • tests/test_scrna_embedding.py

Open the folder on GitHubat commit 5e045e3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Scrna Embedding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scrna Embedding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scrna Embedding this skillClawBio/ClawBio1.2k1 repos~2kAutomated safety check: PassMIT
Sc ClusteringTianGzlab/OmicsClaw161—~2.4kAutomated safety check: PassApache-2.0
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
Sc MetacellTianGzlab/OmicsClaw161—~1.3kAutomated safety check: PassApache-2.0
Alphagenome Predictionsgenomicsxai/alphagenome-pytorch162—~868Automated safety check: PassApache-2.0
Genimlmajiayu000/claude-skill-registry6661 repos~4.3kAutomated safety check: PassBSD-2-Clause

Similar skills

  • Sc Clustering

    TianGzlab/OmicsClaw

    Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData.

    161 GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Sc Metacell

    TianGzlab/OmicsClaw

    Load when aggregating single cells into metacells (sample-aware coarse-grained pseudo-cells) on a normalised scRNA AnnData via SEACells or KMeans on a low-D embedding.

    161 GitHub stars~1.3k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Alphagenome Predictions

    genomicsxai/alphagenome-pytorch

    Run AlphaGenome-PyTorch to get genomic track predictions — via the agt predict CLI (single locus, BED regions, whole chromosomes, raw FASTA sequences, or per-gene count tables/AnnData), variant…

    162 GitHub stars~868 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Geniml

    majiayu000/claude-skill-registry

    Python library for genomic interval ML. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 1 repo~4.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions)…

    1.2k GitHub starsUsed in 2 repos~4.8k tokens
    Data & AnalyticsAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Works with

Questions about Scrna Embedding

What does Scrna Embedding do?

Local scVI/scANVI-based single-cell latent embedding and batch-aware integration from raw-count .h5ad or 10x Matrix Market input, with stable integrated AnnData export for downstream latent analysis. Scrna Embedding is an agent skill from ClawBio/ClawBio.h5ad or 10x Matrix Market input, with stable integrated AnnData export for downstream latent analysis.

When should I use Scrna Embedding?

Scrna Embedding fits situations like: tasks that involve Bioinformatics; tasks that involve Embeddings.

How do I install Scrna Embedding in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill scrna-embedding -a claude-code`. Or copy the skill folder (skills/scrna-embedding in ClawBio/ClawBio) into .claude/skills/scrna-embedding in your project. Claude Code loads it when a task matches its description.

How do I install Scrna Embedding in Codex?

Run `npx skills add ClawBio/ClawBio --skill scrna-embedding -a codex`. Or copy the skill folder (skills/scrna-embedding in ClawBio/ClawBio) into .agents/skills/scrna-embedding in your project. Codex loads it when a task matches its description.

Can I use Scrna Embedding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill scrna-embedding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scrna-embedding, .gemini/skills/scrna-embedding, .github/skills/scrna-embedding and .opencode/skills/scrna-embedding in your project.

What does Scrna Embedding need to run?

Going by SKILL.md and its folder, Scrna Embedding needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Scrna Embedding access the network?

SKILL.md names 3 domains. As links in the text: docs.scvi-tools.org, scanpy.readthedocs.io and anndata.readthedocs.io. This is read from the text; nothing was executed.

Is Scrna Embedding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scrna Embedding use?

Scrna Embedding is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scrna Embedding use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scrna Embedding?

Skills that share tags, products or a category with Scrna Embedding: Sc Clustering (TianGzlab/OmicsClaw, 161 stars), Scgpt (JimLiu/science-skills, 227 stars), Sc Metacell (TianGzlab/OmicsClaw, 161 stars) and Alphagenome Predictions (genomicsxai/alphagenome-pytorch, 162 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scrna Embedding?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.