Scanpy Single-Cell Analysis
davila7/claude-code-templates
Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.
Single-cell RNA-seq data preparation and quality control pipeline.
$ npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install harrisongzhang/TheVirtualBiotech single-cell-data-prep-qc --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/harrisongzhang/TheVirtualBiotech.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/single-cell-data-prep-qc .claude/skills/single-cell-data-prep-qc && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "single-cell-data-prep-qc" agent skill from https://github.com/harrisongzhang/TheVirtualBiotech/tree/main/.claude/skills/single-cell-data-prep-qc into .claude/skills/single-cell-data-prep-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-data-prep-qc", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/harrisongzhang/TheVirtualBiotech/tree/main/.claude/skills/single-cell-data-prep-qcType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install harrisongzhang/TheVirtualBiotech single-cell-data-prep-qc --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harrisongzhang/TheVirtualBiotech.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/single-cell-data-prep-qc .agents/skills/single-cell-data-prep-qc && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "single-cell-data-prep-qc" agent skill from https://github.com/harrisongzhang/TheVirtualBiotech/tree/main/.claude/skills/single-cell-data-prep-qc into .agents/skills/single-cell-data-prep-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-data-prep-qc", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install harrisongzhang/TheVirtualBiotech single-cell-data-prep-qc --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harrisongzhang/TheVirtualBiotech.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/single-cell-data-prep-qc .cursor/skills/single-cell-data-prep-qc && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "single-cell-data-prep-qc" agent skill from https://github.com/harrisongzhang/TheVirtualBiotech/tree/main/.claude/skills/single-cell-data-prep-qc into .cursor/skills/single-cell-data-prep-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-data-prep-qc", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/harrisongzhang/TheVirtualBiotech.git --path .claude/skills/single-cell-data-prep-qc--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install harrisongzhang/TheVirtualBiotech single-cell-data-prep-qc --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harrisongzhang/TheVirtualBiotech.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/single-cell-data-prep-qc .gemini/skills/single-cell-data-prep-qc && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "single-cell-data-prep-qc" agent skill from https://github.com/harrisongzhang/TheVirtualBiotech/tree/main/.claude/skills/single-cell-data-prep-qc into .gemini/skills/single-cell-data-prep-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-data-prep-qc", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install harrisongzhang/TheVirtualBiotech single-cell-data-prep-qcInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/harrisongzhang/TheVirtualBiotech.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/single-cell-data-prep-qc .github/skills/single-cell-data-prep-qc && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "single-cell-data-prep-qc" agent skill from https://github.com/harrisongzhang/TheVirtualBiotech/tree/main/.claude/skills/single-cell-data-prep-qc into .github/skills/single-cell-data-prep-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-data-prep-qc", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install harrisongzhang/TheVirtualBiotech single-cell-data-prep-qc --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harrisongzhang/TheVirtualBiotech.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/single-cell-data-prep-qc .opencode/skills/single-cell-data-prep-qc && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "single-cell-data-prep-qc" agent skill from https://github.com/harrisongzhang/TheVirtualBiotech/tree/main/.claude/skills/single-cell-data-prep-qc into .opencode/skills/single-cell-data-prep-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-data-prep-qc", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
single-cell-data-prep-qcSingle-cell RNA-seq data preparation and quality control pipeline.
Single Cell Data Prep Qc is an agent skill from harrisongzhang/TheVirtualBiotech. Single-cell RNA-seq data preparation and quality control pipeline. Handles data discovery from CELLxGENE Census, quality filtering, cell type harmonization, and batch correction. Outputs clean, integrated AnnData ready for statistical analysis. Use when you need to prepare scRNA-seq data for analysis or create publication-ready integrated datasets.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files (for example `procedures/batch_correction_procedure.md`, `procedures/harmonization_procedure.md` and `procedures/subsampling_procedure.md`).
It sits in Research & Science, covering Bioinformatics. It works with AnnData. The repository describes itself as: Multi-agent AI system for drug-target identification and due diligence. The licence is MIT.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 71f9da6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Single Cell Data Prep Qc loads about 2.9k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 947 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from harrisongzhang/TheVirtualBiotech at commit 71f9da6, republished under its MIT licence (© harrisongzhang). 947 words, ~2,928 tokens.
.claude/skills/single-cell-data-prep-qc/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.This skill prepares single-cell RNA-seq data from CELLxGENE Census for downstream statistical analysis. It handles the full data engineering pipeline from raw data query to batch-corrected, quality-controlled integrated dataset.
Pipeline:
Output: processed/integrated.h5ad - Clean, batch-corrected dataset ready for DE/pathway analysis
Key Features:
Data Acquisition:
mcp__single_cell__get_anndata)Workspace Initialization Pattern:
# Recommended: Initialize workspace before any data operations
python -c "from src.utils.workspace_manager import WorkspaceManager; wm = WorkspaceManager(agent_name='single_cell_analyst'); print(f'DATE={wm.date}'); print(f'RUN_ID={wm.run_id}')"When writing Python scripts:
❌ NEVER write to project root
✅ ALWAYS write to workspace code directory: workspace/{date}/{run_id}/single_cell_analyst/code/scripts/
See reference/workspace_setup.md for complete pattern.
⚠️ CELLxGENE Census Gene Index Fix: AnnData objects from Census use numeric IDs as var_names (e.g., '0', '1', '2'), NOT gene symbols. Gene symbols are in adata.var['feature_name']. Always run this immediately after any Census download, before any other processing:
adata.var.index = adata.var['feature_name'].astype(str)
adata.var_names_make_unique()This skill requires publication-quality visualizations at each major step. All figures should:
dpi=300 for publication qualityworkspace/{date}/{run_id}/single_cell_analyst/results/figures/qc_violin_plots.png, batch_correction_umap.png)Required visualization categories:
See workflow documentation for specific visualization requirements at each step.
This skill contains 3 procedures located in procedures/:
Mandatory Procedures:
You Must NEVER:
See reference/forbidden_actions.md for full details.
See reference/computational_resources.md for details.
Available resources:
No excuses for shortcuts or skipping steps due to computational constraints.
Complete workflow: See workflows/stage1_data_discovery.md
Objectives:
Key outputs:
data/raw/disease_raw_subsampled.h5addata/raw/healthy_raw_subsampled.h5adComplete workflow: See workflows/stage2_qc_integration.md
Objectives:
Key outputs:
data/processed/integrated.h5ad - Ready for analysisBefore declaring data prep complete, verify:
Files present:
# Raw data
ls workspace/{date}/{run_id}/single_cell_analyst/data/raw/*_subsampled.h5ad
# Processed data (MAIN OUTPUT)
ls workspace/{date}/{run_id}/single_cell_analyst/data/processed/integrated.h5ad
# QC figures
ls workspace/{date}/{run_id}/single_cell_analyst/results/figures/qc_*.png
# Integration figures
ls workspace/{date}/{run_id}/single_cell_analyst/results/figures/integration_*.pngSuccess criteria:
Primary output: workspace/{date}/{run_id}/single_cell_analyst/data/processed/integrated.h5ad
This file contains:
X_pca_harmony for neighbors/UMAP).X (preserved for differential expression)var_names (required for pathway enrichment)obs['unified_cell_type']obs['condition'] (disease/healthy)obs['donor_id'] (for pseudobulk DE)Validation before handoff:
import scanpy as sc
# Load integrated data
adata = sc.read_h5ad('workspace/{date}/{run_id}/single_cell_analyst/data/processed/integrated.h5ad')
print("\n[Integration Output Validation]")
print(f" Cells: {adata.n_obs:,}")
print(f" Genes: {adata.n_vars:,}")
print(f" Gene symbols set: {not adata.var_names[0].isdigit()}")
print(f" Has donor_id: {'donor_id' in adata.obs.columns}")
print(f" Has condition: {'condition' in adata.obs.columns}")
print(f" Has unified_cell_type: {'unified_cell_type' in adata.obs.columns}")
print(f" Has Harmony embedding: {'X_pca_harmony' in adata.obsm.keys()}")
# Verify raw counts preserved
import numpy as np
is_integer = np.all(np.equal(np.mod(adata.X.data, 1), 0))
print(f" Raw counts preserved: {is_integer}")
print("\n✅ Data ready for statistical analysis")Metadata summary:
Create results/reports/data_prep_summary.md with:
START: Recognize data prep task
↓
STAGE 1: Data Discovery & Size Optimization
├─ Initialize workspace (print date/run_id)
├─ Query CELLxGENE Census
├─ Download data
├─ ⚠️ CHECKPOINT: Size check → subsample if >100K cells
│ └─ Link: procedures/subsampling_procedure.md
├─ Generate initial QC visualizations
└─ Gate: Verify *_subsampled.h5ad files exist
↓
STAGE 2: QC, Harmonization & Integration
├─ Load subsampled data (≤100K cells guaranteed)
├─ QC and filtering with visualizations
├─ Checkpoint 1: Harmonization → harmonize if needed
│ └─ Link: procedures/harmonization_procedure.md
├─ Checkpoint 2: Batch correction with Harmony + validation
│ └─ Link: procedures/batch_correction_procedure.md
├─ Generate comprehensive visualization suite
└─ Gate: Verify integrated.h5ad exists and validated
↓
COMPLETE: Clean data ready for statistical analysisFollow the workflow sequentially:
Use procedures at checkpoints:
Validate at gates:
Generate publication-quality figures:
Issue: Dataset exceeds 100K cells but subsampling skipped
Issue: Gene symbols still integers after QC
feature_name column exists in varIssue: Batch correction fails validation
Issue: Cell type harmonization unclear
Your data preparation is complete when:
After completing data preparation:
You have the workflow. Execute it completely. Follow every checkpoint. Generate all visualizations. No shortcuts.
© harrisongzhang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files in .claude/skills/single-cell-data-prep-qc of harrisongzhang/TheVirtualBiotech.
Open the folder on GitHubat commit 71f9da6
Single Cell Data Prep Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Single Cell Data Prep Qc this skillharrisongzhang/TheVirtualBiotech | 122 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Scanpy Single-Cell Analysisdavila7/claude-code-templates | 33k | 15 repos | ~2.8k | Automated safety check: Pass | MIT | |
| ScgptJimLiu/science-skills | 228 | 4 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| PyDESeq2 Differential Expressiondavila7/claude-code-templates | 33k | 11 repos | ~4k | Automated safety check: Pass | MIT | |
| Anndatadavila7/claude-code-templates | 33k | 11 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Single-Cell Initial AnalysisLigphiDonk/Oh-my--paper | 739 | 1 repos | ~1.4k | Automated safety check: Pass | MIT |
davila7/claude-code-templates
Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.
JimLiu/science-skills
Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.
davila7/claude-code-templates
Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.
davila7/claude-code-templates
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…
LigphiDonk/Oh-my--paper
Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.
QING1105/ezST
Atera platform branch of the spatial transcriptomics workflow — load and validate Atera cell-level output (AnnData + Zarr segmentation) for downstream analysis.
harrisongzhang/TheVirtualBiotech
How to report findings so they can be cited — writing evidence to files before asserting it, choosing where outputs belong, describing what each artifact shows, and returning findings with…
harrisongzhang/TheVirtualBiotech
How to organise a session run so a human can audit it — directory layout, artifact naming, recording the analysis plan, filing claim-evidence objects, and the end-of-run checklist.
harrisongzhang/TheVirtualBiotech
Statistical analysis and reporting for single-cell RNA-seq data.
Works with
Categories
Single-cell RNA-seq data preparation and quality control pipeline. Single Cell Data Prep Qc is an agent skill from harrisongzhang/TheVirtualBiotech. Single-cell RNA-seq data preparation and quality control pipeline.
Single Cell Data Prep Qc fits situations like: you need to prepare scRNA-seq data for analysis; create publication-ready integrated datasets.
Run `npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a claude-code`. Or copy the skill folder (.claude/skills/single-cell-data-prep-qc in harrisongzhang/TheVirtualBiotech) into .claude/skills/single-cell-data-prep-qc in your project. Claude Code loads it when a task matches its description.
Run `npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a codex`. Or copy the skill folder (.claude/skills/single-cell-data-prep-qc in harrisongzhang/TheVirtualBiotech) into .agents/skills/single-cell-data-prep-qc in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/single-cell-data-prep-qc, .gemini/skills/single-cell-data-prep-qc, .github/skills/single-cell-data-prep-qc and .opencode/skills/single-cell-data-prep-qc in your project.
Going by SKILL.md and its folder, Single Cell Data Prep Qc needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Single Cell Data Prep Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Single Cell Data Prep Qc: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 33k stars), Scgpt (JimLiu/science-skills, 228 stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 33k stars) and Anndata (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
harrisongzhang (a GitHub user) maintains it in harrisongzhang/TheVirtualBiotech, which has 122 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on September 17, 2026.
Source: harrisongzhang/TheVirtualBiotech on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.