Agent skill

Single Cell Data Prep Qc

by harrisongzhang in harrisongzhang/TheVirtualBiotech

Single-cell RNA-seq data preparation and quality control pipeline.

MITAuto-check passedResearch & Science

Install Single Cell Data Prep Qc

skills CLI
$ npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install harrisongzhang/TheVirtualBiotech single-cell-data-prep-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/harrisongzhang/TheVirtualBiotech.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/single-cell-data-prep-qc .claude/skills/single-cell-data-prep-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
single-cell-data-prep-qc
GitHub stars
122
Token cost
~2.9k tokens
SKILL.md length
947 words
Files
9
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Single-cell RNA-seq data preparation and quality control pipeline.

  • Works in 2 steps: Data Discovery & Size Optimization → QC, Harmonization & Integration
  • You need to prepare scRNA-seq data for analysis
  • SKILL.md covers Overview, ⚠️ CRITICAL: Workspace…, Visualization Requirements and Procedure Quick Reference, plus 10 more sections
  • Calls python

What it does

Single Cell Data Prep Qc is an agent skill from harrisongzhang/TheVirtualBiotech. Single-cell RNA-seq data preparation and quality control pipeline. Handles data discovery from CELLxGENE Census, quality filtering, cell type harmonization, and batch correction. Outputs clean, integrated AnnData ready for statistical analysis. Use when you need to prepare scRNA-seq data for analysis or create publication-ready integrated datasets.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files (for example `procedures/batch_correction_procedure.md`, `procedures/harmonization_procedure.md` and `procedures/subsampling_procedure.md`).

It sits in Research & Science, covering Bioinformatics. It works with AnnData. The repository describes itself as: Multi-agent AI system for drug-target identification and due diligence. The licence is MIT.

When your agent uses it

  • You need to prepare scRNA-seq data for analysis
  • Create publication-ready integrated datasets

Example prompts

  • “/single-cell-data-prep-qc”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Data Discovery & Size Optimization
  2. QC, Harmonization & Integration

What it can do on your machine

Read from SKILL.md and the folder at commit 71f9da6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Single Cell Data Prep Qc loads about 2.9k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 947 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from harrisongzhang/TheVirtualBiotech at commit 71f9da6, republished under its MIT licence (© harrisongzhang). 947 words, ~2,928 tokens.

Download SKILL.mdSave it as .claude/skills/single-cell-data-prep-qc/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
single-cell-data-prep-qc
description
Single-cell RNA-seq data preparation and quality control pipeline. Handles data discovery from CELLxGENE Census, quality filtering, cell type harmonization, and batch correction. Outputs clean, integrated AnnData ready for statistical analysis. Use when you need to prepare scRNA-seq data for analysis or create publication-ready integrated datasets.

Single-Cell Data Preparation & QC

Overview

This skill prepares single-cell RNA-seq data from CELLxGENE Census for downstream statistical analysis. It handles the full data engineering pipeline from raw data query to batch-corrected, quality-controlled integrated dataset.

Pipeline:

  1. Data Discovery - Query Census, download disease + healthy data, subsample if needed
  2. Quality Control - Filter low-quality cells, set gene symbols
  3. Harmonization - Standardize cell type annotations across datasets
  4. Integration - Batch correction with Harmony, validation, visualization

Output: processed/integrated.h5ad - Clean, batch-corrected dataset ready for DE/pathway analysis

Key Features:

  • Enforces ~100K cell limit with two-stage stratified subsampling (donor + cell type)
  • Prioritizes donor preservation for robust pseudobulk analysis
  • Conservative cell type harmonization
  • Harmony batch correction with validation metrics
  • Comprehensive QC visualizations at each step

⚠️ CRITICAL: Workspace Initialization

Data Acquisition:

  • Data is downloaded via MCP tools (e.g., mcp__single_cell__get_anndata)
  • WorkspaceManager can be initialized via direct bash command before writing scripts

Workspace Initialization Pattern:

bash
# Recommended: Initialize workspace before any data operations
python -c "from src.utils.workspace_manager import WorkspaceManager; wm = WorkspaceManager(agent_name='single_cell_analyst'); print(f'DATE={wm.date}'); print(f'RUN_ID={wm.run_id}')"

When writing Python scripts: ❌ NEVER write to project root ✅ ALWAYS write to workspace code directory: workspace/{date}/{run_id}/single_cell_analyst/code/scripts/

See reference/workspace_setup.md for complete pattern.

⚠️ CELLxGENE Census Gene Index Fix: AnnData objects from Census use numeric IDs as var_names (e.g., '0', '1', '2'), NOT gene symbols. Gene symbols are in adata.var['feature_name']. Always run this immediately after any Census download, before any other processing:

python
adata.var.index = adata.var['feature_name'].astype(str)
adata.var_names_make_unique()

Visualization Requirements

This skill requires publication-quality visualizations at each major step. All figures should:

  • Use dpi=300 for publication quality
  • Include clear titles and axis labels
  • Save to workspace/{date}/{run_id}/single_cell_analyst/results/figures/
  • Use informative filenames (e.g., qc_violin_plots.png, batch_correction_umap.png)

Required visualization categories:

  1. QC metrics - Distribution plots before/after filtering
  2. Cell type composition - Bar plots comparing disease vs healthy
  3. Batch correction validation - Before/after UMAPs with metrics
  4. Data summary - Cell counts, sample sizes, metadata overview

See workflow documentation for specific visualization requirements at each step.


Procedure Quick Reference

This skill contains 3 procedures located in procedures/:

Mandatory Procedures:


⚠️ Forbidden Actions

You Must NEVER:

  • Skip the subsampling checkpoint when dataset >100K cells
  • Proceed with integration before harmonizing cell types (if needed)
  • Skip batch correction validation metrics
  • Create analysis scripts without gene symbols properly set
  • Write files to project root instead of workspace

See reference/forbidden_actions.md for full details.


Computational Resources

See reference/computational_resources.md for details.

Available resources:

  • 650 GB RAM
  • Up to 30-minute timeouts (1,800,000 ms)
  • Multiple cores for parallel processing

No excuses for shortcuts or skipping steps due to computational constraints.


Two-Stage Workflow

Stage 1: Data Discovery & Size Optimization

Complete workflow: See workflows/stage1_data_discovery.md

Objectives:

  • Initialize WorkspaceManager with date/run_id
  • Query CELLxGENE Census for disease + healthy tissue data
  • Download and save raw data to workspace
  • CHECKPOINT: Evaluate size and subsample to ≤100K cells if needed
  • Generate initial QC visualizations

Key outputs:

  • data/raw/disease_raw_subsampled.h5ad
  • data/raw/healthy_raw_subsampled.h5ad
  • Initial QC plots

Stage 2: QC, Harmonization & Integration

Complete workflow: See workflows/stage2_qc_integration.md

Objectives:

  • Load subsampled data from Stage 1 (≤100K cells guaranteed)
  • Perform QC and filtering
  • Checkpoint 1: Harmonize cell type annotations if needed
  • Checkpoint 2: Perform batch correction with Harmony
  • Validate integration quality
  • Generate comprehensive visualization suite

Key outputs:

  • data/processed/integrated.h5ad - Ready for analysis
  • QC, harmonization, integration figures
  • Metadata summary JSON

Workflow Completion Checklist

Before declaring data prep complete, verify:

Files present:

bash
# Raw data
ls workspace/{date}/{run_id}/single_cell_analyst/data/raw/*_subsampled.h5ad

# Processed data (MAIN OUTPUT)
ls workspace/{date}/{run_id}/single_cell_analyst/data/processed/integrated.h5ad

# QC figures
ls workspace/{date}/{run_id}/single_cell_analyst/results/figures/qc_*.png

# Integration figures
ls workspace/{date}/{run_id}/single_cell_analyst/results/figures/integration_*.png

Success criteria:

  • ✅ Both stages completed in sequence
  • ✅ All mandatory checkpoints completed (subsampling, harmonization, batch correction)
  • ✅ Gene symbols properly set (not integers)
  • ✅ Batch correction validated with metrics
  • ✅ All required visualizations generated
  • ✅ integrated.h5ad file exists and is valid

Show full SKILL.md (375 more words)Show less

Output Handoff to Statistical Analysis Skill

Primary output: workspace/{date}/{run_id}/single_cell_analyst/data/processed/integrated.h5ad

This file contains:

  • Batch-corrected expression data (X_pca_harmony for neighbors/UMAP)
  • Raw counts in .X (preserved for differential expression)
  • Quality-filtered cells
  • Gene symbols as var_names (required for pathway enrichment)
  • Harmonized cell type labels in obs['unified_cell_type']
  • Condition labels in obs['condition'] (disease/healthy)
  • Donor IDs in obs['donor_id'] (for pseudobulk DE)

Validation before handoff:

python
import scanpy as sc

# Load integrated data
adata = sc.read_h5ad('workspace/{date}/{run_id}/single_cell_analyst/data/processed/integrated.h5ad')

print("\n[Integration Output Validation]")
print(f"  Cells: {adata.n_obs:,}")
print(f"  Genes: {adata.n_vars:,}")
print(f"  Gene symbols set: {not adata.var_names[0].isdigit()}")
print(f"  Has donor_id: {'donor_id' in adata.obs.columns}")
print(f"  Has condition: {'condition' in adata.obs.columns}")
print(f"  Has unified_cell_type: {'unified_cell_type' in adata.obs.columns}")
print(f"  Has Harmony embedding: {'X_pca_harmony' in adata.obsm.keys()}")

# Verify raw counts preserved
import numpy as np
is_integer = np.all(np.equal(np.mod(adata.X.data, 1), 0))
print(f"  Raw counts preserved: {is_integer}")

print("\n✅ Data ready for statistical analysis")

Metadata summary: Create results/reports/data_prep_summary.md with:

  • Data sources and queries used
  • Cell counts before/after each QC step
  • Harmonization decisions (if applied)
  • Batch correction validation metrics
  • All figure references

Workflow Summary

START: Recognize data prep task
  ↓
STAGE 1: Data Discovery & Size Optimization
  ├─ Initialize workspace (print date/run_id)
  ├─ Query CELLxGENE Census
  ├─ Download data
  ├─ ⚠️ CHECKPOINT: Size check → subsample if >100K cells
  │   └─ Link: procedures/subsampling_procedure.md
  ├─ Generate initial QC visualizations
  └─ Gate: Verify *_subsampled.h5ad files exist
  ↓
STAGE 2: QC, Harmonization & Integration
  ├─ Load subsampled data (≤100K cells guaranteed)
  ├─ QC and filtering with visualizations
  ├─ Checkpoint 1: Harmonization → harmonize if needed
  │   └─ Link: procedures/harmonization_procedure.md
  ├─ Checkpoint 2: Batch correction with Harmony + validation
  │   └─ Link: procedures/batch_correction_procedure.md
  ├─ Generate comprehensive visualization suite
  └─ Gate: Verify integrated.h5ad exists and validated
  ↓
COMPLETE: Clean data ready for statistical analysis

Key Principles

Follow the workflow sequentially:

  • Complete Stage 1 before Stage 2
  • Complete all checkpoints in order
  • Generate visualizations at each step

Use procedures at checkpoints:

  • When workflow references a procedure file, read it completely
  • Follow the procedure's instructions exactly
  • Return to workflow when procedure is complete

Validate at gates:

  • Run validation checks before progressing
  • Verify required files exist
  • Ensure visualizations are generated
  • Do not skip stages or checkpoints

Generate publication-quality figures:

  • All plots at dpi=300
  • Clear labels and legends
  • Informative titles
  • Save with descriptive filenames

Troubleshooting

Issue: Dataset exceeds 100K cells but subsampling skipped

  • Return to Stage 1 subsampling checkpoint
  • Apply stratified subsampling procedure
  • Regenerate subsampled files

Issue: Gene symbols still integers after QC

  • Check if feature_name column exists in var
  • Apply gene symbol mapping in QC script
  • Make unique if duplicates exist

Issue: Batch correction fails validation

  • Check iLISI and ASW_batch metrics
  • Ensure batch variable is dataset_id (not condition)
  • Try adjusting Harmony parameters or alternative methods

Issue: Cell type harmonization unclear

  • Review harmonization procedure carefully
  • Create decision table for manual adjudication
  • Document all harmonization decisions
  • When in doubt, be conservative (don't force matches)

Success Criteria

Your data preparation is complete when:

  • ✅ Both stages finished in sequence
  • ✅ All checkpoints evaluated and procedures followed
  • ✅ Subsampling performed if needed
  • ✅ Harmonization completed if annotations differed
  • ✅ Batch correction applied and validated
  • ✅ integrated.h5ad file created with all required components
  • ✅ All visualization requirements met
  • ✅ Data prep summary report created
  • ✅ Validation checks pass

Next Steps

After completing data preparation:

  1. Verify all Stage 2 checklist items are checked
  2. Run validation script to confirm integrated.h5ad quality
  3. Review data_prep_summary.md
  4. Ready to invoke single-cell-statistical-analysis skill for DE and pathway enrichment

You have the workflow. Execute it completely. Follow every checkpoint. Generate all visualizations. No shortcuts.

© harrisongzhang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files in .claude/skills/single-cell-data-prep-qc of harrisongzhang/TheVirtualBiotech.

  • SKILL.md
  • procedures/batch_correction_procedure.md
  • procedures/harmonization_procedure.md
  • procedures/subsampling_procedure.md
  • reference/computational_resources.md
  • reference/forbidden_actions.md
  • reference/workspace_setup.md
  • workflows/stage1_data_discovery.md
  • workflows/stage2_qc_integration.md

Open the folder on GitHubat commit 71f9da6

Compare with similar skills

Single Cell Data Prep Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Single Cell Data Prep Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Single Cell Data Prep Qc this skillharrisongzhang/TheVirtualBiotech122—~2.9kAutomated safety check: PassMIT
Scanpy Single-Cell Analysisdavila7/claude-code-templates33k15 repos~2.8kAutomated safety check: PassMIT
ScgptJimLiu/science-skills2284 repos~1.3kAutomated safety check: PassApache-2.0
PyDESeq2 Differential Expressiondavila7/claude-code-templates33k11 repos~4kAutomated safety check: PassMIT
Anndatadavila7/claude-code-templates33k11 repos~2.5kAutomated safety check: PassMIT
Single-Cell Initial AnalysisLigphiDonk/Oh-my--paper7391 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    33k GitHub starsUsed in 15 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    228 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    33k GitHub starsUsed in 11 repos~4k tokens
    Research & ScienceAuto-check passed
  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    33k GitHub starsUsed in 11 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    739 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check passed
  • Spatial Atera

    QING1105/ezST

    Atera platform branch of the spatial transcriptomics workflow — load and validate Atera cell-level output (AnnData + Zarr segmentation) for downstream analysis.

    101 GitHub stars~576 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from harrisongzhang/TheVirtualBiotech

  • Evidence Citation

    harrisongzhang/TheVirtualBiotech

    How to report findings so they can be cited — writing evidence to files before asserting it, choosing where outputs belong, describing what each artifact shows, and returning findings with…

    122 GitHub stars~1.4k tokensUpdated 23 days ago
    Auto-check passed
  • Run Organization

    harrisongzhang/TheVirtualBiotech

    How to organise a session run so a human can audit it — directory layout, artifact naming, recording the analysis plan, filing claim-evidence objects, and the end-of-run checklist.

    122 GitHub stars~1.8k tokensUpdated 23 days ago
    Auto-check passed
  • Single Cell Analysis

    harrisongzhang/TheVirtualBiotech

    Statistical analysis and reporting for single-cell RNA-seq data.

    122 GitHub stars~3.2k tokensUpdated 23 days ago
    Auto-check passed

Works with

Questions about Single Cell Data Prep Qc

What does Single Cell Data Prep Qc do?

Single-cell RNA-seq data preparation and quality control pipeline. Single Cell Data Prep Qc is an agent skill from harrisongzhang/TheVirtualBiotech. Single-cell RNA-seq data preparation and quality control pipeline.

When should I use Single Cell Data Prep Qc?

Single Cell Data Prep Qc fits situations like: you need to prepare scRNA-seq data for analysis; create publication-ready integrated datasets.

How do I install Single Cell Data Prep Qc in Claude Code?

Run `npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a claude-code`. Or copy the skill folder (.claude/skills/single-cell-data-prep-qc in harrisongzhang/TheVirtualBiotech) into .claude/skills/single-cell-data-prep-qc in your project. Claude Code loads it when a task matches its description.

How do I install Single Cell Data Prep Qc in Codex?

Run `npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a codex`. Or copy the skill folder (.claude/skills/single-cell-data-prep-qc in harrisongzhang/TheVirtualBiotech) into .agents/skills/single-cell-data-prep-qc in your project. Codex loads it when a task matches its description.

Can I use Single Cell Data Prep Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add harrisongzhang/TheVirtualBiotech --skill single-cell-data-prep-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/single-cell-data-prep-qc, .gemini/skills/single-cell-data-prep-qc, .github/skills/single-cell-data-prep-qc and .opencode/skills/single-cell-data-prep-qc in your project.

What does Single Cell Data Prep Qc need to run?

Going by SKILL.md and its folder, Single Cell Data Prep Qc needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Single Cell Data Prep Qc access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Single Cell Data Prep Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Single Cell Data Prep Qc use?

Single Cell Data Prep Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Single Cell Data Prep Qc use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Single Cell Data Prep Qc?

Skills that share tags, products or a category with Single Cell Data Prep Qc: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 33k stars), Scgpt (JimLiu/science-skills, 228 stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 33k stars) and Anndata (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Single Cell Data Prep Qc?

harrisongzhang (a GitHub user) maintains it in harrisongzhang/TheVirtualBiotech, which has 122 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on September 17, 2026.

Source: harrisongzhang/TheVirtualBiotech on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.