Agent skill

Bio Single Cell Preprocessing

by GPTomics in GPTomics/bioSkills

Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R).

MITAuto-check passedResearch & Science

Install Bio Single Cell Preprocessing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-single-cell-preprocessing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/single-cell/preprocessing .claude/skills/bio-single-cell-preprocessing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-single-cell-preprocessing
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.1k tokens
SKILL.md length
2,031 words
Files
4
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R).

  • Works in 9 steps: Load the RAW (unfiltered) droplet matrix. → Empty-droplet calling (EmptyDrops,… → Ambient-RNA removal (optional;… → …
  • Filtering low-quality cells with MAD-adaptive thresholds
  • SKILL.md covers Version Compatibility, Governing Principle, Canonical Pipeline Order and Quality Control, plus 7 more sections
  • Runs Python and R scripts from its folder; calls pip

What it does

Bio Single Cell Preprocessing is an agent skill from GPTomics/bioSkills. Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R). Use when filtering low-quality cells with MAD-adaptive thresholds, setting tissue-aware mito cutoffs, removing ambient RNA (SoupX/CellBender/DecontX), choosing a normalization (shifted-log vs scran vs sctransform vs Pearson residuals), selecting highly variable genes, or deciding whether to scale and regress out covariates.

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/preprocess_scanpy.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with Python and Scanpy. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Filtering low-quality cells with MAD-adaptive thresholds
  • Setting tissue-aware mito cutoffs
  • Removing ambient RNA (SoupX/CellBender/DecontX)
  • Choosing a normalization (shifted-log vs scran vs sctransform vs Pearson residuals)

Example prompts

  • “/bio-single-cell-preprocessing”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Load the RAW (unfiltered) droplet matrix.
  2. Empty-droplet calling (EmptyDrops, FDR<=0.001 on raw) or CellBender (folds calling + denoising).
  3. Ambient-RNA removal (optional; SoupX/DecontX/CellBender) - BEFORE QC, because it needs the soup estimate.
  4. QC filtering: cells (MAD on counts/genes/mito) + genes (min_cells).
  5. Doublet detection - per sample, on raw counts (see single-cell/doublet-detection).
  6. Normalization - shifted-log default, or scran/Pearson.
  7. HVG selection - mind the raw-vs-lognorm input per flavor.
  8. Scaling (optional, increasingly skipped).
  9. PCA on HVG (~50 comps), then neighbors/clustering.

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Single Cell Preprocessing loads about 5.1k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 2,031 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,031 words, ~5,090 tokens.

Download SKILL.mdSave it as .claude/skills/bio-single-cell-preprocessing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-single-cell-preprocessing
description
Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R). Use when filtering low-quality cells with MAD-adaptive thresholds, setting tissue-aware mito cutoffs, removing ambient RNA (SoupX/CellBender/DecontX), choosing a normalization (shifted-log vs scran vs sctransform vs Pearson residuals), selecting highly variable genes, or deciding whether to scale and regress out covariates.
tool_type
mixed
primary_tool
Seurat

Version Compatibility

Reference examples tested with: scanpy 1.10+, Seurat 5.0+, scran 1.30+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Single-Cell Preprocessing

"Preprocess my scRNA-seq data" -> Remove bad barcodes, correct technical biases, and select informative genes before dimensionality reduction.

  • Python: calculate_qc_metrics() -> filter -> normalize_total()+log1p() -> highly_variable_genes()
  • R: QC -> NormalizeData() or SCTransform() -> FindVariableFeatures() -> ScaleData()

Governing Principle

Every preprocessing choice propagates to every downstream result, and the two highest-leverage choices both encode a hidden biological assumption.

Normalization assumes near-constant total mRNA per cell. Shifted-log CP10k, scran deconvolution, and sctransform all divide by a per-cell size factor meant to capture only capture-efficiency and sequencing depth. They silently assume total transcriptome size is roughly constant across cell types, so a cell's total UMI count is a pure technical nuisance. This is false for plasma/antibody-secreting cells, secretory epithelia, hepatocytes, large neurons, and S/G2M cells, which carry 2-10x more mRNA. Dividing them to a common total deflates every gene that is not one of their few dominant transcripts (a compositional see-saw), and partially erases real proliferation biology. The honest framing: single-cell measures proportions, not absolute amounts.

QC metrics are biology metrics in disguise. pct_counts_mt conflates apoptosis, genuine metabolic activity (cardiomyocytes/hepatocytes/muscle are constitutively 20-40% mito and healthy), dissociation stress, and technical contamination. A flat global mito cutoff deletes entire healthy parenchymal populations and the survivors still cluster cleanly, so the loss is invisible. Use adaptive, tissue-aware thresholds and treat all three QC covariates jointly.

A beautiful UMAP proves nothing. Compositional normalization bias, deleted high-mito parenchyma, dissociation-stress clusters, ambient-induced co-expression, and residual homotypic doublets are all compatible with tidy clusters. The dangerous artifacts are precisely the ones that do not look like artifacts.

Canonical Pipeline Order

  1. Load the RAW (unfiltered) droplet matrix.
  2. Empty-droplet calling (EmptyDrops, FDR<=0.001 on raw) or CellBender (folds calling + denoising).
  3. Ambient-RNA removal (optional; SoupX/DecontX/CellBender) - BEFORE QC, because it needs the soup estimate.
  4. QC filtering: cells (MAD on counts/genes/mito) + genes (min_cells).
  5. Doublet detection - per sample, on raw counts (see single-cell/doublet-detection).
  6. Normalization - shifted-log default, or scran/Pearson.
  7. HVG selection - mind the raw-vs-lognorm input per flavor.
  8. Scaling (optional, increasingly skipped).
  9. PCA on HVG (~50 comps), then neighbors/clustering.

Ambient correction needs the raw matrix and must precede QC filtering; once subset to cells, the soup estimate is gone. Doublet detection runs on raw counts, so stash counts before normalizing.

Steps 1-5 are per-sample operations performed BEFORE merge or integration: empty-droplet calling, ambient removal (SoupX load10X is inherently per-run), adaptive QC, and doublet detection all reason about one capture's droplet population, so QC-then-merge is correct and merge-then-QC leaks batch effects into every threshold and contaminates the soup and doublet-scoring neighborhoods. Merge only after each sample is cleaned.

Quality Control

Goal: Remove empty/dying/stressed barcodes using data-driven thresholds that do not delete real cell types.

Approach: Annotate mito/ribo/hemoglobin gene sets, compute joint QC metrics, then flag outliers by median absolute deviation (MAD) on the log scale rather than fixed cutoffs.

python
import scanpy as sc
import numpy as np
from scipy.stats import median_abs_deviation

adata.var['mt'] = adata.var_names.str.startswith('MT-')                       # mouse: 'mt-'
adata.var['ribo'] = adata.var_names.str.startswith(('RPS', 'RPL'))
adata.var['hb'] = adata.var_names.str.contains(r'^HB[ABDEGMQZ]\d*(?!\w)')      # explicit subunits, not legacy ^HB[^(P)]
sc.pp.calculate_qc_metrics(adata, qc_vars=['mt', 'ribo', 'hb'], percent_top=[20], log1p=True, inplace=True)
# inplace defaults to False and returns DataFrames; pass inplace=True to write .obs/.var

def is_outlier(adata, metric, nmads):
    M = adata.obs[metric]
    return (M < np.median(M) - nmads * median_abs_deviation(M)) | (np.median(M) + nmads * median_abs_deviation(M) < M)

adata.obs['outlier'] = (is_outlier(adata, 'log1p_total_counts', 5) | is_outlier(adata, 'log1p_n_genes_by_counts', 5)
                        | is_outlier(adata, 'pct_counts_in_top_20_genes', 5))
adata.obs['mt_outlier'] = is_outlier(adata, 'pct_counts_mt', 3) | (adata.obs['pct_counts_mt'] > 8)
adata = adata[~(adata.obs['outlier'] | adata.obs['mt_outlier'])].copy()
sc.pp.filter_genes(adata, min_cells=3)

When samples differ in depth/quality or were sequenced in separate batches, compute MAD thresholds PER SAMPLE, not globally: a single global MAD over-cuts the shallow batch and under-cuts the deep one. Apply is_outlier within each batch_key group (the same per-batch logic the HVG step uses).

python
flags = ['log1p_total_counts', 'log1p_n_genes_by_counts', 'pct_counts_in_top_20_genes']
adata.obs['outlier'] = adata.obs.groupby('sample', observed=True).apply(
    lambda g: (is_outlier(adata[g.index], flags[0], 5) | is_outlier(adata[g.index], flags[1], 5)
               | is_outlier(adata[g.index], flags[2], 5))).droplevel(0)
r
# '^MT-' matches gene SYMBOLS; with Ensembl-ID feature names it matches nothing and the mito filter silently does nothing
seurat_obj[['percent.mt']] <- PercentageFeatureSet(seurat_obj, pattern = '^MT-')
VlnPlot(seurat_obj, features = c('nFeature_RNA', 'nCount_RNA', 'percent.mt'), ncol = 3)
seurat_obj <- subset(seurat_obj, subset = nFeature_RNA > 200 & nFeature_RNA < 5000 & percent.mt < 20)
QC Thresholds and Rationale
MetricReference valueRationale and caveat
min_genes200Below this is mostly empty droplets / debris; raise for deep data
log1p_total_counts / log1p_n_genes_by_counts5 MADsc-best-practices loosens from scater's 3 MAD to avoid cutting real biology; filter on the log scale (depth is right-skewed)
pct_counts_in_top_20_genes5 MADHigh value flags low-complexity / dying cells
pct_counts_mt3 MAD plus hard >8%Tissue-dependent: 5-20% typical, but cardiomyocytes/hepatocytes/muscle are constitutively high; nuclei are ~0-2% and any mito flags ambient
min_cells (genes)3Remove genes seen in too few cells to be informative

Fixed cutoffs are a fast first pass for well-characterized tissue but silently delete valid populations; MAD-adaptive is the modern default; miQC (a mito-vs-detected-genes mixture model) helps when that relationship varies across samples.

Mito and Dissociation Confounds

High mito is ambiguous: apoptosis co-occurs with low gene counts and apoptotic markers, while warm-dissociation stress co-occurs with immediate-early genes (FOS, JUN, JUNB, EGR1) and heat-shock proteins (HSPA1A/B) at normal gene counts. The IEG/HSP program creates a spurious "activated/stressed" cluster that passes every count/mito filter, is cell-type-specific in magnitude (so it does not cancel as a uniform batch effect), and overlaps real immune/stem activation, so naive removal can itself delete biology. Score the dissociation module per cell, then exclude those genes from HVG/clustering or flag and interpret cautiously; cold-protease digestion and single-nucleus assays reduce the artifact. For nuclei, standard mito thresholds are meaningless (baseline near zero) - lean on counts/genes outliers.

Ambient RNA Removal

Goal: Remove cell-free "soup" mRNA that inflates off-target markers (hemoglobin everywhere in PBMCs, hepatocyte genes in non-hepatocytes) and fabricates co-expression.

Approach: Estimate the soup profile and a per-cell contamination fraction, then subtract; pick ONE tool and validate that a known-specific marker survives.

ToolInputNeeds empty droplets?StrengthFails / risk
SoupX (R)Cell Ranger raw+filteredYesFast, interpretable rho, auto-estimateautoEstCont fails on homogeneous data; single global soup wrong when ambient is heterogeneous
CellBender (Python, GPU)RAW h5Yes (core of model)Deep generative; removes ambient + barcode noise; also does cell-calling; strong on nucleiOver-removes real low-abundance genes at high --fpr; black-box; slow
DecontX (R, celda)Filtered cellsNoNo raw needed; easy SCE/Seurat integrationRelies on cluster purity
r
library(SoupX)
sc <- load10X('cellranger_outs/')          # needs BOTH raw and filtered
sc <- autoEstCont(sc)                       # estimates contamination fraction rho
counts_adj <- adjustCounts(sc, roundToInt = TRUE)   # output is non-integer by default; round for NB models

SoupX and CellBender disagree on what "ambient" is: SoupX subtracts a per-cell scalar of a single global soup profile; CellBender learns a probabilistic per-droplet background in a generative model. There is no consensus on which is better - CellBender is more powerful and more dangerous. Subtracting a shared soup vector from every cell can manufacture artificial negative correlations and zero out genes cells genuinely lacked, so validate. Matters most for solid tumors, snRNA-seq, and blood. Do not stack tools; double-correction compounds over-removal.

Normalization

Goal: Remove per-cell depth bias and stabilize variance so high-expression genes do not dominate PCA/kNN distances.

Approach: Default to shifted-log; reach for scran on shallow data and Pearson residuals on UMI count models; never normalize already-normalized data.

python
adata.layers['counts'] = adata.X.copy()                  # stash raw before normalizing (HVG/doublets need it)
sc.pp.normalize_total(adata)                             # target_sum=None scales each cell to the dataset MEDIAN; pass target_sum=1e4 for the historical, arbitrary CP10k
sc.pp.log1p(adata)                                       # natural-log(1+x); the variance-stabilizing transform (no target_sum argument)
r
seurat_obj <- NormalizeData(seurat_obj, normalization.method = 'LogNormalize', scale.factor = 10000)
# or variance-stabilized: seurat_obj <- SCTransform(seurat_obj, verbose = FALSE)
MethodModel / assumptionUse whenFails when
Shifted-log (CP10k / median)Size-factor + log1p; constant total mRNAGeneral default; strong, fast, defensibleComposition-divergent types (plasma, cycling) distort fold-changes
scran deconvolutionPooled size factors robust to compositionLow-depth, high-dropout, plate-basedR-only; needs pre-clustering (quickCluster); factors can go negative
sctransform v1/v2NB regularized regression (Pearson residuals)Seurat depth removal for HVG/vizSlow; off the count scale; v1 overfits theta (use v2)
Analytic Pearson residualsr=(x-mu)/sqrt(mu+mu^2/theta)UMI HVG+PCA without ad-hoc stepsExperimental; residual variance depends on theta and depth; clip to +/-sqrt(n)

Ahlmann-Eltze and Huber 2023 found plain shifted-log + PCA performs as well as or better than sctransform, Pearson residuals, and GLM-PCA on kNN-overlap recovery, so shifted-log is the defensible default and the sophisticated methods are "use if preferred," not mandated. Because methods genuinely compete here, verify current best practice against the installed tool's docs before committing. Normalize raw counts exactly once and keep the transform consistent across HVG, scaling, and PCA.

Show full SKILL.md (826 more words)Show less

Highly Variable Genes

Goal: Restrict PCA/clustering to genes carrying biological signal.

Approach: Select a flavor, then feed it the input type it expects - the single most consequential gotcha is that dispersion flavors want log-normalized data while seurat_v3 and Pearson want RAW COUNTS.

python
# seurat_v3 reads raw counts from a layer and REQUIRES n_top_genes; needs the scikit-misc package
sc.pp.highly_variable_genes(adata, n_top_genes=2000, flavor='seurat_v3', layer='counts')
FlavorFunctionInputn_top_genes required?Extra dependency
seurat (default)sc.pp.highly_variable_geneslog-normalizedNo-
cell_rangersc.pp.highly_variable_geneslog-normalizedNo-
seurat_v3sc.pp.highly_variable_genesRAW countsYesscikit-misc
pearson_residualssc.experimental.pp.highly_variable_genesRAW countsrecommended-

Running seurat_v3 on logged values, or seurat on raw counts, runs silently and yields garbage HVGs. The field is shifting toward binomial-deviance and Pearson-residual feature selection on raw counts because dispersion HVGs are sensitive to the upstream normalization choice. Set batch_key to compute HVGs per batch and avoid batch-specific technical genes.

Scaling and Regressing Out

Goal: Optionally equalize gene weight in PCA, and remove unwanted covariates - both now discouraged as reflexive defaults.

Approach: Prefer PCA on log-normalized HVG without scaling; regress out only a validated, non-confounded covariate.

python
sc.pp.scale(adata, max_value=10)        # max_value default is None (no clipping); 10 is an explicit choice to cap z-scores
# sc.pp.regress_out(adata, ['total_counts', 'pct_counts_mt'])   # scanpy itself warns this overcorrects

"Always regress out mito and total_counts" is folklore: those covariates are confounded with real cell identity and state (cycling cells legitimately have more RNA), so regressing them erases biology and can collapse data into a blob. Modern normalization already stabilizes depth; address unwanted variation with integration (Harmony, scVI) rather than linear regression. Scaling inflates lowly-expressed noisy genes; sc-best-practices runs PCA on the normalized layer directly.

Common Errors

SymptomCauseFix
HVGs look random; clustering is mushseurat_v3 fed log-normalized (or seurat fed raw)Feed each flavor its required input; use layer='counts' for seurat_v3
An entire healthy cell type disappearedFlat mito cutoff deleted high-mito parenchymaUse MAD/tissue-aware thresholds; inspect what was removed
Values inflated ~2x after re-running normalizationNormalized already-normalized dataNormalize raw once; restore from layers['counts']
ModuleNotFoundError: skmiscseurat_v3 needs scikit-miscpip install scikit-misc
QC metrics missing from .obscalculate_qc_metrics inplace defaults to FalsePass inplace=True
Proliferation / activation signal vanishedRegressed out total_counts / cell-cycle confounded with biologyDo not reflexively regress; validate the covariate is not confounded
New "stressed/transitional" clusterWarm-dissociation IEG/HSP artifactScore the dissociation module; exclude those genes from HVG/clustering
Off-target markers everywhere (Hb, Ig)Ambient RNA contaminationRun SoupX/CellBender/DecontX on the raw matrix before QC
Spike to ~2x counts deflated other genesCompositional see-saw from a few dominant genesUse exclude_highly_expressed=True or scran; report relative, not absolute, expression
Almost all cells filtered / tiny survivor countMAD ~ 0 on a low-variance, tiny, or nuclei sample (>50% share a value), so is_outlier flags every non-median cellAssert n_obs > 0 and a sane survival fraction; fall back to fixed cutoffs when MAD is ~0
Mito filter removes nothing (percent.mt all 0)'^MT-' pattern matched against Ensembl-ID feature namesUse gene symbols, or match the mito Ensembl IDs / a mito gene list
Shallow batch over-filtered, deep batch under-filteredGlobal MAD thresholds across samples of differing depthCompute is_outlier per batch_key/sample group
Batch effects baked into QC/soup/doublet callsMerged samples before QC, ambient, and doublet stepsRun steps 1-5 per sample, then merge
  • single-cell/data-io - load the raw matrix before preprocessing
  • single-cell/doublet-detection - per-sample doublet calling around the QC step
  • single-cell/clustering - PCA, neighbors, and clustering after preprocessing
  • single-cell/batch-integration - correct batch effects instead of regressing them out
  • single-cell/markers-annotation - find markers after clustering
  • differential-expression/deseq2-basics - pseudobulk DE across samples (avoids single-cell pseudo-replication)

References

  • Heumos L, Schaar AC, Lance C, et al. (2023) Best practices for single-cell analysis across modalities. Nature Reviews Genetics 24:550-572. DOI 10.1038/s41576-023-00586-w
  • Ahlmann-Eltze C, Huber W (2023) Comparison of transformations for single-cell RNA-seq data. Nature Methods 20:665-672. DOI 10.1038/s41592-023-01814-1
  • Lun ATL, Bach K, Marioni JC (2016) Pooling across cells to normalize single-cell RNA sequencing data (scran). Genome Biology 17:75. DOI 10.1186/s13059-016-0947-7
  • Hafemeister C, Satija R (2019) Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression (sctransform). Genome Biology 20:296. DOI 10.1186/s13059-019-1874-1
  • Choudhary S, Satija R (2022) Comparison and evaluation of statistical error models for scRNA-seq (sctransform v2). Genome Biology 23:27. DOI 10.1186/s13059-021-02584-9
  • Lause J, Berens P, Kobak D (2021) Analytic Pearson residuals for normalization of single-cell RNA-seq UMI data. Genome Biology 22:258. DOI 10.1186/s13059-021-02451-7
  • Vallejos CA, Risso D, Scialdone A, Dudoit S, Marioni JC (2017) Normalizing single-cell RNA sequencing data: challenges and opportunities. Nature Methods 14(6):565-571. DOI 10.1038/nmeth.4292
  • Osorio D, Cai JJ (2021) Systematic determination of the mitochondrial proportion in human and mouse tissues for scRNA-seq quality control. Bioinformatics 37(7):963-967. DOI 10.1093/bioinformatics/btaa751
  • Hippen AA, Falco MM, Weber LM, et al. (2021) miQC: An adaptive probabilistic framework for quality control of single-cell RNA-seq data. PLoS Computational Biology 17(8):e1009290. DOI 10.1371/journal.pcbi.1009290
  • Young MD, Behjati S (2020) SoupX removes ambient RNA contamination from droplet-based single-cell RNA sequencing data. GigaScience 9(12):giaa151. DOI 10.1093/gigascience/giaa151
  • Fleming SJ, Chaffin MD, Arduini A, et al. (2023) Unsupervised removal of systematic background noise (CellBender remove-background). Nature Methods 20:1323-1335. DOI 10.1038/s41592-023-01943-7
  • van den Brink SC, Sage F, Vertesy A, et al. (2017) Single-cell sequencing reveals dissociation-induced gene expression in tissue subpopulations. Nature Methods 14(10):935-936. DOI 10.1038/nmeth.4437
  • Squair JW, Gautier M, Kathe C, et al. (2021) Confronting false discoveries in single-cell differential expression. Nature Communications 12:5692. DOI 10.1038/s41467-021-25960-2

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in single-cell/preprocessing of GPTomics/bioSkills.

  • SKILL.md
  • examples/preprocess_scanpy.py
  • examples/preprocess_seurat.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Single Cell Preprocessing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Single Cell Preprocessing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Single Cell Preprocessing this skillGPTomics/bioSkills1.2k1 repos~5.1kAutomated safety check: PassMIT
Anndatadavila7/claude-code-templates32k12 repos~2.5kAutomated safety check: PassMIT
ScanpyK-Dense-AI/scientific-agent-skills48k1 repos~5.1kAutomated safety check: PassBSD-3-Clause
AnndataK-Dense-AI/scientific-agent-skills48k1 repos~3.9kAutomated safety check: NotesBSD-3-Clause
Bio Single Cell Data IoFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2kAutomated safety check: PassNone
Bio Single Cell PreprocessingFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2.4kAutomated safety check: PassNone

Similar skills

  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    32k GitHub starsUsed in 12 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Scanpy

    K-Dense-AI/scientific-agent-skills

    Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…

    48k GitHub starsUsed in 1 repo~5.1k tokens
    Research & ScienceAuto-check passed
  • Anndata

    K-Dense-AI/scientific-agent-skills

    Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.

    48k GitHub starsUsed in 1 repo~3.9k tokens
    Research & ScienceAuto-check: notes
  • Bio Single Cell Data Io

    FreedomIntelligence/OpenClaw-Medical-Skills

    Read, write, and create single-cell data objects using Seurat (R) and Scanpy (Python).

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Bio Single Cell Preprocessing

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control, filtering, and normalization for single-cell RNA-seq using Seurat (R) and Scanpy (Python).

    3.1k GitHub starsUsed in 1 repo~2.4k tokens
    Research & ScienceAuto-check passed
  • Biopython

    foryourhealth111-pixel/Vibe-Skills

    Primary retained Python toolkit for molecular biology sequence work.

    3.6k GitHub stars~3.5k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Single Cell Preprocessing

What does Bio Single Cell Preprocessing do?

Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R). Bio Single Cell Preprocessing is an agent skill from GPTomics/bioSkills. Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R).

When should I use Bio Single Cell Preprocessing?

Bio Single Cell Preprocessing fits situations like: filtering low-quality cells with MAD-adaptive thresholds; setting tissue-aware mito cutoffs; removing ambient RNA (SoupX/CellBender/DecontX); choosing a normalization (shifted-log vs scran vs sctransform vs Pearson residuals).

How do I install Bio Single Cell Preprocessing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a claude-code`. Or copy the skill folder (single-cell/preprocessing in GPTomics/bioSkills) into .claude/skills/bio-single-cell-preprocessing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Single Cell Preprocessing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a codex`. Or copy the skill folder (single-cell/preprocessing in GPTomics/bioSkills) into .agents/skills/bio-single-cell-preprocessing in your project. Codex loads it when a task matches its description.

Can I use Bio Single Cell Preprocessing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-single-cell-preprocessing, .gemini/skills/bio-single-cell-preprocessing, .github/skills/bio-single-cell-preprocessing and .opencode/skills/bio-single-cell-preprocessing in your project.

What does Bio Single Cell Preprocessing need to run?

Going by SKILL.md and its folder, Bio Single Cell Preprocessing needs Python and R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Single Cell Preprocessing access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Single Cell Preprocessing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Single Cell Preprocessing use?

Bio Single Cell Preprocessing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Single Cell Preprocessing use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Single Cell Preprocessing?

Skills that share tags, products or a category with Bio Single Cell Preprocessing: Anndata (davila7/claude-code-templates, 32k stars), Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars), Anndata (K-Dense-AI/scientific-agent-skills, 48k stars) and Bio Single Cell Data Io (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Single Cell Preprocessing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.