Anndata
davila7/claude-code-templates
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…
Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R).
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-preprocessing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/single-cell/preprocessing .claude/skills/bio-single-cell-preprocessing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-single-cell-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/preprocessing into .claude/skills/bio-single-cell-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-preprocessing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/single-cell/preprocessingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-preprocessing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/single-cell/preprocessing .agents/skills/bio-single-cell-preprocessing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-single-cell-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/preprocessing into .agents/skills/bio-single-cell-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-preprocessing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-preprocessing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/single-cell/preprocessing .cursor/skills/bio-single-cell-preprocessing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-single-cell-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/preprocessing into .cursor/skills/bio-single-cell-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-preprocessing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path single-cell/preprocessing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-preprocessing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/single-cell/preprocessing .gemini/skills/bio-single-cell-preprocessing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-single-cell-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/preprocessing into .gemini/skills/bio-single-cell-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-preprocessing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-single-cell-preprocessingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/single-cell/preprocessing .github/skills/bio-single-cell-preprocessing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-single-cell-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/preprocessing into .github/skills/bio-single-cell-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-preprocessing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-preprocessing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/single-cell/preprocessing .opencode/skills/bio-single-cell-preprocessing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-single-cell-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/preprocessing into .opencode/skills/bio-single-cell-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-preprocessing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-single-cell-preprocessingQuality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R).
Bio Single Cell Preprocessing is an agent skill from GPTomics/bioSkills. Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R). Use when filtering low-quality cells with MAD-adaptive thresholds, setting tissue-aware mito cutoffs, removing ambient RNA (SoupX/CellBender/DecontX), choosing a normalization (shifted-log vs scran vs sctransform vs Pearson residuals), selecting highly variable genes, or deciding whether to scale and regress out covariates.
Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/preprocess_scanpy.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. It works with Python and Scanpy. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
9 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python and R), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Single Cell Preprocessing loads about 5.1k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 2,031 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,031 words, ~5,090 tokens.
.claude/skills/bio-single-cell-preprocessing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: scanpy 1.10+, Seurat 5.0+, scran 1.30+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturespackageVersion('<pkg>') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Preprocess my scRNA-seq data" -> Remove bad barcodes, correct technical biases, and select informative genes before dimensionality reduction.
calculate_qc_metrics() -> filter -> normalize_total()+log1p() -> highly_variable_genes()NormalizeData() or SCTransform() -> FindVariableFeatures() -> ScaleData()Every preprocessing choice propagates to every downstream result, and the two highest-leverage choices both encode a hidden biological assumption.
Normalization assumes near-constant total mRNA per cell. Shifted-log CP10k, scran deconvolution, and sctransform all divide by a per-cell size factor meant to capture only capture-efficiency and sequencing depth. They silently assume total transcriptome size is roughly constant across cell types, so a cell's total UMI count is a pure technical nuisance. This is false for plasma/antibody-secreting cells, secretory epithelia, hepatocytes, large neurons, and S/G2M cells, which carry 2-10x more mRNA. Dividing them to a common total deflates every gene that is not one of their few dominant transcripts (a compositional see-saw), and partially erases real proliferation biology. The honest framing: single-cell measures proportions, not absolute amounts.
QC metrics are biology metrics in disguise. pct_counts_mt conflates apoptosis, genuine metabolic activity (cardiomyocytes/hepatocytes/muscle are constitutively 20-40% mito and healthy), dissociation stress, and technical contamination. A flat global mito cutoff deletes entire healthy parenchymal populations and the survivors still cluster cleanly, so the loss is invisible. Use adaptive, tissue-aware thresholds and treat all three QC covariates jointly.
A beautiful UMAP proves nothing. Compositional normalization bias, deleted high-mito parenchyma, dissociation-stress clusters, ambient-induced co-expression, and residual homotypic doublets are all compatible with tidy clusters. The dangerous artifacts are precisely the ones that do not look like artifacts.
min_cells).Ambient correction needs the raw matrix and must precede QC filtering; once subset to cells, the soup estimate is gone. Doublet detection runs on raw counts, so stash counts before normalizing.
Steps 1-5 are per-sample operations performed BEFORE merge or integration: empty-droplet calling, ambient removal (SoupX load10X is inherently per-run), adaptive QC, and doublet detection all reason about one capture's droplet population, so QC-then-merge is correct and merge-then-QC leaks batch effects into every threshold and contaminates the soup and doublet-scoring neighborhoods. Merge only after each sample is cleaned.
Goal: Remove empty/dying/stressed barcodes using data-driven thresholds that do not delete real cell types.
Approach: Annotate mito/ribo/hemoglobin gene sets, compute joint QC metrics, then flag outliers by median absolute deviation (MAD) on the log scale rather than fixed cutoffs.
import scanpy as sc
import numpy as np
from scipy.stats import median_abs_deviation
adata.var['mt'] = adata.var_names.str.startswith('MT-') # mouse: 'mt-'
adata.var['ribo'] = adata.var_names.str.startswith(('RPS', 'RPL'))
adata.var['hb'] = adata.var_names.str.contains(r'^HB[ABDEGMQZ]\d*(?!\w)') # explicit subunits, not legacy ^HB[^(P)]
sc.pp.calculate_qc_metrics(adata, qc_vars=['mt', 'ribo', 'hb'], percent_top=[20], log1p=True, inplace=True)
# inplace defaults to False and returns DataFrames; pass inplace=True to write .obs/.var
def is_outlier(adata, metric, nmads):
M = adata.obs[metric]
return (M < np.median(M) - nmads * median_abs_deviation(M)) | (np.median(M) + nmads * median_abs_deviation(M) < M)
adata.obs['outlier'] = (is_outlier(adata, 'log1p_total_counts', 5) | is_outlier(adata, 'log1p_n_genes_by_counts', 5)
| is_outlier(adata, 'pct_counts_in_top_20_genes', 5))
adata.obs['mt_outlier'] = is_outlier(adata, 'pct_counts_mt', 3) | (adata.obs['pct_counts_mt'] > 8)
adata = adata[~(adata.obs['outlier'] | adata.obs['mt_outlier'])].copy()
sc.pp.filter_genes(adata, min_cells=3)When samples differ in depth/quality or were sequenced in separate batches, compute MAD thresholds PER SAMPLE, not globally: a single global MAD over-cuts the shallow batch and under-cuts the deep one. Apply is_outlier within each batch_key group (the same per-batch logic the HVG step uses).
flags = ['log1p_total_counts', 'log1p_n_genes_by_counts', 'pct_counts_in_top_20_genes']
adata.obs['outlier'] = adata.obs.groupby('sample', observed=True).apply(
lambda g: (is_outlier(adata[g.index], flags[0], 5) | is_outlier(adata[g.index], flags[1], 5)
| is_outlier(adata[g.index], flags[2], 5))).droplevel(0)# '^MT-' matches gene SYMBOLS; with Ensembl-ID feature names it matches nothing and the mito filter silently does nothing
seurat_obj[['percent.mt']] <- PercentageFeatureSet(seurat_obj, pattern = '^MT-')
VlnPlot(seurat_obj, features = c('nFeature_RNA', 'nCount_RNA', 'percent.mt'), ncol = 3)
seurat_obj <- subset(seurat_obj, subset = nFeature_RNA > 200 & nFeature_RNA < 5000 & percent.mt < 20)| Metric | Reference value | Rationale and caveat |
|---|---|---|
min_genes | 200 | Below this is mostly empty droplets / debris; raise for deep data |
log1p_total_counts / log1p_n_genes_by_counts | 5 MAD | sc-best-practices loosens from scater's 3 MAD to avoid cutting real biology; filter on the log scale (depth is right-skewed) |
pct_counts_in_top_20_genes | 5 MAD | High value flags low-complexity / dying cells |
pct_counts_mt | 3 MAD plus hard >8% | Tissue-dependent: 5-20% typical, but cardiomyocytes/hepatocytes/muscle are constitutively high; nuclei are ~0-2% and any mito flags ambient |
min_cells (genes) | 3 | Remove genes seen in too few cells to be informative |
Fixed cutoffs are a fast first pass for well-characterized tissue but silently delete valid populations; MAD-adaptive is the modern default; miQC (a mito-vs-detected-genes mixture model) helps when that relationship varies across samples.
High mito is ambiguous: apoptosis co-occurs with low gene counts and apoptotic markers, while warm-dissociation stress co-occurs with immediate-early genes (FOS, JUN, JUNB, EGR1) and heat-shock proteins (HSPA1A/B) at normal gene counts. The IEG/HSP program creates a spurious "activated/stressed" cluster that passes every count/mito filter, is cell-type-specific in magnitude (so it does not cancel as a uniform batch effect), and overlaps real immune/stem activation, so naive removal can itself delete biology. Score the dissociation module per cell, then exclude those genes from HVG/clustering or flag and interpret cautiously; cold-protease digestion and single-nucleus assays reduce the artifact. For nuclei, standard mito thresholds are meaningless (baseline near zero) - lean on counts/genes outliers.
Goal: Remove cell-free "soup" mRNA that inflates off-target markers (hemoglobin everywhere in PBMCs, hepatocyte genes in non-hepatocytes) and fabricates co-expression.
Approach: Estimate the soup profile and a per-cell contamination fraction, then subtract; pick ONE tool and validate that a known-specific marker survives.
| Tool | Input | Needs empty droplets? | Strength | Fails / risk |
|---|---|---|---|---|
| SoupX (R) | Cell Ranger raw+filtered | Yes | Fast, interpretable rho, auto-estimate | autoEstCont fails on homogeneous data; single global soup wrong when ambient is heterogeneous |
| CellBender (Python, GPU) | RAW h5 | Yes (core of model) | Deep generative; removes ambient + barcode noise; also does cell-calling; strong on nuclei | Over-removes real low-abundance genes at high --fpr; black-box; slow |
| DecontX (R, celda) | Filtered cells | No | No raw needed; easy SCE/Seurat integration | Relies on cluster purity |
library(SoupX)
sc <- load10X('cellranger_outs/') # needs BOTH raw and filtered
sc <- autoEstCont(sc) # estimates contamination fraction rho
counts_adj <- adjustCounts(sc, roundToInt = TRUE) # output is non-integer by default; round for NB modelsSoupX and CellBender disagree on what "ambient" is: SoupX subtracts a per-cell scalar of a single global soup profile; CellBender learns a probabilistic per-droplet background in a generative model. There is no consensus on which is better - CellBender is more powerful and more dangerous. Subtracting a shared soup vector from every cell can manufacture artificial negative correlations and zero out genes cells genuinely lacked, so validate. Matters most for solid tumors, snRNA-seq, and blood. Do not stack tools; double-correction compounds over-removal.
Goal: Remove per-cell depth bias and stabilize variance so high-expression genes do not dominate PCA/kNN distances.
Approach: Default to shifted-log; reach for scran on shallow data and Pearson residuals on UMI count models; never normalize already-normalized data.
adata.layers['counts'] = adata.X.copy() # stash raw before normalizing (HVG/doublets need it)
sc.pp.normalize_total(adata) # target_sum=None scales each cell to the dataset MEDIAN; pass target_sum=1e4 for the historical, arbitrary CP10k
sc.pp.log1p(adata) # natural-log(1+x); the variance-stabilizing transform (no target_sum argument)seurat_obj <- NormalizeData(seurat_obj, normalization.method = 'LogNormalize', scale.factor = 10000)
# or variance-stabilized: seurat_obj <- SCTransform(seurat_obj, verbose = FALSE)| Method | Model / assumption | Use when | Fails when |
|---|---|---|---|
| Shifted-log (CP10k / median) | Size-factor + log1p; constant total mRNA | General default; strong, fast, defensible | Composition-divergent types (plasma, cycling) distort fold-changes |
| scran deconvolution | Pooled size factors robust to composition | Low-depth, high-dropout, plate-based | R-only; needs pre-clustering (quickCluster); factors can go negative |
| sctransform v1/v2 | NB regularized regression (Pearson residuals) | Seurat depth removal for HVG/viz | Slow; off the count scale; v1 overfits theta (use v2) |
| Analytic Pearson residuals | r=(x-mu)/sqrt(mu+mu^2/theta) | UMI HVG+PCA without ad-hoc steps | Experimental; residual variance depends on theta and depth; clip to +/-sqrt(n) |
Ahlmann-Eltze and Huber 2023 found plain shifted-log + PCA performs as well as or better than sctransform, Pearson residuals, and GLM-PCA on kNN-overlap recovery, so shifted-log is the defensible default and the sophisticated methods are "use if preferred," not mandated. Because methods genuinely compete here, verify current best practice against the installed tool's docs before committing. Normalize raw counts exactly once and keep the transform consistent across HVG, scaling, and PCA.
Goal: Restrict PCA/clustering to genes carrying biological signal.
Approach: Select a flavor, then feed it the input type it expects - the single most consequential gotcha is that dispersion flavors want log-normalized data while seurat_v3 and Pearson want RAW COUNTS.
# seurat_v3 reads raw counts from a layer and REQUIRES n_top_genes; needs the scikit-misc package
sc.pp.highly_variable_genes(adata, n_top_genes=2000, flavor='seurat_v3', layer='counts')| Flavor | Function | Input | n_top_genes required? | Extra dependency |
|---|---|---|---|---|
seurat (default) | sc.pp.highly_variable_genes | log-normalized | No | - |
cell_ranger | sc.pp.highly_variable_genes | log-normalized | No | - |
seurat_v3 | sc.pp.highly_variable_genes | RAW counts | Yes | scikit-misc |
pearson_residuals | sc.experimental.pp.highly_variable_genes | RAW counts | recommended | - |
Running seurat_v3 on logged values, or seurat on raw counts, runs silently and yields garbage HVGs. The field is shifting toward binomial-deviance and Pearson-residual feature selection on raw counts because dispersion HVGs are sensitive to the upstream normalization choice. Set batch_key to compute HVGs per batch and avoid batch-specific technical genes.
Goal: Optionally equalize gene weight in PCA, and remove unwanted covariates - both now discouraged as reflexive defaults.
Approach: Prefer PCA on log-normalized HVG without scaling; regress out only a validated, non-confounded covariate.
sc.pp.scale(adata, max_value=10) # max_value default is None (no clipping); 10 is an explicit choice to cap z-scores
# sc.pp.regress_out(adata, ['total_counts', 'pct_counts_mt']) # scanpy itself warns this overcorrects"Always regress out mito and total_counts" is folklore: those covariates are confounded with real cell identity and state (cycling cells legitimately have more RNA), so regressing them erases biology and can collapse data into a blob. Modern normalization already stabilizes depth; address unwanted variation with integration (Harmony, scVI) rather than linear regression. Scaling inflates lowly-expressed noisy genes; sc-best-practices runs PCA on the normalized layer directly.
| Symptom | Cause | Fix |
|---|---|---|
| HVGs look random; clustering is mush | seurat_v3 fed log-normalized (or seurat fed raw) | Feed each flavor its required input; use layer='counts' for seurat_v3 |
| An entire healthy cell type disappeared | Flat mito cutoff deleted high-mito parenchyma | Use MAD/tissue-aware thresholds; inspect what was removed |
| Values inflated ~2x after re-running normalization | Normalized already-normalized data | Normalize raw once; restore from layers['counts'] |
ModuleNotFoundError: skmisc | seurat_v3 needs scikit-misc | pip install scikit-misc |
QC metrics missing from .obs | calculate_qc_metrics inplace defaults to False | Pass inplace=True |
| Proliferation / activation signal vanished | Regressed out total_counts / cell-cycle confounded with biology | Do not reflexively regress; validate the covariate is not confounded |
| New "stressed/transitional" cluster | Warm-dissociation IEG/HSP artifact | Score the dissociation module; exclude those genes from HVG/clustering |
| Off-target markers everywhere (Hb, Ig) | Ambient RNA contamination | Run SoupX/CellBender/DecontX on the raw matrix before QC |
| Spike to ~2x counts deflated other genes | Compositional see-saw from a few dominant genes | Use exclude_highly_expressed=True or scran; report relative, not absolute, expression |
| Almost all cells filtered / tiny survivor count | MAD ~ 0 on a low-variance, tiny, or nuclei sample (>50% share a value), so is_outlier flags every non-median cell | Assert n_obs > 0 and a sane survival fraction; fall back to fixed cutoffs when MAD is ~0 |
| Mito filter removes nothing (percent.mt all 0) | '^MT-' pattern matched against Ensembl-ID feature names | Use gene symbols, or match the mito Ensembl IDs / a mito gene list |
| Shallow batch over-filtered, deep batch under-filtered | Global MAD thresholds across samples of differing depth | Compute is_outlier per batch_key/sample group |
| Batch effects baked into QC/soup/doublet calls | Merged samples before QC, ambient, and doublet steps | Run steps 1-5 per sample, then merge |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in single-cell/preprocessing of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Single Cell Preprocessing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Single Cell Preprocessing this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.1k | Automated safety check: Pass | MIT | |
| Anndatadavila7/claude-code-templates | 32k | 12 repos | ~2.5k | Automated safety check: Pass | MIT | |
| ScanpyK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~5.1k | Automated safety check: Pass | BSD-3-Clause | |
| AnndataK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.9k | Automated safety check: Notes | BSD-3-Clause | |
| Bio Single Cell Data IoFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | 1 repos | ~2k | Automated safety check: Pass | None | |
| Bio Single Cell PreprocessingFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | 1 repos | ~2.4k | Automated safety check: Pass | None |
davila7/claude-code-templates
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…
K-Dense-AI/scientific-agent-skills
Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…
K-Dense-AI/scientific-agent-skills
Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.
FreedomIntelligence/OpenClaw-Medical-Skills
Read, write, and create single-cell data objects using Seurat (R) and Scanpy (Python).
FreedomIntelligence/OpenClaw-Medical-Skills
Quality control, filtering, and normalization for single-cell RNA-seq using Seurat (R) and Scanpy (Python).
foryourhealth111-pixel/Vibe-Skills
Primary retained Python toolkit for molecular biology sequence work.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Categories
Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R). Bio Single Cell Preprocessing is an agent skill from GPTomics/bioSkills. Quality control, ambient-RNA handling, normalization, and feature selection for single-cell RNA-seq using Scanpy (Python) and Seurat (R).
Bio Single Cell Preprocessing fits situations like: filtering low-quality cells with MAD-adaptive thresholds; setting tissue-aware mito cutoffs; removing ambient RNA (SoupX/CellBender/DecontX); choosing a normalization (shifted-log vs scran vs sctransform vs Pearson residuals).
Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a claude-code`. Or copy the skill folder (single-cell/preprocessing in GPTomics/bioSkills) into .claude/skills/bio-single-cell-preprocessing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a codex`. Or copy the skill folder (single-cell/preprocessing in GPTomics/bioSkills) into .agents/skills/bio-single-cell-preprocessing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-single-cell-preprocessing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-single-cell-preprocessing, .gemini/skills/bio-single-cell-preprocessing, .github/skills/bio-single-cell-preprocessing and .opencode/skills/bio-single-cell-preprocessing in your project.
Going by SKILL.md and its folder, Bio Single Cell Preprocessing needs Python and R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Single Cell Preprocessing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Single Cell Preprocessing: Anndata (davila7/claude-code-templates, 32k stars), Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars), Anndata (K-Dense-AI/scientific-agent-skills, 48k stars) and Bio Single Cell Data Io (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.