Scanpy Single-Cell Analysis
davila7/claude-code-templates
Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.
Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy.
$ npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-spatial-transcriptomics-spatial-preprocessing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/spatial-transcriptomics/spatial-preprocessing .claude/skills/bio-spatial-transcriptomics-spatial-preprocessing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-spatial-transcriptomics-spatial-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/spatial-transcriptomics/spatial-preprocessing into .claude/skills/bio-spatial-transcriptomics-spatial-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-spatial-transcriptomics-spatial-preprocessing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/spatial-transcriptomics/spatial-preprocessingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-spatial-transcriptomics-spatial-preprocessing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/spatial-transcriptomics/spatial-preprocessing .agents/skills/bio-spatial-transcriptomics-spatial-preprocessing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-spatial-transcriptomics-spatial-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/spatial-transcriptomics/spatial-preprocessing into .agents/skills/bio-spatial-transcriptomics-spatial-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-spatial-transcriptomics-spatial-preprocessing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-spatial-transcriptomics-spatial-preprocessing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/spatial-transcriptomics/spatial-preprocessing .cursor/skills/bio-spatial-transcriptomics-spatial-preprocessing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-spatial-transcriptomics-spatial-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/spatial-transcriptomics/spatial-preprocessing into .cursor/skills/bio-spatial-transcriptomics-spatial-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-spatial-transcriptomics-spatial-preprocessing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path spatial-transcriptomics/spatial-preprocessing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-spatial-transcriptomics-spatial-preprocessing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/spatial-transcriptomics/spatial-preprocessing .gemini/skills/bio-spatial-transcriptomics-spatial-preprocessing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-spatial-transcriptomics-spatial-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/spatial-transcriptomics/spatial-preprocessing into .gemini/skills/bio-spatial-transcriptomics-spatial-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-spatial-transcriptomics-spatial-preprocessing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-spatial-transcriptomics-spatial-preprocessingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/spatial-transcriptomics/spatial-preprocessing .github/skills/bio-spatial-transcriptomics-spatial-preprocessing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-spatial-transcriptomics-spatial-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/spatial-transcriptomics/spatial-preprocessing into .github/skills/bio-spatial-transcriptomics-spatial-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-spatial-transcriptomics-spatial-preprocessing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-spatial-transcriptomics-spatial-preprocessing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/spatial-transcriptomics/spatial-preprocessing .opencode/skills/bio-spatial-transcriptomics-spatial-preprocessing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-spatial-transcriptomics-spatial-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/spatial-transcriptomics/spatial-preprocessing into .opencode/skills/bio-spatial-transcriptomics-spatial-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-spatial-transcriptomics-spatial-preprocessing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-spatial-transcriptomics-spatial-preprocessingQuality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy.
Bio Spatial Transcriptomics Spatial Preprocessing is an agent skill from GPTomics/bioSkills. Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy. Use when setting QC floors that do NOT delete real low-count imaging cells (an scRNA mincounts=500 floor deletes nearly every Xenium cell, whose vector is tens-to-low-hundreds of transcripts); deciding whether to normalize at all when library size carries spatial biology rather than pure technical depth; choosing cell-volume/area normalization over…
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/preprocess_spatial.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. It works with Scanpy. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Spatial Transcriptomics Spatial Preprocessing loads about 4.7k tokens when it runs. Until then it costs about 188 tokens; SKILL.md has 1,956 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,956 words, ~4,715 tokens.
.claude/skills/bio-spatial-transcriptomics-spatial-preprocessing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: squidpy 1.5+, scanpy 1.10+, anndata 0.10+, spatialdata 0.2+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturesIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"QC and normalize my spatial data" -> Flag and remove low-quality spots/cells, then put counts on a scale fit for downstream domain and marker analysis -- but the right QC floors and the right normalization both depend on which side of the platform fork the data sits.
The first question on any spatial dataset is which assay family produced it, because it changes every QC threshold and the entire normalization decision.
| Axis | Sequencing/spot (Visium, Slide-seq, Stereo-seq) | Imaging/in-situ (Xenium, MERSCOPE, CosMx) |
|---|---|---|
| Unit | spot/bin = 1-10-cell mixture | segmentation-derived single cell |
| Counts/unit | hundreds-thousands UMI | tens-low hundreds transcripts |
| Gene space | whole-transcriptome (poly-A) or probe panel | TARGETED panel (100-1000), genes/cell ceilinged at panel size |
| Mito-% QC | available (mixed-cell average) | usually impossible (mito off-panel) |
| Specificity metric | none native | negative-control-probe / blank-barcode FDR |
| Library-size meaning | confounds cells-per-spot + cellularity | confounds cell SIZE/AREA + segmentation error |
In single-cell RNA-seq library size is a technical nuisance to divide out. In spatial transcriptomics LIBRARY SIZE CARRIES BIOLOGY, and that single fact governs both QC and normalization. On Visium, total counts per spot are spatially structured and correlated with anatomy because they confound with the number of cells per spot and tissue cellularity (Bhuva 2024 Genome Biol 25:99). On imaging platforms, total counts per cell confound with cell SIZE/AREA -- a physically larger segmented cell holds more molecules for purely geometric reasons -- and with segmentation error itself. Naively dividing library size out (CP10k, log1p, scran pooling) therefore removes real spatially-structured biology and measurably degrades spatial-domain detection.
For imaging the bias is UPSTREAM of any residual model. Because a targeted panel is small, hand-curated, and skewed toward a few high markers, any gene-count-based size factor is dominated by a handful of genes and becomes panel-composition-dependent. Atta and Fan 2024 (Genome Biol 25:153) compared library-size, Pearson/SCTransform, DESeq2, TMM, and volume/area normalization on skewed panels: the four gene-count methods inject region-specific bias of up to ~13% DE error and fold-change SIGN REVERSAL in up to 19% of genes, while volume/area normalization avoids it because its denominator is independent of panel composition. The load-bearing consequence: Pearson residuals and SCTransform, the gold standard for whole-transcriptome scRNA, do NOT rescue a skewed imaging panel -- they are still gene-count-based and the bias sits in the panel design. The fix is non-gene-count normalization (cell volume/area, Moffitt-style) or spatially-aware joint modeling (SpaNorm).
A clean violin plot proves nothing. A count floor that silently deletes a spatial cluster of small cells (lymphocytes), a normalization that erases a cellularity gradient, and a focal hybridization failure all survive a tidy genes-per-cell distribution. The dangerous artifacts are spatial, so QC must be inspected spatially.
Carrying scRNA defaults onto imaging data deletes the data. Imaging cells carry tens-to-low-hundreds of transcripts -- one to two orders of magnitude below droplet scRNA -- so an min_counts=500 floor removes nearly every real cell. Genes/cell can never exceed the panel size, so the "high genes = doublet" heuristic is meaningless (the ceiling is the panel, not a doublet). Mito genes are usually off-panel, so pct_counts_mt QC is often impossible. Worst, an aggressive count floor preferentially deletes the smallest REAL cells (lymphocytes, neutrophils), biasing tissue composition rather than removing noise.
| Metric | Visium (spot) | Imaging (Xenium/MERSCOPE/CosMx) | Rationale / trap |
|---|---|---|---|
| counts/unit | UMI/spot; no universal floor; OSTA DLPFC flags <600 | transcripts/cell tens-low hundreds; community floor ~10 (Squidpy), some 20 | scRNA min_counts=500 deletes nearly every imaging cell; floor confounds cellularity/cell-size |
| genes/unit | genes/spot; OSTA flags <400 | genes/cell CEILING = panel size (100-1000) | "high genes = doublet" meaningless for imaging |
| mito-% | mixed-cell average; OSTA flags >0.28 (brain) | usually off-panel -> impossible | tissue-dependent; brain tolerates higher |
| cell area (um^2) | n/a (spot is fixed) | MAD-based on counts/area | flags over/under-segmentation; no fixed vendor min |
| negative-control FDR | n/a | THE imaging specificity metric | false-discovery proxy; no scRNA analogue |
| cells/spot | nuclei estimate; OSTA flags >10 | n/a | confirms spot is a mixture |
Thresholds are tissue-dependent and "somewhat arbitrary" (the OSTA Visium worked example flags UMI<600, genes<400, mito>0.28, cells/spot>10 on DLPFC, removing 32/3639 spots -- a starting point, not a law). The imaging floor of ~10 transcripts/cell is a Squidpy/community convention, NOT a vendor specification; community CosMx floors run higher (commonly ~20 counts/cell, scaling up with plex), so confirm the cutoff against the panel and tissue rather than copying a number.
Goal: Annotate negative controls and mito genes (where present), compute QC metrics, and set platform-appropriate floors without deleting real low-count cells.
Approach: Branch on the fork. For imaging, identify control-probe prefixes, filter on a low transcript floor and cell area; for spot data, use UMI/genes/mito floors. Always compute metrics, then look at them spatially before cutting.
import squidpy as sq
import scanpy as sc
import numpy as np
# Imaging branch: control features carry platform-specific prefixes -- they are the specificity ruler, not genes
ctrl_prefixes = ('NegControlProbe', 'NegControlCodeword', 'BLANK', 'Blank', 'NegPrb') # Xenium / MERFISH / CosMx
adata.var['control'] = adata.var_names.str.startswith(ctrl_prefixes)
adata.var['mt'] = adata.var_names.str.startswith(('MT-', 'mt-')) # usually empty on imaging panels
sc.pp.calculate_qc_metrics(adata, qc_vars=['control', 'mt'], percent_top=None, inplace=True)
# inplace defaults to False and returns DataFrames; pass inplace=True to write .obs/.varImaging platforms include features that decode to nothing biological: Xenium negative-control PROBES (off-target binding) plus negative-control CODEWORDS (pure optical/decoding error), MERFISH/MERSCOPE blank barcodes (valid codewords with no probe), CosMx NegPrb (alien synthetic sequences). They are the only native false-discovery proxy in spatial data. The canonical metric is FDR = mean counts per control feature / mean counts per real gene; the community-acceptable band is roughly <=1-5% of signal (Xenium typically <0.1%, MERFISH ~4%, CosMx highest). Compute it before trusting any gene-level claim, and treat controls as a panel-wide QC gate -- not as genes to cluster on.
Goal: Quantify the per-feature false-discovery rate and drop controls before normalization and clustering.
Approach: Average per-feature counts within the control set and within real genes, take the ratio, then subset the matrix to real genes only.
ctrl = adata.var['control'].values
mean_ctrl = np.asarray(adata[:, ctrl].X.sum(axis=0)).ravel().mean() if ctrl.any() else 0.0
mean_gene = np.asarray(adata[:, ~ctrl].X.sum(axis=0)).ravel().mean()
fdr = mean_ctrl / mean_gene if mean_gene else float('nan')
print(f'negative-control FDR: {fdr:.4f} (band ~<=0.01-0.05)')
adata = adata[:, ~ctrl].copy() # controls are a QC ruler, never clustering featuresA QC gradient across the section -- counts falling toward one edge, mito rising in a corner -- is a technical artifact (edge effects, uneven permeabilization, focal hybridization failure), not biology, and a violin plot hides it. Always map QC onto tissue coordinates before choosing thresholds, and check WHERE the cells slated for removal actually fall.
Goal: Reveal spatially-structured quality artifacts and confirm a proposed floor is not removing a coherent tissue region.
Approach: Color the spatial scatter by each QC metric; a smooth spatial gradient signals a technical artifact to address (or model) rather than threshold away.
sq.pl.spatial_scatter(adata, color=['total_counts', 'n_genes_by_counts'], shape=None, ncols=2)
# shape=None renders points (imaging/Slide-seq); omit it for Visium hex spots with a tissue imageGoal: Apply platform-appropriate floors that remove debris and segmentation failures without biasing composition.
Approach: Imaging -- low transcript floor plus a cell-area sanity bound. Spot -- UMI/genes/mito floors. Either way, filter genes seen in too few units last.
# Imaging floor: ~10 transcripts/cell is a Squidpy/community convention, NOT a vendor spec
sc.pp.filter_cells(adata, min_counts=10)
if 'cell_area' in adata.obs:
lo, hi = adata.obs['cell_area'].quantile([0.01, 0.99]) # trim segmentation over/under-calls, tissue-dependent
adata = adata[(adata.obs['cell_area'] > lo) & (adata.obs['cell_area'] < hi)].copy()
sc.pp.filter_genes(adata, min_cells=5)
# Spot branch instead (Visium): tissue-dependent floors -- the OSTA DLPFC example, not universal law
# sc.pp.filter_cells(adata, min_counts=600)
# sc.pp.filter_cells(adata, min_genes=400)
# adata = adata[adata.obs['pct_counts_mt'] < 28].copy()Do not reach reflexively for normalize_total + log1p. The shipped Squidpy tutorials run it for both Xenium and MERFISH, so it is the de-facto default -- and it is exactly what the benchmark papers argue is biased for spatial data. Decide deliberately from the table, and because methods compete here, verify current best practice against the installed tool's docs and the latest benchmarks before committing.
| Method | Assumption | Best when | Fails when |
|---|---|---|---|
normalize_total + log1p | library size = pure technical nuisance | cross-platform comparability; quick default; tool tutorials | spatial -- removes spatially-structured biology (Bhuva 2024); imaging skewed panel |
| Analytic Pearson residuals | closed-form NB offset; depth as fixed offset | whole-transcriptome Visium HVG/PCA | imaging targeted panel -- still gene-count-based, inherits panel-skew bias (Atta/Fan 2024) |
| SCTransform v2 | regularized NB GLM, depth slope fixed | whole-transcriptome UMI / Visium | does NOT fix skewed imaging panels (gene-count-based) |
| Cell volume/area | concentration is the biological quantity | IMAGING skewed panel (Moffitt-style) | denominator needs reliable segmentation; cannot fix segmentation error itself |
| SpaNorm (spatially-aware) | library size and biology are entangled; remove only library-size component | spot AND imaging; preserve spatial structure | newer; R/Bioconductor |
Goal (spot, whole-transcriptome): Stabilize depth for HVG/PCA while keeping raw counts, accepting that crude library-size division can blur domains.
Approach: Stash raw counts, then either run the standard log1p pipeline knowingly or prefer analytic Pearson residuals for feature selection on whole-transcriptome Visium.
adata.layers['counts'] = adata.X.copy() # stash raw -- HVG flavors and re-normalization need it
sc.pp.normalize_total(adata) # target_sum=None scales to the dataset MEDIAN, not the arbitrary 1e4
sc.pp.log1p(adata) # library size carries biology -- this can blur spatial domains
# Whole-transcriptome Visium feature selection alternative (gene-count-based, fine here, NOT for imaging panels):
# sc.experimental.pp.normalize_pearson_residuals(adata)Goal (imaging, targeted panel): Normalize without injecting panel-composition bias, using a denominator independent of gene counts.
Approach: Divide each cell's counts by its segmented area/volume (copies per unit area), then log-transform -- the Moffitt-style fix benchmarks favour over Pearson residuals for skewed panels.
adata.layers['counts'] = adata.X.copy()
if 'cell_area' in adata.obs: # area/volume denominator is panel-composition-independent
sf = adata.obs['cell_area'].values / adata.obs['cell_area'].median()
adata.X = adata.X / sf[:, None]
sc.pp.log1p(adata)
# If no segmentation area is available, prefer SpaNorm (R) over reflexive normalize_total on a skewed panel| Symptom | Cause | Fix |
|---|---|---|
| Nearly all imaging cells filtered out | scRNA min_counts=500 floor on tens-of-transcript cells | Use a low floor (~10 transcripts/cell); branch QC on the platform fork |
| Spatial domains blur / merge after normalization | Crude library-size division erased a cellularity/anatomy gradient (library size carries biology) | Prefer SpaNorm or volume/area; if using log1p, know it can blur domains |
| Fold-change sign flips between normalizations | Gene-count size factor is panel-composition-dependent on a skewed imaging panel | Use non-gene-count (cell area/volume) normalization; Pearson/SCT do NOT fix it |
pct_counts_mt is all zero / NaN | Mito genes are not on the targeted imaging panel | Skip mito-% QC for imaging; QC on transcript floor + cell area instead |
| "Doublet" cells flagged by high gene count | genes/cell ceiling IS the panel size -- not a doublet signal | Drop the high-genes heuristic for imaging; use cell area / spatial doublets |
| Control features cluster as their own group | Negative-control probes/codewords left in the matrix | Compute control FDR, then subset to real genes before clustering |
| Smallest cell type vanished after filtering | A count floor preferentially deleted small real cells (lymphocytes) | Inspect spatially where cuts fall; lower the floor; check composition before/after |
| Claimed a novel cell-state signature from imaging marker genes | An imaging panel (even 5,000-plex) is pre-selected for KNOWN biology; off-panel genes are absent by design, not by expression | Treat absence of an off-panel gene as uninformative; de-novo state discovery is bounded by the panel -- corroborate on whole-transcriptome data before claiming novelty |
| QC looks fine in violins but a region is empty | Spatial QC gradient (edge/permeabilization artifact) invisible in violins | Map QC onto tissue with sq.pl.spatial_scatter before thresholding |
| Counts inflated ~2x after re-running normalization | Normalized already-normalized data | Normalize raw once; restore from layers['counts'] |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in spatial-transcriptomics/spatial-preprocessing of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Spatial Transcriptomics Spatial Preprocessing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Spatial Transcriptomics Spatial Preprocessing this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Scanpy Single-Cell Analysisdavila7/claude-code-templates | 32k | 16 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Single Cell Rna AnalysisPKU-YuanGroup/OpenAI4S | 608 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Anndatadavila7/claude-code-templates | 32k | 12 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Cellxgene Censusdavila7/claude-code-templates | 32k | 11 repos | ~3.8k | Automated safety check: Pass | MIT | |
| Bulk RnaseqK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~4.2k | Automated safety check: Pass | MIT |
davila7/claude-code-templates
Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.
PKU-YuanGroup/OpenAI4S
Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…
davila7/claude-code-templates
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…
davila7/claude-code-templates
Query CZ CELLxGENE Census (61M+ cells). An agent skill from davila7/claude-code-templates.
K-Dense-AI/scientific-agent-skills
Prepares bulk RNA-seq FASTQ, Salmon, STAR or featureCounts output for gene-level differential expression.
K-Dense-AI/scientific-agent-skills
Performs pathway and gene-set enrichment analysis on gene lists or ranked gene data and interprets the results.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Works with
Categories
Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy. Bio Spatial Transcriptomics Spatial Preprocessing is an agent skill from GPTomics/bioSkills. Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy.
Bio Spatial Transcriptomics Spatial Preprocessing fits situations like: setting QC floors that do NOT delete real low-count imaging cells (an scRNA mincounts=500 floor deletes nearly every Xenium cell; whose vector is tens-to-low-hundreds of transcripts); deciding whether to normalize at all when library size carries spatial biology rather than pure technical depth; choosing cell-volume/area normalization over Pearson residuals for skewed targeted panels.
Run `npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a claude-code`. Or copy the skill folder (spatial-transcriptomics/spatial-preprocessing in GPTomics/bioSkills) into .claude/skills/bio-spatial-transcriptomics-spatial-preprocessing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a codex`. Or copy the skill folder (spatial-transcriptomics/spatial-preprocessing in GPTomics/bioSkills) into .agents/skills/bio-spatial-transcriptomics-spatial-preprocessing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-spatial-transcriptomics-spatial-preprocessing, .gemini/skills/bio-spatial-transcriptomics-spatial-preprocessing, .github/skills/bio-spatial-transcriptomics-spatial-preprocessing and .opencode/skills/bio-spatial-transcriptomics-spatial-preprocessing in your project.
Going by SKILL.md and its folder, Bio Spatial Transcriptomics Spatial Preprocessing needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Spatial Transcriptomics Spatial Preprocessing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Spatial Transcriptomics Spatial Preprocessing: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), Single Cell Rna Analysis (PKU-YuanGroup/OpenAI4S, 608 stars), Anndata (davila7/claude-code-templates, 32k stars) and Cellxgene Census (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.