Agent skill

Bio Spatial Transcriptomics Spatial Preprocessing

by GPTomics in GPTomics/bioSkills

Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy.

MITAuto-check passedResearch & Science

Install Bio Spatial Transcriptomics Spatial Preprocessing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-spatial-transcriptomics-spatial-preprocessing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/spatial-transcriptomics/spatial-preprocessing .claude/skills/bio-spatial-transcriptomics-spatial-preprocessing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-spatial-transcriptomics-spatial-preprocessing
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.7k tokens
SKILL.md length
1,956 words
Files
3
Skills in repo
553
Repo updated
First seen
Licence
MIT

At a glance

Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy.

  • Setting QC floors that do NOT delete real low-count imaging cells (an scRNA mincounts=500 floor deletes nearly every Xenium cell
  • SKILL.md covers Version Compatibility, The platform-class fork…, Governing Principle and scRNA QC thresholds are wrong…, plus 6 more sections
  • Runs Python scripts from its folder; calls pip
  • Whose vector is tens-to-low-hundreds of transcripts)

What it does

Bio Spatial Transcriptomics Spatial Preprocessing is an agent skill from GPTomics/bioSkills. Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy. Use when setting QC floors that do NOT delete real low-count imaging cells (an scRNA mincounts=500 floor deletes nearly every Xenium cell, whose vector is tens-to-low-hundreds of transcripts); deciding whether to normalize at all when library size carries spatial biology rather than pure technical depth; choosing cell-volume/area normalization over…

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/preprocess_spatial.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with Scanpy. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Setting QC floors that do NOT delete real low-count imaging cells (an scRNA mincounts=500 floor deletes nearly every Xenium cell
  • Whose vector is tens-to-low-hundreds of transcripts)
  • Deciding whether to normalize at all when library size carries spatial biology rather than pure technical depth
  • Choosing cell-volume/area normalization over Pearson residuals for skewed targeted panels

Example prompts

  • “/bio-spatial-transcriptomics-spatial-preprocessing”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Spatial Transcriptomics Spatial Preprocessing loads about 4.7k tokens when it runs. Until then it costs about 188 tokens; SKILL.md has 1,956 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~188
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,956 words, ~4,715 tokens.

Download SKILL.mdSave it as .claude/skills/bio-spatial-transcriptomics-spatial-preprocessing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-spatial-transcriptomics-spatial-preprocessing
description
Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy. Use when setting QC floors that do NOT delete real low-count imaging cells (an scRNA min_counts=500 floor deletes nearly every Xenium cell, whose vector is tens-to-low-hundreds of transcripts); deciding whether to normalize at all when library size carries spatial biology rather than pure technical depth; choosing cell-volume/area normalization over Pearson residuals for skewed targeted panels; reading negative-control-probe / blank-barcode false-discovery rates; and inspecting QC spatially on the tissue rather than only in violins.
tool_type
python
primary_tool
squidpy

Version Compatibility

Reference examples tested with: squidpy 1.5+, scanpy 1.10+, anndata 0.10+, spatialdata 0.2+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Spatial Preprocessing

"QC and normalize my spatial data" -> Flag and remove low-quality spots/cells, then put counts on a scale fit for downstream domain and marker analysis -- but the right QC floors and the right normalization both depend on which side of the platform fork the data sits.

  • Sequencing/spot (Visium, Visium HD, Slide-seq, Stereo-seq): a spot/bin is mini-bulk over 1-10 cells; QC on UMI/spot, genes/spot, mito-%, cells/spot; normalization must respect that library size tracks cellularity.
  • Imaging/in-situ (Xenium, MERSCOPE/MERFISH, CosMx, seqFISH): a cell is segmentation-derived, carries tens-to-low-hundreds of transcripts over a TARGETED panel; QC on a low transcript floor, cell area, and negative-control FDR; normalization must not be gene-count-based.

The platform-class fork (decide this first)

The first question on any spatial dataset is which assay family produced it, because it changes every QC threshold and the entire normalization decision.

AxisSequencing/spot (Visium, Slide-seq, Stereo-seq)Imaging/in-situ (Xenium, MERSCOPE, CosMx)
Unitspot/bin = 1-10-cell mixturesegmentation-derived single cell
Counts/unithundreds-thousands UMItens-low hundreds transcripts
Gene spacewhole-transcriptome (poly-A) or probe panelTARGETED panel (100-1000), genes/cell ceilinged at panel size
Mito-% QCavailable (mixed-cell average)usually impossible (mito off-panel)
Specificity metricnone nativenegative-control-probe / blank-barcode FDR
Library-size meaningconfounds cells-per-spot + cellularityconfounds cell SIZE/AREA + segmentation error

Governing Principle

In single-cell RNA-seq library size is a technical nuisance to divide out. In spatial transcriptomics LIBRARY SIZE CARRIES BIOLOGY, and that single fact governs both QC and normalization. On Visium, total counts per spot are spatially structured and correlated with anatomy because they confound with the number of cells per spot and tissue cellularity (Bhuva 2024 Genome Biol 25:99). On imaging platforms, total counts per cell confound with cell SIZE/AREA -- a physically larger segmented cell holds more molecules for purely geometric reasons -- and with segmentation error itself. Naively dividing library size out (CP10k, log1p, scran pooling) therefore removes real spatially-structured biology and measurably degrades spatial-domain detection.

For imaging the bias is UPSTREAM of any residual model. Because a targeted panel is small, hand-curated, and skewed toward a few high markers, any gene-count-based size factor is dominated by a handful of genes and becomes panel-composition-dependent. Atta and Fan 2024 (Genome Biol 25:153) compared library-size, Pearson/SCTransform, DESeq2, TMM, and volume/area normalization on skewed panels: the four gene-count methods inject region-specific bias of up to ~13% DE error and fold-change SIGN REVERSAL in up to 19% of genes, while volume/area normalization avoids it because its denominator is independent of panel composition. The load-bearing consequence: Pearson residuals and SCTransform, the gold standard for whole-transcriptome scRNA, do NOT rescue a skewed imaging panel -- they are still gene-count-based and the bias sits in the panel design. The fix is non-gene-count normalization (cell volume/area, Moffitt-style) or spatially-aware joint modeling (SpaNorm).

A clean violin plot proves nothing. A count floor that silently deletes a spatial cluster of small cells (lymphocytes), a normalization that erases a cellularity gradient, and a focal hybridization failure all survive a tidy genes-per-cell distribution. The dangerous artifacts are spatial, so QC must be inspected spatially.

scRNA QC thresholds are wrong for imaging

Carrying scRNA defaults onto imaging data deletes the data. Imaging cells carry tens-to-low-hundreds of transcripts -- one to two orders of magnitude below droplet scRNA -- so an min_counts=500 floor removes nearly every real cell. Genes/cell can never exceed the panel size, so the "high genes = doublet" heuristic is meaningless (the ceiling is the panel, not a doublet). Mito genes are usually off-panel, so pct_counts_mt QC is often impossible. Worst, an aggressive count floor preferentially deletes the smallest REAL cells (lymphocytes, neutrophils), biasing tissue composition rather than removing noise.

QC-metric-by-platform table
MetricVisium (spot)Imaging (Xenium/MERSCOPE/CosMx)Rationale / trap
counts/unitUMI/spot; no universal floor; OSTA DLPFC flags <600transcripts/cell tens-low hundreds; community floor ~10 (Squidpy), some 20scRNA min_counts=500 deletes nearly every imaging cell; floor confounds cellularity/cell-size
genes/unitgenes/spot; OSTA flags <400genes/cell CEILING = panel size (100-1000)"high genes = doublet" meaningless for imaging
mito-%mixed-cell average; OSTA flags >0.28 (brain)usually off-panel -> impossibletissue-dependent; brain tolerates higher
cell area (um^2)n/a (spot is fixed)MAD-based on counts/areaflags over/under-segmentation; no fixed vendor min
negative-control FDRn/aTHE imaging specificity metricfalse-discovery proxy; no scRNA analogue
cells/spotnuclei estimate; OSTA flags >10n/aconfirms spot is a mixture

Thresholds are tissue-dependent and "somewhat arbitrary" (the OSTA Visium worked example flags UMI<600, genes<400, mito>0.28, cells/spot>10 on DLPFC, removing 32/3639 spots -- a starting point, not a law). The imaging floor of ~10 transcripts/cell is a Squidpy/community convention, NOT a vendor specification; community CosMx floors run higher (commonly ~20 counts/cell, scaling up with plex), so confirm the cutoff against the panel and tissue rather than copying a number.

Goal: Annotate negative controls and mito genes (where present), compute QC metrics, and set platform-appropriate floors without deleting real low-count cells.

Approach: Branch on the fork. For imaging, identify control-probe prefixes, filter on a low transcript floor and cell area; for spot data, use UMI/genes/mito floors. Always compute metrics, then look at them spatially before cutting.

python
import squidpy as sq
import scanpy as sc
import numpy as np

# Imaging branch: control features carry platform-specific prefixes -- they are the specificity ruler, not genes
ctrl_prefixes = ('NegControlProbe', 'NegControlCodeword', 'BLANK', 'Blank', 'NegPrb')   # Xenium / MERFISH / CosMx
adata.var['control'] = adata.var_names.str.startswith(ctrl_prefixes)
adata.var['mt'] = adata.var_names.str.startswith(('MT-', 'mt-'))                          # usually empty on imaging panels
sc.pp.calculate_qc_metrics(adata, qc_vars=['control', 'mt'], percent_top=None, inplace=True)
# inplace defaults to False and returns DataFrames; pass inplace=True to write .obs/.var

Negative controls -- the imaging specificity metric

Imaging platforms include features that decode to nothing biological: Xenium negative-control PROBES (off-target binding) plus negative-control CODEWORDS (pure optical/decoding error), MERFISH/MERSCOPE blank barcodes (valid codewords with no probe), CosMx NegPrb (alien synthetic sequences). They are the only native false-discovery proxy in spatial data. The canonical metric is FDR = mean counts per control feature / mean counts per real gene; the community-acceptable band is roughly <=1-5% of signal (Xenium typically <0.1%, MERFISH ~4%, CosMx highest). Compute it before trusting any gene-level claim, and treat controls as a panel-wide QC gate -- not as genes to cluster on.

Goal: Quantify the per-feature false-discovery rate and drop controls before normalization and clustering.

Approach: Average per-feature counts within the control set and within real genes, take the ratio, then subset the matrix to real genes only.

python
ctrl = adata.var['control'].values
mean_ctrl = np.asarray(adata[:, ctrl].X.sum(axis=0)).ravel().mean() if ctrl.any() else 0.0
mean_gene = np.asarray(adata[:, ~ctrl].X.sum(axis=0)).ravel().mean()
fdr = mean_ctrl / mean_gene if mean_gene else float('nan')
print(f'negative-control FDR: {fdr:.4f}  (band ~<=0.01-0.05)')
adata = adata[:, ~ctrl].copy()   # controls are a QC ruler, never clustering features

Inspect QC spatially, then filter

A QC gradient across the section -- counts falling toward one edge, mito rising in a corner -- is a technical artifact (edge effects, uneven permeabilization, focal hybridization failure), not biology, and a violin plot hides it. Always map QC onto tissue coordinates before choosing thresholds, and check WHERE the cells slated for removal actually fall.

Goal: Reveal spatially-structured quality artifacts and confirm a proposed floor is not removing a coherent tissue region.

Approach: Color the spatial scatter by each QC metric; a smooth spatial gradient signals a technical artifact to address (or model) rather than threshold away.

python
sq.pl.spatial_scatter(adata, color=['total_counts', 'n_genes_by_counts'], shape=None, ncols=2)
# shape=None renders points (imaging/Slide-seq); omit it for Visium hex spots with a tissue image

Goal: Apply platform-appropriate floors that remove debris and segmentation failures without biasing composition.

Approach: Imaging -- low transcript floor plus a cell-area sanity bound. Spot -- UMI/genes/mito floors. Either way, filter genes seen in too few units last.

python
# Imaging floor: ~10 transcripts/cell is a Squidpy/community convention, NOT a vendor spec
sc.pp.filter_cells(adata, min_counts=10)
if 'cell_area' in adata.obs:
    lo, hi = adata.obs['cell_area'].quantile([0.01, 0.99])     # trim segmentation over/under-calls, tissue-dependent
    adata = adata[(adata.obs['cell_area'] > lo) & (adata.obs['cell_area'] < hi)].copy()
sc.pp.filter_genes(adata, min_cells=5)

# Spot branch instead (Visium): tissue-dependent floors -- the OSTA DLPFC example, not universal law
# sc.pp.filter_cells(adata, min_counts=600)
# sc.pp.filter_cells(adata, min_genes=400)
# adata = adata[adata.obs['pct_counts_mt'] < 28].copy()
Show full SKILL.md (809 more words)Show less

Normalization -- the central decision

Do not reach reflexively for normalize_total + log1p. The shipped Squidpy tutorials run it for both Xenium and MERFISH, so it is the de-facto default -- and it is exactly what the benchmark papers argue is biased for spatial data. Decide deliberately from the table, and because methods compete here, verify current best practice against the installed tool's docs and the latest benchmarks before committing.

Normalization-method table
MethodAssumptionBest whenFails when
normalize_total + log1plibrary size = pure technical nuisancecross-platform comparability; quick default; tool tutorialsspatial -- removes spatially-structured biology (Bhuva 2024); imaging skewed panel
Analytic Pearson residualsclosed-form NB offset; depth as fixed offsetwhole-transcriptome Visium HVG/PCAimaging targeted panel -- still gene-count-based, inherits panel-skew bias (Atta/Fan 2024)
SCTransform v2regularized NB GLM, depth slope fixedwhole-transcriptome UMI / Visiumdoes NOT fix skewed imaging panels (gene-count-based)
Cell volume/areaconcentration is the biological quantityIMAGING skewed panel (Moffitt-style)denominator needs reliable segmentation; cannot fix segmentation error itself
SpaNorm (spatially-aware)library size and biology are entangled; remove only library-size componentspot AND imaging; preserve spatial structurenewer; R/Bioconductor

Goal (spot, whole-transcriptome): Stabilize depth for HVG/PCA while keeping raw counts, accepting that crude library-size division can blur domains.

Approach: Stash raw counts, then either run the standard log1p pipeline knowingly or prefer analytic Pearson residuals for feature selection on whole-transcriptome Visium.

python
adata.layers['counts'] = adata.X.copy()                    # stash raw -- HVG flavors and re-normalization need it
sc.pp.normalize_total(adata)                               # target_sum=None scales to the dataset MEDIAN, not the arbitrary 1e4
sc.pp.log1p(adata)                                         # library size carries biology -- this can blur spatial domains
# Whole-transcriptome Visium feature selection alternative (gene-count-based, fine here, NOT for imaging panels):
# sc.experimental.pp.normalize_pearson_residuals(adata)

Goal (imaging, targeted panel): Normalize without injecting panel-composition bias, using a denominator independent of gene counts.

Approach: Divide each cell's counts by its segmented area/volume (copies per unit area), then log-transform -- the Moffitt-style fix benchmarks favour over Pearson residuals for skewed panels.

python
adata.layers['counts'] = adata.X.copy()
if 'cell_area' in adata.obs:                               # area/volume denominator is panel-composition-independent
    sf = adata.obs['cell_area'].values / adata.obs['cell_area'].median()
    adata.X = adata.X / sf[:, None]
    sc.pp.log1p(adata)
# If no segmentation area is available, prefer SpaNorm (R) over reflexive normalize_total on a skewed panel

Common Errors

SymptomCauseFix
Nearly all imaging cells filtered outscRNA min_counts=500 floor on tens-of-transcript cellsUse a low floor (~10 transcripts/cell); branch QC on the platform fork
Spatial domains blur / merge after normalizationCrude library-size division erased a cellularity/anatomy gradient (library size carries biology)Prefer SpaNorm or volume/area; if using log1p, know it can blur domains
Fold-change sign flips between normalizationsGene-count size factor is panel-composition-dependent on a skewed imaging panelUse non-gene-count (cell area/volume) normalization; Pearson/SCT do NOT fix it
pct_counts_mt is all zero / NaNMito genes are not on the targeted imaging panelSkip mito-% QC for imaging; QC on transcript floor + cell area instead
"Doublet" cells flagged by high gene countgenes/cell ceiling IS the panel size -- not a doublet signalDrop the high-genes heuristic for imaging; use cell area / spatial doublets
Control features cluster as their own groupNegative-control probes/codewords left in the matrixCompute control FDR, then subset to real genes before clustering
Smallest cell type vanished after filteringA count floor preferentially deleted small real cells (lymphocytes)Inspect spatially where cuts fall; lower the floor; check composition before/after
Claimed a novel cell-state signature from imaging marker genesAn imaging panel (even 5,000-plex) is pre-selected for KNOWN biology; off-panel genes are absent by design, not by expressionTreat absence of an off-panel gene as uninformative; de-novo state discovery is bounded by the panel -- corroborate on whole-transcriptome data before claiming novelty
QC looks fine in violins but a region is emptySpatial QC gradient (edge/permeabilization artifact) invisible in violinsMap QC onto tissue with sq.pl.spatial_scatter before thresholding
Counts inflated ~2x after re-running normalizationNormalized already-normalized dataNormalize raw once; restore from layers['counts']
  • spatial-data-io - load Visium/Xenium/MERFISH and reach the molecule table vs the segmentation-derived matrix
  • image-analysis - segment cells from imaging data, the upstream error source that sets imaging QC and cell area
  • spatial-deconvolution - the next step for spot data, where library size and reference choice decide proportions
  • single-cell/preprocessing - the scRNA QC/normalization baseline these thresholds deliberately depart from
  • single-cell/clustering - cluster the QC'd cells; resolution is not a truth knob
  • single-cell/cell-annotation - label-transfer typing for a targeted panel (de-novo marker discovery is panel-bounded)

References

  • Bhuva DD, Tan CW, Salim A, et al. (2024) Library size confounds biology in spatial transcriptomics data. Genome Biology 25:99. DOI 10.1186/s13059-024-03241-7
  • Atta L, Clifton K, Anant M, Aihara G, Fan J (2024) Gene count normalization in single-cell imaging-based spatially resolved transcriptomics. Genome Biology 25:153. DOI 10.1186/s13059-024-03303-w
  • Lause J, Berens P, Kobak D (2021) Analytic Pearson residuals for normalization of single-cell RNA-seq UMI data. Genome Biology 22:258. DOI 10.1186/s13059-021-02451-7
  • Palla G, Spitzer H, Klein M, et al. (2022) Squidpy: a scalable framework for spatial omics analysis. Nature Methods 19:171-178. DOI 10.1038/s41592-021-01358-2
  • Salim A, Bhuva DD, Chen C, et al. (2025) SpaNorm: spatially-aware normalisation for spatial transcriptomics data. Genome Biology 26:109. DOI 10.1186/s13059-025-03565-y
  • Moffitt JR, Bambah-Mukku D, Eichhorn SW, et al. (2018) Molecular, spatial, and functional single-cell profiling of the hypothalamic preoptic region. Science 362:eaau5324. DOI 10.1126/science.aau5324
  • Janesick A, Shelansky R, Gottscho AD, et al. (2023) High resolution mapping of the tumor microenvironment using integrated single-cell, spatial and in situ analysis. Nature Communications 14:8353. DOI 10.1038/s41467-023-43458-x
  • Maynard KR, Collado-Torres L, Weber LM, et al. (2021) Transcriptome-scale spatial gene expression in the human dorsolateral prefrontal cortex. Nature Neuroscience 24:425-436. DOI 10.1038/s41593-020-00787-0

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in spatial-transcriptomics/spatial-preprocessing of GPTomics/bioSkills.

  • SKILL.md
  • examples/preprocess_spatial.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Spatial Transcriptomics Spatial Preprocessing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Spatial Transcriptomics Spatial Preprocessing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Spatial Transcriptomics Spatial Preprocessing this skillGPTomics/bioSkills1.2k1 repos~4.7kAutomated safety check: PassMIT
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
Single Cell Rna AnalysisPKU-YuanGroup/OpenAI4S608—~1.3kAutomated safety check: PassMIT
Anndatadavila7/claude-code-templates32k12 repos~2.5kAutomated safety check: PassMIT
Cellxgene Censusdavila7/claude-code-templates32k11 repos~3.8kAutomated safety check: PassMIT
Bulk RnaseqK-Dense-AI/scientific-agent-skills48k1 repos~4.2kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Single Cell Rna Analysis

    PKU-YuanGroup/OpenAI4S

    Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…

    608 GitHub stars~1.3k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    32k GitHub starsUsed in 12 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Cellxgene Census

    davila7/claude-code-templates

    Query CZ CELLxGENE Census (61M+ cells). An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.8k tokens
    Research & ScienceAuto-check passed
  • Bulk Rnaseq

    K-Dense-AI/scientific-agent-skills

    Prepares bulk RNA-seq FASTQ, Salmon, STAR or featureCounts output for gene-level differential expression.

    48k GitHub starsUsed in 1 repo~4.2k tokens
    Research & ScienceAuto-check passed
  • Pathway Enrichment

    K-Dense-AI/scientific-agent-skills

    Performs pathway and gene-set enrichment analysis on gene lists or ranked gene data and interprets the results.

    48k GitHub starsUsed in 1 repo~4.2k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 553 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Spatial Transcriptomics Spatial Preprocessing

What does Bio Spatial Transcriptomics Spatial Preprocessing do?

Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy. Bio Spatial Transcriptomics Spatial Preprocessing is an agent skill from GPTomics/bioSkills. Quality control, filtering, and normalization for spatial transcriptomics (Visium, Visium HD, Xenium, MERFISH/MERSCOPE, CosMx, Slide-seq) with Squidpy and Scanpy.

When should I use Bio Spatial Transcriptomics Spatial Preprocessing?

Bio Spatial Transcriptomics Spatial Preprocessing fits situations like: setting QC floors that do NOT delete real low-count imaging cells (an scRNA mincounts=500 floor deletes nearly every Xenium cell; whose vector is tens-to-low-hundreds of transcripts); deciding whether to normalize at all when library size carries spatial biology rather than pure technical depth; choosing cell-volume/area normalization over Pearson residuals for skewed targeted panels.

How do I install Bio Spatial Transcriptomics Spatial Preprocessing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a claude-code`. Or copy the skill folder (spatial-transcriptomics/spatial-preprocessing in GPTomics/bioSkills) into .claude/skills/bio-spatial-transcriptomics-spatial-preprocessing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Spatial Transcriptomics Spatial Preprocessing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a codex`. Or copy the skill folder (spatial-transcriptomics/spatial-preprocessing in GPTomics/bioSkills) into .agents/skills/bio-spatial-transcriptomics-spatial-preprocessing in your project. Codex loads it when a task matches its description.

Can I use Bio Spatial Transcriptomics Spatial Preprocessing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-spatial-transcriptomics-spatial-preprocessing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-spatial-transcriptomics-spatial-preprocessing, .gemini/skills/bio-spatial-transcriptomics-spatial-preprocessing, .github/skills/bio-spatial-transcriptomics-spatial-preprocessing and .opencode/skills/bio-spatial-transcriptomics-spatial-preprocessing in your project.

What does Bio Spatial Transcriptomics Spatial Preprocessing need to run?

Going by SKILL.md and its folder, Bio Spatial Transcriptomics Spatial Preprocessing needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Spatial Transcriptomics Spatial Preprocessing access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Spatial Transcriptomics Spatial Preprocessing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Spatial Transcriptomics Spatial Preprocessing use?

Bio Spatial Transcriptomics Spatial Preprocessing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Spatial Transcriptomics Spatial Preprocessing use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Spatial Transcriptomics Spatial Preprocessing?

Skills that share tags, products or a category with Bio Spatial Transcriptomics Spatial Preprocessing: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), Single Cell Rna Analysis (PKU-YuanGroup/OpenAI4S, 608 stars), Anndata (davila7/claude-code-templates, 32k stars) and Cellxgene Census (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Spatial Transcriptomics Spatial Preprocessing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.