Agent skill

Spatial Raw Processing

by TianGzlab in TianGzlab/OmicsClaw

Load when converting spatial transcriptomics raw FASTQ pairs through ST-Pipeline into a rawcounts.h5ad ready for spatial-preprocess.

Apache-2.0Auto-check passedResearch & Science

Install Spatial Raw Processing

skills CLI
$ npx skills add TianGzlab/OmicsClaw --skill spatial-raw-processing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TianGzlab/OmicsClaw spatial-raw-processing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TianGzlab/OmicsClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spatial/spatial-raw-processing .claude/skills/spatial-raw-processing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spatial-raw-processing
GitHub stars
161
Token cost
~1.7k tokens
SKILL.md length
530 words
Files
9 (incl. references)
Skills in repo
88
Repo updated
First seen
Licence
Apache-2.0

At a glance

Load when converting spatial transcriptomics raw FASTQ pairs through ST-Pipeline into a rawcounts.h5ad ready for spatial-preprocess.

  • Works in 6 steps: Parse args (or load bundle JSON / YAML… → _apply_effective_defaults fills missing… → _validate_real_run_bundle: check read1 /… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Use from a step, When to use, Inputs & Outputs and Flow, plus 4 more sections
  • Runs Python and R scripts from its folder; calls python

What it does

Spatial Raw Processing is an agent skill from TianGzlab/OmicsClaw. Load when converting spatial transcriptomics raw FASTQ pairs through ST-Pipeline into a rawcounts.h5ad ready for spatial-preprocess. Skip when input is already a count-matrix AnnData (use spatial-preprocess); non-spatial bulk / scRNA FASTQ (use bulkrna-read-qc).

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `examples/example_step.py`, `r_visualization/README.md` and `references/methodology.md`).

It sits in Research & Science, covering Bioinformatics. It works with AnnData. The repository describes itself as: Conversational & memory-enabled AI research partner for multi-omics analysis. CLI + Desktop App (installers in Releases). From biological idea to full research paper. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/spatial-raw-processing”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Parse args (or load bundle JSON / YAML from positional --input).
  2. _apply_effective_defaults fills missing parameter values (threads, trimming, UMI ranges, etc.).
  3. _validate_real_run_bundle: check read1 / read2 / ids / ref-map exist and are well-typed; reject duplicate read1=read2; verify FASTQ…
  4. Call run_stpipeline(...) which shells out to ST-Pipeline (requires the stpipeline binary on PATH or --stpipeline-repo + --bin-path).
  5. Wrap the resulting count matrix into AnnData with X = raw_counts, layers["counts"], raw = raw_counts_snapshot.
  6. Save raw_counts.h5ad and result.json. Print "next: spatial-preprocess on raw_counts.h5ad".

What it can do on your machine

Read from SKILL.md and the folder at commit 90a3bec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and R), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Spatial Raw Processing loads about 1.7k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 530 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TianGzlab/OmicsClaw at commit 90a3bec, republished under its Apache-2.0 licence (© TianGzlab). 530 words, ~1,748 tokens.

Download SKILL.mdSave it as .claude/skills/spatial-raw-processing/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
spatial-raw-processing
description
Load when converting spatial transcriptomics raw FASTQ pairs through ST-Pipeline into a `raw_counts.h5ad` ready for spatial-preprocess. Skip when input is already a count-matrix AnnData (use spatial-preprocess); non-spatial bulk / scRNA FASTQ (use bulkrna-read-qc).
trigger
spatial raw processing, raw spatial fastq, spatial fastq, st_pipeline, st pipeline, barcode coordinates, ids file, visium raw fastq, slide-seq fastq, slideseq…
tags
spatial, raw-processing, fastq, st-pipeline, visium, slideseq

spatial-raw-processing

Use from a step

This is CLI_ONLY: ST-Pipeline consumes FASTQ files, barcode coordinates and reference indexes. It is not an in-memory analysis function.

python
from skills._sdk.notebook import run_cli

output = run_cli("spatial-raw-processing", "--input", "data/run_bundle.json",
                 inputs=["data/run_bundle.json", "data/R1.fastq.gz",
                         "data/R2.fastq.gz", "data/barcodes.tsv"])

Record the reference index version in the module README. The executable example uses synthetic upstream outputs and does not run ST-Pipeline.

When to use

The user has paired-end spatial-transcriptomics FASTQ files (read1 = spatial barcode + UMI, read2 = cDNA) plus a STAR genome index, and wants the standard ST-Pipeline run that produces a raw_counts.h5ad with one row per spatial spot. Single backend: st_pipeline (calls run_stpipeline from skills/spatial/_lib/stpipeline_adapter.py).

After this skill, chain to spatial-preprocess for QC + normalisation. For non-spatial scRNA FASTQ use sc-fastq-qc. For bulk RNA-seq read QC use bulkrna-read-qc.

Inputs & Outputs

Inputs

  • Input kinds: file, directory
  • Modalities: visium, slideseq
  • File types: .fastq, .fq, .json, .yaml, .yml
  • FASTQ structure: valid first record; paired layout
  • Directory layouts (any): paired-fastq

Outputs

  • tables/gene_qc.csv
  • tables/raw_gene_qc.csv
  • tables/raw_processing_run_summary.csv
  • tables/raw_processing_spatial_points.csv
  • tables/raw_spot_qc.csv
  • tables/raw_top_genes.csv
  • tables/run_summary.csv
  • tables/saturation_curve.csv
  • tables/spatial_coordinates.csv
  • tables/spot_qc.csv
  • tables/stage_summary.csv
  • tables/top_genes.csv
  • figures/raw_detected_genes_spatial.png
  • figures/raw_spot_qc_histograms.png
  • figures/raw_top_genes_barplot.png
  • figures/raw_total_counts_spatial.png
  • figures/st_pipeline_saturation_curve.png
  • figures/st_pipeline_stage_attrition.png
  • omicsclaw_stpipeline_run.json
  • raw_counts.h5ad
  • st_pipeline.stderr.txt
  • st_pipeline.stdout.txt
  • report.md
  • result.json
  • Processed AnnData (saves_h5ad) — adds obs: barcode, x_array, y_array; obsm: spatial

Flow

  1. Parse args (or load bundle JSON / YAML from positional --input).
  2. _apply_effective_defaults fills missing parameter values (threads, trimming, UMI ranges, etc.).
  3. _validate_real_run_bundle: check read1 / read2 / ids / ref-map exist and are well-typed; reject duplicate read1=read2; verify FASTQ extension.
  4. Call run_stpipeline(...) which shells out to ST-Pipeline (requires the stpipeline binary on PATH or --stpipeline-repo + --bin-path).
  5. Wrap the resulting count matrix into AnnData with X = raw_counts, layers["counts"], raw = raw_counts_snapshot.
  6. Save raw_counts.h5ad and result.json. Print "next: spatial-preprocess on raw_counts.h5ad".
Show full SKILL.md (275 more words)Show less

Gotchas

  • All input failures raise typed exceptions wrapped in SystemExit(1). spatial_raw_processing.py catches DataError / DependencyError / ParameterError / ProcessingError and re-raises as SystemExit(1). The originating raises live in _validate_real_run_bundle — it raises ParameterError(f"Missing required parameter: {key}") for missing read1/read2/ids; raises DataError(...) for non-existent files; raises DataError("Resolved read1/read2 inputs must be FASTQ files.") for non-FASTQ extensions; raises ParameterError for read1==read2; raises DataError for missing / wrong-type STAR index dir; raises DataError only when --ref-annotation was provided but the path is missing or not a file (the param itself is optional — omitting it doesn't raise).
  • --read1 / --read2 / --ids / --ref-map are all required for real runs (not enforced by argparse required=True, validated later). Missing any → ParameterError. Demo mode skips this validation entirely.
  • The output filename is always raw_counts.h5ad (spatial_raw_processing.py). It's not configurable — the contract is consumed by spatial-preprocess. Multiple runs to the same --output will overwrite.
  • QC tables and figures are written. tables/spot_qc.csv and figures/raw_total_counts_spatial.png describe the count matrix; upstream stage and saturation plots depend on available pipeline metrics.
  • Demo mode skips ST-Pipeline entirely. spatial_raw_processing.py calls create_demo_upstream_outputs(...) to fabricate a synthetic raw_counts.h5ad. Useful for plumbing checks; does NOT exercise the FASTQ → matrix code path.
  • --platform is a metadata label only. _build_parser documents it as "Label recorded in outputs"; ST-Pipeline doesn't branch on it. Common values: visium, visium_hd, slideseq, custom strings.

Key CLI

bash
# Demo (synthetic raw_counts.h5ad — does NOT run ST-Pipeline)
python skills/spatial/spatial-raw-processing/spatial_raw_processing.py --demo --output /tmp/spatial_raw_demo

# Real run with explicit args
python skills/spatial/spatial-raw-processing/spatial_raw_processing.py \
  --read1 sample_R1.fastq.gz --read2 sample_R2.fastq.gz \
  --ids barcodes.tsv \
  --ref-map /refs/star_index_human \
  --ref-annotation /refs/genes.gtf \
  --exp-name visium_001 --platform visium \
  --threads 16 \
  --output results/

# Real run from bundle JSON
python skills/spatial/spatial-raw-processing/spatial_raw_processing.py \
  --input run_bundle.json --output results/

# Slide-seq with custom UMI range
python skills/spatial/spatial-raw-processing/spatial_raw_processing.py \
  --read1 R1.fq.gz --read2 R2.fq.gz --ids barcodes.tsv \
  --ref-map /refs/star_index --platform slideseq \
  --umi-start-position 1 --umi-end-position 8 \
  --output results/

See also

  • references/parameters.md — every CLI flag, ST-Pipeline option mapping
  • references/methodology.md — when ST-Pipeline wins vs Space Ranger; barcode-ID format
  • references/output_contract.md — raw_counts.h5ad schema
  • Adjacent skills: spatial-preprocess (downstream — required next step; consumes raw_counts.h5ad), bulkrna-read-qc / sc-fastq-qc (parallel — non-spatial FASTQ paths)

Dependencies

Python packages this skill's script needs. They are not installed for you — check before a long run.

anndata, matplotlib, numpy, pandas, PyYAML, scipy, seaborn

© TianGzlab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/spatial/spatial-raw-processing of TianGzlab/OmicsClaw.

  • SKILL.md
  • examples/example_step.py
  • r_visualization/README.md
  • r_visualization/raw_processing_publication_template.R
  • references/methodology.md
  • references/output_contract.md
  • references/parameters.md
  • spatial_raw_processing.py
  • tests/test_spatial_raw_processing.py

Open the folder on GitHubat commit 90a3bec

Compare with similar skills

Spatial Raw Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spatial Raw Processing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spatial Raw Processing this skillTianGzlab/OmicsClaw161—~1.7kAutomated safety check: PassApache-2.0
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
PyDESeq2 Differential Expressiondavila7/claude-code-templates32k12 repos~4kAutomated safety check: PassMIT
Anndatadavila7/claude-code-templates32k12 repos~2.5kAutomated safety check: PassMIT
Single-Cell Initial AnalysisLigphiDonk/Oh-my--paper7381 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    32k GitHub starsUsed in 12 repos~4k tokens
    Research & ScienceAuto-check passed
  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    32k GitHub starsUsed in 12 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    738 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check passed
  • Single Cell Data Prep Qc

    harrisongzhang/TheVirtualBiotech

    Single-cell RNA-seq data preparation and quality control pipeline.

    120 GitHub stars~2.9k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed

More from TianGzlab/OmicsClaw

All 88 skills in this repo
  • Bulkrna De

    TianGzlab/OmicsClaw

    Load when comparing gene expression between two conditions in bulk RNA-seq count data.

    161 GitHub starsUsed in 1 repo~867 tokens
    Auto-check passed
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated yesterday
    Auto-check passed
  • Bulkrna Qc

    TianGzlab/OmicsClaw

    Load when checking a bulk RNA-seq count matrix for library-size outliers, gene detection rates, and sample-sample correlation before DE.

    161 GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Sc Fastq Qc

    TianGzlab/OmicsClaw

    Load when checking raw single-cell FASTQ read quality (Phred / GC / adapter / length) before counting.

    161 GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Sc Filter

    TianGzlab/OmicsClaw

    Load when removing low-quality cells and lowly-detected genes from a single-cell AnnData using QC-derived thresholds or tissue presets.

    161 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Sc Markers

    TianGzlab/OmicsClaw

    Load when ranking cluster-level marker genes from a clustered single-cell AnnData via Scanpy Wilcoxon / t-test / logreg or COSG specificity.

    161 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Works with

Questions about Spatial Raw Processing

What does Spatial Raw Processing do?

Load when converting spatial transcriptomics raw FASTQ pairs through ST-Pipeline into a rawcounts.h5ad ready for spatial-preprocess. Spatial Raw Processing is an agent skill from TianGzlab/OmicsClaw.h5ad ready for spatial-preprocess.

When should I use Spatial Raw Processing?

Spatial Raw Processing fits situations like: tasks that involve Bioinformatics.

How do I install Spatial Raw Processing in Claude Code?

Run `npx skills add TianGzlab/OmicsClaw --skill spatial-raw-processing -a claude-code`. Or copy the skill folder (skills/spatial/spatial-raw-processing in TianGzlab/OmicsClaw) into .claude/skills/spatial-raw-processing in your project. Claude Code loads it when a task matches its description.

How do I install Spatial Raw Processing in Codex?

Run `npx skills add TianGzlab/OmicsClaw --skill spatial-raw-processing -a codex`. Or copy the skill folder (skills/spatial/spatial-raw-processing in TianGzlab/OmicsClaw) into .agents/skills/spatial-raw-processing in your project. Codex loads it when a task matches its description.

Can I use Spatial Raw Processing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TianGzlab/OmicsClaw --skill spatial-raw-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spatial-raw-processing, .gemini/skills/spatial-raw-processing, .github/skills/spatial-raw-processing and .opencode/skills/spatial-raw-processing in your project.

What does Spatial Raw Processing need to run?

Going by SKILL.md and its folder, Spatial Raw Processing needs Python and R for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Spatial Raw Processing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Spatial Raw Processing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Spatial Raw Processing use?

Spatial Raw Processing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spatial Raw Processing use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Spatial Raw Processing?

Skills that share tags, products or a category with Spatial Raw Processing: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), Scgpt (JimLiu/science-skills, 227 stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 32k stars) and Anndata (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spatial Raw Processing?

TianGzlab (a GitHub organization) maintains it in TianGzlab/OmicsClaw, which has 161 GitHub stars. The repository holds 88 skills in this directory. The repository was last updated on October 7, 2026.

Source: TianGzlab/OmicsClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.