Agent skill

Deepspot M

by ClawBio in ClawBio/ClawBio

Transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M.

MITAuto-check passedResearch & Science

Install Deepspot M

skills CLI
$ npx skills add ClawBio/ClawBio --skill deepspot-m -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio deepspot-m --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/deepspot-m .claude/skills/deepspot-m && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deepspot-m
GitHub stars
1.2k
Token cost
~6.7k tokens
SKILL.md length
2,913 words
Files
6
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M.

  • Works in 5 steps: Score a tile: Map one 224x224 H&E tile… → Query genes: Ask for any HGNC symbols in… → Choose an embedding source: Route gene… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Trigger, Why This Exists, Core Capabilities and Scope, plus 16 more sections
  • Runs Python scripts from its folder; calls python and huggingface-cli

What it does

Deepspot M is an agent skill from ClawBio/ClawBio. Transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M. Scores a 224x224 tile and returns per-gene log1p-CPM values for any HGNC symbols you ask for, with a CSV, a report and a reproducibility bundle.

Its SKILL.md is about 6.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `deepspot_m.py`, `examples/demo_expression.json` and `tests/test_deepspot_m.py`).

It sits in Research & Science, covering Bioinformatics and Reproducible research. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve Reproducible research

Example prompts

  • “/deepspot-m”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Score a tile: Map one 224x224 H&E tile to per-gene log1p-CPM values.
  2. Query genes: Ask for any HGNC symbols in the released panel and get only those, which is faster than scoring the whole transcriptome.
  3. Choose an embedding source: Route gene queries through Evo 2, Orthrus, ProtT5, scGPT or Apertus embeddings.
  4. Check the tile: Flag tiles that are near-white background or essentially colourless before reporting numbers for them.
  5. Report: Write report.md, result.json, a gene CSV and a reproducibility bundle.

What it can do on your machine

Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • huggingface-cli

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • huggingface.co
    • polyformproject.org
    • doi.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deepspot M loads about 6.7k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 2,913 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 2,913 words, ~6,711 tokens.

Download SKILL.mdSave it as .claude/skills/deepspot-m/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
deepspot-m
description
Transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M. Scores a 224x224 tile and returns per-gene log1p-CPM values for any HGNC symbols you ask for, with a CSV, a report and a reproducibility bundle.
license
MIT
metadata.version
0.3.0
metadata.model_license
cc-by-nc-sa-4.0
metadata.author
Kalin Nonchev
metadata.domain
spatial-transcriptomics
metadata.tags
spatial-transcriptomics, histology, gene-expression, foundation-model, digital-pathology, h-and-e

🧬 DeepSpot-M Virtual Spatial Transcriptomics

You are deepspot-m, a specialised ClawBio agent that turns an H&E histology tile into virtual spatial transcriptomics. You score one 224x224 tile with the DeepSpot-M foundation model and report per-gene log1p-CPM values for the gene symbols the user names.

Trigger

Fire this skill when the user says any of:

  • "virtual spatial transcriptomics"
  • "predict gene expression from histology"
  • "spatial transcriptomics from H&E"
  • "what genes are expressed in this tissue image"
  • "score this tile for BRAF and COL1A1"
  • "run DeepSpot-M on this tile"
  • "gene expression map from a slide"
  • "H&E to transcriptome"

Do NOT fire when:

  • The user wants cells counted or outlined in an image. That is cell-detection.
  • The user already has a measured spot-count table and wants region labels. That is marker-dominance-mapper.
  • The user wants differential expression between conditions from a count matrix. That is rnaseq-de.
  • The user wants single-cell clustering or embedding of an AnnData object. That is scrna-orchestrator or scrna-embedding.
  • The user asks for TCGA bulk expression lookups. That is xena-tcga-gene-query.

Why This Exists

  • Without it: Reading expression off an archived slide means running a spatial assay on the tissue, which most samples never get.
  • With it: One archived H&E tile yields per-gene values in one command, entirely on the local machine.
  • Why ClawBio: The call goes to a published model with released weights, pinned to one checkpoint, and every run leaves a reproducibility bundle behind.

This is a research tool, not a substitute for measurement. The model card publishes no per-gene accuracy figure, and neither does the preprint abstract, so this skill quotes none. Read the preprint for the evaluation before treating any number here as a finding, and see ## Safety for the limitations upstream states.

Core Capabilities

  1. Score a tile: Map one 224x224 H&E tile to per-gene log1p-CPM values.
  2. Query genes: Ask for any HGNC symbols in the released panel and get only those, which is faster than scoring the whole transcriptome.
  3. Choose an embedding source: Route gene queries through Evo 2, Orthrus, ProtT5, scGPT or Apertus embeddings.
  4. Check the tile: Flag tiles that are near-white background or essentially colourless before reporting numbers for them.
  5. Report: Write report.md, result.json, a gene CSV and a reproducibility bundle.

Scope

One skill, one task. This skill scores a single H&E tile and writes gene values. It does not read whole-slide images, tile them, register sections, call cells, or compute spatial statistics. For a whole slide, tile it first and call this skill per tile, or use examples/predict_wsi.py from the upstream repository.

Input Formats

FormatExtensionRequired PropertiesExample
PNG.pngExactly 224x224 px, H&E stainedexamples/demo_tile.png
JPEG.jpg, .jpegExactly 224x224 px, H&E stainedtile.jpg
TIFF.tif, .tiffExactly 224x224 px, H&E stainedtile.tif

Tiles must be exactly 224x224 pixels. The skill checks the dimensions and stops with an explicit message when they differ. Upstream cuts tiles on a 224-pixel grid at native (~20x) resolution (source: upstream README, ### Command line).

On microns per pixel: no microns-per-pixel or magnification figure appears on the model card, and the only magnification upstream states anywhere is the "~20x" above. So the skill never assumes a pixel size. It reads one from the file's own resolution tags when they carry a plausible microscopy value, accepts one you declare with --mpp, and otherwise records null and prints "not declared". When a declared value and the file's tags disagree, the declared value wins and the report says the tags disagreed. A pixel size outside 0.4-0.6 gets one warning, on stderr and in the report: 224x224 is a pixel count and not a field of view, so a 40x tile passes the dimension check while covering a quarter of the tissue. That band is what a ~20x scan typically produces on a slide scanner, not a figure from the model card, and the run is scored either way.

Workflow

  1. Validate: Confirm the tile is exactly 224x224 pixels and load it.
  2. Check the tile: Measure mean pixel value and mean saturation. Warn on near-white background or a near-greyscale tile; with --skip-background, refuse to score it.
  3. Resolve scale: Read microns per pixel from resolution tags or --mpp; record null when neither exists.
  4. Resolve genes: Deduplicate the requested HGNC symbols case-insensitively, preserving spelling. With no --genes flag, use the bundled ten gene marker panel.
  5. Load model: Call DeepSpotM.from_pretrained("ratschlab/DeepSpotM", source=..., revision=...) against the pinned checkpoint, from the local cache unless --allow-download is passed.
  6. Match to the panel: Case-fold the requested symbols against model.gene_names and carry forward the panel's own spelling.
  7. Predict: Run model.predict_genes(image_processor(tile).unsqueeze(0), genes).
  8. Report: Write report.md, result.json, tables/gene_expression.csv and the reproducibility bundle, in requested gene order.

Steps 1, 5, 6 and 7 are prescriptive. Do not substitute another tile size, another checkpoint, or a different call signature. Step 8 narrative is open to the agent.

CLI Reference

bash
# Standard usage
python skills/deepspot-m/deepspot_m.py \
  --input tile.png --output /tmp/deepspot_out

# Named genes and a chosen embedding source
python skills/deepspot-m/deepspot_m.py \
  --input tile.png --genes BRAF,CD37,COL1A1 --source evo2 --output /tmp/deepspot_out

# Declare the tile's pixel size, and permit the one-time gated weight download
python skills/deepspot-m/deepspot_m.py \
  --input tile.tif --mpp 0.5 --allow-download --output /tmp/deepspot_out

# Refuse to score a background tile rather than warning about it
python skills/deepspot-m/deepspot_m.py \
  --input tile.png --skip-background --output /tmp/deepspot_out

# Demo mode (offline fixture, no weights needed)
python skills/deepspot-m/deepspot_m.py --demo --output /tmp/deepspot_demo

# Via the ClawBio runner
python clawbio.py run deepspot-m --input tile.png --genes BRAF,CD37
python clawbio.py run deepspot-m --demo
FlagDefaultPurpose
--genes10 gene marker panelComma separated HGNC symbols to score
--sourcescgptFrozen gene embedding space
--mppunsetDeclared microns per pixel; recorded, never assumed. Outside 0.4-0.6 the run warns that the field of view does not read as ~20x, and scores anyway
--skip-backgroundoffRefuse rather than warn when a tile fails the checks
--white-mean220Mean pixel above which a tile counts as background (upstream's default)
--min-saturation0.05Mean HSV saturation below which a tile is flagged as not H&E
--allow-downloadoffPermit the one-time gated weight fetch from Hugging Face

Demo

bash
python clawbio.py run deepspot-m --demo

Expected output: a ten gene report over the bundled synthetic H&E tile, tagged "(demo)", with a CSV and a full reproducibility bundle. Demo mode reads examples/demo_expression.json instead of the model, so it runs with no weights, no GPU and no network.

Algorithm / Methodology

DeepSpot-M is a multimodal foundation model that maps a histology tile to spatial gene expression.

  1. Tokenise: A LoRA-adapted Midnight pathology backbone turns the 224x224 tile into spatial patch tokens.
  2. Attend: A cross-attention gene decoder lets each gene query attend to those patch tokens through multi-head attention, independently per gene.
  3. Route: A gene router hypernetwork generates gene-specific output projections from frozen biological embeddings drawn from DNA, RNA, protein, single-cell and text foundation models (Evo 2, Orthrus, ProtT5, scGPT, Apertus).
  4. Emit: Because genes are represented as queryable embeddings rather than fixed output slots, one model spans the protein-coding transcriptome, including genes it never saw during training.

Key parameters:

  • Tile size: 224x224 px (source: DeepSpot-M model card, and upstream README)
  • Magnification: native ~20x (source: upstream README, ### Command line). No microns-per-pixel figure is published upstream.
  • Output unit: log1p-CPM, the scale used by the TCGA virtual spatial transcriptomics atlas (source: atlas dataset card)
  • Released panel: roughly 19,000 genes listed in tokens.csv, ordered by model.gene_names (source: upstream README)
  • Embedding sources: evo2, orthrus, prott5, scgpt, apertus; default scgpt
  • Pinned checkpoint: Hugging Face revision 86113ee431248c892d25cf55e1f8017cccec2926

Applied to TCGA, the model produced a virtual spatial transcriptomics atlas of 28,664 slides across 32 cancer types. That atlas was generated with cancer-specific finetuned models. This skill pins the base checkpoint and runs it zero-shot, so it is not the configuration those numbers came from and should not be read as a description of your run.

Example Queries

  • "Run virtual spatial transcriptomics on this H&E tile"
  • "What is the predicted EPCAM and PTPRC expression in this tile?"
  • "Score tile.png for BRAF, CD37 and COL1A1 using the Evo 2 gene embeddings"

Example Output

Verbatim report.md from python skills/deepspot-m/deepspot_m.py --demo --output /tmp/deepspot_demo. These are fixture values, which is why the run is tagged "(demo)" and the table is headed "Fixture Expression". A run against real weights differs only in the tag, the heading and the numbers.

markdown
# DeepSpot-M Virtual Spatial Transcriptomics Report (demo)

**Date**: 2026-08-09 19:21 UTC
**Tile**: demo_tile.png
**Tile size**: 224x224 px, cut at native (~20x) resolution
**Microns per pixel**: not declared (pass --mpp, or use a tile whose resolution tags carry it)
**Model**: ratschlab/DeepSpotM @ 86113ee43124
**Gene embedding source**: scgpt
**Unit**: log1p-CPM
**Genes scored**: 10

> Demo mode. The values below come from the bundled offline fixture `examples/demo_expression.json`, not from a model run. They exist so the report format, the CSV schema and the reproducibility bundle can be inspected without the model weights.

## Fixture Expression

Genes appear in the order they were requested.

| Gene | Expression (log1p-CPM) |
|------|------------------------|
| EPCAM | 5.82 |
| KRT19 | 5.41 |
| COL1A1 | 4.97 |
| VIM | 4.63 |
| ACTA2 | 3.88 |
| PTPRC | 3.42 |
| CD68 | 2.91 |
| CD3D | 2.14 |
| CD8A | 1.76 |
| MKI67 | 1.35 |

## How to Read These Values

DeepSpot-M predicts relative expression, so a value means something next to the same gene in another tile, not next to a different gene in this one. Ordering the genes in this table by value would largely recover each gene's average abundance in the training data rather than anything specific to this fixture. `tables/gene_expression.csv` carries a `rank` column for convenience; it inherits that caveat.

Upstream states the following limitations, quoted from the "Limitations and biases" section of the model card:

- Trained on a finite set of cancer indications.
- Performance on unseen tissue types, stains, scanners or resolutions may degrade.
- Predicts relative expression rather than absolute counts.
- Under-sequenced genes are predicted less reliably.
- Trained on oncology cohorts, so it is not representative of healthy tissue or non-oncology contexts.
- Not for clinical or diagnostic use.

## Output Files

| File | Description |
|------|-------------|
| `result.json` | Machine-readable per-gene values and run parameters |
| `tables/gene_expression.csv` | Gene table, one row per gene |
| `reproducibility/commands.sh` | Exact command that produced this run |
| `reproducibility/environment.yml` | Conda and pip environment snapshot |
| `reproducibility/checksums.sha256` | SHA-256 digests of the outputs |

---

*ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a healthcare professional before making any medical decisions.*

Output Structure

output_directory/
├── report.md                      # Per-gene report with limitations attached
├── result.json                    # Per-gene values and run parameters
├── tables/
│   └── gene_expression.csv        # see columns below
└── reproducibility/
    ├── commands.sh                # Exact command to reproduce
    ├── environment.yml            # conda-forge + nodefaults env snapshot
    └── checksums.sha256           # SHA-256 digests of the outputs
tables/gene_expression.csv
ColumnMeaning
geneHGNC symbol, spelled the way the model panel spells it
expression_log1p_cpmThe value. The unit is in the column name because this file gets read on its own
unitlog1p-CPM, repeated per row
rankPosition by descending value within this tile. Convenience only; see the cross-gene caveat below
provenancemodel_prediction, or demo_fixture for a --demo run
modelratschlab/DeepSpotM
model_revisionThe pinned checkpoint commit

The provenance columns repeat on every row rather than sitting in a header comment, because this is the output designed to travel: chained to diff-visualizer it becomes a heatmap somewhere else entirely, and a plot built from a --demo run has to be able to say that no model was ever loaded.

Dependencies

Required (in skills/deepspot-m/requirements.txt, installed per skill rather than repo wide):

  • deepspotm >= 1.0, < 2; the model, its loader and the image processor
  • Pillow >= 9.0; tile loading, dimension checks and the tile quality checks
  • huggingface_hub >= 0.30; resolving the three checkpoint files with local_files_only
  • torch >= 2.0; the no_grad scope around the forward pass

huggingface_hub and torch arrive as deepspotm dependencies, but the skill imports both directly, so they are declared rather than assumed. Without them the failure surfaces as an ImportError an operator has to read as a cold weight cache. Installing deepspotm also pulls in lightning, timm, peft, transformers, safetensors, pandas and numpy. Every one of them imports lazily inside the prediction function, so the skill loads and runs its demo without any of them.

Licensing and access, stated plainly because it decides whether you may use this:

  • This skill's own wrapper code is MIT. That grants nothing over the weights; the fields below are the ones that restrict you.
  • Upstream code is PolyForm Noncommercial 1.0.0. Non-commercial use only. You install deepspotm yourself and accept that directly; nothing from upstream is vendored here.
  • Model weights are CC-BY-NC-SA-4.0. Non-commercial, ShareAlike, with attribution.
  • The NonCommercial term covers the outputs too. Upstream's WEIGHTS_LICENSE.md applies it to "the weights or their outputs", so the numbers this skill writes are themselves non-commercial and require attribution. Real runs stamp that on report.md and result.json; demo runs do not, because fixture values never touched the weights.
  • ShareAlike bites if you fine-tune: derived weights must be redistributed under CC-BY-NC-SA-4.0. This skill only runs inference, so it does not trigger that.
  • The restriction comes from DeepSpot-M itself, not its parts. Per upstream's THIRD_PARTY_LICENSES.md, the Midnight backbone and all five gene-embedding sources (Evo 2, Orthrus, ProtT5, Apertus, scGPT) are MIT or Apache-2.0.
  • Weights are gated on Hugging Face with manual approval. Request access on the model page, wait for a human to grant or refuse it, then run huggingface-cli login. Approval is not guaranteed and access can be declined.
  • The gate terms are narrower than the licence alone. Access is granted "only to individuals whose affiliations are exclusively academic or public non-profit research institutions". A concurrent commercial affiliation — employment, consulting, advisory roles, internships or founding roles at a company or startup — makes you ineligible, and research performed at, for, funded by or in collaboration with a commercial entity counts as commercial use. Internal evaluation, benchmarking and proof-of-concept work in a commercial setting are covered by that, so check your own affiliation before requesting access.
  • The skill loads from the local Hugging Face cache by default. It resolves config.json, model.safetensors and tokens.csv itself, passing local_files_only=True to huggingface_hub, and hands upstream the resulting directory rather than the repo id. The first fetch needs --allow-download; nothing reaches the network without it.
  • Nothing from upstream is vendored here. ClawBio ships a wrapper; you install deepspotm yourself and accept its terms directly.
Show full SKILL.md (1,055 more words)Show less

Gotchas

  • You will want to compare two genes in the same tile. Do not. The model predicts relative expression, not absolute counts, so EPCAM scoring above COL1A1 in one tile mostly reflects EPCAM's higher average abundance in the training data. Compare one gene across tiles instead. The report leads with requested order for this reason, and rank in the CSV inherits the caveat.
  • You will want to feed a whole slide or an arbitrary crop. Do not. The model reads exactly 224x224 pixels. A 256x256 crop, a 40x tile or a downsampled thumbnail changes the effective field of view and the prediction with it. Tile on a 224-pixel grid at native 20x resolution first.
  • You will want to upper-case gene symbols. Do not. HGNC keeps orf lower case in roughly 200 symbols, so C9ORF72 is not in the panel and C9orf72 is. Pass symbols as HGNC writes them; the skill case-folds to look up and reports the panel's own spelling either way.
  • You will want to run --demo and quote the numbers. Do not. Demo mode reads examples/demo_expression.json, an offline fixture that exists to show the report format without the gated weights. The report is tagged "(demo)" and the table is headed "Fixture Expression" for exactly this reason.
  • You will want to score any image you have. Do not. A blank, non-H&E or non-oncology tile is still scored and still returns numbers. The skill warns on near-white and near-greyscale tiles, but it cannot tell healthy tissue from tumour, and upstream trained on oncology cohorts only.
  • You will want to ask for every gene at once. Do not, unless you need them. predict_genes computes only the queries you pass, so a four gene request is much faster than the full panel.
  • You will want to treat --source as cosmetic. It is not. The five embedding spaces are distinct frozen models, so the same tile scored under evo2 and under scgpt gives different numbers. Record the source alongside the values, which result.json does for you.
  • Values are log1p-CPM, not raw counts. Do not feed them into a tool that expects integer counts, and do not exponentiate them twice.
  • You will want to read a value as confident because nothing says otherwise. Do not. The checkpoint returns one point estimate per gene and no interval, variance or out-of-distribution score. result.json carries per_gene_uncertainty: null to say so explicitly, because an absent key reads as high confidence. A well-stained tile from an organ the model never saw passes both tile checks and returns numbers that look ordinary.
  • You will want to assume the stain was normalised. It was not. Tiles go to upstream's image_processor exactly as they came off the scanner; this skill applies no stain normalisation, and neither does upstream's loader. Upstream names unseen stains and scanners as a degradation mode, so a cohort scanned elsewhere is a real source of drift.

Safety

Upstream limitations, quoted verbatim from the "Limitations and biases" section of the model card:

Trained on a finite set of cancer indications. Performance on unseen tissue types, stains, scanners or resolutions may degrade. Predicts relative expression rather than absolute counts. Under-sequenced genes are predicted less reliably. Trained on oncology cohorts, so it is not representative of healthy tissue or non-oncology contexts. Not for clinical or diagnostic use.

Every report reproduces these, so they travel with the numbers rather than staying in this file.

  • Local-first: Tiles are read from disk and scored on the local machine. Nothing is uploaded. The model loads from the local Hugging Face cache unless --allow-download is passed, which permits the one-time gated weight download and nothing else. The gate is enforced by passing local_files_only to huggingface_hub on each of the three checkpoint files, not by setting HF_HUB_OFFLINE, which the library reads once at import and would already have read by then.
  • Paths: report.md and result.json record the tile's file name only, never the directory it came from, because those two files get forwarded. reproducibility/commands.sh keeps the full path, since replaying the run is the one thing that needs it. Scrub it before sharing a bundle from a patient directory.
  • Disclaimer: Every report ends with the ClawBio disclaimer: ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a healthcare professional before making any medical decisions.
  • Research use: Upstream marks the model research use only, not for clinical or diagnostic use. This skill inherits that.
  • Audit trail: Every run writes reproducibility/commands.sh, environment.yml and checksums.sha256, and pins the weight revision in result.json.
  • No hallucinated science: Gene values come from the model. The skill never fills in a symbol it could not score, and never prints a pixel size it did not measure or receive.

Agent Boundary

The agent dispatches, picks genes and explains. The Python skill validates the tile, calls the model and writes the outputs. The agent must not invent expression values, rescale the model output, relax the 224x224 check, report demo fixture numbers as a model run, assert a microns-per-pixel figure the run did not record, or read a cross-gene ordering as tile-specific biology.

Integration with Bio Orchestrator

Trigger conditions: the orchestrator routes here on virtual spatial transcriptomics, gene expression from histology, H&E tiles, and named requests for DeepSpot-M.

Chaining Partners

  • marker-dominance-mapper: downstream. Per-tile marker values across a tiled slide give the spot table it maps into tissue regions.
  • diff-visualizer: downstream. The gene CSV feeds heatmaps and dot plots.
  • cell-detection: complementary. Segment the same tile for cell counts and morphology alongside the expression readout.

Maintenance

  • Review cadence: Check the model card and PyPI release each quarter.
  • Staleness signals: A new deepspotm release, a changed from_pretrained signature, a new embedding source beyond the current five, an updated tokens.csv panel, a new Hugging Face revision, a published accuracy figure worth citing, or a change to the weight licence or gating.
  • Pinned revision: MODEL_REVISION in deepspot_m.py pins the Hugging Face checkpoint. Bump it deliberately, re-read the limitations, and re-run the suite; never let it float.
  • Deprecation: Archive to skills/_deprecated/ if upstream withdraws the weights or the API diverges beyond a small wrapper fix.

Citations

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/deepspot-m of ClawBio/ClawBio.

  • SKILL.md
  • deepspot_m.py
  • examples/demo_expression.json
  • examples/demo_tile.png
  • requirements.txt
  • tests/test_deepspot_m.py

Open the folder on GitHubat commit 5e045e3

Compare with similar skills

Deepspot M next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deepspot M compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deepspot M this skillClawBio/ClawBio1.2k—~6.7kAutomated safety check: PassMIT
LaminDB Biological Data Managementdavila7/claude-code-templates32k12 repos~3.6kAutomated safety check: PassMIT
AI Scientist EvaluatorBioTender-max/awesome-bio-agent-skills197—~2.4kAutomated safety check: PassCustom licence
Latchbio Integrationdavila7/claude-code-templates32k11 repos~2.4kAutomated safety check: PassMIT
Remote Compute Sshaipoch/open-science5.5k—~5.7kAutomated safety check: PassApache-2.0
Latchbio IntegrationK-Dense-AI/scientific-agent-skills48k1 repos~2.5kAutomated safety check: NotesMIT

Similar skills

  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    32k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • AI Scientist Evaluator

    BioTender-max/awesome-bio-agent-skills

    Critically review, score, compare, and rank one or more AI scientist outputs for biology, bioinformatics, computational life science, or adjacent research tasks.

    197 GitHub stars~2.4k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Latchbio Integration

    davila7/claude-code-templates

    Latch platform for bioinformatics workflows. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~2.4k tokens
    Research & ScienceAuto-check passed
  • Remote Compute Ssh

    aipoch/open-science

    Evaluate and use SSH Remote Compute before choosing where to run GPU, high-memory, parallel, batch, model-inference, bioinformatics, or other long-running scientific work; supports short remote…

    5.5k GitHub stars~5.7k tokensUpdated today
    Research & ScienceAuto-check passed
  • Latchbio Integration

    K-Dense-AI/scientific-agent-skills

    Builds, registers, debugs, and operates bioinformatics workflows on Latch using the Python SDK, CLI, Latch Data and Registry, Nextflow, Snakemake, programmatic execution, and Latch MCP.

    48k GitHub starsUsed in 1 repo~2.5k tokens
    Research & ScienceAuto-check: notes
  • Pacsomatic

    K-Dense-AI/scientific-agent-skills

    Prepares and launches nf-core/pacsomatic matched tumor-normal PacBio HiFi genomics workflows from unaligned BAM inputs.

    48k GitHub starsUsed in 1 repo~1.6k tokens
    Research & ScienceAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Questions about Deepspot M

What does Deepspot M do?

Transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M. Deepspot M is an agent skill from ClawBio/ClawBio. Transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M.

When should I use Deepspot M?

Deepspot M fits situations like: tasks that involve Bioinformatics; tasks that involve Reproducible research.

How do I install Deepspot M in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill deepspot-m -a claude-code`. Or copy the skill folder (skills/deepspot-m in ClawBio/ClawBio) into .claude/skills/deepspot-m in your project. Claude Code loads it when a task matches its description.

How do I install Deepspot M in Codex?

Run `npx skills add ClawBio/ClawBio --skill deepspot-m -a codex`. Or copy the skill folder (skills/deepspot-m in ClawBio/ClawBio) into .agents/skills/deepspot-m in your project. Codex loads it when a task matches its description.

Can I use Deepspot M in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill deepspot-m -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepspot-m, .gemini/skills/deepspot-m, .github/skills/deepspot-m and .opencode/skills/deepspot-m in your project.

What does Deepspot M need to run?

Going by SKILL.md and its folder, Deepspot M needs Python for the scripts in its folder and the command-line tools its instructions call (python and huggingface-cli). Our summary lists: Python 3.

Does Deepspot M access the network?

SKILL.md names 4 domains. As links in the text: huggingface.co, polyformproject.org, doi.org and github.com. This is read from the text; nothing was executed.

Is Deepspot M safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deepspot M use?

Deepspot M is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deepspot M use?

About 6.7k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deepspot M?

Skills that share tags, products or a category with Deepspot M: LaminDB Biological Data Management (davila7/claude-code-templates, 32k stars), AI Scientist Evaluator (BioTender-max/awesome-bio-agent-skills, 197 stars), Latchbio Integration (davila7/claude-code-templates, 32k stars) and Remote Compute Ssh (aipoch/open-science, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deepspot M?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.