Agent skill

Sc Clustering

by TianGzlab in TianGzlab/OmicsClaw

Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Sc Clustering

skills CLI
$ npx skills add TianGzlab/OmicsClaw --skill sc-clustering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TianGzlab/OmicsClaw sc-clustering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TianGzlab/OmicsClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/singlecell/scrna/sc-clustering .claude/skills/sc-clustering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sc-clustering
GitHub stars
161
Token cost
~2.4k tokens
SKILL.md length
1,077 words
Files
10 (incl. references)
Skills in repo
88
Repo updated
First seen
Licence
Apache-2.0

At a glance

Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData.

  • Tasks that involve Bioinformatics
  • SKILL.md covers When to use, Use from a step, API and Methods and parameters, plus 5 more sections
  • Runs Python scripts from its folder; calls python
  • Tasks that involve Embeddings

What it does

Sc Clustering is an agent skill from TianGzlab/OmicsClaw. Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData. Skip when QC/normalisation/HVG/PCA have not run yet (use sc-preprocessing); marker ranking after clustering (use sc-markers).

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `_api.py`, `examples/example_step.py` and `references/methodology.md`).

It sits in AI & LLM Engineering, covering Bioinformatics and Embeddings. It works with UMAP and AnnData. The repository describes itself as: Conversational & memory-enabled AI research partner for multi-omics analysis. CLI + Desktop App (installers in Releases). From biological idea to full research paper. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve Embeddings

Example prompts

  • “/sc-clustering”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 90a3bec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sc Clustering loads about 2.4k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 1,077 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TianGzlab/OmicsClaw at commit 90a3bec, republished under its Apache-2.0 licence (© TianGzlab). 1,077 words, ~2,434 tokens.

Download SKILL.mdSave it as .claude/skills/sc-clustering/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
sc-clustering
description
Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData. Skip when QC/normalisation/HVG/PCA have not run yet (use sc-preprocessing); marker ranking after clustering (use sc-markers).
tags
singlecell, scrna, clustering, leiden, louvain, umap, tsne, phate

sc-clustering

When to use

The user has a normalised AnnData with PCA / integrated embedding already populated and wants the standard scRNA neighbour-graph → embedding → cluster workflow. Combinable in one call: pick an embedding method (umap default, also tsne / diffmap / phate) and a clustering method (leiden default, also louvain), with an explicit resolution or auto-resolution search. Reads from obsm["X_pca"] / obsm["X_harmony"] / etc. via use_rep.

Use from a step

python
clustering = load_skill("sc-clustering")
adata = read_input("results/02_preprocess/intermediate/adata_preprocessed.h5ad")
# resolution 0.8: the user wants broad cell types; sc-clustering's default is 1.0.
adata = clustering.cluster(adata, resolution=0.8)
write_output(clustering.cluster_summary(adata, key="leiden"), "tables/cluster_summary.csv")
write_output(clustering.embedding_figure(adata, color="leiden"), "figures/umap_leiden.png")
write_output(adata, "intermediate/adata_clustered.h5ad")

A complete step that runs on demo data: examples/example_step.py.

API

<!-- api:begin generated from _api.py; regenerate with run.py api <skill dir> --write -->
cluster(adata, *, use_rep: str | None=None, use_existing_graph: bool=False, n_neighbors: int=15, n_pcs: int=50, embedding: str='umap', method: str='leiden', resolution: float=1.0, random_state: int=0, umap_min_dist: float=0.5, umap_spread: float=1.0, tsne_perplexity: float=30.0, tsne_metric: str='euclidean', diffmap_n_comps: int=15, phate_knn: int=15, phate_decay: int=40)

Build the neighbour graph, cluster the cells and compute a 2-D embedding, in place.

Writes obs[method] (the cluster labels, a categorical of strings), obsm['X_<embedding>'], the graph in obsp and uns['neighbors'], and records the cluster key as the primary one in the matrix contract.

:param use_rep: The obsm key the graph is built on. Default None: the first of X_pca, X_harmony, X_scvi, X_scanvi, X_scanorama present. Pass X_harmony or X_scvi after batch integration. :param n_neighbors: Neighbours per cell in the graph. Default 15, scanpy's default. Larger values give smoother, coarser structure. :param use_existing_graph: Keep the graph in uns['neighbors'] and obsp instead of rebuilding it. Default False; set True after BBKNN. Graph construction parameters are then ignored; UMAP and diffusion maps use the retained graph, while t-SNE and PHATE still use use_rep. :param n_pcs: Components of use_rep to use. Default 50; pick it from sc-preprocessing's pca_variance_table (the elbow) when the data are small. :param embedding: "umap" (default), "tsne", "diffmap" or "phate". :param method: "leiden" (default) or "louvain"; also the obs column written. :param resolution: Clustering resolution. Default 1.0, scanpy's default. Lower it (0.3 to 0.6) for broad cell types, raise it for finer subtypes; or use :func:auto_resolution. :param random_state: Seed for the graph, the clustering and the embedding. Default 0, scanpy's default, which makes reruns reproduce the labels. :param umap_min_dist: UMAP min_dist. Default 0.5, scanpy's default. :param umap_spread: UMAP spread. Default 1.0, scanpy's default. :param tsne_perplexity: t-SNE perplexity. Default 30. :param tsne_metric: t-SNE distance metric. Default "euclidean". :param diffmap_n_comps: Diffusion-map components. Default 15. :param phate_knn: PHATE neighbours. Default 15. :param phate_decay: PHATE alpha decay. Default 40. :returns: The same AnnData. :raises ValueError: no embedding is available, or an unknown method or embedding.

auto_resolution(adata, *, use_rep: str | None=None, n_neighbors: int=15, n_pcs: int=50, method: str='leiden', resolutions: list[float] | None=None, n_subsample_reps: int=5, subsample_fraction: float=0.8, random_state: int=1) -> pd.DataFrame

Choose a clustering resolution by how stably cells co-cluster under subsampling.

Builds the neighbour graph, then for every candidate resolution clusters n_subsample_reps random subsamples, turns how often two cells share a cluster into a distance, and scores the full-data clustering with the silhouette on that distance. The best-scoring clustering is left in obs[method]; pass the selected resolution to :func:cluster. The cost grows with the square of the number of cells.

:param use_rep: As in :func:cluster. :param n_neighbors: As in :func:cluster. :param n_pcs: As in :func:cluster. :param method: "leiden" (default) or "louvain". :param resolutions: Candidates. Default [0.4, 0.6, 0.8, 1.0, 1.2, 1.4]. :param n_subsample_reps: Subsamples per resolution. Default 5. :param subsample_fraction: Fraction of cells in each subsample. Default 0.8. :param random_state: Seed for the subsampling. Default 1, the CLI's value. :returns: Columns resolution, silhouette_score and selected (true on the chosen row).

run_info(adata, *, keep: bool=True) -> dict

What :func:cluster recorded: use_rep, cluster_key, embedding_key and the contracts.

:param keep: Leave the record in adata.uns; False removes it. :returns: The record, or an empty dict when cluster has not run on adata.

cluster_summary(adata, *, key: str='leiden') -> pd.DataFrame

Cells per cluster, largest first.

:param key: The obs column holding the labels. Default "leiden". :returns: Columns cluster, n_cells and proportion_pct (rounded to 2 decimals).

embedding_figure(adata, *, color: str='leiden', basis: str | None=None)

A scatter plot of the 2-D embedding, coloured by an obs column.

:param color: The obs column to colour by. Default "leiden". :param basis: The obsm key to plot. Default: the first of X_umap, X_tsne, X_diffmap, X_phate present. :returns: A matplotlib Figure. :raises KeyError: no 2-D embedding is present.

<!-- api:end -->
Show full SKILL.md (410 more words)Show less

Methods and parameters

  • Graph: scanpy neighbors on use_rep with n_neighbors=15 and n_pcs=50, scanpy's defaults. Pick n_pcs from the elbow of sc-preprocessing's pca_variance_table when it is clearly below 50; raise n_neighbors (30 to 50) for smoother, coarser structure on large datasets.
  • Clustering: leiden (default; Traag et al. 2019, guarantees connected communities) or louvain. resolution=1.0 is scanpy's default. It is the parameter the user most often has an opinion on: broad cell types usually come out around 0.3 to 0.6, subtypes at 1.0 to 2.0. Ask, or try two values in variant steps (02a_leiden_r05.py, 02b_leiden_r10.py) and compare.
  • auto_resolution: picks the resolution whose clustering is most stable under subsampling (silhouette on a co-clustering distance). It costs time quadratic in the cell count; on more than a few thousand cells subsample first or choose by hand.
  • Embedding: umap (default; min_dist=0.5, spread=1.0, scanpy's defaults), tsne (perplexity 30), diffmap (15 components, for trajectories), phate (needs the phate package).
  • random_state=0: scanpy's default seed for the graph, clustering and UMAP; keep it fixed so reruns and the standalone run of the step give the same labels.

Gotchas

  • No embedding → ValueError. cluster needs obsm["X_pca"] or the use_rep you name. Run sc-preprocessing (or sc-batch-integration) first.
  • The label column is named after the method. cluster(method="leiden") writes obs["leiden"] and overwrites an existing one; copy the old labels to another column first when you need to compare.
  • cluster works in place. It returns the same AnnData; write it out with write_output to keep the labels.
  • Use the integrated embedding after batch correction. With X_harmony or X_scvi present, pass it as use_rep; the default picks X_pca first.

Inputs and outputs

  • Reads obsm[use_rep] (default X_pca) and the matrix contract in uns.
  • Writes obs[method] (labels as strings), obsm['X_<embedding>'], obsp['connectivities'] / obsp['distances'], uns['neighbors'], and the contract's primary cluster key. auto_resolution also leaves its best clustering in obs[method].
  • cluster_summary and auto_resolution return DataFrames; embedding_figure returns a matplotlib Figure.

CLI

sc_cluster.py runs the same functions outside a project and writes a report, figures, tables and processed.h5ad: python <skill directory>/sc_cluster.py --help. --demo runs it on PBMC3k.

See also

  • references/parameters.md — every CLI flag and per-method tuning hint
  • references/methodology.md — embedding choice guide, auto-resolution heuristic
  • references/output_contract.md — obs / obsm keys + the CLI's table schemas
  • Adjacent skills: sc-preprocessing (upstream — normalise/HVG/PCA before this), sc-batch-integration (parallel — produces the integrated embedding use_rep reads from), sc-markers (downstream — rank cluster markers from obs["leiden"])

Dependencies

Python packages this skill's script needs. They are not installed for you — check before a long run.

anndata, louvain, matplotlib, numpy, pandas, phate, scanpy, scikit-learn, scipy, seaborn

© TianGzlab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in skills/singlecell/scrna/sc-clustering of TianGzlab/OmicsClaw.

  • SKILL.md
  • _api.py
  • examples/example_step.py
  • references/methodology.md
  • references/output_contract.md
  • references/parameters.md
  • references/r_visualization.md
  • sc_cluster.py
  • tests/test_sc_cluster.py
  • tests/test_sc_cluster_methods.py

Open the folder on GitHubat commit 90a3bec

Compare with similar skills

Sc Clustering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sc Clustering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sc Clustering this skillTianGzlab/OmicsClaw161—~2.4kAutomated safety check: PassApache-2.0
Scrna EmbeddingClawBio/ClawBio1.2k1 repos~2kAutomated safety check: PassMIT
Bio Data Visualization Dimensionality Reduction PlotsGPTomics/bioSkills1.2k2 repos~4.8kAutomated safety check: PassMIT
Alphagenome Predictionsgenomicsxai/alphagenome-pytorch162—~868Automated safety check: PassApache-2.0
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
Umap Tsne Analysisaipoch/medical-research-skills2k—~2.7kAutomated safety check: PassMIT

Similar skills

  • Scrna Embedding

    ClawBio/ClawBio

    Local scVI/scANVI-based single-cell latent embedding and batch-aware integration from raw-count .h5ad or 10x Matrix Market input, with stable integrated AnnData export for downstream latent analysis.

    1.2k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions)…

    1.2k GitHub starsUsed in 2 repos~4.8k tokens
    Data & AnalyticsAuto-check passed
  • Alphagenome Predictions

    genomicsxai/alphagenome-pytorch

    Run AlphaGenome-PyTorch to get genomic track predictions — via the agt predict CLI (single locus, BED regions, whole chromosomes, raw FASTA sequences, or per-gene count tables/AnnData), variant…

    162 GitHub stars~868 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Umap Tsne Analysis

    aipoch/medical-research-skills

    A skill your agent uses when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE…

    2k GitHub stars~2.7k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • Scanpy

    K-Dense-AI/scientific-agent-skills

    Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…

    48k GitHub starsUsed in 1 repo~5.1k tokens
    Research & ScienceAuto-check passed

More from TianGzlab/OmicsClaw

All 88 skills in this repo
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated yesterday
    Auto-check passed
  • Bulkrna Batch Correction

    TianGzlab/OmicsClaw

    Load when correcting batch effects in bulk expression using R sva ComBat or the legacy Python parametric approximation.

    161 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Bulkrna Coexpression

    TianGzlab/OmicsClaw

    Load when discovering bulk gene co-expression modules and hub genes with R WGCNA.

    161 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Bulkrna De

    TianGzlab/OmicsClaw

    Load when comparing gene expression between two conditions in bulk RNA-seq count data.

    161 GitHub stars~867 tokensUpdated yesterday
    Auto-check passed
  • Bulkrna Deconvolution

    TianGzlab/OmicsClaw

    Load when estimating cell-type proportions in bulk RNA-seq samples from a single-cell or signature-matrix reference.

    161 GitHub stars~757 tokensUpdated yesterday
    Auto-check passed
  • Bulkrna Enrichment

    TianGzlab/OmicsClaw

    Load when running pathway / GO term enrichment on a bulk RNA-seq DE result list.

    161 GitHub stars~860 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Sc Clustering

What does Sc Clustering do?

Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData. Sc Clustering is an agent skill from TianGzlab/OmicsClaw. Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData.

When should I use Sc Clustering?

Sc Clustering fits situations like: tasks that involve Bioinformatics; tasks that involve Embeddings.

How do I install Sc Clustering in Claude Code?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-clustering -a claude-code`. Or copy the skill folder (skills/singlecell/scrna/sc-clustering in TianGzlab/OmicsClaw) into .claude/skills/sc-clustering in your project. Claude Code loads it when a task matches its description.

How do I install Sc Clustering in Codex?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-clustering -a codex`. Or copy the skill folder (skills/singlecell/scrna/sc-clustering in TianGzlab/OmicsClaw) into .agents/skills/sc-clustering in your project. Codex loads it when a task matches its description.

Can I use Sc Clustering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TianGzlab/OmicsClaw --skill sc-clustering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sc-clustering, .gemini/skills/sc-clustering, .github/skills/sc-clustering and .opencode/skills/sc-clustering in your project.

What does Sc Clustering need to run?

Going by SKILL.md and its folder, Sc Clustering needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Sc Clustering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sc Clustering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sc Clustering use?

Sc Clustering is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sc Clustering use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Sc Clustering?

Skills that share tags, products or a category with Sc Clustering: Scrna Embedding (ClawBio/ClawBio, 1.2k stars), Bio Data Visualization Dimensionality Reduction Plots (GPTomics/bioSkills, 1.2k stars), Alphagenome Predictions (genomicsxai/alphagenome-pytorch, 162 stars) and Scgpt (JimLiu/science-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sc Clustering?

TianGzlab (a GitHub organization) maintains it in TianGzlab/OmicsClaw, which has 161 GitHub stars. The repository holds 88 skills in this directory. The repository was last updated on October 7, 2026.

Source: TianGzlab/OmicsClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.