Agent skill

Bio Flow Cytometry Clustering Phenotyping

by GPTomics in GPTomics/bioSkills

Unsupervised clustering and cell-type identification for high-dimensional flow, spectral, and mass cytometry - FlowSOM, PhenoGraph, FlowSOM-via-CATALYST, with UMAP/tSNE for visualization.

MITAuto-check passedAI & LLM Engineering

Install Bio Flow Cytometry Clustering Phenotyping

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-flow-cytometry-clustering-phenotyping -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-flow-cytometry-clustering-phenotyping --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/flow-cytometry/clustering-phenotyping .claude/skills/bio-flow-cytometry-clustering-phenotyping && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-flow-cytometry-clustering-phenotyping
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
846 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Unsupervised clustering and cell-type identification for high-dimensional flow, spectral, and mass cytometry - FlowSOM, PhenoGraph, FlowSOM-via-CATALYST, with UMAP/tSNE for visualization.

  • Discovering populations without predefined gates
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithm Taxonomy and Why Over-Provision the Grid,…, plus 8 more sections
  • Runs R scripts from its folder
  • Choosing a clustering algorithm

What it does

Bio Flow Cytometry Clustering Phenotyping is an agent skill from GPTomics/bioSkills. Unsupervised clustering and cell-type identification for high-dimensional flow, spectral, and mass cytometry - FlowSOM, PhenoGraph, FlowSOM-via-CATALYST, with UMAP/tSNE for visualization. Covers the type-vs-state marker distinction (cluster on lineage, test state within clusters), over-provision-then-metacluster, the Weber-Robinson benchmark, seed dependence and metacluster stability, why embeddings are for looking not measuring, and median-heatmap annotation/merging. Use when discovering populations without…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in AI & LLM Engineering, covering Embeddings. It works with UMAP and GitHub. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Discovering populations without predefined gates
  • Choosing a clustering algorithm
  • Selecting the number of metaclusters
  • Annotating clusters into cell types

Example prompts

  • “/bio-flow-cytometry-clustering-phenotyping”

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Flow Cytometry Clustering Phenotyping loads about 2.3k tokens when it runs. Until then it costs about 171 tokens; SKILL.md has 846 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~171
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 846 words, ~2,279 tokens.

Download SKILL.mdSave it as .claude/skills/bio-flow-cytometry-clustering-phenotyping/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-flow-cytometry-clustering-phenotyping
description
Unsupervised clustering and cell-type identification for high-dimensional flow, spectral, and mass cytometry - FlowSOM, PhenoGraph, FlowSOM-via-CATALYST, with UMAP/tSNE for visualization. Covers the type-vs-state marker distinction (cluster on lineage, test state within clusters), over-provision-then-metacluster, the Weber-Robinson benchmark, seed dependence and metacluster stability, why embeddings are for looking not measuring, and median-heatmap annotation/merging. Use when discovering populations without predefined gates, choosing a clustering algorithm, selecting the number of metaclusters, or annotating clusters into cell types.
tool_type
r
primary_tool
CATALYST

Version Compatibility

Reference examples tested with: CATALYST 1.26+, FlowSOM 2.10+, flowCore 2.14+; Rphenograph (GitHub: JinmiaoChenLab/Rphenograph).

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

Rphenograph is GitHub-only (remotes::install_github('JinmiaoChenLab/Rphenograph')) and returns a list - membership is igraph::membership(out[[2]]), not a vector. Adapt rather than retrying.

Clustering and Phenotyping

"Cluster my cytometry data to find cell types" -> Discover populations in high-dimensional data without gates, then annotate them by marker expression.

  • R: CATALYST::cluster() (wraps FlowSOM + ConsensusClusterPlus) - the field default
  • R: FlowSOM::FlowSOM() directly, or Rphenograph() for graph-based clustering

The Single Most Important Modern Insight -- Cluster on Type Markers; the Embedding Is for Looking, Not Measuring

Two rules carry most of the correctness here. First, the type-vs-state distinction: LINEAGE/type markers (CD3, CD4, CD8, CD19) DEFINE clusters; functional/STATE markers (phospho-epitopes, cytokines, Ki-67, activation markers) must be WITHHELD from clustering and tested within clusters instead (the DA/DS framework, Nowicka 2017 F1000Res 6:748). Clustering on state markers splits "activated CD4" from "resting CD4" and confounds abundance with activation - a classic, silent design error. Second, t-SNE/UMAP embeddings do NOT preserve inter-cluster distances, cluster sizes, or densities (the apparent "UMAP preserves global structure" edge over tSNE is largely an initialization artifact - Kobak & Linderman 2021 Nat Biotechnol 39:156). Define populations by clustering in the HIGH-DIMENSIONAL space and COLOR the embedding by cluster; never gate on the embedding or read biology off blob distances.

Algorithm Taxonomy

AlgorithmCitationMechanismSpeedRare-popDeterminism
FlowSOMVan Gassen 2015 Cytometry A 87:636SOM grid -> MST (viz) -> consensus metaclusteringfastestgood if grid over-provisionedstochastic; seed-controllable
PhenoGraphLevine 2015 Cell 162:184kNN graph (Jaccard) + Louvainmoderatestrong (no preset k)seed-fragile (>40% reassignment reported)
X-shiftSamusik 2016 Nat Methods 13:493weighted kNN density + auto cluster #slowexcellentmore deterministic
flowMeansAghaeepour 2011 Cytometry A 79:6k-means multi-cluster + change-point kfastmoderatestochastic

Benchmark: Weber & Robinson 2016 Cytometry A 89:1084 tested 18 methods - FlowSOM (with metaclustering) was a top performer AND by far fastest, hence the field default; but its accuracy depends on supplying the right number of metaclusters.

Why Over-Provision the Grid, Then Metacluster

Set the SOM grid (e.g. 10x10 = 100 nodes) MUCH larger than the number of populations expected, then metacluster down. The asymmetry: metaclustering can MERGE over-fine nodes into a real population, but can NEVER SPLIT a node that erroneously fused two cell types. Too coarse commits the unrecoverable error; too fine commits only the recoverable one. So over-cluster, then merge by hand off the median heatmap.

CATALYST Clustering Pipeline

Goal: Cluster on type markers and prepare for annotation.

Approach: prepData builds the SCE (panel marker_class flags type vs state); cluster() wraps FlowSOM+ConsensusClusterPlus. Defaults xdim=ydim=10, maxK=20 (the metacluster cap people forget); set seed on the function.

r
library(CATALYST)

sce <- prepData(fs, panel, md, transform = TRUE, cofactor = 5)   # cofactor 5 = CyTOF; ~150 for fluorescence
sce <- cluster(sce, features = 'type',                            # type markers only
               xdim = 10, ydim = 10, maxK = 20, seed = 42)        # maxK caps metaclusters at 20 by default
plotExprHeatmap(sce, features = 'type', by = 'cluster_id', k = 'meta20', scale = 'last')

PhenoGraph (graph-based alternative)

Goal: Cluster with a kNN graph when a data-driven cluster count is wanted.

Approach: Rphenograph on the type-marker matrix (cells x markers); extract membership from the list.

r
library(Rphenograph)
type_expr <- t(assay(sce, 'exprs')[rowData(sce)$marker_class == 'type', ])
out <- Rphenograph(type_expr, k = 30)                  # only knob: k (neighbors)
sce$phenograph <- factor(igraph::membership(out[[2]]))  # list -> membership, not a vector

Dimensionality Reduction (visualization only) and Annotation

Goal: Visualize structure and assign cell-type labels.

Approach: runDR subsamples per sample (cells=); color by cluster, never gate on it. Annotate from the median heatmap, then mergeClusters with a curated table.

r
sce <- runDR(sce, dr = 'UMAP', features = 'type', cells = 2000)   # subsampled embedding
plotDR(sce, 'UMAP', color_by = 'meta20')

merging <- data.frame(old_cluster = 1:20,
                      new_cluster = c('CD4 T','CD4 T','CD8 T', '...'))   # curated from the heatmap
sce <- mergeClusters(sce, k = 'meta20', table = merging, id = 'annotated')
Show full SKILL.md (330 more words)Show less

Per-Method Failure Modes

Clustering on state markers

Trigger: activation/phospho markers in the clustering feature set. Mechanism: state contaminates lineage identity. Symptom: "activated" and "resting" versions of a type split as separate clusters. Fix: cluster on type only; test state markers within clusters (differential-analysis).

Seed-dependent "novel populations"

Trigger: a population that appears at one seed and vanishes at another. Mechanism: FlowSOM init / Louvain are stochastic. Symptom: non-reproducible clusters. Fix: set + report the seed; check multi-seed stability; treat unstable clusters as hypotheses.

Reading biology off the embedding

Trigger: "cluster A is closer to B than C." Mechanism: UMAP/tSNE distances are non-metric. Symptom: false developmental/relatedness claims. Fix: quantify in marker space; embedding for display only.

Clustering uncompensated/untransformed data

Trigger: raw linear input to FlowSOM. Mechanism: spillover + scale dominate Euclidean distance. Symptom: clusters track intensity, not biology. Fix: compensate + transform first.

Quantitative Thresholds

ThresholdSourceRationale
over-provision grid (10x10) >> expected popsVan Gassen 2015metacluster can merge, never split
maxK = 20 defaultCATALYSTmetacluster cap; raise if expecting more
FlowSOM needs correct KWeber & Robinson 2016accuracy depends on metacluster number
use median (not mean) per clusterBendall 2011 Science 332:687robust to doublet/spillover contamination

Common Errors

Error / symptomCauseSolution
clustering uses scatter/Time/statefeatures not restrictedfeatures='type' / colsToUse= lineage markers
Rphenograph result unusableit returns a listigraph::membership(out[[2]])
set.seed doesn't make FlowSOM reproducibleinternal reseedingpass seed= to cluster()
only 20 clusters no matter whatmaxK defaultraise maxK

References

  • Van Gassen 2015 Cytometry A 87(7):636-645 — FlowSOM.
  • Levine 2015 Cell 162(1):184-197 — PhenoGraph.
  • Samusik 2016 Nat Methods 13(6):493-496 — X-shift.
  • Weber & Robinson 2016 Cytometry A 89(12):1084-1096 — clustering benchmark (FlowSOM top + fastest).
  • Nowicka 2017 F1000Research 6:748 — CyTOF workflow; type-vs-state markers.
  • Kobak & Linderman 2021 Nat Biotechnol 39:156-157 — embedding initialization artifact.
  • Bendall 2011 Science 332(6030):687-696 — arcsinh-median analysis of CyTOF data.
  • compensation-transformation - Compensate/transform before clustering
  • gating-analysis - Supervised alternative; needed for rare populations
  • differential-analysis - Test abundance/state of clusters between conditions
  • cytometry-qc - Cluster only QC-passed events
  • single-cell/clustering - Leiden/Louvain on scRNA-seq (shared graph-clustering ideas)
  • imaging-mass-cytometry/phenotyping - Same CATALYST/FlowSOM conventions for imaging

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in flow-cytometry/clustering-phenotyping of GPTomics/bioSkills.

  • SKILL.md
  • examples/cluster_cytof.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Flow Cytometry Clustering Phenotyping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Flow Cytometry Clustering Phenotyping compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Flow Cytometry Clustering Phenotyping this skillGPTomics/bioSkills1.2k1 repos~2.3kAutomated safety check: PassMIT
Embeddings via 9Routerdecolua/9router31k—~604Automated safety check: PassMIT
Esmfold2JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.0
PR Demomikeyobrien/ralph-orchestrator3.2k—~1.3kAutomated safety check: PassMIT
Yas Demo Texttmck-code/yet-another-statusline242—~649Automated safety check: PassBSD-3-Clause
Michel CLI Demo RecorderPackmindHub/packmind318—~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    31k GitHub stars~604 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • PR Demo

    mikeyobrien/ralph-orchestrator

    A skill your agent uses when creating animated demos (GIFs) for pull requests or documentation.

    3.2k GitHub stars~1.3k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Yas Demo Text

    tmck-code/yet-another-statusline

    Convert make demo/img statusline snapshots into ANSI-stripped plain text for diffing and PR embedding.

    242 GitHub stars~649 tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Michel CLI Demo Recorder

    PackmindHub/packmind

    Produce proof-of-execution demos of the Packmind CLI (packmind-cli) as terminal-styled images (colors and formatting preserved exactly), for embedding in a GitHub PR.

    318 GitHub stars~3.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Copilot SDK

    github/awesome-copilot

    Official

    Build agentic applications with GitHub Copilot SDK. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 4 repos~6.3k tokens
    AI & LLM EngineeringAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Flow Cytometry Clustering Phenotyping

What does Bio Flow Cytometry Clustering Phenotyping do?

Unsupervised clustering and cell-type identification for high-dimensional flow, spectral, and mass cytometry - FlowSOM, PhenoGraph, FlowSOM-via-CATALYST, with UMAP/tSNE for visualization. Bio Flow Cytometry Clustering Phenotyping is an agent skill from GPTomics/bioSkills. Unsupervised clustering and cell-type identification for high-dimensional flow, spectral, and mass cytometry - FlowSOM, PhenoGraph, FlowSOM-via-CATALYST, with UMAP/tSNE for visualization.

When should I use Bio Flow Cytometry Clustering Phenotyping?

Bio Flow Cytometry Clustering Phenotyping fits situations like: discovering populations without predefined gates; choosing a clustering algorithm; selecting the number of metaclusters; annotating clusters into cell types.

How do I install Bio Flow Cytometry Clustering Phenotyping in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-flow-cytometry-clustering-phenotyping -a claude-code`. Or copy the skill folder (flow-cytometry/clustering-phenotyping in GPTomics/bioSkills) into .claude/skills/bio-flow-cytometry-clustering-phenotyping in your project. Claude Code loads it when a task matches its description.

How do I install Bio Flow Cytometry Clustering Phenotyping in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-flow-cytometry-clustering-phenotyping -a codex`. Or copy the skill folder (flow-cytometry/clustering-phenotyping in GPTomics/bioSkills) into .agents/skills/bio-flow-cytometry-clustering-phenotyping in your project. Codex loads it when a task matches its description.

Can I use Bio Flow Cytometry Clustering Phenotyping in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-flow-cytometry-clustering-phenotyping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-flow-cytometry-clustering-phenotyping, .gemini/skills/bio-flow-cytometry-clustering-phenotyping, .github/skills/bio-flow-cytometry-clustering-phenotyping and .opencode/skills/bio-flow-cytometry-clustering-phenotyping in your project.

What does Bio Flow Cytometry Clustering Phenotyping need to run?

Going by SKILL.md and its folder, Bio Flow Cytometry Clustering Phenotyping needs R for the scripts in its folder.

Does Bio Flow Cytometry Clustering Phenotyping access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Flow Cytometry Clustering Phenotyping safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Flow Cytometry Clustering Phenotyping use?

Bio Flow Cytometry Clustering Phenotyping is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Flow Cytometry Clustering Phenotyping use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Flow Cytometry Clustering Phenotyping?

Skills that share tags, products or a category with Bio Flow Cytometry Clustering Phenotyping: Embeddings via 9Router (decolua/9router, 31k stars), Esmfold2 (JimLiu/science-skills, 228 stars), PR Demo (mikeyobrien/ralph-orchestrator, 3.2k stars) and Yas Demo Text (tmck-code/yet-another-statusline, 242 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Flow Cytometry Clustering Phenotyping?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.