Agent skill

Bio Data Visualization Dimensionality Reduction Plots

by GPTomics in GPTomics/bioSkills

Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions)…

MITAuto-check passedData & Analytics

Install Bio Data Visualization Dimensionality Reduction Plots

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-data-visualization-dimensionality-reduction-plots -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-data-visualization-dimensionality-reduction-plots --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-visualization/dimensionality-reduction-plots .claude/skills/bio-data-visualization-dimensionality-reduction-plots && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-data-visualization-dimensionality-reduction-plots
GitHub stars
1.2k
Used in
2 other repos
Token cost
~4.8k tokens
SKILL.md length
1,931 words
Files
3
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions)…

  • Works in 3 steps: Initialize with PCA, not random —… → High learning rate — learning_rate =… → Early exaggeration — exaggeration=12,…
  • Visualizing high-dimensional data — bulk PCA
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithmic Taxonomy and Decision Tree by Scenario, plus 11 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Data Visualization Dimensionality Reduction Plots is an agent skill from GPTomics/bioSkills. Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions), hyperparameter sensitivity, and the well-documented limits of 2D embeddings. Covers PCA biplot/scree/loadings, t-SNE PCA initialization (Kobak-Berens 2019), UMAP nneighbors/mindist trade-offs, and the Chari-Pachter 2023 critique. Use when visualizing high-dimensional data — bulk PCA, single-cell embeddings, multi-omics integration…

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/embedding_phd.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Data visualization, Embeddings and Bioinformatics. It works with UMAP and Scanpy. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Visualizing high-dimensional data — bulk PCA
  • Single-cell embeddings
  • Multi-omics integration projections

Example prompts

  • “/bio-data-visualization-dimensionality-reduction-plots”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Initialize with PCA, not random — init='pca' (openTSNE) or pre-compute PCA scores as init
  2. High learning rate — learning_rate = n/12 (n = number of points), not the default 200
  3. Early exaggeration — exaggeration=12, early_exaggeration_iter=250 for large data

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Data Visualization Dimensionality Reduction Plots loads about 4.8k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 1,931 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,931 words, ~4,838 tokens.

Download SKILL.mdSave it as .claude/skills/bio-data-visualization-dimensionality-reduction-plots/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-data-visualization-dimensionality-reduction-plots
description
Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions), hyperparameter sensitivity, and the well-documented limits of 2D embeddings. Covers PCA biplot/scree/loadings, t-SNE PCA initialization (Kobak-Berens 2019), UMAP n_neighbors/min_dist trade-offs, and the Chari-Pachter 2023 critique. Use when visualizing high-dimensional data — bulk PCA, single-cell embeddings, multi-omics integration projections.
tool_type
mixed
primary_tool
scanpy

Version Compatibility

Reference examples tested with: scanpy 1.10+, anndata 0.10+, scikit-learn 1.4+, umap-learn 0.5+, openTSNE 1.0+, phate 1.0+, ggplot2 3.5+, PCAtools 2.16+, matplotlib 3.8+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Dimensionality-Reduction Plots

"Make a PCA / UMAP / t-SNE plot" -> Choose a projection method aligned with what the plot must reveal — variance explained (PCA), local neighborhood structure (t-SNE), manifold approximation with some global structure (UMAP), or continuous transitions (PHATE). Set hyperparameters deliberately. Communicate the projection's limits and refuse to over-interpret 2D distances.

  • Python: sklearn.decomposition.PCA, openTSNE, umap-learn, phate, scanpy.tl.umap / scanpy.tl.tsne / scanpy.tl.pca
  • R: prcomp, PCAtools::pca, Seurat::RunPCA / RunUMAP / RunTSNE, phateR

The Single Most Important Modern Insight -- 2D Embeddings Distort

Chari & Pachter 2023 PLOS Comp Biol 19:e1011288 demonstrated that 2D embeddings of single-cell data lose >95% of the high-dimensional geometry — local neighborhoods are preserved by construction, but distances between distant cells, density estimates, and global topology are NOT preserved. The "specious art" of single-cell genomics is the practice of reading 2D layout as biology.

Practical consequence: a UMAP plot communicates "these cells are similar locally" and nothing more. Distance between clusters is meaningless. Density of points within a cluster is dominated by the embedding's repulsion parameter, not the underlying biology. A trajectory inferred from "the gap" between two clusters in UMAP space is an artifact unless validated against the high-dimensional data (RNA velocity, diffusion pseudotime, PHATE).

A second foundational paper is Kobak & Berens 2019 Nat Commun 10:5416 on t-SNE for single-cell: PCA initialization + early-exaggeration + multi-scale similarity kernels recover more global structure than default t-SNE settings. The same logic applies to UMAP via init='spectral' (default) and min_dist.

Algorithmic Taxonomy

MethodPreservesHyperparametersStrengthFails when
PCALinear variance (orthogonal, ordered)n_components, scalingInterpretable via loadings; deterministic; variance % per axisNon-linear manifolds; high-dim data with few effective dims
t-SNE (van der Maaten 2008)Local neighborhoods (Student-t similarity)perplexity (typ. 30-50), learning_rate, n_iter, initCrisp cluster separationGlobal distances meaningless; cluster sizes deceptive; non-deterministic
UMAP (McInnes 2018, Becht 2018)Manifold local + partial globaln_neighbors (typ. 15-50), min_dist (typ. 0.1-0.5), spreadFaster than t-SNE; better global preservation than default t-SNE; deterministic given seedStill distorts; n_neighbors small -> shattered; large -> homogenized
PHATE (Moon 2019)Continuous transitions, branching trajectoriesk (knn), t (diffusion power)Best for developmental trajectories; preserves transition geometrySlower; less canonical for clustering display
Diffusion mapDiffusion distanceepsilon, n_componentsTheoretically motivated; supports pseudotimeLess visually striking; less commonly used
MDS / classical MDSGlobal Euclidean distancesn_components, dissimilarity matrixHonest about distance preservationComputationally expensive >5000 points
IsomapGeodesic distance on knn graphn_neighbors, n_componentsCaptures non-linear manifoldSensitive to k; less popular than UMAP
Force-directed (PAGA, ForceAtlas2)Graph topologyLayout-specificBest for connectivity (PAGA cluster graph)Not for dense cells; aesthetic

Decision Tree by Scenario

ScenarioRecommendedWhy
Bulk RNA-seq sample QCPCA on log-vst counts; show PC1 vs PC2 with metadata colorVariance explained is meaningful for batch detection
Single-cell broad cluster overviewUMAP n_neighbors=30, min_dist=0.3 after PCA(50)Standard; preserves clusters; faster than t-SNE
Single-cell with delicate trajectoriesPHATE OR diffusion mapPreserves continuous transitions
Cluster cardinality / boundary visualizationt-SNE with PCA init, perplexity=50 (Kobak-Berens)Crisper cluster separation than UMAP
Multi-omics integration projectionMOFA factors + PCA, or UMAP of joint embeddingPer-omics projection often misleading
Spatial transcriptomics with histologyUMAP for transcriptional axis; SEPARATE spatial scatterUMAP collapses physical space
Identify which genes drive variationPCA biplot with loadings as arrowsLoadings are interpretable; UMAP/t-SNE has no loadings
Demonstrating batch confoundPCA color by batch -- if PC1/PC2 separates batches, batch is the dominant varianceUMAP can hide batch effect via local neighborhood preservation
Visualizing 50 conditionsUMAP/t-SNE for nuance; faceted PCA for interpretabilityMethod choice depends on question

PCA -- The Underused Workhorse

PCA is interpretable, deterministic, and the loadings explain WHY samples cluster — UMAP/t-SNE cannot do this. For bulk RNA-seq sample QC, PCA is the right answer 90% of the time.

Goal: Project samples into a low-dim space whose axes are linear combinations of features ordered by variance explained, then visualize PC1 vs PC2 colored by metadata.

Approach: Variance-stabilize counts (DESeq2 vst() / rlog()); run PCA on transposed expression matrix; annotate axes with variance-explained percentages; layer screeplot and loadings plot to support interpretation.

r
library(DESeq2)
library(PCAtools)
library(ggplot2)
vsd <- vst(dds, blind = FALSE)
p <- pca(assay(vsd), metadata = as.data.frame(colData(dds)))
biplot(p, colby = 'condition', shape = 'batch', lab = NULL,
       hline = 0, vline = 0,
       legendPosition = 'right',
       title = paste0('PCA: PC1 (', round(p$variance[1], 1), '%) vs PC2 (', round(p$variance[2], 1), '%)'))
screeplot(p, components = 1:10)
loadings_plot <- plotloadings(p, components = 1, rangeRetain = 0.05)
python
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt

pca = PCA(n_components=10)
X_pca = pca.fit_transform(X)
var = pca.explained_variance_ratio_

fig, axes = plt.subplots(1, 2, figsize=(10, 4))
axes[0].scatter(X_pca[:, 0], X_pca[:, 1], c=labels, alpha=0.7)
axes[0].set_xlabel(f'PC1 ({var[0]*100:.1f}%)')
axes[0].set_ylabel(f'PC2 ({var[1]*100:.1f}%)')
axes[1].plot(range(1, 11), var, 'o-')
axes[1].set_xlabel('PC')
axes[1].set_ylabel('Variance explained')

Always label axes with variance explained. A PCA plot without PC1 (45%) annotation is unreadable. If PC1 = 5% and PC2 = 4%, apparent "clusters" may be noise.

t-SNE -- Kobak-Berens Modern Defaults

Default t-SNE (Maaten 2008) loses global structure. Kobak-Berens 2019 demonstrated that three changes recover it:

  1. Initialize with PCA, not random — init='pca' (openTSNE) or pre-compute PCA scores as init
  2. High learning rate — learning_rate = n/12 (n = number of points), not the default 200
  3. Early exaggeration — exaggeration=12, early_exaggeration_iter=250 for large data
python
import openTSNE
import numpy as np

# Kobak-Berens defaults
embedding = openTSNE.TSNE(
    perplexity=30,                          # 30-50 typical
    n_iter=750,
    initialization='pca',                   # NOT random
    learning_rate=X.shape[0] / 12,          # scales with n
    n_jobs=-1,
    random_state=42).fit(X)
r
library(Rtsne)
set.seed(42)
ts <- Rtsne(X, perplexity = 30, theta = 0.5, pca_scale = TRUE,
            initial_dims = 50, max_iter = 750)
# Rtsne does not natively support PCA initialization; use external init via Y_init=

Perplexity is the local-vs-global trade-off. Low (5) -> local; high (100) -> global. 30-50 is standard for >1000 points.

UMAP -- Modern Defaults and the Random-Seed Trap

python
import umap
reducer = umap.UMAP(
    n_neighbors=30,        # local-global balance; 15-50 typical
    min_dist=0.3,          # tightness of clusters; 0.1-0.5 typical
    n_components=2,
    metric='euclidean',
    random_state=42)       # reproducibility
embedding = reducer.fit_transform(X)
r
library(uwot)
set.seed(42)
um <- umap(X, n_neighbors = 30, min_dist = 0.3, metric = 'euclidean')
python
# scanpy convention -- after sc.tl.pca, sc.pp.neighbors
sc.pp.neighbors(adata, n_neighbors=30, n_pcs=50)
sc.tl.umap(adata, min_dist=0.3, random_state=42)
sc.pl.umap(adata, color='leiden', palette='tab20', frameon=False,
           legend_loc='on data', legend_fontsize=7,
           save='_clusters.pdf')

min_dist controls tightness, NOT separation. Smaller min_dist = tighter clusters. Does not change which cells cluster together — only how dense the rendering is.

n_neighbors controls local-vs-global. Small n_neighbors = local fragmentation; large n_neighbors = clusters merge.

Random seed matters. UMAP is deterministic given seed; without setting seed, results vary across runs. Always set random_state (umap-learn) or seed= (uwot).

scanpy.pl.umap save trap: save='_x.pdf' writes to sc.settings.figdir (default ./figures/) with prefix umap, producing figures/umap_x.pdf — not the path specified. Default dpi_save = 150 is below journal requirements.

PHATE -- For Continuous Trajectories

python
import phate
phate_op = phate.PHATE(knn=10, decay=40, t='auto', n_jobs=-1, random_state=42)
emb = phate_op.fit_transform(X)
plt.scatter(emb[:, 0], emb[:, 1], c=pseudotime, cmap='viridis', s=5)

PHATE preserves transition geometry — for embryonic development, differentiation trajectories, or any continuous-state biology, PHATE is more faithful than UMAP. For discrete cell types, UMAP is fine.

Per-Method Failure Modes

Over-interpreting UMAP distances

Trigger: Reading "cluster A is closer to cluster B than to C" as biological similarity.

Mechanism: UMAP preserves local neighborhoods; global distances are NOT preserved (Chari-Pachter 2023).

Symptom: Conclusion contradicts hierarchical clustering / RNA velocity / known biology.

Fix: Validate inter-cluster relationships against high-dimensional metrics (correlation, distance in PCA space, RNA velocity).

t-SNE / UMAP without random seed

Trigger: Reproducibility request; figure differs between runs.

Mechanism: Both methods use stochastic optimization; default seed varies.

Symptom: Re-running the script produces visibly different layouts.

Fix: Set random_state=42 (umap-learn, sklearn) or seed=42 (R uwot/Rtsne).

Perplexity too low for the data

Trigger: Default t-SNE perplexity (30) on small dataset (<500 points).

Mechanism: Perplexity > n/3 fails; cells artificially fragment.

Symptom: Plot shows "shattered" small clusters that don't correspond to biology.

Fix: For small n: perplexity = max(5, n/30). For very large n: perplexity 50-100.

PCA without scaling

Trigger: prcomp(X) or PCA().fit(X) without scaling rows/columns first.

Mechanism: Genes with high absolute expression dominate variance; PCA captures library-size effect rather than biological variation.

Symptom: PC1 perfectly correlates with library size or with total expression.

Fix: Use vst() / rlog() (DESeq2) or log + scale (prcomp(X, scale.=TRUE)). For single-cell, normalize then sc.pp.scale_data.

Show full SKILL.md (793 more words)Show less
UMAP cluster shapes "interpreted" as biological signal

Trigger: Reporting that a cluster is "elongated" or "round" as biological observation.

Mechanism: UMAP cluster shape is an artifact of min_dist and n_neighbors, not biology.

Symptom: Reviewer asks "why is the immune cluster elongated?"; no answer except the embedding.

Fix: Do not interpret cluster shape. Report cluster membership and validate biology via marker genes.

scanpy.pl.umap save writes to figures/ subdirectory

Trigger: sc.pl.umap(adata, save='myplot.pdf') with the expectation that myplot.pdf will land in the current directory.

Mechanism: save= is concatenated with sc.settings.figdir (default ./figures/) and prefixed with umap.

Symptom: File not at the requested path; actually at figures/umapmyplot.pdf.

Fix: Set sc.settings.figdir='/abs/path/' AND save='_descriptive.pdf' so result is figures/umap_descriptive.pdf. For full path control use matplotlib.savefig after sc.pl.umap(show=False).

Default DPI 150 below journal requirements

Trigger: scanpy sc.settings.set_figure_params() default dpi_save=150.

Mechanism: Nature/Cell require 300+ DPI for raster.

Symptom: Figure looks fine on screen, rejected at submission.

Fix: sc.set_figure_params(dpi_save=300, figsize=(4, 4)).

PCA loadings interpreted on UMAP/t-SNE coordinates

Trigger: "PC1 axis on UMAP" — projecting loadings onto UMAP.

Mechanism: UMAP/t-SNE coordinates have no linear interpretation; loadings are PCA-specific.

Symptom: Conclusion about "what UMAP-x means" that has no foundation.

Fix: Use PCA when loadings are needed. Show UMAP for visualization and PCA for axis-driving gene identification, separately.

Reconciliation: When Methods Disagree

PatternLikely causeAction
t-SNE shows distinct clusters; UMAP merges themt-SNE over-emphasizes local structure; UMAP n_neighbors too largeBoth views valid; check Leiden cluster assignments rather than embedding
PCA shows batch on PC1; UMAP hides itUMAP preserves local neighborhood within each batchRun UMAP only after batch correction; PCA is the canonical batch-effect diagnostic
PHATE shows continuous trajectory; UMAP shows discrete clustersPHATE preserves transitions; UMAP "blobifies" continuous dataUse PHATE for trajectory display; UMAP for discrete cell-type display
Reproducibility breaks across re-runsRandom seed not setSet seed; document version of umap-learn/openTSNE
Cluster boundaries differ between Seurat/scanpy UMAPDifferent defaults for n_neighbors, min_dist, initStandardize hyperparameters; report explicitly

Operational rule: report ALL hyperparameters used (perplexity, n_neighbors, min_dist, random_state). State the embedding's interpretation limit ("local neighborhood; distances between clusters not meaningful"). For trajectory claims, validate with RNA velocity, pseudotime, or PHATE.

Quantitative Thresholds

ThresholdValueSource
t-SNE perplexity30-50 for n>1000; max(5, n/30) for n<500Maaten 2008; Kobak-Berens 2019
UMAP n_neighbors15-50 default rangeumap-learn docs; Becht 2018
UMAP min_dist0.1-0.5Tighter for crisp clusters, looser for continua
t-SNE learning raten / 12 (Kobak-Berens)Default 200 over-shrinks large data
PCA n_components for downstream UMAP30-50Standard scanpy workflow
Single-cell n_neighbors for sc.pp.neighbors15-30Wolf 2018 Scanpy paper
Save DPI for publication300+Nature/Cell figure guidelines
Random seed42 (or any fixed integer)Reproducibility

Common Errors

Error / symptomCauseSolution
Plot differs between runsRandom seed not setrandom_state=42 always
PC1 = library sizeNo scaling/normalizationvst() or log + scale before PCA
Cluster shapes "interpreted" biologicallyUMAP artifactDo not interpret shape; report membership
scanpy save writes to wrong pathfigdir + prefix concatenationSet figdir explicitly OR use matplotlib.savefig
t-SNE fragments small datasetPerplexity too high for nUse perplexity = max(5, n/30)
Inter-cluster "distance" used in trajectory claimUMAP distances not meaningfulValidate with RNA velocity / PHATE
Loadings interpreted on UMAP axesUMAP has no loadingsUse PCA for loading-driven interpretation

Anticipated Reviewer Pushback

PushbackStandard response
"Why UMAP and not t-SNE?"UMAP for cluster overview (faster, better global preservation given Becht 2018); t-SNE in supplementary if cluster boundaries are the focus
"What hyperparameters?"Explicit n_neighbors, min_dist, n_pcs, random_state in caption AND methods
"Why is cluster X shaped this way?"UMAP/t-SNE cluster shape is an embedding artifact; cluster membership is the biological observation
"Are these trajectories real?"Validated via RNA velocity / PHATE / diffusion pseudotime (NOT inferred from UMAP layout alone)
"Why PCA?"Variance explained per axis is interpretable for sample QC; loadings identify driving genes (uniquely PCA, NOT UMAP)

References

  • Becht E, McInnes L, Healy J, et al. 2019. Dimensionality reduction for visualizing single-cell data using UMAP. Nat Biotechnol 37(1):38-44.
  • Chari T, Pachter L. 2023. The specious art of single-cell genomics. PLOS Comp Biol 19(8):e1011288.
  • Kobak D, Berens P. 2019. The art of using t-SNE for single-cell transcriptomics. Nat Commun 10:5416.
  • McInnes L, Healy J, Melville J. 2018. UMAP: Uniform Manifold Approximation and Projection for dimension reduction. arXiv:1802.03426.
  • Moon KR, van Dijk D, Wang Z, et al. 2019. Visualizing structure and transitions in high-dimensional biological data. Nat Biotechnol 37(12):1482-1492.
  • van der Maaten L, Hinton G. 2008. Visualizing data using t-SNE. J Mach Learn Res 9:2579-2605.
  • Wolf FA, Angerer P, Theis FJ. 2018. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol 19:15.
  • single-cell/preprocessing - PCA / neighbor-graph computation before embedding
  • single-cell/clustering - Leiden / Louvain cluster assignments visualized in UMAP
  • single-cell/trajectory-inference - PHATE, diffusion pseudotime, RNA velocity for trajectory claims
  • data-visualization/color-palettes - Categorical palette for cluster labels
  • data-visualization/distribution-plots - Per-cluster gene-expression follow-up
  • data-visualization/heatmaps-clustering - Alternative view of the same high-dim matrix

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in data-visualization/dimensionality-reduction-plots of GPTomics/bioSkills.

  • SKILL.md
  • examples/embedding_phd.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Data Visualization Dimensionality Reduction Plots next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Data Visualization Dimensionality Reduction Plots compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Data Visualization Dimensionality Reduction Plots this skillGPTomics/bioSkills1.2k2 repos~4.8kAutomated safety check: PassMIT
Umap Tsne Analysisaipoch/medical-research-skills2k—~2.7kAutomated safety check: PassMIT
Genimlaipoch/medical-research-skills2k—~1.9kAutomated safety check: PassMIT
Sc ClusteringTianGzlab/OmicsClaw161—~2.4kAutomated safety check: PassApache-2.0
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
Scholar Lingjoshzyj/open-scholar-skill168—~6.7kAutomated safety check: PassCustom licence

Similar skills

  • Umap Tsne Analysis

    aipoch/medical-research-skills

    A skill your agent uses when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE…

    2k GitHub stars~2.7k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed
  • Geniml

    aipoch/medical-research-skills

    Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run…

    2k GitHub stars~1.9k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed
  • Sc Clustering

    TianGzlab/OmicsClaw

    Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData.

    161 GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Scholar Ling

    joshzyj/open-scholar-skill

    Design and analyze studies in sociolinguistics, language variation, acoustic phonetics, discourse analysis, language contact, and computational linguistics.

    168 GitHub stars~6.7k tokensUpdated 19 days ago
    Data & AnalyticsAuto-check passed
  • Metagenomic Krona Chart

    aipoch/medical-research-skills

    Analyze data with metagenomic-krona-chart using a reproducible workflow, explicit validation, and structured outputs for review-ready interpretation.

    2k GitHub stars~2.5k tokensUpdated 21 days ago
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Data Visualization Dimensionality Reduction Plots

What does Bio Data Visualization Dimensionality Reduction Plots do?

Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions)…. Bio Data Visualization Dimensionality Reduction Plots is an agent skill from GPTomics/bioSkills. Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions), hyperparameter sensitivity, and the well-documented limits of 2D embeddings.

When should I use Bio Data Visualization Dimensionality Reduction Plots?

Bio Data Visualization Dimensionality Reduction Plots fits situations like: visualizing high-dimensional data — bulk PCA; single-cell embeddings; multi-omics integration projections.

How do I install Bio Data Visualization Dimensionality Reduction Plots in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-data-visualization-dimensionality-reduction-plots -a claude-code`. Or copy the skill folder (data-visualization/dimensionality-reduction-plots in GPTomics/bioSkills) into .claude/skills/bio-data-visualization-dimensionality-reduction-plots in your project. Claude Code loads it when a task matches its description.

How do I install Bio Data Visualization Dimensionality Reduction Plots in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-data-visualization-dimensionality-reduction-plots -a codex`. Or copy the skill folder (data-visualization/dimensionality-reduction-plots in GPTomics/bioSkills) into .agents/skills/bio-data-visualization-dimensionality-reduction-plots in your project. Codex loads it when a task matches its description.

Can I use Bio Data Visualization Dimensionality Reduction Plots in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-data-visualization-dimensionality-reduction-plots -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-data-visualization-dimensionality-reduction-plots, .gemini/skills/bio-data-visualization-dimensionality-reduction-plots, .github/skills/bio-data-visualization-dimensionality-reduction-plots and .opencode/skills/bio-data-visualization-dimensionality-reduction-plots in your project.

What does Bio Data Visualization Dimensionality Reduction Plots need to run?

Going by SKILL.md and its folder, Bio Data Visualization Dimensionality Reduction Plots needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Data Visualization Dimensionality Reduction Plots access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Data Visualization Dimensionality Reduction Plots safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Data Visualization Dimensionality Reduction Plots use?

Bio Data Visualization Dimensionality Reduction Plots is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Data Visualization Dimensionality Reduction Plots use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Data Visualization Dimensionality Reduction Plots?

Skills that share tags, products or a category with Bio Data Visualization Dimensionality Reduction Plots: Umap Tsne Analysis (aipoch/medical-research-skills, 2k stars), Geniml (aipoch/medical-research-skills, 2k stars), Sc Clustering (TianGzlab/OmicsClaw, 161 stars) and Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Data Visualization Dimensionality Reduction Plots?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.