Agent skill

Umap Tsne Analysis

by aipoch in aipoch/medical-research-skills

A skill your agent uses when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE…

MITAuto-check passedResearch & Science

Install Umap Tsne Analysis

skills CLI
$ npx skills add aipoch/medical-research-skills --skill umap-tsne-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills umap-tsne-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'awesome-med-research-skills/Data Analysis/umap-tsne-analysis' .claude/skills/umap-tsne-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
umap-tsne-analysis
GitHub stars
1.9k
Token cost
~2.7k tokens
SKILL.md length
1,072 words
Files
18 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE…

  • Works in 4 steps: Validate Input → Prepare Matrix → Run Dimensionality Reduction → …
  • Performing sample-level dimensionality reduction and visualization on abundance
  • SKILL.md covers Prerequisites, When to Read External Files, Usage and Arguments, plus 10 more sections
  • Runs R scripts from its folder; reaches cloud.r-project.org

What it does

Umap Tsne Analysis is an agent skill from aipoch/medical-research-skills. Use when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE coordinates and plots for group separation assessment. NOT for: differential expression testing, single-cell workflows requiring dedicated embeddings pipelines, or analyses without a sample grouping file.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including scripts and reference files (for example `eval_report_umap-tsne-analysis_result.json`, `references/algorithm.md` and `references/cli-guide.md`).

It sits in Research & Science, covering Bioinformatics and Embeddings. It works with UMAP. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Performing sample-level dimensionality reduction and visualization on abundance
  • OTU-style matrices with a companion group file
  • Generating UMAP and/or t-SNE coordinates and plots for group separation assessment

Example prompts

  • “/umap-tsne-analysis”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Validate Input
  2. Prepare Matrix
  3. Run Dimensionality Reduction
  4. Generate Visualizations

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (R, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • cloud.r-project.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Umap Tsne Analysis loads about 2.7k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 98 tokens; SKILL.md has 1,072 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 1,072 words, ~2,712 tokens.

Download SKILL.mdSave it as .claude/skills/umap-tsne-analysis/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
umap-tsne-analysis
description
Use when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE coordinates and plots for group separation assessment. NOT for: differential expression testing, single-cell workflows requiring dedicated embeddings pipelines, or analyses without a sample grouping file.
license
MIT
skill-author
AIPOCH

UMAP and t-SNE Analysis

Prerequisites

Run the following before the first analysis to install all required R packages:

bash
Rscript scripts/install_dependencies.R

Alternative manual installation:

bash
Rscript -e "install.packages(c('optparse','data.table','Rtsne','umap','ggplot2','vegan','R.utils'), repos='https://cloud.r-project.org')"

Note: R.utils is only required when --timeout > 0, but pre-installing it avoids environment drift across runs. testthat is installed by scripts/install_dependencies.R as the development test dependency.

The skill cannot run until these packages are installed. In new or bare R environments, always run the prerequisite step first.


When to Read External Files

SituationFile to ReadPurpose
Need algorithm detailsreferences/algorithm.mdDimensionality reduction methods, assumptions, parameter interpretation
Need to run analysisscripts/main.RExecute: Rscript scripts/main.R --input_file ... --group_file ...
Encounter errorsreferences/troubleshooting.mdCommon errors and solutions
Need CLI examplesreferences/cli-guide.mdDetailed CLI usage examples
Need test datatests/data/Sample input files for testing

Usage

bash
Rscript scripts/main.R \
  --input_file ./otu_table.csv \
  --group_file ./group_info.csv \
  --output_dir ./output/ \
  --method both \
  --seed 42

Arguments

ShortLongTypeDefaultDescription
-i--input_filecharacterrequiredAbundance / OTU matrix file
-g--group_filecharacterrequiredGroup information file
-o--output_dircharacter./output/Output directory
-m--methodcharacterbothMethod: tsne, umap, or both
--sample_id_colcharacterfirst columnSample ID column in group file
--group_colcharactersecond columnGroup column in group file
--perplexitynumeric25t-SNE perplexity
--thetanumeric0.0t-SNE theta
--pcalogicalFALSEWhether to use PCA before t-SNE
--check_duplicateslogicalFALSEWhether t-SNE should check duplicated rows
--normalizelogicalTRUEWhether to normalize data before UMAP
--norm_methodcharacterhellingerNormalization method for vegan::decostand()
--n_neighborsinteger10UMAP neighborhood size
-s--seedinteger42Random seed for reproducibility
-t--timeoutinteger0Timeout in seconds; 0 disables timeout

Dependency Baseline

The skill was validated with the exact package baseline recorded in dependencies.lock.tsv.

PackageTested Version
optparse1.7.5
data.table1.15.4
Rtsne0.17
umap0.2.10.0
ggplot23.4.0
vegan2.7.3
R.utils2.13.0
testthat3.1.2

Use this file as the reproducibility baseline when validating a new environment.


Input Format

Abundance / OTU Matrix (input_file)

Features as rows, samples as columns, CSV/TSV-like tabular file with feature ID in the first column.

csv
OTU_ID,S1,S2,S3,S4
OTU_1,10,3,0,5
OTU_2,2,8,1,0
OTU_3,0,0,6,9
Group File (group_file)

Tabular file with at least two columns: sample ID and group label.

csv
SampleID,Group
S1,Control
S2,Control
S3,Treatment
S4,Treatment

Requirements:

  • At least 2 groups with at least 2 samples per group are required.
  • All sample IDs in the group file must exist in the matrix columns.
  • Single-group inputs will produce a SKILL_INVALID_PARAMETER error because dimensionality reduction without group contrast produces uninterpretable plots.

Output Files

FileDescription
table/tsne_coordinates.csvt-SNE coordinates with sample and group annotations
table/umap_coordinates.csvUMAP coordinates with sample and group annotations
plot/tsne_plot.pdft-SNE scatter plot with group colors and ellipses
plot/umap_plot.pdfUMAP scatter plot with group colors and ellipses
data/session_info.txtR session and package version info
data/analysis_data.rdaSaved analysis object with aligned matrix, metadata, colors, and runtime parameters

Workflow

Step 1: Validate Input
  • Check that matrix file and group file exist
  • Resolve sample ID and group columns
  • Validate at least 2 groups and at least 2 samples per group
  • Ensure all group-file sample IDs exist in the matrix
  • Remove samples with zero total abundance after alignment if needed
Step 2: Prepare Matrix
  • Convert input table into numeric matrix
  • Align matrix columns to sample order from the group file
  • Transpose matrix so rows become samples and columns become features
Step 3: Run Dimensionality Reduction
  • Run t-SNE if --method tsne or --method both
  • Run UMAP if --method umap or --method both
  • Apply fixed random seed for reproducibility
Step 4: Generate Visualizations
  • Plot sample embeddings
  • Color points by group
  • Draw group ellipses when enabled
  • Save PDF outputs

Methods

t-SNE

t-SNE is a non-linear dimensionality reduction method that preserves local neighborhood structure. It is useful for identifying local sample clustering patterns.

UMAP

UMAP is a manifold learning method that aims to preserve both local and some global structure. It is often faster than t-SNE and can produce stable low-dimensional embeddings when parameters are chosen appropriately.

Normalization

When --normalize TRUE, the script uses vegan::decostand() with the selected --norm_method before UMAP. This is helpful for abundance-style ecological matrices.


Show full SKILL.md (440 more words)Show less

Agent Response Contract

After a successful run, report:

  1. Method(s) run (tsne, umap, or both)
  2. Sample count and group count processed
  3. Key parameters used (perplexity for t-SNE, n_neighbors for UMAP)
  4. Group separation quality (describe visible clustering from coordinate ranges if accessible)
  5. Artifact paths: coordinate CSV(s) and plot PDF(s) produced

Examples

Basic Usage
bash
Rscript scripts/main.R \
  -i otu_table.csv \
  -g group_info.csv \
  -o ./output \
  -m both
Only t-SNE
bash
Rscript scripts/main.R \
  -i otu_table.csv \
  -g group_info.csv \
  -o ./output \
  -m tsne \
  --perplexity 10
Only UMAP with Custom Group Column
bash
Rscript scripts/main.R \
  -i otu_table.csv \
  -g metadata.csv \
  -o ./output \
  -m umap \
  --sample_id_col SampleID \
  --group_col Treatment \
  --n_neighbors 15

Error Handling

Common Errors
ErrorCauseSolution
SKILL_FILE_NOT_FOUNDInput file does not existCheck file path
SKILL_MISSING_COLUMNSGroup or matrix file lacks required columnsVerify file format
SKILL_SAMPLE_MISMATCHSample IDs in group file do not match matrix columnsCheck sample naming consistency
SKILL_EMPTY_DATAMatrix becomes empty after preprocessingCheck input values and filtering
SKILL_INVALID_PARAMETERInvalid method, invalid parameter value, or single-group inputAdjust CLI arguments; ensure at least 2 groups are present
SKILL_PACKAGE_NOT_FOUNDRequired R package is missingRun Rscript scripts/install_dependencies.R; note that file errors will only surface after packages are installed
SKILL_TIMEOUTAnalysis exceeded the configured timeoutIncrease --timeout or set --timeout 0

IF error persists, READ: references/troubleshooting.md

Troubleshooting note: In environments where packages are not yet installed, SKILL_PACKAGE_NOT_FOUND will fire before file-validation errors. Install dependencies first, then re-run to expose any file-related errors.


Input Validation

This skill accepts:

  1. An abundance or OTU-style feature matrix (CSV/TSV, features as rows, samples as columns)
  2. A group file with at least two groups (CSV/TSV, sample IDs and group labels)

If the user's request does not involve UMAP or t-SNE dimensionality reduction for group separation visualization — for example, asking to run differential expression testing, process single-cell RNA-seq with specialized pipelines, perform clustering without a group file, or impute missing values — do not proceed with the workflow. Instead respond:

"UMAP and t-SNE Analysis is designed to perform sample-level dimensionality reduction and visualization on abundance or OTU-style matrices. Your request appears to be outside this scope. Please provide a feature matrix and group file for UMAP/t-SNE, or use a more appropriate tool for differential expression testing, single-cell analysis, or clustering."


Testing

Test with Sample Data
bash
Rscript scripts/install_dependencies.R

Rscript scripts/main.R --help

Rscript scripts/main.R \
  -i tests/data/otu_table.csv \
  -g tests/data/group_info.csv \
  -o tests/output/ \
  -m both

Rscript tests/test_skill.R

Rscript tests/run_smoke_test.R
Validation Commands
bash
ls -la tests/output/
ls -la tests/output/table
ls -la tests/output/plot
ls -la tests/output/data
wc -l tests/output/table/tsne_coordinates.csv
wc -l tests/output/table/umap_coordinates.csv

The canonical sample data live in tests/data/. Use those files for examples, smoke tests, and regression checks. The canonical output layout is output_dir/table, output_dir/plot, and output_dir/data.


Implementation Checklist

  • CLI parsing with optparse
  • set.seed() for reproducibility
  • requireNamespace() dependency checks
  • Dependency bootstrap script
  • Session info recording
  • File reading instructions in SKILL.md
  • Modular script structure
  • Error handling with SKILL_* codes
  • Test data provided in tests/data/
  • Version-pinned dependency baseline in dependencies.lock.tsv
  • Automated testthat coverage for validation and plotting edge cases
  • Scripts in scripts/ directory
  • References in references/ directory

Last updated: 2026-04-27 | Version: 1.1.0

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts, references) in awesome-med-research-skills/Data Analysis/umap-tsne-analysis of aipoch/medical-research-skills.

  • SKILL.md
  • dependencies.lock.tsv
  • eval_report_umap-tsne-analysis_result.json
  • references/algorithm.md
  • references/cli-guide.md
  • references/troubleshooting.md
  • scripts/dim_reduction_methods.R
  • scripts/functions.R
  • scripts/install_dependencies.R
  • scripts/main.R
  • scripts/run_analysis.R
  • scripts/utils.R
  • scripts/visualization.R
  • tests/data/group_info.csv
  • tests/data/otu_table.csv
  • tests/run_smoke_test.R
  • tests/testthat
  • … and 1 more

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Umap Tsne Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Umap Tsne Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Umap Tsne Analysis this skillaipoch/medical-research-skills1.9k—~2.7kAutomated safety check: PassMIT
Bio Data Visualization Dimensionality Reduction PlotsGPTomics/bioSkills1.2k2 repos~4.8kAutomated safety check: PassMIT
Sc ClusteringTianGzlab/OmicsClaw161—~2.4kAutomated safety check: PassApache-2.0
Evo2JimLiu/science-skills2284 repos~1.3kAutomated safety check: PassApache-2.0
ScgptJimLiu/science-skills2284 repos~1.3kAutomated safety check: PassApache-2.0
Genimldavila7/claude-code-templates33k11 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions)…

    1.2k GitHub starsUsed in 2 repos~4.8k tokens
    Data & AnalyticsAuto-check passed
  • Sc Clustering

    TianGzlab/OmicsClaw

    Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData.

    161 GitHub stars~2.4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Evo2

    JimLiu/science-skills

    Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.

    228 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    228 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Geniml

    davila7/claude-code-templates

    This skill should be used when working with genomic interval data (BED files) for machine learning tasks.

    33k GitHub starsUsed in 11 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Given a gene and a single-cell atlas, compute how cell-type-specific its expression is — the tau specificity index, Sarle's expression bimodality coefficient, and the cell types that drive the…

    1.2k GitHub stars~4.3k tokensUpdated yesterday
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    1.9k GitHub stars~2.2k tokensUpdated 23 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    1.9k GitHub stars~1.4k tokensUpdated 23 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    1.9k GitHub stars~3.7k tokensUpdated 23 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    1.9k GitHub stars~1.8k tokensUpdated 23 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    1.9k GitHub stars~1.7k tokensUpdated 23 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    1.9k GitHub stars~1.3k tokensUpdated 23 days ago
    Auto-check passed

Works with

Questions about Umap Tsne Analysis

What does Umap Tsne Analysis do?

A skill your agent uses when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE…. Umap Tsne Analysis is an agent skill from aipoch/medical-research-skills. Use when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE coordinates and plots for group separation assessment.

When should I use Umap Tsne Analysis?

Umap Tsne Analysis fits situations like: performing sample-level dimensionality reduction and visualization on abundance; OTU-style matrices with a companion group file; generating UMAP and/or t-SNE coordinates and plots for group separation assessment.

How do I install Umap Tsne Analysis in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill umap-tsne-analysis -a claude-code`. Or copy the skill folder (awesome-med-research-skills/Data Analysis/umap-tsne-analysis in aipoch/medical-research-skills) into .claude/skills/umap-tsne-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Umap Tsne Analysis in Codex?

Run `npx skills add aipoch/medical-research-skills --skill umap-tsne-analysis -a codex`. Or copy the skill folder (awesome-med-research-skills/Data Analysis/umap-tsne-analysis in aipoch/medical-research-skills) into .agents/skills/umap-tsne-analysis in your project. Codex loads it when a task matches its description.

Can I use Umap Tsne Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill umap-tsne-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/umap-tsne-analysis, .gemini/skills/umap-tsne-analysis, .github/skills/umap-tsne-analysis and .opencode/skills/umap-tsne-analysis in your project.

What does Umap Tsne Analysis need to run?

Going by SKILL.md and its folder, Umap Tsne Analysis needs R for the scripts in its folder.

Does Umap Tsne Analysis access the network?

SKILL.md names 1 domain. In commands or code: cloud.r-project.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Umap Tsne Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Umap Tsne Analysis use?

Umap Tsne Analysis is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Umap Tsne Analysis use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Umap Tsne Analysis?

Skills that share tags, products or a category with Umap Tsne Analysis: Bio Data Visualization Dimensionality Reduction Plots (GPTomics/bioSkills, 1.2k stars), Sc Clustering (TianGzlab/OmicsClaw, 161 stars), Evo2 (JimLiu/science-skills, 228 stars) and Scgpt (JimLiu/science-skills, 228 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Umap Tsne Analysis?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,937 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.