Agent skill

Umap Learn

by jaechang-hits in jaechang-hits/SciAgent-Skills

UMAP dimensionality reduction for visualization, clustering prep, and feature engineering.

BSD-3-ClauseAuto-check passedData & Analytics

Install Umap Learn

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills umap-learn --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scientific-computing/umap-learn .claude/skills/umap-learn && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
umap-learn
GitHub stars
374
Token cost
~4.7k tokens
SKILL.md length
1,014 words
Files
2 (incl. references)
Skills in repo
169
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

UMAP dimensionality reduction for visualization, clustering prep, and feature engineering.

  • Works in 6 steps: Standard UMAP → Supervised & Semi-Supervised UMAP → Transform New Data → …
  • Tasks that involve Machine learning
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 8 more sections
  • Calls pip

What it does

Umap Learn is an agent skill from jaechang-hits/SciAgent-Skills. UMAP dimensionality reduction for visualization, clustering prep, and feature engineering. Fast nonlinear manifold learning preserving local and global structure. Standard UMAP (fit/transform, sklearn-compatible), supervised/semi-supervised, Parametric UMAP (NN encoder/decoder, TensorFlow), DensMAP (density), AlignedUMAP (temporal/batch). 15+ distance metrics, custom Numba metrics, precomputed distances. For linear reduction use PCA; for neighborhood graphs use sklearn NearestNeighbors.

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/api_reference.md`).

It sits in Data & Analytics, covering Machine learning and Deep learning. It works with UMAP, scikit-learn and TensorFlow. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.

When your agent uses it

  • Tasks that involve Machine learning
  • Tasks that involve Deep learning

Example prompts

  • “/umap-learn”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Standard UMAP
  2. Supervised & Semi-Supervised UMAP
  3. Transform New Data
  4. Parametric UMAP
  5. DensMAP
  6. AlignedUMAP

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • umap-learn.readthedocs.io
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Umap Learn loads about 4.7k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 1,014 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,014 words, ~4,745 tokens.

Download SKILL.mdSave it as .claude/skills/umap-learn/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
umap-learn
description
UMAP dimensionality reduction for visualization, clustering prep, and feature engineering. Fast nonlinear manifold learning preserving local and global structure. Standard UMAP (fit/transform, sklearn-compatible), supervised/semi-supervised, Parametric UMAP (NN encoder/decoder, TensorFlow), DensMAP (density), AlignedUMAP (temporal/batch). 15+ distance metrics, custom Numba metrics, precomputed distances. For linear reduction use PCA; for neighborhood graphs use sklearn NearestNeighbors.
license
BSD-3-Clause

UMAP-Learn

Overview

UMAP (Uniform Manifold Approximation and Projection) is a dimensionality reduction algorithm for visualization and general non-linear dimensionality reduction. It is faster than t-SNE, scales to larger datasets, preserves both local and global structure, and supports supervised learning and embedding of new data points.

When to Use

  • Reducing high-dimensional data to 2D/3D for visualization
  • Preprocessing for density-based clustering (HDBSCAN, DBSCAN)
  • Feature engineering in ML pipelines (transform new data into learned embedding)
  • Supervised/semi-supervised embedding with partial labels
  • Tracking embeddings across time points or batches (AlignedUMAP)
  • Density-preserving embeddings (DensMAP)
  • Neural network-based embedding with custom architectures (Parametric UMAP)
  • For linear dimensionality reduction use PCA (scikit-learn)
  • For neighborhood-graph construction without embedding use scikit-learn NearestNeighbors

Prerequisites

bash
pip install umap-learn

# For Parametric UMAP (neural network variant)
pip install umap-learn[parametric_umap]  # requires TensorFlow 2.x

Critical: Always standardize features before applying UMAP to ensure equal weighting across dimensions.

Quick Start

python
import umap
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import load_digits

# Load and scale data
X, y = load_digits(return_X_y=True)
X_scaled = StandardScaler().fit_transform(X)

# Fit and transform
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled)
print(f"Input: {X_scaled.shape}, Output: {embedding.shape}")
# Input: (1797, 64), Output: (1797, 2)

Core API

1. Standard UMAP

Basic dimensionality reduction following scikit-learn conventions.

python
import umap
from sklearn.preprocessing import StandardScaler

X_scaled = StandardScaler().fit_transform(data)

# Method 1: fit_transform (single step)
embedding = umap.UMAP(
    n_neighbors=15,     # local neighborhood size (2-200)
    min_dist=0.1,       # min distance between embedded points (0.0-0.99)
    n_components=2,     # output dimensions
    metric='euclidean', # distance metric
    random_state=42,    # reproducibility
).fit_transform(X_scaled)
print(f"Embedding shape: {embedding.shape}")

# Method 2: fit + access (for reuse)
reducer = umap.UMAP(random_state=42)
reducer.fit(X_scaled)
embedding = reducer.embedding_  # trained embedding
graph = reducer.graph_          # fuzzy simplicial set (sparse matrix)
python
# Visualization
import matplotlib.pyplot as plt

plt.figure(figsize=(8, 6))
plt.scatter(embedding[:, 0], embedding[:, 1], c=labels, cmap='Spectral', s=5)
plt.colorbar()
plt.title('UMAP Embedding')
plt.tight_layout()
plt.savefig('umap_embedding.png', dpi=150)
2. Supervised & Semi-Supervised UMAP

Incorporate label information to guide embedding via the y parameter.

python
import umap

# Supervised — all labels known
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled, y=labels)

# Semi-supervised — partial labels (mark unlabeled as -1)
semi_labels = labels.copy()
semi_labels[unlabeled_indices] = -1
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled, y=semi_labels)

# Control label influence with target_weight (0.0=unsupervised, 1.0=fully supervised)
reducer = umap.UMAP(
    target_weight=0.7,               # emphasize labels
    target_metric='categorical',     # for classification; use distance metric for regression
    random_state=42
)
embedding = reducer.fit_transform(X_scaled, y=labels)
print(f"Supervised embedding: {embedding.shape}")
3. Transform New Data

Project unseen data into the trained embedding space.

python
import umap
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

# Fit on training data
reducer = umap.UMAP(n_components=10, random_state=42)
X_train_emb = reducer.fit_transform(X_train_scaled)

# Transform test data
X_test_emb = reducer.transform(X_test_scaled)
print(f"Train: {X_train_emb.shape}, Test: {X_test_emb.shape}")

# Works in sklearn Pipelines
from sklearn.pipeline import Pipeline
from sklearn.svm import SVC

pipeline = Pipeline([
    ('scaler', StandardScaler()),
    ('umap', umap.UMAP(n_components=10, random_state=42)),
    ('classifier', SVC())
])
pipeline.fit(X_train, y_train)
accuracy = pipeline.score(X_test, y_test)
print(f"Pipeline accuracy: {accuracy:.3f}")
4. Parametric UMAP

Neural network-based embedding via TensorFlow/Keras. Enables efficient transform, reconstruction, and custom architectures.

python
from umap.parametric_umap import ParametricUMAP

# Default architecture (3-layer, 100-neuron FC network)
embedder = ParametricUMAP(n_components=2, random_state=42)
embedding = embedder.fit_transform(X_scaled)
new_emb = embedder.transform(new_data)  # fast neural network inference
print(f"Parametric embedding: {embedding.shape}")
python
import tensorflow as tf
from umap.parametric_umap import ParametricUMAP

# Custom encoder/decoder for autoencoder mode
input_dim = X_scaled.shape[1]
encoder = tf.keras.Sequential([
    tf.keras.layers.InputLayer(input_shape=(input_dim,)),
    tf.keras.layers.Dense(128, activation='relu'),
    tf.keras.layers.Dense(64, activation='relu'),
    tf.keras.layers.Dense(2),
])
decoder = tf.keras.Sequential([
    tf.keras.layers.InputLayer(input_shape=(2,)),
    tf.keras.layers.Dense(64, activation='relu'),
    tf.keras.layers.Dense(128, activation='relu'),
    tf.keras.layers.Dense(input_dim),
])

embedder = ParametricUMAP(
    encoder=encoder, decoder=decoder, dims=(input_dim,),
    parametric_reconstruction=True, autoencoder_loss=True,
    n_training_epochs=10, batch_size=128,
    n_neighbors=15, min_dist=0.1, random_state=42
)
embedding = embedder.fit_transform(X_scaled)
reconstructed = embedder.inverse_transform(embedding)
print(f"Reconstruction error: {np.mean((X_scaled - reconstructed)**2):.4f}")
5. DensMAP

Variant preserving local density information in the embedding.

python
import umap

reducer = umap.UMAP(
    densmap=True,          # enable DensMAP
    dens_lambda=2.0,       # density preservation weight
    dens_frac=0.3,         # fraction for density estimation
    output_dens=True,      # output density estimates
    n_neighbors=15,
    min_dist=0.1,
    random_state=42
)
embedding = reducer.fit_transform(X_scaled)

# Access density estimates
original_density = reducer.rad_orig_  # density in original space
embedded_density = reducer.rad_emb_   # density in embedded space
print(f"DensMAP embedding: {embedding.shape}")
print(f"Density correlation: {np.corrcoef(original_density, embedded_density)[0,1]:.3f}")
6. AlignedUMAP

Align embeddings across multiple related datasets (time points, batches).

python
from umap import AlignedUMAP

# Multiple related datasets
datasets = [day1_data, day2_data, day3_data]

mapper = AlignedUMAP(
    n_neighbors=15,
    alignment_regularisation=1e-2,  # alignment strength
    alignment_window_size=2,        # align with N adjacent datasets
    n_components=2,
    random_state=42
)
mapper.fit(datasets)

aligned_embeddings = mapper.embeddings_  # list of aligned embedding arrays
print(f"Aligned {len(aligned_embeddings)} datasets")
for i, emb in enumerate(aligned_embeddings):
    print(f"  Dataset {i}: {emb.shape}")

Key Concepts

Parameter Tuning Guide
ParameterLowMedium (default)HighEffect
n_neighbors2-51550-200Local detail vs global structure
min_dist0.00.10.5-0.99Tight clusters vs spread out
n_components225-50Visualization vs ML/clustering
spread0.51.02.0Embedding scale (with min_dist)
Configuration by Use-Case
Use-Casen_neighborsmin_distn_componentsmetric
Visualization150.12euclidean
Clustering (HDBSCAN)300.05-10euclidean
Text/document embedding150.12cosine
Global structure1000.52euclidean
ML feature engineering15-300.110-50euclidean
Binary/set data150.12hamming/jaccard
Supported Metrics

Minkowski family: euclidean, manhattan, chebyshev, minkowski. Spatial: canberra, braycurtis, haversine. Correlation: cosine, correlation. Binary: hamming, jaccard, dice, russellrao, rogerstanimoto, sokalmichener, sokalsneath, yule. Special: precomputed (distance matrix), custom Numba-compiled callables.

Standard UMAP vs Parametric UMAP
FeatureStandardParametric
BackendDirect optimizationTensorFlow neural network
Transform speedModerateFast (neural net inference)
Inverse transformApproximate, expensiveDecoder network, fast
Custom architectureNoYes (CNNs, RNNs, etc.)
Requirementsumap-learnumap-learn + TensorFlow 2.x
Best forQuick explorationProduction pipelines, reconstruction

Common Workflows

Workflow 1: UMAP + HDBSCAN Clustering Pipeline
python
import umap
import hdbscan
import numpy as np
import matplotlib.pyplot as plt
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import adjusted_rand_score

# Step 1: Preprocess
X_scaled = StandardScaler().fit_transform(data)
print(f"Input shape: {X_scaled.shape}")

# Step 2: UMAP for clustering (NOT visualization parameters)
reducer = umap.UMAP(
    n_neighbors=30,     # more global structure for clustering
    min_dist=0.0,       # allow tight packing
    n_components=10,    # higher dims preserve density better than 2D
    metric='euclidean',
    random_state=42
)
embedding = reducer.fit_transform(X_scaled)

# Step 3: HDBSCAN clustering
clusterer = hdbscan.HDBSCAN(min_cluster_size=15, min_samples=5)
cluster_labels = clusterer.fit_predict(embedding)

n_clusters = len(set(cluster_labels)) - (1 if -1 in cluster_labels else 0)
noise = sum(cluster_labels == -1)
print(f"Clusters: {n_clusters}, Noise: {noise}")

# Step 4: Separate 2D embedding for visualization
vis_emb = umap.UMAP(n_neighbors=15, min_dist=0.1, random_state=42).fit_transform(X_scaled)
plt.scatter(vis_emb[:, 0], vis_emb[:, 1], c=cluster_labels, cmap='Spectral', s=5)
plt.colorbar()
plt.title(f'HDBSCAN Clusters (n={n_clusters})')
plt.tight_layout()
plt.savefig('umap_clusters.png', dpi=150)
Workflow 2: Supervised Embedding for Classification
python
import umap
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.metrics import classification_report

# Split and scale
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
scaler = StandardScaler()
X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)

# Supervised UMAP for feature engineering
reducer = umap.UMAP(n_components=10, random_state=42)
X_train_emb = reducer.fit_transform(X_train_s, y=y_train)
X_test_emb = reducer.transform(X_test_s)

# Downstream classifier
clf = SVC(kernel='rbf')
clf.fit(X_train_emb, y_train)
y_pred = clf.predict(X_test_emb)

print(classification_report(y_test, y_pred))
Workflow 3: Exploring Embedding Space with Inverse Transform

Text-only — combines Core API modules 1 and 3 (inverse_transform on standard UMAP):

  1. Fit standard UMAP on data (Core API: Standard UMAP)
  2. Create a grid of points spanning the embedding space
  3. Apply reducer.inverse_transform(grid_points) to reconstruct high-dimensional data
  4. Visualize reconstructed samples to understand embedding regions

Note: inverse transform is approximate; works poorly outside the convex hull of the training embedding.

Key Parameters

ParameterModuleDefaultRangeEffect
n_neighborsUMAP152-200Local vs global structure balance
min_distUMAP0.10.0-0.99Cluster tightness
n_componentsUMAP22-100Output dimensionality
metricUMAP'euclidean'See metrics listDistance calculation method
spreadUMAP1.0>0Embedding scale (with min_dist)
n_epochsUMAPNone (auto)50-500+Training iterations
learning_rateUMAP1.0>0SGD step size
initUMAP'spectral'spectral/random/pcaEmbedding initialization
random_stateUMAPNoneintReproducibility seed
target_weightUMAP0.50.0-1.0Label influence (supervised)
densmapUMAPFalseboolEnable DensMAP
dens_lambdaUMAP2.0>0DensMAP density weight
low_memoryUMAPTrueboolMemory-efficient mode
encoderParametricUMAPNoneKeras modelCustom encoder network
decoderParametricUMAPNoneKeras modelCustom decoder network
n_training_epochsParametricUMAP11-100Neural network training epochs
alignment_regularisationAlignedUMAP0.01>0Alignment strength
alignment_window_sizeAlignedUMAP31-NAdjacent datasets to align
Show full SKILL.md (429 more words)Show less

Best Practices

  1. Always standardize features: Use StandardScaler before UMAP — unscaled features with different ranges will dominate the embedding.

  2. Set random_state for reproducibility: UMAP uses stochastic optimization; results vary between runs without a fixed seed.

  3. Use different parameters for clustering vs visualization: Clustering needs n_neighbors=30, min_dist=0.0, n_components=5-10. Visualization needs n_neighbors=15, min_dist=0.1, n_components=2.

  4. Anti-pattern — interpreting distances literally: UMAP preserves topology, not precise distances. Cluster separations and point distances in the embedding are not proportional to original distances.

  5. Anti-pattern — using 2D embeddings for clustering: 2D projections lose density information. Use 5-10 components for HDBSCAN input.

  6. Consider PCA preprocessing for very high dimensions: For data with >1000 features, reducing to 50-100 PCA components first can speed up UMAP without losing quality.

  7. Use Parametric UMAP for production: When you need fast transform on new data or reconstruction capabilities, Parametric UMAP's neural network provides consistent, fast inference.

Common Recipes

Recipe: Custom Numba Distance Metric
python
from numba import njit
import umap

@njit()
def weighted_euclidean(x, y):
    """Custom distance with feature weights."""
    result = 0.0
    for i in range(x.shape[0]):
        result += (x[i] - y[i]) ** 2 * (1.0 + i * 0.01)  # increasing weight
    return np.sqrt(result)

embedding = umap.UMAP(metric=weighted_euclidean, random_state=42).fit_transform(data)
Recipe: Precomputed Distance Matrix
python
import umap
from scipy.spatial.distance import pdist, squareform

# Compute custom distance matrix
dist_matrix = squareform(pdist(data, metric='correlation'))

# Use precomputed distances
embedding = umap.UMAP(
    metric='precomputed', random_state=42
).fit_transform(dist_matrix)
print(f"Embedding from precomputed: {embedding.shape}")
Recipe: Metric Learning Pipeline
python
import umap
from sklearn.svm import SVC

# Train supervised embedding on labeled data
mapper = umap.UMAP(n_components=10, random_state=42)
train_emb = mapper.fit_transform(X_train, y=y_train)

# Transform unlabeled test data using learned metric
test_emb = mapper.transform(X_test)

# Downstream classifier
clf = SVC().fit(train_emb, y_train)
predictions = clf.predict(test_emb)
print(f"Accuracy: {(predictions == y_test).mean():.3f}")

Troubleshooting

ProblemCauseSolution
Disconnected/fragmented clustersn_neighbors too lowIncrease n_neighbors (try 30-50)
Clusters too spread outmin_dist too highDecrease min_dist (try 0.0-0.05)
All points collapsedBad preprocessing or min_dist too lowCheck StandardScaler; increase min_dist
Poor clustering resultsUsing visualization parameters for clusteringSet n_neighbors=30, min_dist=0.0, n_components=5-10
Transform results differ from trainingDistribution shiftEnsure test data matches training distribution; use Parametric UMAP
Slow on large datasets (>100k)Default settingsSet low_memory=True; preprocess with PCA to 50-100 dims
First run very slowNumba JIT compilationExpected — subsequent runs are fast (compiled cache)
ImportError: umapName conflict with umap packagepip install umap-learn (not pip install umap)
Parametric UMAP import errorMissing TensorFlowpip install umap-learn[parametric_umap]
Non-reproducible resultsMissing random_stateAlways set random_state=42 (or any int)

Bundled Resources

references/api_reference.md

Complete UMAP constructor parameter reference (60+ parameters organized by category: core, training, advanced structural, supervised, transform, performance, DensMAP), all methods and attributes, ParametricUMAP class with autoencoder parameters, AlignedUMAP class, utility functions (nearest_neighbors, fuzzy_simplicial_set). Core parameter tuning guidance was relocated to SKILL.md Key Concepts and Core API modules. Usage examples duplicating SKILL.md workflows omitted.

  • scikit-learn-machine-learning — ML classifiers, preprocessing, pipelines for downstream tasks
  • matplotlib-scientific-plotting — Visualization of UMAP embeddings
  • scikit-bio — Biological distance matrices that can feed into UMAP via metric='precomputed'

References

  • McInnes L, Healy J, Melville J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv:1802.03426
  • Sainburg T, McInnes L, Gentner TQ. Parametric UMAP Embeddings for Representation and Semisupervised Learning. Neural Computation (2021)
  • Narayan A, Berger B, Cho H. Assessing single-cell transcriptomic variability through density-preserving data visualization. Nature Biotechnology (2021) — DensMAP
  • Official docs: https://umap-learn.readthedocs.io/
  • GitHub: https://github.com/lmcinnes/umap

© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/scientific-computing/umap-learn of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/api_reference.md

Open the folder on GitHubat commit 82c862c

Compare with similar skills

Umap Learn next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Umap Learn compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Umap Learn this skilljaechang-hits/SciAgent-Skills374—~4.7kAutomated safety check: PassBSD-3-Clause
ML Model Trainingsecondsky/claude-skills227—~1.7kAutomated safety check: PassMIT
Editomegaml/omegaml108—~206Automated safety check: PassApache-2.0
ML EngineerRightNow-AI/openfang18k—~987Automated safety check: PassApache-2.0
Mlflow OnboardingKilo-Org/kilo-marketplace190—~3.2kAutomated safety check: PassApache-2.0
Molfeatdavila7/claude-code-templates33k9 repos~3.7kAutomated safety check: PassMIT

Similar skills

  • ML Model Training

    secondsky/claude-skills

    Train ML models with scikit-learn, PyTorch, TensorFlow. An agent skill from secondsky/claude-skills.

    227 GitHub stars~1.7k tokensUpdated 12 days ago
    Data & AnalyticsAuto-check passed
  • Edit

    omegaml/omegaml

    how to use the edit command properly

    108 GitHub stars~206 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • ML Engineer

    RightNow-AI/openfang

    Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps

    18k GitHub stars~987 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Mlflow Onboarding

    Kilo-Org/kilo-marketplace

    Onboards users to MLflow by determining their use case (GenAI agents/apps or traditional ML/deep learning) and guiding them through relevant quickstart tutorials and initial integration.

    190 GitHub stars~3.2k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Molfeat

    davila7/claude-code-templates

    Molecular featurization for ML (100+ featurizers). An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 9 repos~3.7k tokens
    Data & AnalyticsAuto-check passed
  • Databricks ML Training

    databricks/databricks-agent-skills

    Official

    Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills.

    345 GitHub stars~4.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 11 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 11 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 11 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 11 days ago
    Auto-check passed

Questions about Umap Learn

What does Umap Learn do?

UMAP dimensionality reduction for visualization, clustering prep, and feature engineering. Umap Learn is an agent skill from jaechang-hits/SciAgent-Skills. UMAP dimensionality reduction for visualization, clustering prep, and feature engineering.

When should I use Umap Learn?

Umap Learn fits situations like: tasks that involve Machine learning; tasks that involve Deep learning.

How do I install Umap Learn in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a claude-code`. Or copy the skill folder (skills/scientific-computing/umap-learn in jaechang-hits/SciAgent-Skills) into .claude/skills/umap-learn in your project. Claude Code loads it when a task matches its description.

How do I install Umap Learn in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a codex`. Or copy the skill folder (skills/scientific-computing/umap-learn in jaechang-hits/SciAgent-Skills) into .agents/skills/umap-learn in your project. Codex loads it when a task matches its description.

Can I use Umap Learn in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/umap-learn, .gemini/skills/umap-learn, .github/skills/umap-learn and .opencode/skills/umap-learn in your project.

What does Umap Learn need to run?

Going by SKILL.md and its folder, Umap Learn needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Umap Learn access the network?

SKILL.md names 2 domains. As links in the text: umap-learn.readthedocs.io and github.com. This is read from the text; nothing was executed.

Is Umap Learn safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Umap Learn use?

Umap Learn is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Umap Learn use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Umap Learn?

Skills that share tags, products or a category with Umap Learn: ML Model Training (secondsky/claude-skills, 227 stars), Edit (omegaml/omegaml, 108 stars), ML Engineer (RightNow-AI/openfang, 18k stars) and Mlflow Onboarding (Kilo-Org/kilo-marketplace, 190 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Umap Learn?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.