ML Model Training
secondsky/claude-skills
Train ML models with scikit-learn, PyTorch, TensorFlow. An agent skill from secondsky/claude-skills.
UMAP dimensionality reduction for visualization, clustering prep, and feature engineering.
$ npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills umap-learn --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scientific-computing/umap-learn .claude/skills/umap-learn && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "umap-learn" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/umap-learn into .claude/skills/umap-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "umap-learn", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/umap-learnType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills umap-learn --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/scientific-computing/umap-learn .agents/skills/umap-learn && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "umap-learn" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/umap-learn into .agents/skills/umap-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "umap-learn", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills umap-learn --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/scientific-computing/umap-learn .cursor/skills/umap-learn && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "umap-learn" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/umap-learn into .cursor/skills/umap-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "umap-learn", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/scientific-computing/umap-learn--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills umap-learn --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/scientific-computing/umap-learn .gemini/skills/umap-learn && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "umap-learn" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/umap-learn into .gemini/skills/umap-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "umap-learn", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills umap-learnInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/scientific-computing/umap-learn .github/skills/umap-learn && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "umap-learn" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/umap-learn into .github/skills/umap-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "umap-learn", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills umap-learn --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/scientific-computing/umap-learn .opencode/skills/umap-learn && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "umap-learn" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/umap-learn into .opencode/skills/umap-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "umap-learn", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
umap-learnUMAP dimensionality reduction for visualization, clustering prep, and feature engineering.
Umap Learn is an agent skill from jaechang-hits/SciAgent-Skills. UMAP dimensionality reduction for visualization, clustering prep, and feature engineering. Fast nonlinear manifold learning preserving local and global structure. Standard UMAP (fit/transform, sklearn-compatible), supervised/semi-supervised, Parametric UMAP (NN encoder/decoder, TensorFlow), DensMAP (density), AlignedUMAP (temporal/batch). 15+ distance metrics, custom Numba metrics, precomputed distances. For linear reduction use PCA; for neighborhood graphs use sklearn NearestNeighbors.
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/api_reference.md`).
It sits in Data & Analytics, covering Machine learning and Deep learning. It works with UMAP, scikit-learn and TensorFlow. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
umap-learn.readthedocs.iogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Umap Learn loads about 4.7k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 1,014 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,014 words, ~4,745 tokens.
.claude/skills/umap-learn/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.UMAP (Uniform Manifold Approximation and Projection) is a dimensionality reduction algorithm for visualization and general non-linear dimensionality reduction. It is faster than t-SNE, scales to larger datasets, preserves both local and global structure, and supports supervised learning and embedding of new data points.
pip install umap-learn
# For Parametric UMAP (neural network variant)
pip install umap-learn[parametric_umap] # requires TensorFlow 2.xCritical: Always standardize features before applying UMAP to ensure equal weighting across dimensions.
import umap
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import load_digits
# Load and scale data
X, y = load_digits(return_X_y=True)
X_scaled = StandardScaler().fit_transform(X)
# Fit and transform
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled)
print(f"Input: {X_scaled.shape}, Output: {embedding.shape}")
# Input: (1797, 64), Output: (1797, 2)Basic dimensionality reduction following scikit-learn conventions.
import umap
from sklearn.preprocessing import StandardScaler
X_scaled = StandardScaler().fit_transform(data)
# Method 1: fit_transform (single step)
embedding = umap.UMAP(
n_neighbors=15, # local neighborhood size (2-200)
min_dist=0.1, # min distance between embedded points (0.0-0.99)
n_components=2, # output dimensions
metric='euclidean', # distance metric
random_state=42, # reproducibility
).fit_transform(X_scaled)
print(f"Embedding shape: {embedding.shape}")
# Method 2: fit + access (for reuse)
reducer = umap.UMAP(random_state=42)
reducer.fit(X_scaled)
embedding = reducer.embedding_ # trained embedding
graph = reducer.graph_ # fuzzy simplicial set (sparse matrix)# Visualization
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 6))
plt.scatter(embedding[:, 0], embedding[:, 1], c=labels, cmap='Spectral', s=5)
plt.colorbar()
plt.title('UMAP Embedding')
plt.tight_layout()
plt.savefig('umap_embedding.png', dpi=150)Incorporate label information to guide embedding via the y parameter.
import umap
# Supervised — all labels known
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled, y=labels)
# Semi-supervised — partial labels (mark unlabeled as -1)
semi_labels = labels.copy()
semi_labels[unlabeled_indices] = -1
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled, y=semi_labels)
# Control label influence with target_weight (0.0=unsupervised, 1.0=fully supervised)
reducer = umap.UMAP(
target_weight=0.7, # emphasize labels
target_metric='categorical', # for classification; use distance metric for regression
random_state=42
)
embedding = reducer.fit_transform(X_scaled, y=labels)
print(f"Supervised embedding: {embedding.shape}")Project unseen data into the trained embedding space.
import umap
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
# Fit on training data
reducer = umap.UMAP(n_components=10, random_state=42)
X_train_emb = reducer.fit_transform(X_train_scaled)
# Transform test data
X_test_emb = reducer.transform(X_test_scaled)
print(f"Train: {X_train_emb.shape}, Test: {X_test_emb.shape}")
# Works in sklearn Pipelines
from sklearn.pipeline import Pipeline
from sklearn.svm import SVC
pipeline = Pipeline([
('scaler', StandardScaler()),
('umap', umap.UMAP(n_components=10, random_state=42)),
('classifier', SVC())
])
pipeline.fit(X_train, y_train)
accuracy = pipeline.score(X_test, y_test)
print(f"Pipeline accuracy: {accuracy:.3f}")Neural network-based embedding via TensorFlow/Keras. Enables efficient transform, reconstruction, and custom architectures.
from umap.parametric_umap import ParametricUMAP
# Default architecture (3-layer, 100-neuron FC network)
embedder = ParametricUMAP(n_components=2, random_state=42)
embedding = embedder.fit_transform(X_scaled)
new_emb = embedder.transform(new_data) # fast neural network inference
print(f"Parametric embedding: {embedding.shape}")import tensorflow as tf
from umap.parametric_umap import ParametricUMAP
# Custom encoder/decoder for autoencoder mode
input_dim = X_scaled.shape[1]
encoder = tf.keras.Sequential([
tf.keras.layers.InputLayer(input_shape=(input_dim,)),
tf.keras.layers.Dense(128, activation='relu'),
tf.keras.layers.Dense(64, activation='relu'),
tf.keras.layers.Dense(2),
])
decoder = tf.keras.Sequential([
tf.keras.layers.InputLayer(input_shape=(2,)),
tf.keras.layers.Dense(64, activation='relu'),
tf.keras.layers.Dense(128, activation='relu'),
tf.keras.layers.Dense(input_dim),
])
embedder = ParametricUMAP(
encoder=encoder, decoder=decoder, dims=(input_dim,),
parametric_reconstruction=True, autoencoder_loss=True,
n_training_epochs=10, batch_size=128,
n_neighbors=15, min_dist=0.1, random_state=42
)
embedding = embedder.fit_transform(X_scaled)
reconstructed = embedder.inverse_transform(embedding)
print(f"Reconstruction error: {np.mean((X_scaled - reconstructed)**2):.4f}")Variant preserving local density information in the embedding.
import umap
reducer = umap.UMAP(
densmap=True, # enable DensMAP
dens_lambda=2.0, # density preservation weight
dens_frac=0.3, # fraction for density estimation
output_dens=True, # output density estimates
n_neighbors=15,
min_dist=0.1,
random_state=42
)
embedding = reducer.fit_transform(X_scaled)
# Access density estimates
original_density = reducer.rad_orig_ # density in original space
embedded_density = reducer.rad_emb_ # density in embedded space
print(f"DensMAP embedding: {embedding.shape}")
print(f"Density correlation: {np.corrcoef(original_density, embedded_density)[0,1]:.3f}")Align embeddings across multiple related datasets (time points, batches).
from umap import AlignedUMAP
# Multiple related datasets
datasets = [day1_data, day2_data, day3_data]
mapper = AlignedUMAP(
n_neighbors=15,
alignment_regularisation=1e-2, # alignment strength
alignment_window_size=2, # align with N adjacent datasets
n_components=2,
random_state=42
)
mapper.fit(datasets)
aligned_embeddings = mapper.embeddings_ # list of aligned embedding arrays
print(f"Aligned {len(aligned_embeddings)} datasets")
for i, emb in enumerate(aligned_embeddings):
print(f" Dataset {i}: {emb.shape}")| Parameter | Low | Medium (default) | High | Effect |
|---|---|---|---|---|
n_neighbors | 2-5 | 15 | 50-200 | Local detail vs global structure |
min_dist | 0.0 | 0.1 | 0.5-0.99 | Tight clusters vs spread out |
n_components | 2 | 2 | 5-50 | Visualization vs ML/clustering |
spread | 0.5 | 1.0 | 2.0 | Embedding scale (with min_dist) |
| Use-Case | n_neighbors | min_dist | n_components | metric |
|---|---|---|---|---|
| Visualization | 15 | 0.1 | 2 | euclidean |
| Clustering (HDBSCAN) | 30 | 0.0 | 5-10 | euclidean |
| Text/document embedding | 15 | 0.1 | 2 | cosine |
| Global structure | 100 | 0.5 | 2 | euclidean |
| ML feature engineering | 15-30 | 0.1 | 10-50 | euclidean |
| Binary/set data | 15 | 0.1 | 2 | hamming/jaccard |
Minkowski family: euclidean, manhattan, chebyshev, minkowski. Spatial: canberra, braycurtis, haversine. Correlation: cosine, correlation. Binary: hamming, jaccard, dice, russellrao, rogerstanimoto, sokalmichener, sokalsneath, yule. Special: precomputed (distance matrix), custom Numba-compiled callables.
| Feature | Standard | Parametric |
|---|---|---|
| Backend | Direct optimization | TensorFlow neural network |
| Transform speed | Moderate | Fast (neural net inference) |
| Inverse transform | Approximate, expensive | Decoder network, fast |
| Custom architecture | No | Yes (CNNs, RNNs, etc.) |
| Requirements | umap-learn | umap-learn + TensorFlow 2.x |
| Best for | Quick exploration | Production pipelines, reconstruction |
import umap
import hdbscan
import numpy as np
import matplotlib.pyplot as plt
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import adjusted_rand_score
# Step 1: Preprocess
X_scaled = StandardScaler().fit_transform(data)
print(f"Input shape: {X_scaled.shape}")
# Step 2: UMAP for clustering (NOT visualization parameters)
reducer = umap.UMAP(
n_neighbors=30, # more global structure for clustering
min_dist=0.0, # allow tight packing
n_components=10, # higher dims preserve density better than 2D
metric='euclidean',
random_state=42
)
embedding = reducer.fit_transform(X_scaled)
# Step 3: HDBSCAN clustering
clusterer = hdbscan.HDBSCAN(min_cluster_size=15, min_samples=5)
cluster_labels = clusterer.fit_predict(embedding)
n_clusters = len(set(cluster_labels)) - (1 if -1 in cluster_labels else 0)
noise = sum(cluster_labels == -1)
print(f"Clusters: {n_clusters}, Noise: {noise}")
# Step 4: Separate 2D embedding for visualization
vis_emb = umap.UMAP(n_neighbors=15, min_dist=0.1, random_state=42).fit_transform(X_scaled)
plt.scatter(vis_emb[:, 0], vis_emb[:, 1], c=cluster_labels, cmap='Spectral', s=5)
plt.colorbar()
plt.title(f'HDBSCAN Clusters (n={n_clusters})')
plt.tight_layout()
plt.savefig('umap_clusters.png', dpi=150)import umap
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.metrics import classification_report
# Split and scale
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
scaler = StandardScaler()
X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)
# Supervised UMAP for feature engineering
reducer = umap.UMAP(n_components=10, random_state=42)
X_train_emb = reducer.fit_transform(X_train_s, y=y_train)
X_test_emb = reducer.transform(X_test_s)
# Downstream classifier
clf = SVC(kernel='rbf')
clf.fit(X_train_emb, y_train)
y_pred = clf.predict(X_test_emb)
print(classification_report(y_test, y_pred))Text-only — combines Core API modules 1 and 3 (inverse_transform on standard UMAP):
reducer.inverse_transform(grid_points) to reconstruct high-dimensional dataNote: inverse transform is approximate; works poorly outside the convex hull of the training embedding.
| Parameter | Module | Default | Range | Effect |
|---|---|---|---|---|
n_neighbors | UMAP | 15 | 2-200 | Local vs global structure balance |
min_dist | UMAP | 0.1 | 0.0-0.99 | Cluster tightness |
n_components | UMAP | 2 | 2-100 | Output dimensionality |
metric | UMAP | 'euclidean' | See metrics list | Distance calculation method |
spread | UMAP | 1.0 | >0 | Embedding scale (with min_dist) |
n_epochs | UMAP | None (auto) | 50-500+ | Training iterations |
learning_rate | UMAP | 1.0 | >0 | SGD step size |
init | UMAP | 'spectral' | spectral/random/pca | Embedding initialization |
random_state | UMAP | None | int | Reproducibility seed |
target_weight | UMAP | 0.5 | 0.0-1.0 | Label influence (supervised) |
densmap | UMAP | False | bool | Enable DensMAP |
dens_lambda | UMAP | 2.0 | >0 | DensMAP density weight |
low_memory | UMAP | True | bool | Memory-efficient mode |
encoder | ParametricUMAP | None | Keras model | Custom encoder network |
decoder | ParametricUMAP | None | Keras model | Custom decoder network |
n_training_epochs | ParametricUMAP | 1 | 1-100 | Neural network training epochs |
alignment_regularisation | AlignedUMAP | 0.01 | >0 | Alignment strength |
alignment_window_size | AlignedUMAP | 3 | 1-N | Adjacent datasets to align |
Always standardize features: Use StandardScaler before UMAP — unscaled features with different ranges will dominate the embedding.
Set random_state for reproducibility: UMAP uses stochastic optimization; results vary between runs without a fixed seed.
Use different parameters for clustering vs visualization: Clustering needs n_neighbors=30, min_dist=0.0, n_components=5-10. Visualization needs n_neighbors=15, min_dist=0.1, n_components=2.
Anti-pattern — interpreting distances literally: UMAP preserves topology, not precise distances. Cluster separations and point distances in the embedding are not proportional to original distances.
Anti-pattern — using 2D embeddings for clustering: 2D projections lose density information. Use 5-10 components for HDBSCAN input.
Consider PCA preprocessing for very high dimensions: For data with >1000 features, reducing to 50-100 PCA components first can speed up UMAP without losing quality.
Use Parametric UMAP for production: When you need fast transform on new data or reconstruction capabilities, Parametric UMAP's neural network provides consistent, fast inference.
from numba import njit
import umap
@njit()
def weighted_euclidean(x, y):
"""Custom distance with feature weights."""
result = 0.0
for i in range(x.shape[0]):
result += (x[i] - y[i]) ** 2 * (1.0 + i * 0.01) # increasing weight
return np.sqrt(result)
embedding = umap.UMAP(metric=weighted_euclidean, random_state=42).fit_transform(data)import umap
from scipy.spatial.distance import pdist, squareform
# Compute custom distance matrix
dist_matrix = squareform(pdist(data, metric='correlation'))
# Use precomputed distances
embedding = umap.UMAP(
metric='precomputed', random_state=42
).fit_transform(dist_matrix)
print(f"Embedding from precomputed: {embedding.shape}")import umap
from sklearn.svm import SVC
# Train supervised embedding on labeled data
mapper = umap.UMAP(n_components=10, random_state=42)
train_emb = mapper.fit_transform(X_train, y=y_train)
# Transform unlabeled test data using learned metric
test_emb = mapper.transform(X_test)
# Downstream classifier
clf = SVC().fit(train_emb, y_train)
predictions = clf.predict(test_emb)
print(f"Accuracy: {(predictions == y_test).mean():.3f}")| Problem | Cause | Solution |
|---|---|---|
| Disconnected/fragmented clusters | n_neighbors too low | Increase n_neighbors (try 30-50) |
| Clusters too spread out | min_dist too high | Decrease min_dist (try 0.0-0.05) |
| All points collapsed | Bad preprocessing or min_dist too low | Check StandardScaler; increase min_dist |
| Poor clustering results | Using visualization parameters for clustering | Set n_neighbors=30, min_dist=0.0, n_components=5-10 |
| Transform results differ from training | Distribution shift | Ensure test data matches training distribution; use Parametric UMAP |
| Slow on large datasets (>100k) | Default settings | Set low_memory=True; preprocess with PCA to 50-100 dims |
| First run very slow | Numba JIT compilation | Expected — subsequent runs are fast (compiled cache) |
ImportError: umap | Name conflict with umap package | pip install umap-learn (not pip install umap) |
| Parametric UMAP import error | Missing TensorFlow | pip install umap-learn[parametric_umap] |
| Non-reproducible results | Missing random_state | Always set random_state=42 (or any int) |
Complete UMAP constructor parameter reference (60+ parameters organized by category: core, training, advanced structural, supervised, transform, performance, DensMAP), all methods and attributes, ParametricUMAP class with autoencoder parameters, AlignedUMAP class, utility functions (nearest_neighbors, fuzzy_simplicial_set). Core parameter tuning guidance was relocated to SKILL.md Key Concepts and Core API modules. Usage examples duplicating SKILL.md workflows omitted.
metric='precomputed'© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/scientific-computing/umap-learn of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
Umap Learn next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Umap Learn this skilljaechang-hits/SciAgent-Skills | 374 | — | ~4.7k | Automated safety check: Pass | BSD-3-Clause | |
| ML Model Trainingsecondsky/claude-skills | 227 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Editomegaml/omegaml | 108 | — | ~206 | Automated safety check: Pass | Apache-2.0 | |
| ML EngineerRightNow-AI/openfang | 18k | — | ~987 | Automated safety check: Pass | Apache-2.0 | |
| Mlflow OnboardingKilo-Org/kilo-marketplace | 190 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Molfeatdavila7/claude-code-templates | 33k | 9 repos | ~3.7k | Automated safety check: Pass | MIT |
secondsky/claude-skills
Train ML models with scikit-learn, PyTorch, TensorFlow. An agent skill from secondsky/claude-skills.
omegaml/omegaml
how to use the edit command properly
RightNow-AI/openfang
Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps
Kilo-Org/kilo-marketplace
Onboards users to MLflow by determining their use case (GenAI agents/apps or traditional ML/deep learning) and guiding them through relevant quickstart tutorials and initial integration.
davila7/claude-code-templates
Molecular featurization for ML (100+ featurizers). An agent skill from davila7/claude-code-templates.
databricks/databricks-agent-skills
Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
UMAP dimensionality reduction for visualization, clustering prep, and feature engineering. Umap Learn is an agent skill from jaechang-hits/SciAgent-Skills. UMAP dimensionality reduction for visualization, clustering prep, and feature engineering.
Umap Learn fits situations like: tasks that involve Machine learning; tasks that involve Deep learning.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a claude-code`. Or copy the skill folder (skills/scientific-computing/umap-learn in jaechang-hits/SciAgent-Skills) into .claude/skills/umap-learn in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a codex`. Or copy the skill folder (skills/scientific-computing/umap-learn in jaechang-hits/SciAgent-Skills) into .agents/skills/umap-learn in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/umap-learn, .gemini/skills/umap-learn, .github/skills/umap-learn and .opencode/skills/umap-learn in your project.
Going by SKILL.md and its folder, Umap Learn needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: umap-learn.readthedocs.io and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Umap Learn is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Umap Learn: ML Model Training (secondsky/claude-skills, 227 stars), Edit (omegaml/omegaml, 108 stars), ML Engineer (RightNow-AI/openfang, 18k stars) and Mlflow Onboarding (Kilo-Org/kilo-marketplace, 190 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.