Senior Data Scientist
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
Classical ML in Python: classification, regression, clustering, dim reduction, evaluation, tuning, preprocessing pipelines.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scikit-learn-machine-learning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scientific-computing/scikit-learn-machine-learning .claude/skills/scikit-learn-machine-learning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scikit-learn-machine-learning" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/scikit-learn-machine-learning into .claude/skills/scikit-learn-machine-learning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn-machine-learning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/scikit-learn-machine-learningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scikit-learn-machine-learning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/scientific-computing/scikit-learn-machine-learning .agents/skills/scikit-learn-machine-learning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scikit-learn-machine-learning" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/scikit-learn-machine-learning into .agents/skills/scikit-learn-machine-learning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn-machine-learning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scikit-learn-machine-learning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/scientific-computing/scikit-learn-machine-learning .cursor/skills/scikit-learn-machine-learning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scikit-learn-machine-learning" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/scikit-learn-machine-learning into .cursor/skills/scikit-learn-machine-learning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn-machine-learning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/scientific-computing/scikit-learn-machine-learning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scikit-learn-machine-learning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/scientific-computing/scikit-learn-machine-learning .gemini/skills/scikit-learn-machine-learning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scikit-learn-machine-learning" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/scikit-learn-machine-learning into .gemini/skills/scikit-learn-machine-learning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn-machine-learning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills scikit-learn-machine-learningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/scientific-computing/scikit-learn-machine-learning .github/skills/scikit-learn-machine-learning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scikit-learn-machine-learning" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/scikit-learn-machine-learning into .github/skills/scikit-learn-machine-learning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn-machine-learning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scikit-learn-machine-learning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/scientific-computing/scikit-learn-machine-learning .opencode/skills/scikit-learn-machine-learning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scikit-learn-machine-learning" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/scikit-learn-machine-learning into .opencode/skills/scikit-learn-machine-learning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn-machine-learning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scikit-learn-machine-learningClassical ML in Python: classification, regression, clustering, dim reduction, evaluation, tuning, preprocessing pipelines.
Scikit Learn Machine Learning is an agent skill from jaechang-hits/SciAgent-Skills. Classical ML in Python: classification, regression, clustering, dim reduction, evaluation, tuning, preprocessing pipelines. Linear models, tree ensembles, SVMs, K-Means, PCA, t-SNE. Use PyTorch/TF for deep learning; XGBoost/LightGBM for scale.
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Machine learning and Deep learning. It works with scikit-learn, Python, PyTorch and NumPy. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
scikit-learn.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scikit Learn Machine Learning loads about 4k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 540 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 540 words, ~4,009 tokens.
.claude/skills/scikit-learn-machine-learning/SKILL.md (or your agent's skills folder).scikit-learn is the standard Python library for classical machine learning. It provides consistent APIs for supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, and preprocessing, with seamless integration into NumPy/pandas workflows.
pytorch or transformers insteadxgboost or lightgbm insteadscikit-learn, numpy, pandasmatplotlib, seaborn for visualizationpip install scikit-learn numpy pandas matplotlib seabornfrom sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, classification_report
from sklearn.datasets import load_breast_cancer
# Load dataset, split, train, evaluate in 10 lines
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
clf = RandomForestClassifier(n_estimators=100, random_state=42)
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
print(f"Accuracy: {accuracy_score(y_test, y_pred):.3f}")
print(classification_report(y_test, y_pred, target_names=["malignant", "benign"]))Scaling, encoding, imputation, and feature engineering.
from sklearn.preprocessing import StandardScaler, MinMaxScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
import numpy as np
# Scaling: zero mean, unit variance
X = np.array([[1, 2], [3, 4], [5, 6]])
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
print(f"Mean: {X_scaled.mean(axis=0)}, Std: {X_scaled.std(axis=0)}")
# Mean: [0. 0.], Std: [1. 1.]
# Imputation: fill missing values
X_missing = np.array([[1, np.nan], [3, 4], [np.nan, 6]])
imputer = SimpleImputer(strategy="median")
X_filled = imputer.fit_transform(X_missing)
print(f"Filled:\n{X_filled}")from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder, LabelEncoder
# One-hot encoding for nominal categories
enc = OneHotEncoder(sparse_output=False, handle_unknown="ignore")
X_cat = np.array([["red"], ["blue"], ["green"], ["red"]])
X_encoded = enc.fit_transform(X_cat)
print(f"Categories: {enc.categories_}")
print(f"Encoded shape: {X_encoded.shape}") # (4, 3)Classifiers for discrete target prediction.
from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
# Compare classifiers
classifiers = {
"LogisticRegression": LogisticRegression(max_iter=200),
"RandomForest": RandomForestClassifier(n_estimators=100, random_state=42),
"SVM": SVC(kernel="rbf", C=1.0),
"GradientBoosting": GradientBoostingClassifier(n_estimators=100, random_state=42),
}
for name, clf in classifiers.items():
clf.fit(X_train, y_train)
print(f"{name}: accuracy = {clf.score(X_test, y_test):.3f}")Regressors for continuous target prediction.
from sklearn.linear_model import LinearRegression, Ridge, Lasso, ElasticNet
from sklearn.ensemble import RandomForestRegressor
from sklearn.datasets import make_regression
from sklearn.metrics import mean_squared_error, r2_score
X, y = make_regression(n_samples=200, n_features=10, noise=10, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
models = {
"Linear": LinearRegression(),
"Ridge": Ridge(alpha=1.0),
"Lasso": Lasso(alpha=0.1),
"RandomForest": RandomForestRegressor(n_estimators=100, random_state=42),
}
for name, model in models.items():
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print(f"{name}: RMSE={mean_squared_error(y_test, y_pred, squared=False):.2f}, R²={r2_score(y_test, y_pred):.3f}")Clustering algorithms for unlabeled data.
from sklearn.cluster import KMeans, DBSCAN, AgglomerativeClustering
from sklearn.metrics import silhouette_score
from sklearn.datasets import make_blobs
X, y_true = make_blobs(n_samples=300, centers=4, random_state=42)
# K-Means with elbow method
for k in [2, 3, 4, 5, 6]:
km = KMeans(n_clusters=k, random_state=42, n_init=10)
labels = km.fit_predict(X)
sil = silhouette_score(X, labels)
print(f"k={k}: silhouette={sil:.3f}, inertia={km.inertia_:.1f}")# DBSCAN — no need to specify k
from sklearn.cluster import DBSCAN
db = DBSCAN(eps=0.5, min_samples=5)
labels = db.fit_predict(X)
n_clusters = len(set(labels)) - (1 if -1 in labels else 0)
n_noise = (labels == -1).sum()
print(f"DBSCAN: {n_clusters} clusters, {n_noise} noise points")PCA, t-SNE, and other methods for visualization and feature reduction.
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.datasets import load_digits
X, y = load_digits(return_X_y=True)
print(f"Original shape: {X.shape}") # (1797, 64)
# PCA — preserve 95% variance
pca = PCA(n_components=0.95)
X_pca = pca.fit_transform(X)
print(f"PCA: {X_pca.shape[1]} components, explained variance: {pca.explained_variance_ratio_.sum():.3f}")
# t-SNE — 2D visualization
tsne = TSNE(n_components=2, perplexity=30, random_state=42)
X_tsne = tsne.fit_transform(X)
print(f"t-SNE shape: {X_tsne.shape}") # (1797, 2)Cross-validation, metrics, hyperparameter tuning.
from sklearn.model_selection import cross_val_score, GridSearchCV, StratifiedKFold
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report, confusion_matrix
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
# Cross-validation
clf = RandomForestClassifier(n_estimators=100, random_state=42)
scores = cross_val_score(clf, X, y, cv=StratifiedKFold(5), scoring="accuracy")
print(f"CV accuracy: {scores.mean():.3f} ± {scores.std():.3f}")# Hyperparameter tuning with GridSearchCV
param_grid = {
"n_estimators": [50, 100, 200],
"max_depth": [5, 10, None],
"min_samples_split": [2, 5]
}
grid = GridSearchCV(
RandomForestClassifier(random_state=42),
param_grid, cv=5, scoring="accuracy", n_jobs=-1
)
grid.fit(X, y)
print(f"Best params: {grid.best_params_}")
print(f"Best score: {grid.best_score_:.3f}")Chain preprocessing and models; prevent data leakage.
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.ensemble import GradientBoostingClassifier
# Mixed-type preprocessing
numeric_features = ["age", "income"]
categorical_features = ["gender", "occupation"]
preprocessor = ColumnTransformer([
("num", Pipeline([
("imputer", SimpleImputer(strategy="median")),
("scaler", StandardScaler())
]), numeric_features),
("cat", Pipeline([
("imputer", SimpleImputer(strategy="most_frequent")),
("onehot", OneHotEncoder(handle_unknown="ignore"))
]), categorical_features),
])
pipe = Pipeline([
("preprocessor", preprocessor),
("classifier", GradientBoostingClassifier(random_state=42))
])
# pipe.fit(X_train, y_train); pipe.predict(X_test)
print("Pipeline steps:", [name for name, _ in pipe.steps])Goal: Complete classification workflow from data loading to evaluation.
import pandas as pd
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report
from sklearn.datasets import load_breast_cancer
# Load data
X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
# Build pipeline
pipe = Pipeline([
("scaler", StandardScaler()),
("clf", RandomForestClassifier(n_estimators=200, random_state=42))
])
# Cross-validate
cv_scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring="f1")
print(f"CV F1: {cv_scores.mean():.3f} ± {cv_scores.std():.3f}")
# Final evaluation
pipe.fit(X_train, y_train)
y_pred = pipe.predict(X_test)
print(classification_report(y_test, y_pred))Goal: Cluster data and visualize with dimensionality reduction.
from sklearn.datasets import make_blobs
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
from sklearn.metrics import silhouette_score
import matplotlib.pyplot as plt
# Generate and scale data
X, _ = make_blobs(n_samples=500, centers=4, random_state=42)
X_scaled = StandardScaler().fit_transform(X)
# Cluster
km = KMeans(n_clusters=4, random_state=42, n_init=10)
labels = km.fit_predict(X_scaled)
print(f"Silhouette: {silhouette_score(X_scaled, labels):.3f}")
# Visualize
X_2d = PCA(n_components=2).fit_transform(X_scaled)
plt.scatter(X_2d[:, 0], X_2d[:, 1], c=labels, cmap="viridis", s=20, alpha=0.7)
plt.title("K-Means Clustering (PCA projection)")
plt.savefig("clustering_result.png", dpi=150, bbox_inches="tight")
print("Saved clustering_result.png")Goal: Select best features and build a tuned model.
from sklearn.datasets import make_classification
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.model_selection import GridSearchCV
X, y = make_classification(n_samples=500, n_features=50, n_informative=10, random_state=42)
pipe = Pipeline([
("scaler", StandardScaler()),
("selector", SelectKBest(f_classif)),
("svm", SVC(kernel="rbf"))
])
param_grid = {
"selector__k": [5, 10, 20],
"svm__C": [0.1, 1, 10],
"svm__gamma": ["scale", "auto"]
}
grid = GridSearchCV(pipe, param_grid, cv=5, scoring="accuracy", n_jobs=-1)
grid.fit(X, y)
print(f"Best params: {grid.best_params_}")
print(f"Best accuracy: {grid.best_score_:.3f}")| Parameter | Module | Default | Range / Options | Effect |
|---|---|---|---|---|
n_estimators | RandomForest, GradientBoosting | 100 | 50-1000 | Number of trees; higher = better but slower |
max_depth | Tree-based models | None | 1-50, None | Tree depth; None = no limit (can overfit) |
C | SVM, LogisticRegression | 1.0 | 0.001-1000 | Regularization strength (inverse); lower = more regularization |
alpha | Ridge, Lasso | 1.0 | 0.001-100 | Regularization strength; higher = more regularization |
n_clusters | KMeans | required | 2-N | Number of clusters to form |
eps | DBSCAN | 0.5 | 0.01-10 | Neighborhood radius; smaller = more clusters |
n_components | PCA | required | 1-N or 0.0-1.0 | Components to keep; float = variance ratio |
perplexity | t-SNE | 30 | 5-50 | Balance local/global structure |
cv | GridSearchCV | 5 | 2-10 | Cross-validation folds |
scoring | GridSearchCV, cross_val_score | varies | accuracy, f1, roc_auc, etc. | Evaluation metric |
When to use: Understanding which features drive model predictions.
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
clf = RandomForestClassifier(n_estimators=200, random_state=42).fit(X, y)
importances = clf.feature_importances_
indices = np.argsort(importances)[::-1]
feature_names = load_iris().feature_names
for i in range(X.shape[1]):
print(f"{feature_names[indices[i]]}: {importances[indices[i]]:.4f}")When to use: Diagnosing overfitting vs underfitting.
from sklearn.model_selection import learning_curve
import matplotlib.pyplot as plt
import numpy as np
train_sizes, train_scores, val_scores = learning_curve(
clf, X, y, cv=5, train_sizes=np.linspace(0.1, 1.0, 10), scoring="accuracy"
)
plt.plot(train_sizes, train_scores.mean(axis=1), label="Train")
plt.plot(train_sizes, val_scores.mean(axis=1), label="Validation")
plt.xlabel("Training size"); plt.ylabel("Accuracy"); plt.legend()
plt.savefig("learning_curve.png", dpi=150, bbox_inches="tight")
print("Saved learning_curve.png")When to use: Persisting trained models for later use.
import joblib
# Save
joblib.dump(pipe, "model_pipeline.joblib")
print("Model saved to model_pipeline.joblib")
# Load
loaded_pipe = joblib.load("model_pipeline.joblib")
y_pred = loaded_pipe.predict(X_test)
print(f"Loaded model predictions: {y_pred[:5]}")| Problem | Cause | Solution |
|---|---|---|
ConvergenceWarning | Model didn't converge | Increase max_iter (e.g., 1000) or scale features with StandardScaler |
| High train accuracy, low test accuracy | Overfitting | Add regularization, reduce max_depth, use cross-validation |
ValueError: unknown categories | New categories in test data | Use OneHotEncoder(handle_unknown='ignore') |
MemoryError with large data | Full dataset in memory | Use SGDClassifier/MiniBatchKMeans for incremental learning |
| Poor clustering results | Unscaled features or wrong k | Scale features first; use silhouette score to find optimal k |
NotFittedError | Predict before fit | Call model.fit(X_train, y_train) first |
| Different results each run | Missing random_state | Set random_state=42 in model and train_test_split |
| Slow GridSearchCV | Large parameter grid | Use RandomizedSearchCV or HalvingGridSearchCV; add n_jobs=-1 |
© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/scientific-computing/scikit-learn-machine-learning of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Scikit Learn Machine Learning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scikit Learn Machine Learning this skilljaechang-hits/SciAgent-Skills | 370 | 1 repos | ~4k | Automated safety check: Pass | BSD-3-Clause | |
| Senior Data ScientistRaidriar7170/hermes-skilleval | 125 | 6 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Machine Learning Trading StrategyHKUDS/Vibe-Trading | 35k | — | ~3.2k | Automated safety check: Pass | MIT | |
| Senior Data Scientistalirezarezvani/claude-skills | 28k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Statistical Data Analysislingzhi227/agent-research-skills | 384 | — | ~886 | Automated safety check: Pass | None | |
| Optimize For GPUK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.4k | Automated safety check: Pass | MIT |
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
HKUDS/Vibe-Trading
Trains scikit-learn models with walk-forward validation on features from OHLCV data to predict return direction and turn the predictions into trading signals.
alirezarezvani/claude-skills
World-class senior data scientist skill specialising in statistical modeling, experiment design, causal inference, and predictive analytics.
lingzhi227/agent-research-skills
Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.
K-Dense-AI/scientific-agent-skills
GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.
RightNow-AI/openfang
Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Classical ML in Python: classification, regression, clustering, dim reduction, evaluation, tuning, preprocessing pipelines. Scikit Learn Machine Learning is an agent skill from jaechang-hits/SciAgent-Skills. Classical ML in Python: classification, regression, clustering, dim reduction, evaluation, tuning, preprocessing pipelines.
Scikit Learn Machine Learning fits situations like: tasks that involve Machine learning; tasks that involve Deep learning.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a claude-code`. Or copy the skill folder (skills/scientific-computing/scikit-learn-machine-learning in jaechang-hits/SciAgent-Skills) into .claude/skills/scikit-learn-machine-learning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a codex`. Or copy the skill folder (skills/scientific-computing/scikit-learn-machine-learning in jaechang-hits/SciAgent-Skills) into .agents/skills/scikit-learn-machine-learning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill scikit-learn-machine-learning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scikit-learn-machine-learning, .gemini/skills/scikit-learn-machine-learning, .github/skills/scikit-learn-machine-learning and .opencode/skills/scikit-learn-machine-learning in your project.
Going by SKILL.md and its folder, Scikit Learn Machine Learning needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: scikit-learn.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Scikit Learn Machine Learning is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Scikit Learn Machine Learning: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Machine Learning Trading Strategy (HKUDS/Vibe-Trading, 35k stars), Senior Data Scientist (alirezarezvani/claude-skills, 28k stars) and Statistical Data Analysis (lingzhi227/agent-research-skills, 384 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 163 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.