Senior Data Scientist
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
$ npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install zLanqing/codex-claude-academic-skills scikit-learn --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/zLanqing/codex-claude-academic-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/scientific-toolkit-skill/references/scientific-skills/scikit-learn .claude/skills/scikit-learn && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scikit-learn" agent skill from https://github.com/zLanqing/codex-claude-academic-skills/tree/main/scientific-toolkit-skill/references/scientific-skills/scikit-learn into .claude/skills/scikit-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/zLanqing/codex-claude-academic-skills/tree/main/scientific-toolkit-skill/references/scientific-skills/scikit-learnType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install zLanqing/codex-claude-academic-skills scikit-learn --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zLanqing/codex-claude-academic-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/scientific-toolkit-skill/references/scientific-skills/scikit-learn .agents/skills/scikit-learn && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scikit-learn" agent skill from https://github.com/zLanqing/codex-claude-academic-skills/tree/main/scientific-toolkit-skill/references/scientific-skills/scikit-learn into .agents/skills/scikit-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install zLanqing/codex-claude-academic-skills scikit-learn --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zLanqing/codex-claude-academic-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/scientific-toolkit-skill/references/scientific-skills/scikit-learn .cursor/skills/scikit-learn && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scikit-learn" agent skill from https://github.com/zLanqing/codex-claude-academic-skills/tree/main/scientific-toolkit-skill/references/scientific-skills/scikit-learn into .cursor/skills/scikit-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/zLanqing/codex-claude-academic-skills.git --path scientific-toolkit-skill/references/scientific-skills/scikit-learn--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install zLanqing/codex-claude-academic-skills scikit-learn --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zLanqing/codex-claude-academic-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/scientific-toolkit-skill/references/scientific-skills/scikit-learn .gemini/skills/scikit-learn && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scikit-learn" agent skill from https://github.com/zLanqing/codex-claude-academic-skills/tree/main/scientific-toolkit-skill/references/scientific-skills/scikit-learn into .gemini/skills/scikit-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install zLanqing/codex-claude-academic-skills scikit-learnInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/zLanqing/codex-claude-academic-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/scientific-toolkit-skill/references/scientific-skills/scikit-learn .github/skills/scikit-learn && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scikit-learn" agent skill from https://github.com/zLanqing/codex-claude-academic-skills/tree/main/scientific-toolkit-skill/references/scientific-skills/scikit-learn into .github/skills/scikit-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install zLanqing/codex-claude-academic-skills scikit-learn --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zLanqing/codex-claude-academic-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/scientific-toolkit-skill/references/scientific-skills/scikit-learn .opencode/skills/scikit-learn && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scikit-learn" agent skill from https://github.com/zLanqing/codex-claude-academic-skills/tree/main/scientific-toolkit-skill/references/scientific-skills/scikit-learn into .opencode/skills/scikit-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scikit-learn", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scikit-learnMachine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Scikit Learn is an agent skill from zLanqing/codex-claude-academic-skills. Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.
Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `references/model_evaluation.md`, `references/pipelines_and_composition.md` and `references/preprocessing.md`).
It sits in Data & Analytics, covering Machine learning. It works with scikit-learn and Python. The repository describes itself as: 本仓库包含三个面向学术科研人员的Skills,覆盖从文献阅读、论文写作到科学计算的完整研究工作流。office-academic-skill 负责论文阅读报告与学术 PPT/Word 文档生成;research-writing-skill 提供论文写作、润色与审稿回复辅助;scientific-toolkit-skill 整合 MATLAB/Python… The licence is BSD-3-Clause.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7ed6377. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvpythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
scikit-learn.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scikit Learn loads about 3.9k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 982 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from zLanqing/codex-claude-academic-skills at commit 7ed6377, republished under its BSD-3-Clause licence (© zLanqing). 982 words, ~3,881 tokens.
.claude/skills/scikit-learn/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.This skill provides comprehensive guidance for machine learning tasks using scikit-learn, the industry-standard Python library for classical machine learning. Use this skill for classification, regression, clustering, dimensionality reduction, preprocessing, model evaluation, and building production-ready ML pipelines.
# Install scikit-learn using uv
uv pip install scikit-learn
# Optional: Install visualization dependencies
uv pip install matplotlib seaborn
# Commonly used with
uv pip install pandas numpyUse the scikit-learn skill when:
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report
# Split data
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42
)
# Preprocess
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
# Train model
model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train_scaled, y_train)
# Evaluate
y_pred = model.predict(X_test_scaled)
print(classification_report(y_test, y_pred))from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.ensemble import GradientBoostingClassifier
# Define feature types
numeric_features = ['age', 'income']
categorical_features = ['gender', 'occupation']
# Create preprocessing pipelines
numeric_transformer = Pipeline([
('imputer', SimpleImputer(strategy='median')),
('scaler', StandardScaler())
])
categorical_transformer = Pipeline([
('imputer', SimpleImputer(strategy='most_frequent')),
('onehot', OneHotEncoder(handle_unknown='ignore'))
])
# Combine transformers
preprocessor = ColumnTransformer([
('num', numeric_transformer, numeric_features),
('cat', categorical_transformer, categorical_features)
])
# Full pipeline
model = Pipeline([
('preprocessor', preprocessor),
('classifier', GradientBoostingClassifier(random_state=42))
])
# Fit and predict
model.fit(X_train, y_train)
y_pred = model.predict(X_test)Comprehensive algorithms for classification and regression tasks.
Key algorithms:
When to use:
See: references/supervised_learning.md for detailed algorithm documentation, parameters, and usage examples.
Discover patterns in unlabeled data through clustering and dimensionality reduction.
Clustering algorithms:
Dimensionality reduction:
When to use:
See: references/unsupervised_learning.md for detailed documentation.
Tools for robust model evaluation, cross-validation, and hyperparameter tuning.
Cross-validation strategies:
Hyperparameter tuning:
Metrics:
When to use:
See: references/model_evaluation.md for comprehensive metrics and tuning strategies.
Transform raw data into formats suitable for machine learning.
Scaling and normalization:
Encoding categorical variables:
Handling missing values:
Feature engineering:
When to use:
See: references/preprocessing.md for detailed preprocessing techniques.
Build reproducible, production-ready ML workflows.
Key components:
Benefits:
When to use:
See: references/pipelines_and_composition.md for comprehensive pipeline patterns.
Run a complete classification workflow with preprocessing, model comparison, hyperparameter tuning, and evaluation:
python scripts/classification_pipeline.pyThis script demonstrates:
Perform clustering analysis with algorithm comparison and visualization:
python scripts/clustering_analysis.pyThis script demonstrates:
This skill includes comprehensive reference files for deep dives into specific topics:
File: references/quick_reference.md
File: references/supervised_learning.md
File: references/unsupervised_learning.md
File: references/model_evaluation.md
File: references/preprocessing.md
File: references/pipelines_and_composition.md
Load and explore data
import pandas as pd
df = pd.read_csv('data.csv')
X = df.drop('target', axis=1)
y = df['target']Split data with stratification
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42
)Create preprocessing pipeline
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.compose import ColumnTransformer
# Handle numeric and categorical features separately
preprocessor = ColumnTransformer([
('num', StandardScaler(), numeric_features),
('cat', OneHotEncoder(), categorical_features)
])Build complete pipeline
model = Pipeline([
('preprocessor', preprocessor),
('classifier', RandomForestClassifier(random_state=42))
])Tune hyperparameters
from sklearn.model_selection import GridSearchCV
param_grid = {
'classifier__n_estimators': [100, 200],
'classifier__max_depth': [10, 20, None]
}
grid_search = GridSearchCV(model, param_grid, cv=5)
grid_search.fit(X_train, y_train)Evaluate on test set
from sklearn.metrics import classification_report
best_model = grid_search.best_estimator_
y_pred = best_model.predict(X_test)
print(classification_report(y_test, y_pred))Preprocess data
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)Find optimal number of clusters
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
scores = []
for k in range(2, 11):
kmeans = KMeans(n_clusters=k, random_state=42)
labels = kmeans.fit_predict(X_scaled)
scores.append(silhouette_score(X_scaled, labels))
optimal_k = range(2, 11)[np.argmax(scores)]Apply clustering
model = KMeans(n_clusters=optimal_k, random_state=42)
labels = model.fit_predict(X_scaled)Visualize with dimensionality reduction
from sklearn.decomposition import PCA
pca = PCA(n_components=2)
X_2d = pca.fit_transform(X_scaled)
plt.scatter(X_2d[:, 0], X_2d[:, 1], c=labels, cmap='viridis')Pipelines prevent data leakage and ensure consistency:
# Good: Preprocessing in pipeline
pipeline = Pipeline([
('scaler', StandardScaler()),
('model', LogisticRegression())
])
# Bad: Preprocessing outside (can leak information)
X_scaled = StandardScaler().fit_transform(X)Never fit on test data:
# Good
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test) # Only transform
# Bad
scaler = StandardScaler()
X_all_scaled = scaler.fit_transform(np.vstack([X_train, X_test]))Preserve class distribution:
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42
)model = RandomForestClassifier(n_estimators=100, random_state=42)Algorithms requiring feature scaling:
Algorithms not requiring scaling:
Issue: Model didn't converge
Solution: Increase max_iter or scale features
model = LogisticRegression(max_iter=1000)Issue: Overfitting Solution: Use regularization, cross-validation, or simpler model
# Add regularization
model = Ridge(alpha=1.0)
# Use cross-validation
scores = cross_val_score(model, X, y, cv=5)Solution: Use algorithms designed for large data
# Use SGD for large datasets
from sklearn.linear_model import SGDClassifier
model = SGDClassifier()
# Or MiniBatchKMeans for clustering
from sklearn.cluster import MiniBatchKMeans
model = MiniBatchKMeans(n_clusters=8, batch_size=100)© zLanqing, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references) in scientific-toolkit-skill/references/scientific-skills/scikit-learn of zLanqing/codex-claude-academic-skills.
Open the folder on GitHubat commit 7ed6377
We found 43 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 17 other GitHub owners. This page covers the copy in zLanqing/codex-claude-academic-skills, which our catalogue first saw on October 7, 2026.
Scikit Learn next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scikit Learn this skillzLanqing/codex-claude-academic-skills | 4.6k | 17 repos | ~3.9k | Automated safety check: Pass | BSD-3-Clause | |
| Senior Data ScientistRaidriar7170/hermes-skilleval | 125 | 6 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Time Series Analytics Useropen-edge-platform/edge-ai-libraries | 168 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Aeon Time Series Machine Learningdavila7/claude-code-templates | 32k | 14 repos | ~2.6k | Automated safety check: Pass | MIT | |
| Precisemicroprediction/precise | 336 | — | ~782 | Automated safety check: Pass | MIT | |
| scikit-survival Time-to-Event Modelingdavila7/claude-code-templates | 32k | 12 repos | ~3.7k | Automated safety check: Pass | MIT |
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
open-edge-platform/edge-ai-libraries
Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…
davila7/claude-code-templates
Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.
microprediction/precise
Online (incremental) covariance, correlation, and precision estimation in Python — the streaming complement to sklearn.covariance.
davila7/claude-code-templates
Fits and evaluates survival models with scikit-survival: Cox models, Random Survival Forests, boosting, survival SVMs, concordance index, Brier score and competing risks.
HKUDS/Vibe-Trading
Trains scikit-learn models with walk-forward validation on features from OHLCV data to predict return direction and turn the predictions into trading signals.
zLanqing/codex-claude-academic-skills
Low-level plotting library for full customization. An agent skill from zLanqing/codex-claude-academic-skills.
zLanqing/codex-claude-academic-skills
Process-based discrete-event simulation framework in Python.
zLanqing/codex-claude-academic-skills
Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs in Python.
zLanqing/codex-claude-academic-skills
Statistical visualization with pandas integration. An agent skill from zLanqing/codex-claude-academic-skills.
zLanqing/codex-claude-academic-skills
Zero-shot time series forecasting with Google's TimesFM foundation model.
zLanqing/codex-claude-academic-skills
Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.
Works with
Categories
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills. Scikit Learn is an agent skill from zLanqing/codex-claude-academic-skills. Machine learning in Python with scikit-learn.
Scikit Learn fits situations like: working with supervised learning (classification; unsupervised learning (clustering; dimensionality reduction); model evaluation.
Run `npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a claude-code`. Or copy the skill folder (scientific-toolkit-skill/references/scientific-skills/scikit-learn in zLanqing/codex-claude-academic-skills) into .claude/skills/scikit-learn in your project. Claude Code loads it when a task matches its description.
Run `npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a codex`. Or copy the skill folder (scientific-toolkit-skill/references/scientific-skills/scikit-learn in zLanqing/codex-claude-academic-skills) into .agents/skills/scikit-learn in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scikit-learn, .gemini/skills/scikit-learn, .github/skills/scikit-learn and .opencode/skills/scikit-learn in your project.
Going by SKILL.md and its folder, Scikit Learn needs Python for the scripts in its folder and the command-line tools its instructions call (uv and python). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: scikit-learn.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Scikit Learn is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Scikit Learn: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 168 stars), Aeon Time Series Machine Learning (davila7/claude-code-templates, 32k stars) and Precise (microprediction/precise, 336 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
zLanqing (a GitHub user) maintains it in zLanqing/codex-claude-academic-skills, which has 4,578 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on May 14, 2026.
Source: zLanqing/codex-claude-academic-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.