Agent skill

Scikit Learn

by zLanqing in zLanqing/codex-claude-academic-skills

Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

BSD-3-ClauseAuto-check passedData & Analytics

Install Scikit Learn

skills CLI
$ npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zLanqing/codex-claude-academic-skills scikit-learn --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zLanqing/codex-claude-academic-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/scientific-toolkit-skill/references/scientific-skills/scikit-learn .claude/skills/scikit-learn && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scikit-learn
GitHub stars
4.6k
Used in
17 other repos
Token cost
~3.9k tokens
SKILL.md length
982 words
Files
9 (incl. scripts, references)
Skills in repo
17
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

  • Works in 5 steps: Supervised Learning → Unsupervised Learning → Model Evaluation and Selection → …
  • Working with supervised learning (classification
  • SKILL.md covers Overview, Installation, When to Use This Skill and Quick Start, plus 5 more sections
  • Runs Python scripts from its folder; calls uv and python

What it does

Scikit Learn is an agent skill from zLanqing/codex-claude-academic-skills. Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `references/model_evaluation.md`, `references/pipelines_and_composition.md` and `references/preprocessing.md`).

It sits in Data & Analytics, covering Machine learning. It works with scikit-learn and Python. The repository describes itself as: 本仓库包含三个面向学术科研人员的Skills,覆盖从文献阅读、论文写作到科学计算的完整研究工作流。office-academic-skill 负责论文阅读报告与学术 PPT/Word 文档生成;research-writing-skill 提供论文写作、润色与审稿回复辅助;scientific-toolkit-skill 整合 MATLAB/Python… The licence is BSD-3-Clause.

When your agent uses it

  • Working with supervised learning (classification
  • Unsupervised learning (clustering
  • Dimensionality reduction)
  • Model evaluation

Example prompts

  • “/scikit-learn”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Supervised Learning
  2. Unsupervised Learning
  3. Model Evaluation and Selection
  4. Data Preprocessing
  5. Pipelines and Composition

What it can do on your machine

Read from SKILL.md and the folder at commit 7ed6377. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • scikit-learn.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scikit Learn loads about 3.9k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 982 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zLanqing/codex-claude-academic-skills at commit 7ed6377, republished under its BSD-3-Clause licence (© zLanqing). 982 words, ~3,881 tokens.

Download SKILL.mdSave it as .claude/skills/scikit-learn/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
scikit-learn
description
Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.
license
BSD-3-Clause license
metadata.skill-author
K-Dense Inc.

Scikit-learn

Overview

This skill provides comprehensive guidance for machine learning tasks using scikit-learn, the industry-standard Python library for classical machine learning. Use this skill for classification, regression, clustering, dimensionality reduction, preprocessing, model evaluation, and building production-ready ML pipelines.

Installation

bash
# Install scikit-learn using uv
uv pip install scikit-learn

# Optional: Install visualization dependencies
uv pip install matplotlib seaborn

# Commonly used with
uv pip install pandas numpy

When to Use This Skill

Use the scikit-learn skill when:

  • Building classification or regression models
  • Performing clustering or dimensionality reduction
  • Preprocessing and transforming data for machine learning
  • Evaluating model performance with cross-validation
  • Tuning hyperparameters with grid or random search
  • Creating ML pipelines for production workflows
  • Comparing different algorithms for a task
  • Working with both structured (tabular) and text data
  • Need interpretable, classical machine learning approaches

Quick Start

Classification Example
python
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report

# Split data
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, stratify=y, random_state=42
)

# Preprocess
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

# Train model
model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train_scaled, y_train)

# Evaluate
y_pred = model.predict(X_test_scaled)
print(classification_report(y_test, y_pred))
Complete Pipeline with Mixed Data
python
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.ensemble import GradientBoostingClassifier

# Define feature types
numeric_features = ['age', 'income']
categorical_features = ['gender', 'occupation']

# Create preprocessing pipelines
numeric_transformer = Pipeline([
    ('imputer', SimpleImputer(strategy='median')),
    ('scaler', StandardScaler())
])

categorical_transformer = Pipeline([
    ('imputer', SimpleImputer(strategy='most_frequent')),
    ('onehot', OneHotEncoder(handle_unknown='ignore'))
])

# Combine transformers
preprocessor = ColumnTransformer([
    ('num', numeric_transformer, numeric_features),
    ('cat', categorical_transformer, categorical_features)
])

# Full pipeline
model = Pipeline([
    ('preprocessor', preprocessor),
    ('classifier', GradientBoostingClassifier(random_state=42))
])

# Fit and predict
model.fit(X_train, y_train)
y_pred = model.predict(X_test)

Core Capabilities

1. Supervised Learning

Comprehensive algorithms for classification and regression tasks.

Key algorithms:

  • Linear models: Logistic Regression, Linear Regression, Ridge, Lasso, ElasticNet
  • Tree-based: Decision Trees, Random Forest, Gradient Boosting
  • Support Vector Machines: SVC, SVR with various kernels
  • Ensemble methods: AdaBoost, Voting, Stacking
  • Neural Networks: MLPClassifier, MLPRegressor
  • Others: Naive Bayes, K-Nearest Neighbors

When to use:

  • Classification: Predicting discrete categories (spam detection, image classification, fraud detection)
  • Regression: Predicting continuous values (price prediction, demand forecasting)

See: references/supervised_learning.md for detailed algorithm documentation, parameters, and usage examples.

2. Unsupervised Learning

Discover patterns in unlabeled data through clustering and dimensionality reduction.

Clustering algorithms:

  • Partition-based: K-Means, MiniBatchKMeans
  • Density-based: DBSCAN, HDBSCAN, OPTICS
  • Hierarchical: AgglomerativeClustering
  • Probabilistic: Gaussian Mixture Models
  • Others: MeanShift, SpectralClustering, BIRCH

Dimensionality reduction:

  • Linear: PCA, TruncatedSVD, NMF
  • Manifold learning: t-SNE, UMAP, Isomap, LLE
  • Feature extraction: FastICA, LatentDirichletAllocation

When to use:

  • Customer segmentation, anomaly detection, data visualization
  • Reducing feature dimensions, exploratory data analysis
  • Topic modeling, image compression

See: references/unsupervised_learning.md for detailed documentation.

3. Model Evaluation and Selection

Tools for robust model evaluation, cross-validation, and hyperparameter tuning.

Cross-validation strategies:

  • KFold, StratifiedKFold (classification)
  • TimeSeriesSplit (temporal data)
  • GroupKFold (grouped samples)

Hyperparameter tuning:

  • GridSearchCV (exhaustive search)
  • RandomizedSearchCV (random sampling)
  • HalvingGridSearchCV (successive halving)

Metrics:

  • Classification: accuracy, precision, recall, F1-score, ROC AUC, confusion matrix
  • Regression: MSE, RMSE, MAE, R², MAPE
  • Clustering: silhouette score, Calinski-Harabasz, Davies-Bouldin

When to use:

  • Comparing model performance objectively
  • Finding optimal hyperparameters
  • Preventing overfitting through cross-validation
  • Understanding model behavior with learning curves

See: references/model_evaluation.md for comprehensive metrics and tuning strategies.

4. Data Preprocessing

Transform raw data into formats suitable for machine learning.

Scaling and normalization:

  • StandardScaler (zero mean, unit variance)
  • MinMaxScaler (bounded range)
  • RobustScaler (robust to outliers)
  • Normalizer (sample-wise normalization)

Encoding categorical variables:

  • OneHotEncoder (nominal categories)
  • OrdinalEncoder (ordered categories)
  • LabelEncoder (target encoding)

Handling missing values:

  • SimpleImputer (mean, median, most frequent)
  • KNNImputer (k-nearest neighbors)
  • IterativeImputer (multivariate imputation)

Feature engineering:

  • PolynomialFeatures (interaction terms)
  • KBinsDiscretizer (binning)
  • Feature selection (RFE, SelectKBest, SelectFromModel)

When to use:

  • Before training any algorithm that requires scaled features (SVM, KNN, Neural Networks)
  • Converting categorical variables to numeric format
  • Handling missing data systematically
  • Creating non-linear features for linear models

See: references/preprocessing.md for detailed preprocessing techniques.

5. Pipelines and Composition

Build reproducible, production-ready ML workflows.

Key components:

  • Pipeline: Chain transformers and estimators sequentially
  • ColumnTransformer: Apply different preprocessing to different columns
  • FeatureUnion: Combine multiple transformers in parallel
  • TransformedTargetRegressor: Transform target variable

Benefits:

  • Prevents data leakage in cross-validation
  • Simplifies code and improves maintainability
  • Enables joint hyperparameter tuning
  • Ensures consistency between training and prediction

When to use:

  • Always use Pipelines for production workflows
  • When mixing numerical and categorical features (use ColumnTransformer)
  • When performing cross-validation with preprocessing steps
  • When hyperparameter tuning includes preprocessing parameters

See: references/pipelines_and_composition.md for comprehensive pipeline patterns.

Example Scripts

Classification Pipeline

Run a complete classification workflow with preprocessing, model comparison, hyperparameter tuning, and evaluation:

bash
python scripts/classification_pipeline.py

This script demonstrates:

  • Handling mixed data types (numeric and categorical)
  • Model comparison using cross-validation
  • Hyperparameter tuning with GridSearchCV
  • Comprehensive evaluation with multiple metrics
  • Feature importance analysis
Clustering Analysis

Perform clustering analysis with algorithm comparison and visualization:

bash
python scripts/clustering_analysis.py

This script demonstrates:

  • Finding optimal number of clusters (elbow method, silhouette analysis)
  • Comparing multiple clustering algorithms (K-Means, DBSCAN, Agglomerative, Gaussian Mixture)
  • Evaluating clustering quality without ground truth
  • Visualizing results with PCA projection

Reference Documentation

This skill includes comprehensive reference files for deep dives into specific topics:

Quick Reference

File: references/quick_reference.md

  • Common import patterns and installation instructions
  • Quick workflow templates for common tasks
  • Algorithm selection cheat sheets
  • Common patterns and gotchas
  • Performance optimization tips
Show full SKILL.md (454 more words)Show less
Supervised Learning

File: references/supervised_learning.md

  • Linear models (regression and classification)
  • Support Vector Machines
  • Decision Trees and ensemble methods
  • K-Nearest Neighbors, Naive Bayes, Neural Networks
  • Algorithm selection guide
Unsupervised Learning

File: references/unsupervised_learning.md

  • All clustering algorithms with parameters and use cases
  • Dimensionality reduction techniques
  • Outlier and novelty detection
  • Gaussian Mixture Models
  • Method selection guide
Model Evaluation

File: references/model_evaluation.md

  • Cross-validation strategies
  • Hyperparameter tuning methods
  • Classification, regression, and clustering metrics
  • Learning and validation curves
  • Best practices for model selection
Preprocessing

File: references/preprocessing.md

  • Feature scaling and normalization
  • Encoding categorical variables
  • Missing value imputation
  • Feature engineering techniques
  • Custom transformers
Pipelines and Composition

File: references/pipelines_and_composition.md

  • Pipeline construction and usage
  • ColumnTransformer for mixed data types
  • FeatureUnion for parallel transformations
  • Complete end-to-end examples
  • Best practices

Common Workflows

Building a Classification Model
  1. Load and explore data

    python
    import pandas as pd
    df = pd.read_csv('data.csv')
    X = df.drop('target', axis=1)
    y = df['target']
  2. Split data with stratification

    python
    from sklearn.model_selection import train_test_split
    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.2, stratify=y, random_state=42
    )
  3. Create preprocessing pipeline

    python
    from sklearn.pipeline import Pipeline
    from sklearn.preprocessing import StandardScaler
    from sklearn.compose import ColumnTransformer
    
    # Handle numeric and categorical features separately
    preprocessor = ColumnTransformer([
        ('num', StandardScaler(), numeric_features),
        ('cat', OneHotEncoder(), categorical_features)
    ])
  4. Build complete pipeline

    python
    model = Pipeline([
        ('preprocessor', preprocessor),
        ('classifier', RandomForestClassifier(random_state=42))
    ])
  5. Tune hyperparameters

    python
    from sklearn.model_selection import GridSearchCV
    
    param_grid = {
        'classifier__n_estimators': [100, 200],
        'classifier__max_depth': [10, 20, None]
    }
    
    grid_search = GridSearchCV(model, param_grid, cv=5)
    grid_search.fit(X_train, y_train)
  6. Evaluate on test set

    python
    from sklearn.metrics import classification_report
    
    best_model = grid_search.best_estimator_
    y_pred = best_model.predict(X_test)
    print(classification_report(y_test, y_pred))
Performing Clustering Analysis
  1. Preprocess data

    python
    from sklearn.preprocessing import StandardScaler
    
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
  2. Find optimal number of clusters

    python
    from sklearn.cluster import KMeans
    from sklearn.metrics import silhouette_score
    
    scores = []
    for k in range(2, 11):
        kmeans = KMeans(n_clusters=k, random_state=42)
        labels = kmeans.fit_predict(X_scaled)
        scores.append(silhouette_score(X_scaled, labels))
    
    optimal_k = range(2, 11)[np.argmax(scores)]
  3. Apply clustering

    python
    model = KMeans(n_clusters=optimal_k, random_state=42)
    labels = model.fit_predict(X_scaled)
  4. Visualize with dimensionality reduction

    python
    from sklearn.decomposition import PCA
    
    pca = PCA(n_components=2)
    X_2d = pca.fit_transform(X_scaled)
    
    plt.scatter(X_2d[:, 0], X_2d[:, 1], c=labels, cmap='viridis')

Best Practices

Always Use Pipelines

Pipelines prevent data leakage and ensure consistency:

python
# Good: Preprocessing in pipeline
pipeline = Pipeline([
    ('scaler', StandardScaler()),
    ('model', LogisticRegression())
])

# Bad: Preprocessing outside (can leak information)
X_scaled = StandardScaler().fit_transform(X)
Fit on Training Data Only

Never fit on test data:

python
# Good
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)  # Only transform

# Bad
scaler = StandardScaler()
X_all_scaled = scaler.fit_transform(np.vstack([X_train, X_test]))
Use Stratified Splitting for Classification

Preserve class distribution:

python
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, stratify=y, random_state=42
)
Set Random State for Reproducibility
python
model = RandomForestClassifier(n_estimators=100, random_state=42)
Choose Appropriate Metrics
  • Balanced data: Accuracy, F1-score
  • Imbalanced data: Precision, Recall, ROC AUC, Balanced Accuracy
  • Cost-sensitive: Define custom scorer
Scale Features When Required

Algorithms requiring feature scaling:

  • SVM, KNN, Neural Networks
  • PCA, Linear/Logistic Regression with regularization
  • K-Means clustering

Algorithms not requiring scaling:

  • Tree-based models (Decision Trees, Random Forest, Gradient Boosting)
  • Naive Bayes

Troubleshooting Common Issues

ConvergenceWarning

Issue: Model didn't converge Solution: Increase max_iter or scale features

python
model = LogisticRegression(max_iter=1000)
Poor Performance on Test Set

Issue: Overfitting Solution: Use regularization, cross-validation, or simpler model

python
# Add regularization
model = Ridge(alpha=1.0)

# Use cross-validation
scores = cross_val_score(model, X, y, cv=5)
Memory Error with Large Datasets

Solution: Use algorithms designed for large data

python
# Use SGD for large datasets
from sklearn.linear_model import SGDClassifier
model = SGDClassifier()

# Or MiniBatchKMeans for clustering
from sklearn.cluster import MiniBatchKMeans
model = MiniBatchKMeans(n_clusters=8, batch_size=100)

Additional Resources

© zLanqing, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in scientific-toolkit-skill/references/scientific-skills/scikit-learn of zLanqing/codex-claude-academic-skills.

  • SKILL.md
  • references/model_evaluation.md
  • references/pipelines_and_composition.md
  • references/preprocessing.md
  • references/quick_reference.md
  • references/supervised_learning.md
  • references/unsupervised_learning.md
  • scripts/classification_pipeline.py
  • scripts/clustering_analysis.py

Open the folder on GitHubat commit 7ed6377

Used in 17 other repositories

We found 43 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 17 other GitHub owners. This page covers the copy in zLanqing/codex-claude-academic-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Scikit Learn next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scikit Learn compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scikit Learn this skillzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries168—~3.1kAutomated safety check: PassApache-2.0
Aeon Time Series Machine Learningdavila7/claude-code-templates32k14 repos~2.6kAutomated safety check: PassMIT
Precisemicroprediction/precise336—~782Automated safety check: PassMIT
scikit-survival Time-to-Event Modelingdavila7/claude-code-templates32k12 repos~3.7kAutomated safety check: PassMIT

Similar skills

  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    168 GitHub stars~3.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Aeon Time Series Machine Learning

    davila7/claude-code-templates

    Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.

    32k GitHub starsUsed in 14 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Precise

    microprediction/precise

    Online (incremental) covariance, correlation, and precision estimation in Python — the streaming complement to sklearn.covariance.

    336 GitHub stars~782 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • scikit-survival Time-to-Event Modeling

    davila7/claude-code-templates

    Fits and evaluates survival models with scikit-survival: Cox models, Random Survival Forests, boosting, survival SVMs, concordance index, Brier score and competing risks.

    32k GitHub starsUsed in 12 repos~3.7k tokens
    Data & AnalyticsAuto-check passed
  • Trains scikit-learn models with walk-forward validation on features from OHLCV data to predict return direction and turn the predictions into trading signals.

    35k GitHub stars~3.2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from zLanqing/codex-claude-academic-skills

All 17 skills in this repo
  • Matplotlib

    zLanqing/codex-claude-academic-skills

    Low-level plotting library for full customization. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~2.9k tokens
    Auto-check passed
  • Simpy

    zLanqing/codex-claude-academic-skills

    Process-based discrete-event simulation framework in Python.

    4.6k GitHub starsUsed in 11 repos~3k tokens
    Auto-check passed
  • Networkx

    zLanqing/codex-claude-academic-skills

    Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs in Python.

    4.6k GitHub starsUsed in 16 repos~3.2k tokens
    Auto-check passed
  • Seaborn

    zLanqing/codex-claude-academic-skills

    Statistical visualization with pandas integration. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 16 repos~4.9k tokens
    Auto-check passed
  • Timesfm Forecasting

    zLanqing/codex-claude-academic-skills

    Zero-shot time series forecasting with Google's TimesFM foundation model.

    4.6k GitHub starsUsed in 6 repos~7.5k tokens
    Auto-check: notes
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 16 repos~4.9k tokens
    Auto-check passed

Questions about Scikit Learn

What does Scikit Learn do?

Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills. Scikit Learn is an agent skill from zLanqing/codex-claude-academic-skills. Machine learning in Python with scikit-learn.

When should I use Scikit Learn?

Scikit Learn fits situations like: working with supervised learning (classification; unsupervised learning (clustering; dimensionality reduction); model evaluation.

How do I install Scikit Learn in Claude Code?

Run `npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a claude-code`. Or copy the skill folder (scientific-toolkit-skill/references/scientific-skills/scikit-learn in zLanqing/codex-claude-academic-skills) into .claude/skills/scikit-learn in your project. Claude Code loads it when a task matches its description.

How do I install Scikit Learn in Codex?

Run `npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a codex`. Or copy the skill folder (scientific-toolkit-skill/references/scientific-skills/scikit-learn in zLanqing/codex-claude-academic-skills) into .agents/skills/scikit-learn in your project. Codex loads it when a task matches its description.

Can I use Scikit Learn in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zLanqing/codex-claude-academic-skills --skill scikit-learn -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scikit-learn, .gemini/skills/scikit-learn, .github/skills/scikit-learn and .opencode/skills/scikit-learn in your project.

What does Scikit Learn need to run?

Going by SKILL.md and its folder, Scikit Learn needs Python for the scripts in its folder and the command-line tools its instructions call (uv and python). Our summary lists: Python 3.

Does Scikit Learn access the network?

SKILL.md names 1 domain. As links in the text: scikit-learn.org. This is read from the text; nothing was executed.

Is Scikit Learn safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Scikit Learn use?

Scikit Learn is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scikit Learn use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.

What are the alternatives to Scikit Learn?

Skills that share tags, products or a category with Scikit Learn: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 168 stars), Aeon Time Series Machine Learning (davila7/claude-code-templates, 32k stars) and Precise (microprediction/precise, 336 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scikit Learn?

zLanqing (a GitHub user) maintains it in zLanqing/codex-claude-academic-skills, which has 4,578 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on May 14, 2026.

Source: zLanqing/codex-claude-academic-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.