Agent skill

Python Data Scientist

by FerroxLabs in FerroxLabs/wayland

Guide for applied machine learning with scikit-learn covering feature engineering, model selection, pipeline construction, evaluation, hyperparameter tuning, and production-ready model patterns.

Apache-2.0Auto-check passedData & Analytics

Install Python Data Scientist

skills CLI
$ npx skills add FerroxLabs/wayland --skill python-data-scientist -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FerroxLabs/wayland python-data-scientist --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/python-data-scientist .claude/skills/python-data-scientist && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
python-data-scientist
GitHub stars
608
Token cost
~4.2k tokens
SKILL.md length
479 words
Files
1
Skills in repo
1,194
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide for applied machine learning with scikit-learn covering feature engineering, model selection, pipeline construction, evaluation, hyperparameter tuning, and production-ready model patterns.

  • Works in 5 steps: Gather information. Ask the user… → Analyze context. Review the information… → Develop recommendations. Apply domain… → …
  • The user asks about python data scientist
  • SKILL.md covers When to Use, Project Structure, Feature Engineering and Pipeline Construction, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Python Data Scientist is an agent skill from FerroxLabs/wayland. Guide for applied machine learning with scikit-learn covering feature engineering, model selection, pipeline construction, evaluation, hyperparameter tuning, and production-ready model patterns. Use when the user asks about python data scientist, related techniques, best practices, or needs guidance in this domain. Do NOT use when the request is outside the scope of python data scientist or requires a different specialized skill.

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Machine learning. It works with Python and scikit-learn. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.

When your agent uses it

  • The user asks about python data scientist
  • Related techniques
  • Needs guidance in this domain
  • The request is outside the scope of python data scientist

Example prompts

  • “/python-data-scientist”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Gather information. Ask the user clarifying questions to understand their specific situation, goals, and constraints
  2. Analyze context. Review the information provided and identify key factors relevant to python data scientist
  3. Develop recommendations. Apply domain expertise to create actionable guidance tailored to the user's needs
  4. Present structured output. Deliver findings in the output format below with clear next steps
  5. Address follow-ups. Answer additional questions and refine recommendations based on feedback

What it can do on your machine

Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and template).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Python Data Scientist loads about 4.2k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 479 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 479 words, ~4,185 tokens.

Download SKILL.mdSave it as .claude/skills/python-data-scientist/SKILL.md (or your agent's skills folder).
name
python-data-scientist
description
Guide for applied machine learning with scikit-learn covering feature engineering, model selection, pipeline construction, evaluation, hyperparameter tuning, and production-ready model patterns. Use when the user asks about python data scientist, related techniques, best practices, or needs guidance in this domain. Do NOT use when the request is outside the scope of python data scientist or requires a different specialized skill.
license
Apache-2.0
metadata.author
foundry-skills
metadata.version
1.0.0
metadata.tags
data-science statistics guide python testing performing-arts
metadata.category
data-analysis
metadata.subcategory
statistics-modeling
metadata.disclaimer
none
metadata.difficulty
intermediate

Python Data Scientist

You are an expert applied data scientist who builds robust machine learning pipelines with scikit-learn, engineering features methodically, selecting models systematically, and evaluating results rigorously.

When to Use

Use this skill when:

  • User asks about python data scientist techniques or best practices
  • User needs guidance on python data scientist concepts
  • User wants to implement or improve their approach to python data scientist

Do NOT use when:

  • The request falls outside the scope of python data scientist
  • User needs a different specialized skill for their specific situation
  • The topic requires professional consultation beyond general guidance

Project Structure

ml-project/
├── data/
│   ├── raw/               # Immutable original data
│   ├── processed/          # Cleaned, feature-engineered data
│   └── external/           # Third-party reference data
├── src/
│   ├── data/
│   │   ├── ingestion.py    # Data loading
│   │   └── validation.py   # Schema checks
│   ├── features/
│   │   ├── engineering.py  # Feature transforms
│   │   └── selection.py    # Feature selection
│   ├── models/
│   │   ├── train.py        # Training pipeline
│   │   ├── evaluate.py     # Metrics and reports
│   │   └── predict.py      # Inference
│   └── utils/
│       └── config.py       # Hyperparameters, paths
├── notebooks/
│   ├── 01-exploration.ipynb
│   └── 02-modeling.ipynb
├── models/                 # Serialized model artifacts
├── tests/
└── pyproject.toml

Feature Engineering

Numeric Features
python
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import (
    StandardScaler, MinMaxScaler, RobustScaler,
    PowerTransformer, QuantileTransformer
)
from sklearn.impute import SimpleImputer
import numpy as np

numeric_pipeline = Pipeline([
    ('imputer', SimpleImputer(strategy='median')),
    ('scaler', RobustScaler()),  # Robust to outliers
])

# When to use each scaler:
# StandardScaler    - Normal-ish data, linear models, SVMs
# MinMaxScaler      - Neural networks, bounded features
# RobustScaler      - Data with outliers
# PowerTransformer  - Skewed distributions (Box-Cox, Yeo-Johnson)
# QuantileTransformer - Force uniform or normal distribution
Categorical Features
python
from sklearn.preprocessing import (
    OneHotEncoder, OrdinalEncoder, TargetEncoder
)

# One-hot: low cardinality (<15 categories)
ohe = OneHotEncoder(
    drop='if_binary',           # Drop redundant column for binary
    handle_unknown='ignore',     # Handle unseen categories at predict time
    sparse_output=True,          # Memory efficient for high-dim
    min_frequency=0.01,          # Group rare categories
)

# Ordinal: ordered categories
ordinal = OrdinalEncoder(
    categories=[['low', 'medium', 'high', 'critical']],
    handle_unknown='use_encoded_value',
    unknown_value=-1,
)

# Target encoding: high cardinality (cities, zip codes)
target_enc = TargetEncoder(
    smooth='auto',               # Regularization
    target_type='continuous',
)
Date/Time Features
python
def extract_datetime_features(df, col):
    """Extract useful features from a datetime column."""
    df = df.copy()
    dt = df[col]

    df[f'{col}_year'] = dt.dt.year
    df[f'{col}_month'] = dt.dt.month
    df[f'{col}_day_of_week'] = dt.dt.dayofweek
    df[f'{col}_hour'] = dt.dt.hour
    df[f'{col}_is_weekend'] = dt.dt.dayofweek.isin([5, 6]).astype(int)
    df[f'{col}_quarter'] = dt.dt.quarter
    df[f'{col}_day_of_year'] = dt.dt.dayofyear

    # Cyclical encoding for periodic features
    df[f'{col}_month_sin'] = np.sin(2 * np.pi * dt.dt.month / 12)
    df[f'{col}_month_cos'] = np.cos(2 * np.pi * dt.dt.month / 12)
    df[f'{col}_hour_sin'] = np.sin(2 * np.pi * dt.dt.hour / 24)
    df[f'{col}_hour_cos'] = np.cos(2 * np.pi * dt.dt.hour / 24)

    return df
Custom Transformer
python
from sklearn.base import BaseEstimator, TransformerMixin

class InteractionFeatures(BaseEstimator, TransformerMixin):
    """Create interaction features between specified column pairs."""

    def __init__(self, interaction_pairs):
        self.interaction_pairs = interaction_pairs

    def fit(self, X, y=None):
        return self

    def transform(self, X):
        X = X.copy()
        for col_a, col_b in self.interaction_pairs:
            X[f'{col_a}_x_{col_b}'] = X[col_a] * X[col_b]
            X[f'{col_a}_div_{col_b}'] = X[col_a] / X[col_b].replace(0, np.nan)
        return X

    def get_feature_names_out(self, input_features=None):
        names = list(input_features) if input_features else []
        for col_a, col_b in self.interaction_pairs:
            names.extend([f'{col_a}_x_{col_b}', f'{col_a}_div_{col_b}'])
        return names

Pipeline Construction

Full ML Pipeline
python
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.ensemble import GradientBoostingClassifier

# Define column groups
numeric_features = ['age', 'income', 'tenure_months', 'num_products']
categorical_features = ['region', 'plan_type', 'channel']

# Preprocessing
preprocessor = ColumnTransformer(
    transformers=[
        ('num', Pipeline([
            ('imputer', SimpleImputer(strategy='median')),
            ('scaler', StandardScaler()),
        ]), numeric_features),
        ('cat', Pipeline([
            ('imputer', SimpleImputer(strategy='constant', fill_value='unknown')),
            ('encoder', OneHotEncoder(handle_unknown='ignore', sparse_output=False)),
        ]), categorical_features),
    ],
    remainder='drop',
    verbose_feature_names_out=False,
)

# Full pipeline
pipeline = Pipeline([
    ('preprocessor', preprocessor),
    ('classifier', GradientBoostingClassifier(
        n_estimators=200,
        max_depth=5,
        learning_rate=0.1,
        random_state=42,
    )),
])

# Fit
pipeline.fit(X_train, y_train)

# Predict
y_pred = pipeline.predict(X_test)
y_proba = pipeline.predict_proba(X_test)[:, 1]

Model Selection Guide

By Problem Type
ProblemStart WithThen TryWhen to Use
Binary classificationLogistic RegressionGBM, Random ForestChurn, fraud, click
Multi-classRandom ForestGBM, SVMCategory prediction
RegressionRidge/LassoGBM, Random ForestRevenue, pricing
RankingLambdaMART (LightGBM)XGBoost rankerSearch, recommendations
Anomaly detectionIsolation ForestOne-class SVM, LOFFraud, outliers
ClusteringK-MeansDBSCAN, HDBSCANSegmentation
Time seriesARIMA/ProphetLSTM, LightGBMForecasting
Model Comparison
python
from sklearn.model_selection import cross_val_score
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import (
    RandomForestClassifier, GradientBoostingClassifier
)

models = {
    'Logistic Regression': LogisticRegression(max_iter=1000, random_state=42),
    'Random Forest': RandomForestClassifier(n_estimators=200, random_state=42),
    'Gradient Boosting': GradientBoostingClassifier(n_estimators=200, random_state=42),
}

results = {}
for name, model in models.items():
    pipe = Pipeline([
        ('preprocessor', preprocessor),
        ('model', model),
    ])
    scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring='roc_auc')
    results[name] = {
        'mean_auc': scores.mean(),
        'std_auc': scores.std(),
        'scores': scores,
    }
    print(f"{name}: AUC = {scores.mean():.4f} (+/- {scores.std():.4f})")

Model Evaluation

Classification Metrics
python
from sklearn.metrics import (
    classification_report, confusion_matrix,
    roc_auc_score, average_precision_score,
    roc_curve, precision_recall_curve
)

def evaluate_classifier(y_true, y_pred, y_proba, model_name='Model'):
    """Comprehensive classification evaluation."""
    print(f"=== {model_name} Evaluation ===\n")

    # Classification report
    print(classification_report(y_true, y_pred, digits=3))

    # AUC metrics
    roc_auc = roc_auc_score(y_true, y_proba)
    avg_precision = average_precision_score(y_true, y_proba)
    print(f"ROC AUC:            {roc_auc:.4f}")
    print(f"Average Precision:  {avg_precision:.4f}")

    # Confusion matrix
    cm = confusion_matrix(y_true, y_pred)
    print(f"\nConfusion Matrix:")
    print(f"  TN={cm[0,0]:5d}  FP={cm[0,1]:5d}")
    print(f"  FN={cm[1,0]:5d}  TP={cm[1,1]:5d}")

    return {'roc_auc': roc_auc, 'avg_precision': avg_precision}
Regression Metrics
python
from sklearn.metrics import (
    mean_absolute_error, mean_squared_error,
    r2_score, mean_absolute_percentage_error
)

def evaluate_regressor(y_true, y_pred, model_name='Model'):
    """Comprehensive regression evaluation."""
    mae = mean_absolute_error(y_true, y_pred)
    rmse = mean_squared_error(y_true, y_pred, squared=False)
    r2 = r2_score(y_true, y_pred)
    mape = mean_absolute_percentage_error(y_true, y_pred)

    print(f"=== {model_name} Evaluation ===")
    print(f"MAE:   {mae:.4f}")
    print(f"RMSE:  {rmse:.4f}")
    print(f"R2:    {r2:.4f}")
    print(f"MAPE:  {mape:.2%}")

    return {'mae': mae, 'rmse': rmse, 'r2': r2, 'mape': mape}

Hyperparameter Tuning

python
from sklearn.model_selection import RandomizedSearchCV
from scipy.stats import randint, uniform

param_distributions = {
    'classifier__n_estimators': randint(100, 500),
    'classifier__max_depth': randint(3, 10),
    'classifier__learning_rate': uniform(0.01, 0.3),
    'classifier__subsample': uniform(0.6, 0.4),
    'classifier__min_samples_leaf': randint(5, 50),
}

search = RandomizedSearchCV(
    pipeline,
    param_distributions,
    n_iter=50,
    cv=5,
    scoring='roc_auc',
    random_state=42,
    n_jobs=-1,
    verbose=1,
)

search.fit(X_train, y_train)

print(f"Best AUC: {search.best_score_:.4f}")
print(f"Best params: {search.best_params_}")
best_model = search.best_estimator_
Optuna Integration
python
import optuna
from sklearn.model_selection import cross_val_score

def objective(trial):
    params = {
        'classifier__n_estimators': trial.suggest_int('n_estimators', 100, 500),
        'classifier__max_depth': trial.suggest_int('max_depth', 3, 10),
        'classifier__learning_rate': trial.suggest_float('learning_rate', 0.01, 0.3, log=True),
        'classifier__subsample': trial.suggest_float('subsample', 0.6, 1.0),
        'classifier__min_samples_leaf': trial.suggest_int('min_samples_leaf', 5, 50),
    }

    pipe = pipeline.set_params(**params)
    scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring='roc_auc')
    return scores.mean()

study = optuna.create_study(direction='maximize')
study.optimize(objective, n_trials=100)

print(f"Best AUC: {study.best_value:.4f}")
print(f"Best params: {study.best_params}")

Feature Selection

python
from sklearn.feature_selection import (
    SelectKBest, f_classif, mutual_info_classif
)
from sklearn.inspection import permutation_importance

# Method 1: Statistical tests
selector = SelectKBest(score_func=f_classif, k=20)
X_selected = selector.fit_transform(X_train_processed, y_train)

# Method 2: Model-based importance
model.fit(X_train_processed, y_train)
importances = pd.DataFrame({
    'feature': feature_names,
    'importance': model.feature_importances_,
}).sort_values('importance', ascending=False)

# Method 3: Permutation importance (model-agnostic)
perm_importance = permutation_importance(
    model, X_test_processed, y_test,
    n_repeats=10, random_state=42, scoring='roc_auc'
)

perm_df = pd.DataFrame({
    'feature': feature_names,
    'importance_mean': perm_importance.importances_mean,
    'importance_std': perm_importance.importances_std,
}).sort_values('importance_mean', ascending=False)

Cross-Validation Strategies

python
from sklearn.model_selection import (
    StratifiedKFold, TimeSeriesSplit, GroupKFold,
    RepeatedStratifiedKFold
)

# Imbalanced classification
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

# Time series (no data leakage)
cv = TimeSeriesSplit(n_splits=5, gap=7)  # 7-day gap between train/test

# Grouped data (e.g., same user in train and test)
cv = GroupKFold(n_splits=5)
scores = cross_val_score(pipeline, X, y, cv=cv, groups=user_ids)

# More robust estimate
cv = RepeatedStratifiedKFold(n_splits=5, n_repeats=3, random_state=42)

Model Serialization

python
import joblib
from pathlib import Path
from datetime import datetime

def save_model(pipeline, metrics, model_dir='models'):
    """Save model with metadata."""
    model_dir = Path(model_dir)
    model_dir.mkdir(exist_ok=True)

    timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
    model_path = model_dir / f'model_{timestamp}.joblib'
    meta_path = model_dir / f'model_{timestamp}_meta.json'

    # Save model
    joblib.dump(pipeline, model_path)

    # Save metadata
    import json
    metadata = {
        'timestamp': timestamp,
        'metrics': metrics,
        'features': list(pipeline.named_steps['preprocessor'].get_feature_names_out()),
        'model_type': type(pipeline.named_steps['classifier']).__name__,
        'model_path': str(model_path),
    }
    with open(meta_path, 'w') as f:
        json.dump(metadata, f, indent=2, default=str)

    print(f"Model saved to {model_path}")
    return model_path

Common Pitfalls

PitfallSymptomFix
Data leakageUnrealistically high test performanceFit preprocessing only on train data
Target leakageFeature contains future informationAudit feature timestamps
Class imbalanceHigh accuracy, low recallUse stratified CV, SMOTE, class weights
OverfittingTrain >> test performanceRegularization, simpler model, more data
Feature scale issuesLinear model ignores some featuresScale all numeric features
Missing value patternsModel fails on new dataHandle unknowns in encoders
Train/serve skewGood offline, bad onlineUse same pipeline for train and predict
Temporal leakageRandom CV on time seriesUse TimeSeriesSplit
Show full SKILL.md (187 more words)Show less

Process

  1. Gather information. Ask the user clarifying questions to understand their specific situation, goals, and constraints
  2. Analyze context. Review the information provided and identify key factors relevant to python data scientist
  3. Develop recommendations. Apply domain expertise to create actionable guidance tailored to the user's needs
  4. Present structured output. Deliver findings in the output format below with clear next steps
  5. Address follow-ups. Answer additional questions and refine recommendations based on feedback

Output Format

template
## Python Data Scientist Analysis

### Assessment
[Key findings and observations]

### Recommendations
1. [Primary recommendation]
2. [Secondary recommendation]
3. [Additional suggestions]

### Action Items
- [ ] [First action step]
- [ ] [Second action step]
- [ ] [Follow-up task]

Edge Cases

  • Incomplete information: Ask clarifying questions before proceeding with recommendations
  • Conflicting requirements: Prioritize the most critical constraint and note trade-offs
  • Out of scope requests: Redirect to appropriate specialized skill or professional resource
  • Beginner vs advanced: Adjust depth and terminology based on user's experience level

Example

Input: "Help me with python data scientist for my current situation"

Output:

Based on your situation, here is a structured approach to python data scientist:

  1. Assessment: Evaluate your current state and identify key areas for improvement
  2. Strategy: Develop a targeted plan based on best practices
  3. Implementation: Execute the plan with specific, measurable steps
  4. Review: Monitor progress and adjust as needed

© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/process/resources/skills-library/bodies/skills/data-analysis/python-data-scientist of FerroxLabs/wayland.

Open the folder on GitHubat commit 4c030c7

Compare with similar skills

Python Data Scientist next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Python Data Scientist compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Python Data Scientist this skillFerroxLabs/wayland608—~4.2kAutomated safety check: PassApache-2.0
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries168—~3.1kAutomated safety check: PassApache-2.0
Aeon Time Series Machine Learningdavila7/claude-code-templates32k14 repos~2.6kAutomated safety check: PassMIT
Precisemicroprediction/precise336—~782Automated safety check: PassMIT

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    168 GitHub stars~3.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Aeon Time Series Machine Learning

    davila7/claude-code-templates

    Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.

    32k GitHub starsUsed in 14 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Precise

    microprediction/precise

    Online (incremental) covariance, correlation, and precision estimation in Python — the streaming complement to sklearn.covariance.

    336 GitHub stars~782 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • scikit-survival Time-to-Event Modeling

    davila7/claude-code-templates

    Fits and evaluates survival models with scikit-survival: Cox models, Random Survival Forests, boosting, survival SVMs, concordance index, Brier score and competing risks.

    32k GitHub starsUsed in 12 repos~3.7k tokens
    Data & AnalyticsAuto-check passed

More from FerroxLabs/wayland

All 1,194 skills in this repo
  • Star Office Helper

    FerroxLabs/wayland

    Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.

    608 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Openclaw Setup

    FerroxLabs/wayland

    OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.

    608 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Tvcontrol Setup

    FerroxLabs/wayland

    Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.

    608 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Ab Testing Specialist

    FerroxLabs/wayland

    End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

    608 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Academic Writer

    FerroxLabs/wayland

    Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…

    608 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Accessibility Auditor

    FerroxLabs/wayland

    Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…

    608 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed

Questions about Python Data Scientist

What does Python Data Scientist do?

Guide for applied machine learning with scikit-learn covering feature engineering, model selection, pipeline construction, evaluation, hyperparameter tuning, and production-ready model patterns. Python Data Scientist is an agent skill from FerroxLabs/wayland. Guide for applied machine learning with scikit-learn covering feature engineering, model selection, pipeline construction, evaluation, hyperparameter tuning, and production-ready model patterns.

When should I use Python Data Scientist?

Python Data Scientist fits situations like: the user asks about python data scientist; related techniques; needs guidance in this domain; the request is outside the scope of python data scientist.

How do I install Python Data Scientist in Claude Code?

Run `npx skills add FerroxLabs/wayland --skill python-data-scientist -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/python-data-scientist in FerroxLabs/wayland) into .claude/skills/python-data-scientist in your project. Claude Code loads it when a task matches its description.

How do I install Python Data Scientist in Codex?

Run `npx skills add FerroxLabs/wayland --skill python-data-scientist -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/python-data-scientist in FerroxLabs/wayland) into .agents/skills/python-data-scientist in your project. Codex loads it when a task matches its description.

Can I use Python Data Scientist in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill python-data-scientist -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/python-data-scientist, .gemini/skills/python-data-scientist, .github/skills/python-data-scientist and .opencode/skills/python-data-scientist in your project.

What does Python Data Scientist need to run?

SKILL.md names no scripts, command-line tools or credentials: Python Data Scientist is instructions for the agent only. Our summary lists: Python 3.

Does Python Data Scientist access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Python Data Scientist safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Python Data Scientist use?

Python Data Scientist is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Python Data Scientist use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Python Data Scientist?

Skills that share tags, products or a category with Python Data Scientist: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 168 stars) and Aeon Time Series Machine Learning (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Python Data Scientist?

FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.

Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.