Agent skill

Data Scientist

by borghei in borghei/Claude-Skills

Data science across machine learning, statistical modeling, and experimentation.

MITAuto-check passedData & Analytics

Install Data Scientist

skills CLI
$ npx skills add borghei/Claude-Skills --skill data-scientist -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills data-scientist --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-analytics/data-scientist .claude/skills/data-scientist && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-scientist
GitHub stars
874
Token cost
~3.3k tokens
SKILL.md length
1,002 words
Files
4 (incl. scripts)
Skills in repo
364
Repo updated
First seen
Licence
MIT

At a glance

Data science across machine learning, statistical modeling, and experimentation.

  • Works in 6 steps: Define the problem -- Restate the… → Collect and profile data -- Identify… → Engineer features -- Create numerical… → …
  • Selecting ML algorithms
  • SKILL.md covers Clarify First, Workflow, Algorithm Selection Matrix and Feature Engineering Examples, plus 9 more sections
  • Runs Python scripts from its folder; calls python

What it does

Data Scientist is an agent skill from borghei/Claude-Skills. Data science across machine learning, statistical modeling, and experimentation. Use when selecting ML algorithms, engineering features, designing A/B tests, evaluating model performance, or building predictive pipelines.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/experiment_tracker.py`, `scripts/feature_selector.py` and `scripts/hypothesis_tester.py`).

It sits in Data & Analytics, covering A/B testing and Machine learning. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Selecting ML algorithms
  • Engineering features
  • Designing A/B tests
  • Evaluating model performance

Example prompts

  • “/data-scientist”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Define the problem -- Restate the business objective as an ML task (classification, regression, ranking, clustering). Define the primary…
  2. Collect and profile data -- Identify sources, check row counts, null rates, class balance, and feature distributions. Flag data-quality…
  3. Engineer features -- Create numerical transforms (log, binning), encode categoricals (one-hot, target, frequency), extract time components…
  4. Select and train models -- Use the algorithm selection matrix below. Start simple (logistic/linear regression), then add complexity…
  5. Evaluate rigorously -- Report classification metrics (accuracy, precision, recall, F1, AUC-ROC) or regression metrics (MAE, RMSE…
  6. Communicate results -- Present business impact (e.g., "model reduces false positives by 30%, saving $500K/yr"). Recommend deployment path…

What it can do on your machine

Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Scientist loads about 3.3k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 1,002 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 1,002 words, ~3,299 tokens.

Download SKILL.mdSave it as .claude/skills/data-scientist/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
data-scientist
description
Data science across machine learning, statistical modeling, and experimentation. Use when selecting ML algorithms, engineering features, designing A/B tests, evaluating model performance, or building predictive pipelines.
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
data-analytics
metadata.updated
2026-03-31
metadata.tags
data-science, machine-learning, statistics, modeling, analytics

Data Scientist

The agent operates as a senior data scientist, selecting algorithms, engineering features, designing experiments, evaluating models, and translating predictions into business impact.

Clarify First

Before modeling, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • ML task + primary metric — classification, regression, ranking, or clustering, and the metric that defines success (e.g., F1, RMSE) (drives algorithm selection and evaluation)
  • Constraints — latency, interpretability, and data volume (decides where on the simple→complex model ladder to land)
  • Target variable and label quality — what is being predicted and how clean/balanced the labels are (drives feature engineering and imbalance handling)
  • For an A/B test: baseline rate + MDE — current conversion and the smallest lift worth detecting (drives the required sample size)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Workflow

  1. Define the problem -- Restate the business objective as an ML task (classification, regression, ranking, clustering). Define the primary evaluation metric (e.g., F1 for imbalanced classification, RMSE for regression). Document constraints (latency, interpretability, data volume).
  2. Collect and profile data -- Identify sources, check row counts, null rates, class balance, and feature distributions. Flag data-quality issues before modeling.
  3. Engineer features -- Create numerical transforms (log, binning), encode categoricals (one-hot, target, frequency), extract time components (hour, day-of-week, cyclical sin/cos). Select top features via importance, mutual information, or RFE.
  4. Select and train models -- Use the algorithm selection matrix below. Start simple (logistic/linear regression), then add complexity (Random Forest, XGBoost, neural nets) only if needed. Use cross-validation.
  5. Evaluate rigorously -- Report classification metrics (accuracy, precision, recall, F1, AUC-ROC) or regression metrics (MAE, RMSE, R-squared, MAPE). Compare against a baseline. Check for overfitting (train vs. test gap).
  6. Communicate results -- Present business impact (e.g., "model reduces false positives by 30%, saving $500K/yr"). Recommend deployment path or next experiment.

Algorithm Selection Matrix

ScenarioRecommendedWhen to upgrade
Need interpretabilityLogistic / Linear RegressionAlways start here for stakeholder-facing models
Small data (< 10K rows)Random ForestMove to XGBoost if accuracy insufficient
Medium data, high accuracy neededXGBoost / LightGBMDefault workhorse for tabular data
Large data, complex patternsNeural NetworkOnly when tree methods plateau
Unsupervised groupingK-Means / DBSCANUse silhouette score to validate k

Feature Engineering Examples

Numerical transforms:

python
import numpy as np, pandas as pd

def engineer_numerical(df: pd.DataFrame, col: str) -> pd.DataFrame:
    return pd.DataFrame({
        f'{col}_log':     np.log1p(df[col]),
        f'{col}_sqrt':    np.sqrt(df[col].clip(lower=0)),
        f'{col}_squared': df[col] ** 2,
        f'{col}_binned':  pd.cut(df[col], bins=5, labels=False),
    })

Time-based features with cyclical encoding:

python
def engineer_time(df: pd.DataFrame, col: str) -> pd.DataFrame:
    dt = pd.to_datetime(df[col])
    return pd.DataFrame({
        f'{col}_hour':      dt.dt.hour,
        f'{col}_dayofweek': dt.dt.dayofweek,
        f'{col}_month':     dt.dt.month,
        f'{col}_is_weekend': dt.dt.dayofweek.isin([5, 6]).astype(int),
        f'{col}_hour_sin':  np.sin(2 * np.pi * dt.dt.hour / 24),
        f'{col}_hour_cos':  np.cos(2 * np.pi * dt.dt.hour / 24),
    })

Feature selection (importance-based):

python
from sklearn.ensemble import RandomForestClassifier

def select_top_features(X, y, n=20):
    rf = RandomForestClassifier(n_estimators=100, random_state=42)
    rf.fit(X, y)
    importance = pd.Series(rf.feature_importances_, index=X.columns)
    return importance.nlargest(n).index.tolist()

Model Evaluation

Classification:

python
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score

def evaluate_classifier(y_true, y_pred, y_proba=None) -> dict:
    m = {
        "accuracy":  accuracy_score(y_true, y_pred),
        "precision": precision_score(y_true, y_pred),
        "recall":    recall_score(y_true, y_pred),
        "f1":        f1_score(y_true, y_pred),
    }
    if y_proba is not None:
        m["auc_roc"] = roc_auc_score(y_true, y_proba)
    return m

Regression:

python
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np

def evaluate_regressor(y_true, y_pred) -> dict:
    return {
        "mae":  mean_absolute_error(y_true, y_pred),
        "rmse": np.sqrt(mean_squared_error(y_true, y_pred)),
        "r2":   r2_score(y_true, y_pred),
    }

A/B Test Design and Analysis

Sample size calculation:

python
from scipy import stats
import numpy as np

def required_sample_size(baseline_rate: float, mde: float, alpha: float = 0.05, power: float = 0.8) -> int:
    """Return required N per variant. mde is relative (e.g., 0.10 = 10% lift)."""
    effect = baseline_rate * mde
    z_a = stats.norm.ppf(1 - alpha / 2)
    z_b = stats.norm.ppf(power)
    p = baseline_rate
    return int(np.ceil(2 * p * (1 - p) * (z_a + z_b) ** 2 / effect ** 2))

# Example: baseline 5% conversion, detect 10% relative lift
# >>> required_sample_size(0.05, 0.10)  -> ~62,214 per variant

Result analysis:

python
def analyze_ab(control: np.ndarray, treatment: np.ndarray, alpha: float = 0.05) -> dict:
    """Analyze A/B test with proportions z-test."""
    n_c, n_t = len(control), len(treatment)
    p_c, p_t = control.mean(), treatment.mean()
    p_pool = (control.sum() + treatment.sum()) / (n_c + n_t)
    se = np.sqrt(p_pool * (1 - p_pool) * (1/n_c + 1/n_t))
    z = (p_t - p_c) / se
    p_val = 2 * (1 - stats.norm.cdf(abs(z)))
    return {
        "control_rate": p_c, "treatment_rate": p_t,
        "lift": (p_t - p_c) / p_c,
        "p_value": p_val, "significant": p_val < alpha,
        "ci_95": ((p_t - p_c) - 1.96 * se, (p_t - p_c) + 1.96 * se),
    }

Project Template

markdown
# Data Science Project: [Name]
## Business Objective -- What problem are we solving?
## Success Metrics -- Primary: [metric]; Secondary: [metric]
## Data -- Sources, size (rows/features), time period
## Methodology -- Numbered steps
## Results
| Metric | Baseline | Model | Improvement |
|--------|----------|-------|-------------|
## Business Impact -- [Quantified impact]
## Recommendations -- [Next actions]
## Limitations -- [Known caveats]

Scripts

bash
python scripts/experiment_tracker.py log --name "xgb_v2" --params '{"lr":0.1,"depth":6}' --metrics '{"f1":0.87,"auc":0.92}'
python scripts/experiment_tracker.py list --sort-by f1 --top 5
python scripts/experiment_tracker.py compare --ids 1 3 5 --json
python scripts/hypothesis_tester.py ttest --file data.csv --col-a group_a --col-b group_b
python scripts/hypothesis_tester.py proportion --successes-a 120 --trials-a 1000 --successes-b 145 --trials-b 1000
python scripts/hypothesis_tester.py chi-square --file contingency.csv --json
python scripts/feature_selector.py --file dataset.csv --target churn --top 10
python scripts/feature_selector.py --file dataset.csv --target revenue --method correlation --json

Tool Reference

ToolPurposeKey Flags
experiment_tracker.pyLog, list, and compare experiments with parameters, metrics, and tags in a local JSON filelog --name --params --metrics --tags, list --sort-by --top, compare --ids, --json
hypothesis_tester.pyRun statistical tests: Welch's t-test, paired t-test, proportion z-test, chi-square independencettest --file --col-a --col-b [--paired], proportion --successes-a --trials-a ..., chi-square --file, --json
feature_selector.pyRank features by composite score (variance, correlation, mutual information, null rate) for a target column--file <csv>, --target <col>, --top <n>, --method all/correlation/mutual_info, --json

Troubleshooting

ProblemLikely CauseResolution
Model overfits (large train-test gap in metrics)Too many features, insufficient regularization, or data leakageReduce feature count with feature_selector.py, add regularization, and audit feature engineering for temporal leakage
A/B test shows significant result but tiny effect sizeLarge sample size makes small differences statistically significantAlways report effect size (Cohen's d) alongside p-value; use practical significance thresholds
hypothesis_tester.py p-value differs from scipyThe tool uses normal/t-distribution approximations (standard library only)For publication-grade analysis, validate with scipy.stats; the tool is designed for fast directional estimates
Feature importance scores are near-zero for all featuresTarget variable has extremely low variance or the feature set lacks predictive signalCheck target distribution; consider feature engineering or collecting additional data sources
experiment_tracker.py shows experiment IDs out of orderExperiments were logged non-sequentially or the log file was manually editedIDs are auto-incremented; use --sort-by on a metric for meaningful ordering
Chi-square test fails with "table must be at least 2x2"CSV contingency table has fewer than 2 rows or 2 columns of numeric dataEnsure the CSV has a header row and at least 2x2 numeric cells; verify the format matches expectations
Class imbalance causes misleading accuracyAccuracy inflated by majority class predictionsUse F1, precision-recall, or AUC-ROC instead; apply SMOTE or class weights during training
Show full SKILL.md (296 more words)Show less

Success Criteria

  • Every ML project follows the Define-Collect-Engineer-Train-Evaluate-Communicate workflow before deployment.
  • Feature selection is documented: feature_selector.py output is saved with the experiment record.
  • All experiments are tracked with experiment_tracker.py including parameters, metrics, and a descriptive name.
  • Model evaluation reports include at least 3 metrics (e.g., F1, AUC-ROC, precision) and comparison against a baseline.
  • A/B tests pre-register the hypothesis, sample size calculation, and primary metric before data collection begins.
  • Statistical tests report effect size and confidence intervals, not just p-values.
  • Business impact is quantified in dollar terms or user-metric terms (e.g., "reduces false positives by 30%, saving $500K/yr").

Scope & Limitations

In scope: Machine learning algorithm selection, feature engineering, model training and evaluation, A/B test design and analysis, statistical hypothesis testing, experiment tracking, and communicating results to stakeholders.

Out of scope: Model deployment to production (see ml-ops-engineer), data pipeline infrastructure, dashboard development, and real-time serving architecture.

Limitations: The Python tools use only the Python standard library. hypothesis_tester.py uses normal and t-distribution approximations that are accurate for moderate sample sizes but should be validated with scipy for edge cases (very small n, extreme skew). feature_selector.py computes approximate mutual information using binned discretization -- for high-precision feature selection, use sklearn's mutual_info_classif or permutation importance. All tools process local files and do not integrate with MLflow, W&B, or other tracking platforms.

Integration Points

  • MLOps Engineer (data-analytics/ml-ops-engineer): Trained models are handed off for production deployment, monitoring, and registry management.
  • Data Analyst (data-analytics/data-analyst): Complex analytical questions requiring predictive modeling are escalated from the analyst to the data scientist.
  • Analytics Engineer (data-analytics/analytics-engineer): Feature engineering pipelines may depend on mart models as upstream data sources.
  • Product Team (product-team/): Experiment results inform product decisions; A/B test designs are co-created with product managers.
  • Engineering (engineering/senior-ml-engineer): Algorithm implementation details and model architecture decisions bridge data science and ML engineering.

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in data-analytics/data-scientist of borghei/Claude-Skills.

  • SKILL.md
  • scripts/experiment_tracker.py
  • scripts/feature_selector.py
  • scripts/hypothesis_tester.py

Open the folder on GitHubat commit c9a1487

Compare with similar skills

Data Scientist next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Scientist compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Scientist this skillborghei/Claude-Skills874—~3.3kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Automl SkillLeoYeAI/openclaw-master-skills2.2k—~3.6kAutomated safety check: PassMIT
Data Sciencemajiayu000/claude-skill-registry6661 repos~4.3kAutomated safety check: PassMIT
ML Experiment Evaluationhashgraph-online/awesome-codex-plugins1.2k—~801Automated safety check: PassMIT
Data Scientisttheneoai/awesome-skills183—~2kAutomated safety check: PassMIT

Similar skills

  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Automl Skill

    LeoYeAI/openclaw-master-skills

    AutoML 自动化机器学习技能 | Automated Machine Learning Skill. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~3.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Data Science

    majiayu000/claude-skill-registry

    A skill your agent uses when performing exploratory data analysis, statistical testing, data visualization, or building predictive models.

    666 GitHub starsUsed in 1 repo~4.3k tokens
    Data & AnalyticsAuto-check passed
  • ML Experiment Evaluation

    hashgraph-online/awesome-codex-plugins

    Plan evaluation strategies for machine-learning product changes.

    1.2k GitHub stars~801 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Data Scientist

    theneoai/awesome-skills

    Elite Data Scientist skill with expertise in statistical analysis, predictive modeling, experimental design (A/B testing), feature engineering, and data visualization.

    183 GitHub stars~2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Model Evaluator

    FerroxLabs/wayland

    ML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2), confusion matrix analysis, cross-validation strategies, bias detection…

    608 GitHub stars~4.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from borghei/Claude-Skills

All 364 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    874 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    874 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    874 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    874 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Data Scientist

What does Data Scientist do?

Data science across machine learning, statistical modeling, and experimentation. Data Scientist is an agent skill from borghei/Claude-Skills. Data science across machine learning, statistical modeling, and experimentation.

When should I use Data Scientist?

Data Scientist fits situations like: selecting ML algorithms; engineering features; designing A/B tests; evaluating model performance.

How do I install Data Scientist in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill data-scientist -a claude-code`. Or copy the skill folder (data-analytics/data-scientist in borghei/Claude-Skills) into .claude/skills/data-scientist in your project. Claude Code loads it when a task matches its description.

How do I install Data Scientist in Codex?

Run `npx skills add borghei/Claude-Skills --skill data-scientist -a codex`. Or copy the skill folder (data-analytics/data-scientist in borghei/Claude-Skills) into .agents/skills/data-scientist in your project. Codex loads it when a task matches its description.

Can I use Data Scientist in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill data-scientist -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-scientist, .gemini/skills/data-scientist, .github/skills/data-scientist and .opencode/skills/data-scientist in your project.

What does Data Scientist need to run?

Going by SKILL.md and its folder, Data Scientist needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Data Scientist access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Scientist safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Scientist use?

Data Scientist is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Scientist use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Scientist?

Skills that share tags, products or a category with Data Scientist: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Automl Skill (LeoYeAI/openclaw-master-skills, 2.2k stars), Data Science (majiayu000/claude-skill-registry, 666 stars) and ML Experiment Evaluation (hashgraph-online/awesome-codex-plugins, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Scientist?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.