Senior Data Scientist
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
Data science across machine learning, statistical modeling, and experimentation.
$ npx skills add borghei/Claude-Skills --skill data-scientist -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install borghei/Claude-Skills data-scientist --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-analytics/data-scientist .claude/skills/data-scientist && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-scientist" agent skill from https://github.com/borghei/Claude-Skills/tree/main/data-analytics/data-scientist into .claude/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/borghei/Claude-Skills/tree/main/data-analytics/data-scientistType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add borghei/Claude-Skills --skill data-scientist -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install borghei/Claude-Skills data-scientist --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/data-analytics/data-scientist .agents/skills/data-scientist && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-scientist" agent skill from https://github.com/borghei/Claude-Skills/tree/main/data-analytics/data-scientist into .agents/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add borghei/Claude-Skills --skill data-scientist -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install borghei/Claude-Skills data-scientist --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/data-analytics/data-scientist .cursor/skills/data-scientist && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-scientist" agent skill from https://github.com/borghei/Claude-Skills/tree/main/data-analytics/data-scientist into .cursor/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/borghei/Claude-Skills.git --path data-analytics/data-scientist--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add borghei/Claude-Skills --skill data-scientist -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install borghei/Claude-Skills data-scientist --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/data-analytics/data-scientist .gemini/skills/data-scientist && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-scientist" agent skill from https://github.com/borghei/Claude-Skills/tree/main/data-analytics/data-scientist into .gemini/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install borghei/Claude-Skills data-scientistInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add borghei/Claude-Skills --skill data-scientist -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/data-analytics/data-scientist .github/skills/data-scientist && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-scientist" agent skill from https://github.com/borghei/Claude-Skills/tree/main/data-analytics/data-scientist into .github/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add borghei/Claude-Skills --skill data-scientist -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install borghei/Claude-Skills data-scientist --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/data-analytics/data-scientist .opencode/skills/data-scientist && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-scientist" agent skill from https://github.com/borghei/Claude-Skills/tree/main/data-analytics/data-scientist into .opencode/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-scientistData science across machine learning, statistical modeling, and experimentation.
Data Scientist is an agent skill from borghei/Claude-Skills. Data science across machine learning, statistical modeling, and experimentation. Use when selecting ML algorithms, engineering features, designing A/B tests, evaluating model performance, or building predictive pipelines.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/experiment_tracker.py`, `scripts/feature_selector.py` and `scripts/hypothesis_tester.py`).
It sits in Data & Analytics, covering A/B testing and Machine learning. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Scientist loads about 3.3k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 1,002 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 1,002 words, ~3,299 tokens.
.claude/skills/data-scientist/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.The agent operates as a senior data scientist, selecting algorithms, engineering features, designing experiments, evaluating models, and translating predictions into business impact.
Before modeling, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
| Scenario | Recommended | When to upgrade |
|---|---|---|
| Need interpretability | Logistic / Linear Regression | Always start here for stakeholder-facing models |
| Small data (< 10K rows) | Random Forest | Move to XGBoost if accuracy insufficient |
| Medium data, high accuracy needed | XGBoost / LightGBM | Default workhorse for tabular data |
| Large data, complex patterns | Neural Network | Only when tree methods plateau |
| Unsupervised grouping | K-Means / DBSCAN | Use silhouette score to validate k |
Numerical transforms:
import numpy as np, pandas as pd
def engineer_numerical(df: pd.DataFrame, col: str) -> pd.DataFrame:
return pd.DataFrame({
f'{col}_log': np.log1p(df[col]),
f'{col}_sqrt': np.sqrt(df[col].clip(lower=0)),
f'{col}_squared': df[col] ** 2,
f'{col}_binned': pd.cut(df[col], bins=5, labels=False),
})Time-based features with cyclical encoding:
def engineer_time(df: pd.DataFrame, col: str) -> pd.DataFrame:
dt = pd.to_datetime(df[col])
return pd.DataFrame({
f'{col}_hour': dt.dt.hour,
f'{col}_dayofweek': dt.dt.dayofweek,
f'{col}_month': dt.dt.month,
f'{col}_is_weekend': dt.dt.dayofweek.isin([5, 6]).astype(int),
f'{col}_hour_sin': np.sin(2 * np.pi * dt.dt.hour / 24),
f'{col}_hour_cos': np.cos(2 * np.pi * dt.dt.hour / 24),
})Feature selection (importance-based):
from sklearn.ensemble import RandomForestClassifier
def select_top_features(X, y, n=20):
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X, y)
importance = pd.Series(rf.feature_importances_, index=X.columns)
return importance.nlargest(n).index.tolist()Classification:
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
def evaluate_classifier(y_true, y_pred, y_proba=None) -> dict:
m = {
"accuracy": accuracy_score(y_true, y_pred),
"precision": precision_score(y_true, y_pred),
"recall": recall_score(y_true, y_pred),
"f1": f1_score(y_true, y_pred),
}
if y_proba is not None:
m["auc_roc"] = roc_auc_score(y_true, y_proba)
return mRegression:
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np
def evaluate_regressor(y_true, y_pred) -> dict:
return {
"mae": mean_absolute_error(y_true, y_pred),
"rmse": np.sqrt(mean_squared_error(y_true, y_pred)),
"r2": r2_score(y_true, y_pred),
}Sample size calculation:
from scipy import stats
import numpy as np
def required_sample_size(baseline_rate: float, mde: float, alpha: float = 0.05, power: float = 0.8) -> int:
"""Return required N per variant. mde is relative (e.g., 0.10 = 10% lift)."""
effect = baseline_rate * mde
z_a = stats.norm.ppf(1 - alpha / 2)
z_b = stats.norm.ppf(power)
p = baseline_rate
return int(np.ceil(2 * p * (1 - p) * (z_a + z_b) ** 2 / effect ** 2))
# Example: baseline 5% conversion, detect 10% relative lift
# >>> required_sample_size(0.05, 0.10) -> ~62,214 per variantResult analysis:
def analyze_ab(control: np.ndarray, treatment: np.ndarray, alpha: float = 0.05) -> dict:
"""Analyze A/B test with proportions z-test."""
n_c, n_t = len(control), len(treatment)
p_c, p_t = control.mean(), treatment.mean()
p_pool = (control.sum() + treatment.sum()) / (n_c + n_t)
se = np.sqrt(p_pool * (1 - p_pool) * (1/n_c + 1/n_t))
z = (p_t - p_c) / se
p_val = 2 * (1 - stats.norm.cdf(abs(z)))
return {
"control_rate": p_c, "treatment_rate": p_t,
"lift": (p_t - p_c) / p_c,
"p_value": p_val, "significant": p_val < alpha,
"ci_95": ((p_t - p_c) - 1.96 * se, (p_t - p_c) + 1.96 * se),
}# Data Science Project: [Name]
## Business Objective -- What problem are we solving?
## Success Metrics -- Primary: [metric]; Secondary: [metric]
## Data -- Sources, size (rows/features), time period
## Methodology -- Numbered steps
## Results
| Metric | Baseline | Model | Improvement |
|--------|----------|-------|-------------|
## Business Impact -- [Quantified impact]
## Recommendations -- [Next actions]
## Limitations -- [Known caveats]python scripts/experiment_tracker.py log --name "xgb_v2" --params '{"lr":0.1,"depth":6}' --metrics '{"f1":0.87,"auc":0.92}'
python scripts/experiment_tracker.py list --sort-by f1 --top 5
python scripts/experiment_tracker.py compare --ids 1 3 5 --json
python scripts/hypothesis_tester.py ttest --file data.csv --col-a group_a --col-b group_b
python scripts/hypothesis_tester.py proportion --successes-a 120 --trials-a 1000 --successes-b 145 --trials-b 1000
python scripts/hypothesis_tester.py chi-square --file contingency.csv --json
python scripts/feature_selector.py --file dataset.csv --target churn --top 10
python scripts/feature_selector.py --file dataset.csv --target revenue --method correlation --json| Tool | Purpose | Key Flags |
|---|---|---|
experiment_tracker.py | Log, list, and compare experiments with parameters, metrics, and tags in a local JSON file | log --name --params --metrics --tags, list --sort-by --top, compare --ids, --json |
hypothesis_tester.py | Run statistical tests: Welch's t-test, paired t-test, proportion z-test, chi-square independence | ttest --file --col-a --col-b [--paired], proportion --successes-a --trials-a ..., chi-square --file, --json |
feature_selector.py | Rank features by composite score (variance, correlation, mutual information, null rate) for a target column | --file <csv>, --target <col>, --top <n>, --method all/correlation/mutual_info, --json |
| Problem | Likely Cause | Resolution |
|---|---|---|
| Model overfits (large train-test gap in metrics) | Too many features, insufficient regularization, or data leakage | Reduce feature count with feature_selector.py, add regularization, and audit feature engineering for temporal leakage |
| A/B test shows significant result but tiny effect size | Large sample size makes small differences statistically significant | Always report effect size (Cohen's d) alongside p-value; use practical significance thresholds |
hypothesis_tester.py p-value differs from scipy | The tool uses normal/t-distribution approximations (standard library only) | For publication-grade analysis, validate with scipy.stats; the tool is designed for fast directional estimates |
| Feature importance scores are near-zero for all features | Target variable has extremely low variance or the feature set lacks predictive signal | Check target distribution; consider feature engineering or collecting additional data sources |
experiment_tracker.py shows experiment IDs out of order | Experiments were logged non-sequentially or the log file was manually edited | IDs are auto-incremented; use --sort-by on a metric for meaningful ordering |
| Chi-square test fails with "table must be at least 2x2" | CSV contingency table has fewer than 2 rows or 2 columns of numeric data | Ensure the CSV has a header row and at least 2x2 numeric cells; verify the format matches expectations |
| Class imbalance causes misleading accuracy | Accuracy inflated by majority class predictions | Use F1, precision-recall, or AUC-ROC instead; apply SMOTE or class weights during training |
feature_selector.py output is saved with the experiment record.experiment_tracker.py including parameters, metrics, and a descriptive name.In scope: Machine learning algorithm selection, feature engineering, model training and evaluation, A/B test design and analysis, statistical hypothesis testing, experiment tracking, and communicating results to stakeholders.
Out of scope: Model deployment to production (see ml-ops-engineer), data pipeline infrastructure, dashboard development, and real-time serving architecture.
Limitations: The Python tools use only the Python standard library. hypothesis_tester.py uses normal and t-distribution approximations that are accurate for moderate sample sizes but should be validated with scipy for edge cases (very small n, extreme skew). feature_selector.py computes approximate mutual information using binned discretization -- for high-precision feature selection, use sklearn's mutual_info_classif or permutation importance. All tools process local files and do not integrate with MLflow, W&B, or other tracking platforms.
data-analytics/ml-ops-engineer): Trained models are handed off for production deployment, monitoring, and registry management.data-analytics/data-analyst): Complex analytical questions requiring predictive modeling are escalated from the analyst to the data scientist.data-analytics/analytics-engineer): Feature engineering pipelines may depend on mart models as upstream data sources.product-team/): Experiment results inform product decisions; A/B test designs are co-created with product managers.engineering/senior-ml-engineer): Algorithm implementation details and model architecture decisions bridge data science and ML engineering.© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in data-analytics/data-scientist of borghei/Claude-Skills.
Open the folder on GitHubat commit c9a1487
Data Scientist next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Scientist this skillborghei/Claude-Skills | 874 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Senior Data ScientistRaidriar7170/hermes-skilleval | 125 | 6 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Automl SkillLeoYeAI/openclaw-master-skills | 2.2k | — | ~3.6k | Automated safety check: Pass | MIT | |
| Data Sciencemajiayu000/claude-skill-registry | 666 | 1 repos | ~4.3k | Automated safety check: Pass | MIT | |
| ML Experiment Evaluationhashgraph-online/awesome-codex-plugins | 1.2k | — | ~801 | Automated safety check: Pass | MIT | |
| Data Scientisttheneoai/awesome-skills | 183 | — | ~2k | Automated safety check: Pass | MIT |
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
LeoYeAI/openclaw-master-skills
AutoML 自动化机器学习技能 | Automated Machine Learning Skill. An agent skill from LeoYeAI/openclaw-master-skills.
majiayu000/claude-skill-registry
A skill your agent uses when performing exploratory data analysis, statistical testing, data visualization, or building predictive models.
hashgraph-online/awesome-codex-plugins
Plan evaluation strategies for machine-learning product changes.
theneoai/awesome-skills
Elite Data Scientist skill with expertise in statistical analysis, predictive modeling, experimental design (A/B testing), feature engineering, and data visualization.
FerroxLabs/wayland
ML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2), confusion matrix analysis, cross-validation strategies, bias detection…
borghei/Claude-Skills
Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.
borghei/Claude-Skills
Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.
borghei/Claude-Skills
Idea to AI-generated prototype to customer validation to engineering handoff.
borghei/Claude-Skills
Analytics engineering across data modeling, dbt, transformation, and semantic layers.
borghei/Claude-Skills
Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.
borghei/Claude-Skills
OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.
Categories
Data science across machine learning, statistical modeling, and experimentation. Data Scientist is an agent skill from borghei/Claude-Skills. Data science across machine learning, statistical modeling, and experimentation.
Data Scientist fits situations like: selecting ML algorithms; engineering features; designing A/B tests; evaluating model performance.
Run `npx skills add borghei/Claude-Skills --skill data-scientist -a claude-code`. Or copy the skill folder (data-analytics/data-scientist in borghei/Claude-Skills) into .claude/skills/data-scientist in your project. Claude Code loads it when a task matches its description.
Run `npx skills add borghei/Claude-Skills --skill data-scientist -a codex`. Or copy the skill folder (data-analytics/data-scientist in borghei/Claude-Skills) into .agents/skills/data-scientist in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill data-scientist -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-scientist, .gemini/skills/data-scientist, .github/skills/data-scientist and .opencode/skills/data-scientist in your project.
Going by SKILL.md and its folder, Data Scientist needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Data Scientist is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Data Scientist: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Automl Skill (LeoYeAI/openclaw-master-skills, 2.2k stars), Data Science (majiayu000/claude-skill-registry, 666 stars) and ML Experiment Evaluation (hashgraph-online/awesome-codex-plugins, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.
Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.