Agent skill

Senior Data Scientist

by borghei in borghei/Claude-Skills

A skill your agent uses when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model…

MITAuto-check passedData & Analytics

Install Senior Data Scientist

skills CLI
$ npx skills add borghei/Claude-Skills --skill senior-data-scientist -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills senior-data-scientist --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/senior-data-scientist .claude/skills/senior-data-scientist && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
senior-data-scientist
GitHub stars
881
Token cost
~1.7k tokens
SKILL.md length
619 words
Files
9 (incl. scripts, references)
Skills in repo
349
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model…

  • The user asks to design an experiment
  • SKILL.md covers Keywords, Core Capabilities, When to Use and Clarify First, plus 4 more sections
  • Runs Python scripts from its folder
  • Build a predictive model

What it does

Senior Data Scientist is an agent skill from borghei/Claude-Skills. Use when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model performance", "set up MLOps pipeline", "analyze time series", "calculate sample size", or "deploy a model to production". Expert data science covering statistical modeling, experimentation, causal inference, feature engineering, ML deployment, and advanced analytics with Python, R, and SQL.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `references/ds-operations.md`, `references/ds-workflows.md` and `references/experiment_design_frameworks.md`).

It sits in Data & Analytics, covering A/B testing, Machine learning and Experimental design. It works with Python and SQL. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • The user asks to design an experiment
  • Build a predictive model
  • Run A/B test analysis
  • Perform causal inference

Example prompts

  • “design an experiment”
  • “build a predictive model”
  • “run A/B test analysis”
  • “/senior-data-scientist”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Senior Data Scientist loads about 1.7k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 619 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 619 words, ~1,718 tokens.

Download SKILL.mdSave it as .claude/skills/senior-data-scientist/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
senior-data-scientist
description
Use when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model performance", "set up MLOps pipeline", "analyze time series", "calculate sample size", or "deploy a model to production". Expert data science covering statistical modeling, experimentation, causal inference, feature engineering, ML deployment, and advanced analytics with Python, R, and SQL.
license
MIT + Commons Clause
metadata.version
1.1.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
data-science
metadata.updated
2026-06-17
metadata.tags
data-science, ml, statistics, experimentation, python, mlops

Senior Data Scientist

Expert data science for statistical modeling, experimentation, ML deployment, and data-driven decision making — A/B test design and analysis, feature engineering, model training/evaluation, production deployment, and causal inference.

Keywords

data-science, machine-learning, statistics, a-b-testing, causal-inference, feature-engineering, mlops, experiment-design, model-deployment, python, scikit-learn, pytorch, tensorflow, spark, airflow

Core Capabilities

  • Experiment design & analysis — hypothesis framing, power analysis and sample sizing, randomization, SRM monitoring, and post-hoc significance testing.
  • Feature engineering — profiling, candidate generation (temporal/aggregation/interaction/text), selection (variance, correlation, SHAP/RFE), and leakage validation.
  • Model training & evaluation — stratified/temporal splits, baselines, hyperparameter tuning, cross-validation, calibration, and fairness checks.
  • Production deployment — containerized serving, input/output drift monitoring (KS/PSI), canary rollouts, and latency/error SLAs.
  • Causal inference — propensity score matching, difference-in-differences, regression discontinuity, instrumental variables, and assumption/placebo testing.

When to Use

  • Designing or analyzing an A/B test.
  • Building a feature engineering pipeline.
  • Training, evaluating, or deploying an ML model.
  • Estimating treatment effects from observational data.

Clarify First

Before running an analysis or pipeline, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Task — A/B test design / feature engineering / model evaluation / causal inference (selects the script and workflow)
  • Dataset & target variable — what you are modeling or measuring (drives feature generation and leakage validation)
  • Decision metric & minimum effect — the metric and the smallest effect worth detecting (drives power analysis and sample size)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Tools

ScriptPurpose
scripts/experiment_designer.pyA/B test design, power analysis, sample size calculation
scripts/feature_engineering_pipeline.pyAutomated feature generation, correlation analysis, feature selection
scripts/model_evaluation_suite.pyModel comparison, cross-validation, deployment readiness checks

For hypothesis testing, causal inference, and regression analysis, use inline scipy/statsmodels — see references/statistical_methods_advanced.md.

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

  • references/ds-workflows.md — quick-start commands, tech stack, the five end-to-end workflows (A/B testing, feature pipeline, train/evaluate, deploy, causal inference) with Python snippets, performance targets, and common commands. Read when executing any data-science task.
  • references/ds-operations.md — troubleshooting table, success criteria, and the full CLI flag reference for each script. Read when diagnosing issues or running the tools.
  • references/statistical_methods_advanced.md — advanced statistical methods reference (hypothesis testing, causal inference, regression). Read for statistical depth.
  • references/experiment_design_frameworks.md — experiment design frameworks and power-analysis foundations. Read when designing rigorous experiments.
  • references/feature_engineering_patterns.md — feature engineering patterns and selection techniques. Read when building features.
Show full SKILL.md (230 more words)Show less

Scope & Limitations

This skill covers:

  • End-to-end experiment design including power analysis, randomization, and post-hoc analysis
  • Feature engineering pipelines with profiling, generation, selection, and validation
  • Model training evaluation including cross-validation, calibration, and fairness checks
  • Production model deployment with monitoring, drift detection, and canary rollouts

This skill does NOT cover:

  • Data engineering infrastructure (ETL orchestration, pipeline scheduling, data lake management) -- see senior-data-engineer
  • Deep learning model architecture design and training at scale (distributed GPU training, custom layers) -- see senior-ml-engineer
  • Prompt engineering, RAG systems, and LLM fine-tuning workflows -- see senior-prompt-engineer
  • Computer vision pipelines (object detection, segmentation, video processing) -- see senior-computer-vision

Integration Points

SkillIntegrationData Flow
senior-data-engineerFeature pipeline ingests data from ETL outputs; shares data quality validation patternsRaw data stores --> feature engineering pipeline --> feature store
senior-ml-engineerTrained models handed off for MLOps deployment; shares model registry and serving configsEvaluated model artifacts --> deployment pipeline --> production serving
senior-prompt-engineerEmbedding features from LLMs feed into ML pipelines; experiment frameworks apply to prompt A/B testsLLM embeddings --> feature vectors; experiment designs --> prompt evaluation
senior-architectModel serving architecture reviewed for scalability; data platform design aligned with training infrastructureArchitecture specs --> deployment topology --> monitoring dashboards
senior-backendModel inference endpoints integrated into backend services; API contracts defined for prediction requestsREST/gRPC model API --> backend service layer --> client applications
senior-devopsCI/CD pipelines extended for model retraining triggers; containerized model images deployed via infrastructure-as-codeDocker images --> Kubernetes manifests --> production clusters

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in engineering/senior-data-scientist of borghei/Claude-Skills.

  • SKILL.md
  • references/ds-operations.md
  • references/ds-workflows.md
  • references/experiment_design_frameworks.md
  • references/feature_engineering_patterns.md
  • references/statistical_methods_advanced.md
  • scripts/experiment_designer.py
  • scripts/feature_engineering_pipeline.py
  • scripts/model_evaluation_suite.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Senior Data Scientist next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Senior Data Scientist compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Senior Data Scientist this skillborghei/Claude-Skills881—~1.7kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Senior Data Scientistalirezarezvani/claude-skills28k2 repos~2.3kAutomated safety check: PassMIT
Mle Workflowaffaan-m/ECC275k1 repos~5.6kAutomated safety check: PassMIT
Power Analysisgaasher/Agent-Loop-Skills174—~2.2kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries169—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    alirezarezvani/claude-skills

    World-class senior data scientist skill specialising in statistical modeling, experiment design, causal inference, and predictive analytics.

    28k GitHub starsUsed in 2 repos~2.3k tokens
    Data & AnalyticsAuto-check passed
  • Mle Workflow

    affaan-m/ECC

    Production machine-learning engineering workflow for data contracts, reproducible training, model evaluation, deployment, monitoring, and rollback.

    275k GitHub starsUsed in 1 repo~5.6k tokens
    Data & AnalyticsAuto-check passed
  • Power Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…

    174 GitHub stars~2.2k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    169 GitHub stars~3.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Aeon Time Series Machine Learning

    davila7/claude-code-templates

    Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.

    32k GitHub starsUsed in 14 repos~2.6k tokens
    Data & AnalyticsAuto-check passed

More from borghei/Claude-Skills

All 349 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    881 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    881 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    881 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Senior Data Scientist

What does Senior Data Scientist do?

A skill your agent uses when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model…. Senior Data Scientist is an agent skill from borghei/Claude-Skills. Use when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model performance", "set up MLOps pipeline", "analyze time series", "calculate sample size", or "deploy a model to production".

When should I use Senior Data Scientist?

Senior Data Scientist fits situations like: the user asks to design an experiment; build a predictive model; run A/B test analysis; perform causal inference.

How do I install Senior Data Scientist in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill senior-data-scientist -a claude-code`. Or copy the skill folder (engineering/senior-data-scientist in borghei/Claude-Skills) into .claude/skills/senior-data-scientist in your project. Claude Code loads it when a task matches its description.

How do I install Senior Data Scientist in Codex?

Run `npx skills add borghei/Claude-Skills --skill senior-data-scientist -a codex`. Or copy the skill folder (engineering/senior-data-scientist in borghei/Claude-Skills) into .agents/skills/senior-data-scientist in your project. Codex loads it when a task matches its description.

Can I use Senior Data Scientist in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill senior-data-scientist -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/senior-data-scientist, .gemini/skills/senior-data-scientist, .github/skills/senior-data-scientist and .opencode/skills/senior-data-scientist in your project.

What does Senior Data Scientist need to run?

Going by SKILL.md and its folder, Senior Data Scientist needs Python for the scripts in its folder. Our summary lists: Python 3; Docker.

Does Senior Data Scientist access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Senior Data Scientist safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Senior Data Scientist use?

Senior Data Scientist is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Senior Data Scientist use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.3k tokens, read only when the agent opens those files.

What are the alternatives to Senior Data Scientist?

Skills that share tags, products or a category with Senior Data Scientist: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Senior Data Scientist (alirezarezvani/claude-skills, 28k stars), Mle Workflow (affaan-m/ECC, 275k stars) and Power Analysis (gaasher/Agent-Loop-Skills, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Senior Data Scientist?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 881 GitHub stars. The repository holds 349 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.