Agent skill

Data Scientist

by majiayu000 in majiayu000/claude-skill-registry

Expert in statistical analysis, predictive modeling, machine learning, and data storytelling to drive business insights.

MITAuto-check passedData & Analytics

Install Data Scientist

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill data-scientist -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry data-scientist --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-ml/data-scientist-skill .claude/skills/data-scientist && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-scientist
GitHub stars
666
Used in
1 other repo
Token cost
~3.5k tokens
SKILL.md length
1,372 words
Files
2
Skills in repo
971
Repo updated
First seen
Licence
MIT

At a glance

Expert in statistical analysis, predictive modeling, machine learning, and data storytelling to drive business insights.

  • Works in 3 steps: Core Workflows → Anti-Patterns & Gotchas → Quality Checklist
  • Tasks that involve Machine learning
  • SKILL.md covers Purpose, When to Use, Core Capabilities and 3. Core Workflows, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Scientist is an agent skill from majiayu000/claude-skill-registry. Expert in statistical analysis, predictive modeling, machine learning, and data storytelling to drive business insights.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in Data & Analytics, covering Machine learning, Storytelling and Data analysis. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Tasks that involve Machine learning
  • Tasks that involve Storytelling
  • Tasks that involve Data analysis

Example prompts

  • “/data-scientist”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Core Workflows
  2. Anti-Patterns & Gotchas
  3. Quality Checklist

What it can do on your machine

Read from SKILL.md and the folder at commit 000116a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Scientist loads about 3.5k tokens when it runs. Until then it costs about 34 tokens; SKILL.md has 1,372 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~34
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 000116a, republished under its MIT licence (© majiayu000). 1,372 words, ~3,531 tokens.

Download SKILL.mdSave it as .claude/skills/data-scientist/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
data-scientist
description
Expert in statistical analysis, predictive modeling, machine learning, and data storytelling to drive business insights.

Data Scientist

Purpose

Provides statistical analysis and predictive modeling expertise specializing in machine learning, experimental design, and causal inference. Builds rigorous models and translates complex statistical findings into actionable business insights with proper validation and uncertainty quantification.

When to Use

  • Performing exploratory data analysis (EDA) to find patterns and anomalies
  • Building predictive models (classification, regression, forecasting)
  • Designing and analyzing A/B tests or experiments
  • Conducting rigorous statistical hypothesis testing
  • Creating advanced visualizations and data narratives
  • Defining metrics and KPIs for business problems


Core Capabilities

Statistical Modeling
  • Building predictive models using regression, classification, and clustering
  • Implementing time series forecasting and causal inference
  • Designing and analyzing A/B tests and experiments
  • Performing feature engineering and selection
Machine Learning
  • Training and evaluating supervised and unsupervised learning models
  • Implementing deep learning models for complex patterns
  • Performing hyperparameter tuning and model optimization
  • Validating models with cross-validation and holdout sets
Data Exploration
  • Conducting exploratory data analysis (EDA) to discover patterns
  • Identifying anomalies and outliers in datasets
  • Creating advanced visualizations for insight discovery
  • Generating hypotheses from data exploration
Communication and Storytelling
  • Translating statistical findings into business language
  • Creating compelling data narratives for stakeholders
  • Building interactive notebooks and reports
  • Presenting findings with uncertainty quantification


3. Core Workflows

Workflow 1: Exploratory Data Analysis (EDA) & Cleaning

Goal: Understand data distribution, quality, and relationships before modeling.

Steps:

  1. Load and Profile Data

    python
    import pandas as pd
    import numpy as np
    import seaborn as sns
    import matplotlib.pyplot as plt
    
    # Load data
    df = pd.read_csv("customer_data.csv")
    
    # Basic profiling
    print(df.info())
    print(df.describe())
    
    # Missing values analysis
    missing = df.isnull().sum() / len(df)
    print(missing[missing > 0].sort_values(ascending=False))
  2. Univariate Analysis (Distributions)

    python
    # Numerical features
    num_cols = df.select_dtypes(include=[np.number]).columns
    for col in num_cols:
        plt.figure(figsize=(10, 4))
        plt.subplot(1, 2, 1)
        sns.histplot(df[col], kde=True)
        plt.subplot(1, 2, 2)
        sns.boxplot(x=df[col])
        plt.show()
    
    # Categorical features
    cat_cols = df.select_dtypes(exclude=[np.number]).columns
    for col in cat_cols:
        print(df[col].value_counts(normalize=True))
  3. Bivariate Analysis (Relationships)

    python
    # Correlation matrix
    corr = df.corr()
    sns.heatmap(corr, annot=True, cmap='coolwarm')
    
    # Target vs Features
    target = 'churn'
    sns.boxplot(x=target, y='tenure', data=df)
  4. Data Cleaning

    python
    # Impute missing values
    df['age'].fillna(df['age'].median(), inplace=True)
    df['category'].fillna('Unknown', inplace=True)
    
    # Handle outliers (Example: Cap at 99th percentile)
    cap = df['income'].quantile(0.99)
    df['income'] = np.where(df['income'] > cap, cap, df['income'])

Verification:

  • No missing values in critical columns.
  • Distributions understood (normal vs skewed).
  • Target variable balance checked.


Workflow 3: A/B Test Analysis

Goal: Analyze results of a website conversion experiment.

Steps:

  1. Define Hypothesis

    • H0: Conversion Rate B <= Conversion Rate A
    • H1: Conversion Rate B > Conversion Rate A
    • Alpha: 0.05
  2. Load and Aggregate Data

    python
    # data: ['user_id', 'group', 'converted']
    results = df.groupby('group')['converted'].agg(['count', 'sum', 'mean'])
    results.columns = ['n_users', 'conversions', 'conversion_rate']
    print(results)
  3. Statistical Test (Proportions Z-test)

    python
    from statsmodels.stats.proportion import proportions_ztest
    
    control = results.loc['A']
    treatment = results.loc['B']
    
    count = np.array([treatment['conversions'], control['conversions']])
    nobs = np.array([treatment['n_users'], control['n_users']])
    
    stat, p_value = proportions_ztest(count, nobs, alternative='larger')
    
    print(f"Z-statistic: {stat:.4f}")
    print(f"P-value: {p_value:.4f}")
  4. Confidence Intervals

    python
    from statsmodels.stats.proportion import proportion_confint
    
    (lower_con, lower_treat), (upper_con, upper_treat) = proportion_confint(count, nobs, alpha=0.05)
    
    print(f"Control CI: [{lower_con:.4f}, {upper_con:.4f}]")
    print(f"Treatment CI: [{lower_treat:.4f}, {upper_treat:.4f}]")
  5. Conclusion

    • If p-value < 0.05: Reject H0. Variation B is statistically significantly better.
    • Check practical significance (Lift magnitude).


Workflow 5: Causal Inference (Propensity Score Matching)

Goal: Estimate impact of a "Premium Membership" on "Spend" when A/B test isn't possible (observational data).

Steps:

  1. Problem Setup

    • Treatment: Premium Member (1) vs Free (0)
    • Outcome: Annual Spend ($)
    • Confounders: Age, Income, Location, Tenure (Factors affecting both membership and spend)
  2. Calculate Propensity Scores

    python
    from sklearn.linear_model import LogisticRegression
    
    # P(Treatment=1 | Confounders)
    confounders = ['age', 'income', 'tenure']
    logit = LogisticRegression()
    logit.fit(df[confounders], df['is_premium'])
    
    df['propensity_score'] = logit.predict_proba(df[confounders])[:, 1]
    
    # Check overlap (Common Support)
    sns.histplot(data=df, x='propensity_score', hue='is_premium', element='step')
  3. Matching (Nearest Neighbor)

    python
    from sklearn.neighbors import NearestNeighbors
    
    # Separate groups
    treatment = df[df['is_premium'] == 1]
    control = df[df['is_premium'] == 0]
    
    # Find neighbors for treatment group in control group
    nn = NearestNeighbors(n_neighbors=1, algorithm='ball_tree')
    nn.fit(control[['propensity_score']])
    
    distances, indices = nn.kneighbors(treatment[['propensity_score']])
    
    # Create matched dataframe
    matched_control = control.iloc[indices.flatten()]
    
    # Compare outcomes
    ate = treatment['spend'].mean() - matched_control['spend'].mean()
    print(f"Average Treatment Effect (ATE): ${ate:.2f}")
  4. Validation (Balance Check)

    • Check if confounders are balanced after matching (e.g., Mean Age of Treatment vs Matched Control should be similar).
    • abs(mean_diff) / pooled_std < 0.1 (Standardized Mean Difference).


5. Anti-Patterns & Gotchas

❌ Anti-Pattern 1: Data Leakage

What it looks like:

  • Scaling/Standardizing the entire dataset before train/test split.
  • Using future information (e.g., "next_month_churn") as a feature.
  • Including target-derived features (e.g., mean target encoding) calculated on the whole set.

Why it fails:

  • Model performance is artificially inflated during training/validation.
  • Fails completely in production on new, unseen data.

Correct approach:

  • Split FIRST, then transform.
  • Fit scalers/encoders ONLY on X_train, then transform X_test.
  • Use Pipeline objects to ensure safety.
❌ Anti-Pattern 2: P-Hacking (Data Dredging)

What it looks like:

  • Testing 50 different hypotheses or subgroups.
  • Reporting only the one result with p < 0.05.
  • Stopping an A/B test exactly when significance is reached (peeking).

Why it fails:

  • High probability of False Positives (Type I error).
  • Findings are random noise, not reproducible effects.

Correct approach:

  • Pre-register hypotheses.
  • Apply Bonferroni correction or False Discovery Rate (FDR) control for multiple comparisons.
  • Determine sample size before the experiment and stick to it.
❌ Anti-Pattern 3: Ignoring Imbalanced Classes

What it looks like:

  • Training a fraud detection model on data with 0.1% fraud.
  • Reporting 99.9% Accuracy as "Success".

Why it fails:

  • The model simply predicts "No Fraud" for everyone.
  • Fails to detect the actual class of interest.

Correct approach:

  • Use appropriate metrics: Precision-Recall AUC, F1-Score.
  • Resampling techniques: SMOTE (Synthetic Minority Over-sampling Technique), Random Undersampling.
  • Class weights: scale_pos_weight in XGBoost, class_weight='balanced' in Sklearn.


Show full SKILL.md (529 more words)Show less

7. Quality Checklist

Methodology & Rigor:

  • Hypothesis defined clearly before analysis.
  • Assumptions checked (normality, independence, homoscedasticity) for statistical tests.
  • Train/Test/Validation split performed correctly (no leakage).
  • Imbalanced classes handled appropriate (metrics, resampling).
  • Cross-validation used for model assessment.

Code & Reproducibility:

  • Code stored in git with requirements.txt or environment.yml.
  • Random seeds set for reproducibility (random_state=42).
  • Hardcoded paths replaced with relative paths or config variables.
  • Complex logic wrapped in functions/classes with docstrings.

Interpretation & Communication:

  • Results interpreted in business terms (e.g., "Revenue lift" vs "Log-loss decrease").
  • Confidence intervals provided for estimates.
  • "Black box" models explained using SHAP or LIME if needed.
  • Caveats and limitations explicitly stated.

Performance:

  • EDA performed on sampled data if dataset > 10GB.
  • Vectorized operations used (pandas/numpy) instead of loops.
  • Query optimized (filtering early, selecting only needed columns).

Examples

Example 1: A/B Test Analysis for Feature Launch

Scenario: Product team wants to know if a new recommendation algorithm increases user engagement.

Analysis Approach:

  1. Experimental Design: Random assignment (50/50), minimum sample size calculation
  2. Data Collection: Tracked click-through rate, time on page, conversion
  3. Statistical Testing: Two-sample t-test with bootstrapped confidence intervals
  4. Results: Significant improvement in CTR (p < 0.01), 12% lift

Key Analysis:

python
# Bootstrap confidence interval for difference in means
from scipy import stats
diff = treatment_means - control_means
ci = np.percentile(bootstrap_diffs, [2.5, 97.5])

Outcome: Feature launched with 95% probability of positive impact

Example 2: Time Series Forecasting for Demand Planning

Scenario: Retail chain needs to forecast next-quarter sales for inventory planning.

Modeling Approach:

  1. Exploratory Analysis: Identified trends, seasonality (weekly, holiday)
  2. Feature Engineering: Promotions, weather, economic indicators
  3. Model Selection: Compared ARIMA, Prophet, and gradient boosting
  4. Validation: Walk-forward validation on last 12 months

Results:

ModelMAPE90% CI Width
ARIMA12.3%±15%
Prophet9.8%±12%
XGBoost7.2%±9%

Deliverable: Production model with automated retraining pipeline

Example 3: Causal Attribution Analysis

Scenario: Marketing wants to understand which channels drive actual conversions vs. appear correlated.

Causal Methods:

  1. Propensity Score Matching: Match users with similar characteristics
  2. Difference-in-Differences: Compare changes before/after campaigns
  3. Instrumental Variables: Address selection bias in observational data

Key Findings:

  • TV ads: 3.2x ROAS (strongest attribution)
  • Social media: 1.1x ROAS (attribution unclear)
  • Email: 5.8x ROAS (highest efficiency)

Best Practices

Experimental Design
  • Randomization: Ensure true random assignment to treatment/control
  • Sample Size Calculation: Power analysis before starting experiments
  • Multiple Testing: Adjust significance levels when testing multiple hypotheses
  • Control Variables: Include relevant covariates to reduce variance
  • Duration Planning: Run experiments long enough for stable results
Model Development
  • Feature Engineering: Create interpretable, predictive features
  • Cross-Validation: Use time-aware splits for time series data
  • Model Interpretability: Use SHAP/LIME to explain predictions
  • Validation Metrics: Choose metrics aligned with business objectives
  • Overfitting Prevention: Regularization, early stopping, held-out data
Statistical Rigor
  • Uncertainty Quantification: Always report confidence intervals
  • Significance Interpretation: P-value is not effect size
  • Assumption Checking: Validate statistical test assumptions
  • Sensitivity Analysis: Test robustness to modeling choices
  • Pre-registration: Document analysis plan before seeing results
Communication and Impact
  • Business Translation: Convert statistical terms to business impact
  • Actionable Recommendations: Tie findings to specific decisions
  • Visual Storytelling: Create compelling narratives from data
  • Stakeholder Communication: Tailor level of technical detail
  • Documentation: Maintain reproducible analysis records
Ethical Data Science
  • Fairness Considerations: Check for bias across protected groups
  • Privacy Protection: Anonymize sensitive data appropriately
  • Transparency: Document data sources and methodology
  • Responsible AI: Consider societal impact of models
  • Data Quality: Acknowledge limitations and potential biases

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ai-ml/data-scientist-skill of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 000116a

Used in 1 other repository

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Data Scientist next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Scientist compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Scientist this skillmajiayu000/claude-skill-registry6661 repos~3.5kAutomated safety check: PassMIT
Data Sciencetravisjneuman/.claude1011 repos~2.3kAutomated safety check: PassMIT
Data Scientistdavila7/claude-code-templates32k8 repos~2.6kAutomated safety check: PassMIT
Scientific Toolkit SkillzLanqing/codex-claude-academic-skills4.6k—~1.2kAutomated safety check: PassMIT
Power Analysisgaasher/Agent-Loop-Skills174—~2.2kAutomated safety check: PassMIT
Automl SkillLeoYeAI/openclaw-master-skills2.2k—~3.6kAutomated safety check: PassMIT

Similar skills

  • Data Science

    travisjneuman/.claude

    Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy.

    101 GitHub starsUsed in 1 repo~2.3k tokens
    Data & AnalyticsAuto-check passed
  • Data Scientist

    davila7/claude-code-templates

    Expert data scientist for advanced analytics, machine learning, and statistical modeling.

    32k GitHub starsUsed in 8 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Toolkit Skill

    zLanqing/codex-claude-academic-skills

    Research computing toolkit for optoelectronic information science and engineering, MATLAB/Octave, Python scientific analysis, signal processing, image processing, statistics, simulation…

    4.6k GitHub stars~1.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Power Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…

    174 GitHub stars~2.2k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Automl Skill

    LeoYeAI/openclaw-master-skills

    AutoML 自动化机器学习技能 | Automated Machine Learning Skill. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~3.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Fcr Data Analysis

    franklee16/academic-research-skills

    A skill your agent uses when executing and reporting the statistical analysis for a Field Crops Research (FCR) manuscript — mixed models for multi-environment and blocked/split-plot designs…

    223 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed

More from majiayu000/claude-skill-registry

All 971 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed
  • Open Notebook

    majiayu000/claude-skill-registry

    Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis.

    666 GitHub starsUsed in 4 repos~2.4k tokens
    Auto-check passed

Questions about Data Scientist

What does Data Scientist do?

Expert in statistical analysis, predictive modeling, machine learning, and data storytelling to drive business insights. Data Scientist is an agent skill from majiayu000/claude-skill-registry. Expert in statistical analysis, predictive modeling, machine learning, and data storytelling to drive business insights.

When should I use Data Scientist?

Data Scientist fits situations like: tasks that involve Machine learning; tasks that involve Storytelling; tasks that involve Data analysis.

How do I install Data Scientist in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill data-scientist -a claude-code`. Or copy the skill folder (skills/ai-ml/data-scientist-skill in majiayu000/claude-skill-registry) into .claude/skills/data-scientist in your project. Claude Code loads it when a task matches its description.

How do I install Data Scientist in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill data-scientist -a codex`. Or copy the skill folder (skills/ai-ml/data-scientist-skill in majiayu000/claude-skill-registry) into .agents/skills/data-scientist in your project. Codex loads it when a task matches its description.

Can I use Data Scientist in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill data-scientist -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-scientist, .gemini/skills/data-scientist, .github/skills/data-scientist and .opencode/skills/data-scientist in your project.

What does Data Scientist need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Scientist is instructions for the agent only. Our summary lists: Python 3.

Does Data Scientist access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Scientist safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Scientist use?

Data Scientist is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Scientist use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Scientist?

Skills that share tags, products or a category with Data Scientist: Data Science (travisjneuman/.claude, 101 stars), Data Scientist (davila7/claude-code-templates, 32k stars), Scientific Toolkit Skill (zLanqing/codex-claude-academic-skills, 4.6k stars) and Power Analysis (gaasher/Agent-Loop-Skills, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Scientist?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 971 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.