Agent skill

Data Science

by travisjneuman in travisjneuman/.claude

Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy.

MITAuto-check passedData & Analytics

Install Data Science

skills CLI
$ npx skills add travisjneuman/.claude --skill data-science -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install travisjneuman/.claude data-science --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/travisjneuman/.claude.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-science .claude/skills/data-science && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-science
GitHub stars
101
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
382 words
Files
4 (incl. references)
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy.

  • Building ML models
  • SKILL.md covers Data Strategy, Statistical Analysis, Machine Learning and Data Governance, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Creating dashboards

What it does

Data Science is an agent skill from travisjneuman/.claude. Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy. Use when building ML models, analyzing data, creating dashboards, or designing data architectures.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/data-science-expert.md`, `references/ml-pipelines.md` and `references/statistical-methods.md`).

It sits in Data & Analytics, covering Machine learning, Statistics and Data governance. The repository describes itself as: The Ultimate Claude Code Toolkit: 180 skills, 10 agents, 29 commands, 7 hooks, and 81 marketplace repos (11,000+ community skills). Drop-in ~/.claude config with a generated… The licence is MIT.

When your agent uses it

  • Building ML models
  • Creating dashboards
  • Designing data architectures

Example prompts

  • “/data-science”

What it can do on your machine

Read from SKILL.md and the folder at commit 139c07b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Science loads about 2.3k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 382 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from travisjneuman/.claude at commit 139c07b, republished under its MIT licence (© travisjneuman). 382 words, ~2,267 tokens.

Download SKILL.mdSave it as .claude/skills/data-science/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
data-science
description
Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy. Use when building ML models, analyzing data, creating dashboards, or designing data architectures.

Data Science Expert

Comprehensive data science frameworks for analytics, machine learning, and data-driven decision making.

Data Strategy

Data Maturity Model
LevelNameCharacteristics
1Ad HocManual, inconsistent, siloed
2OpportunisticSome automation, point solutions
3SystematicDefined processes, governance emerging
4DifferentiatingData-driven decisions, advanced analytics
5TransformativeAI-first, competitive advantage
Analytics Value Chain
DATA → INFORMATION → INSIGHT → ACTION → VALUE

PROGRESSION:
Descriptive: What happened?
Diagnostic: Why did it happen?
Predictive: What will happen?
Prescriptive: What should we do?
Autonomous: Self-optimizing systems

Statistical Analysis

Descriptive Statistics
CENTRAL TENDENCY:
- Mean: Sum / Count (sensitive to outliers)
- Median: Middle value (robust to outliers)
- Mode: Most frequent value

DISPERSION:
- Range: Max - Min
- Variance: Average squared deviation
- Standard Deviation: √Variance
- IQR: Q3 - Q1 (robust)

DISTRIBUTION SHAPE:
- Skewness: Asymmetry (0 = symmetric)
- Kurtosis: Tail heaviness (3 = normal)

For detailed inferential statistics and hypothesis testing, see Statistical Methods Reference.

Machine Learning

Algorithm Selection
TaskAlgorithmsWhen to Use
ClassificationLogistic Regression, Random Forest, XGBoost, Neural NetworksCategorical outcomes
RegressionLinear Regression, Ridge/Lasso, Random Forest, XGBoostContinuous outcomes
ClusteringK-Means, Hierarchical, DBSCANGroup discovery
Dimensionality ReductionPCA, t-SNE, UMAPFeature reduction, visualization
Anomaly DetectionIsolation Forest, One-Class SVM, AutoencodersOutlier detection
Time SeriesARIMA, Prophet, LSTMSequential data
RecommendationCollaborative Filtering, Content-Based, Matrix FactorizationPersonalization
NLPTransformers, BERT, GPTText understanding/generation

For detailed ML pipelines, feature engineering, and model monitoring, see ML Pipelines Reference.

Data Governance

Data Governance Framework
GOVERNANCE PILLARS:

POLICIES:
- Data ownership
- Data classification
- Data retention
- Data access
- Data quality standards

ROLES:
- Data Owner: Accountable for data domain
- Data Steward: Day-to-day quality management
- Data Custodian: Technical implementation
- Data Consumer: End user

PROCESSES:
- Data cataloging
- Metadata management
- Data lineage
- Issue resolution
- Change management

METRICS:
- Data quality scores
- Policy compliance
- Data access requests
- Issue resolution time
Data Quality Dimensions
DimensionDefinitionMeasurement
AccuracyCorrect representation of reality% records matching source
CompletenessAll required data present% non-null values
ConsistencySame across systems% matching across sources
TimelinessAvailable when neededLatency, freshness
ValidityConforms to format/rules% passing validation
UniquenessNo unwanted duplicatesDuplicate rate

Business Intelligence

BI Architecture
ARCHITECTURE LAYERS:

DATA SOURCES:
- Operational systems
- External data
- IoT/streaming

DATA INTEGRATION:
- ETL/ELT pipelines
- Data lakes
- Data warehouses

SEMANTIC LAYER:
- Business definitions
- Calculated metrics
- Hierarchies
- Relationships

PRESENTATION:
- Dashboards
- Reports
- Ad-hoc analysis
- Embedded analytics
Dashboard Design Principles
DESIGN PRINCIPLES:

PURPOSE:
- One clear objective per dashboard
- Know your audience
- Enable decisions

LAYOUT:
- Most important top-left
- Related items grouped
- Progressive disclosure
- Whitespace for clarity

VISUALS:
- Right chart for data type
- Consistent formatting
- Minimal decoration
- Color with purpose

INTERACTIVITY:
- Filters for exploration
- Drill-down capability
- Cross-filtering
- Tooltip details
Metric Design
METRIC DEFINITION TEMPLATE:

NAME: [Metric name]
DEFINITION: [Clear business definition]
FORMULA: [Precise calculation]
OWNER: [Responsible person]
DATA SOURCE: [Where it comes from]
GRAIN: [Level of detail]
FREQUENCY: [Update cadence]
DIMENSIONS: [Slicing attributes]
TARGETS: [Goals/benchmarks]
RELATED: [Related metrics]

Predictive Modeling

Show full SKILL.md (166 more words)Show less
Use Case Framework
Use CaseBusiness ApplicationApproach
Churn PredictionRetention programsClassification
Demand ForecastingInventory planningTime series
Lead ScoringSales prioritizationClassification
Price OptimizationRevenue managementRegression/RL
Fraud DetectionRisk mitigationAnomaly detection
RecommendationPersonalizationCollaborative filtering
Customer SegmentationMarketing targetingClustering
Lifetime ValueCustomer investmentRegression

Data Ethics & Privacy

Ethical AI Framework
PRINCIPLES:

FAIRNESS:
- No discriminatory outcomes
- Bias testing across groups
- Regular auditing

ACCOUNTABILITY:
- Clear ownership
- Decision audit trails
- Escalation process

TRANSPARENCY:
- Explainable decisions
- Clear documentation
- User communication

PRIVACY:
- Data minimization
- Consent management
- Security controls
Bias Detection
BIAS TYPES:

HISTORICAL: Reflects past discrimination
REPRESENTATION: Training data not representative
MEASUREMENT: Proxy variables correlate with protected attributes
AGGREGATION: Single model for diverse populations
EVALUATION: Inappropriate benchmarks

FAIRNESS METRICS:
- Demographic Parity: Equal positive rates
- Equalized Odds: Equal TPR and FPR
- Individual Fairness: Similar inputs, similar outputs
- Calibration: Equal accuracy across groups

Analytics Team Structure

Team Roles
RoleFocusSkills
Data EngineerPipelines, infrastructureSQL, Python, Spark, Cloud
Data AnalystReporting, ad-hoc analysisSQL, BI tools, Statistics
Data ScientistModeling, MLPython/R, ML, Statistics
ML EngineerModel deploymentMLOps, Software Engineering
Analytics EngineerData modelingdbt, SQL, Data Modeling
Operating Models
ModelDescriptionBest For
CentralizedSingle analytics teamConsistency, efficiency
DecentralizedEmbedded in business unitsBusiness alignment
Hub & SpokeCentral CoE + embeddedBalance of both
FederatedShared platform, domain teamsScale with autonomy

References

See Also

© travisjneuman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/data-science of travisjneuman/.claude.

  • SKILL.md
  • references/data-science-expert.md
  • references/ml-pipelines.md
  • references/statistical-methods.md

Open the folder on GitHubat commit 139c07b

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in travisjneuman/.claude, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Data Science next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Science compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Science this skilltravisjneuman/.claude1011 repos~2.3kAutomated safety check: PassMIT
Data Scientistdavila7/claude-code-templates32k9 repos~2.6kAutomated safety check: PassMIT
Scientific Toolkit SkillzLanqing/codex-claude-academic-skills4.6k—~1.2kAutomated safety check: PassMIT
Data Scientistmajiayu000/claude-skill-registry6661 repos~3.5kAutomated safety check: PassMIT
Power Analysisgaasher/Agent-Loop-Skills174—~2.2kAutomated safety check: PassMIT
Automl SkillLeoYeAI/openclaw-master-skills2.2k—~3.6kAutomated safety check: PassMIT

Similar skills

  • Data Scientist

    davila7/claude-code-templates

    Expert data scientist for advanced analytics, machine learning, and statistical modeling.

    32k GitHub starsUsed in 9 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Toolkit Skill

    zLanqing/codex-claude-academic-skills

    Research computing toolkit for optoelectronic information science and engineering, MATLAB/Octave, Python scientific analysis, signal processing, image processing, statistics, simulation…

    4.6k GitHub stars~1.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Data Scientist

    majiayu000/claude-skill-registry

    Expert in statistical analysis, predictive modeling, machine learning, and data storytelling to drive business insights.

    666 GitHub starsUsed in 1 repo~3.5k tokens
    Data & AnalyticsAuto-check passed
  • Power Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…

    174 GitHub stars~2.2k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Automl Skill

    LeoYeAI/openclaw-master-skills

    AutoML 自动化机器学习技能 | Automated Machine Learning Skill. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~3.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Fcr Data Analysis

    franklee16/academic-research-skills

    A skill your agent uses when executing and reporting the statistical analysis for a Field Crops Research (FCR) manuscript — mixed models for multi-environment and blocked/split-plot designs…

    223 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed

More from travisjneuman/.claude

All 21 skills in this repo
  • Test Specialist

    travisjneuman/.claude

    Test-writing patterns for JS/TS, Python, Go, and Rust (unit, integration, E2E, visual regression).

    101 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Codebase Documenter

    travisjneuman/.claude

    Write codebase documentation: READMEs, architecture docs, getting-started guides, API docs, and code comments.

    101 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Business Strategy

    travisjneuman/.claude

    Business strategy expertise for strategic planning, competitive analysis, market entry, M&A strategy, portfolio management, and strategic decision-making.

    101 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Finance

    travisjneuman/.claude

    Financial analysis expertise for financial modeling (DCF, LBO, M&A), valuation, financial statement analysis, capital allocation, treasury management, and corporate finance decisions.

    101 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Health Wellness

    travisjneuman/.claude

    Workplace health and wellness expertise for employee wellness programs, mental health initiatives, ergonomics and safety, healthcare benefits strategy, and health analytics.

    101 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Hr Talent

    travisjneuman/.claude

    HR and talent management expertise for talent acquisition, performance management, compensation strategy, organizational design, culture building, succession planning, and D&I programs.

    101 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed

Questions about Data Science

What does Data Science do?

Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy. claude. Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy.

When should I use Data Science?

Data Science fits situations like: building ML models; creating dashboards; designing data architectures.

How do I install Data Science in Claude Code?

Run `npx skills add travisjneuman/.claude --skill data-science -a claude-code`. Or copy the skill folder (skills/data-science in travisjneuman/.claude) into .claude/skills/data-science in your project. Claude Code loads it when a task matches its description.

How do I install Data Science in Codex?

Run `npx skills add travisjneuman/.claude --skill data-science -a codex`. Or copy the skill folder (skills/data-science in travisjneuman/.claude) into .agents/skills/data-science in your project. Codex loads it when a task matches its description.

Can I use Data Science in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add travisjneuman/.claude --skill data-science -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-science, .gemini/skills/data-science, .github/skills/data-science and .opencode/skills/data-science in your project.

What does Data Science need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Science is instructions for the agent only.

Does Data Science access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Science safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Science use?

Data Science is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Science use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Data Science?

Skills that share tags, products or a category with Data Science: Data Scientist (davila7/claude-code-templates, 32k stars), Scientific Toolkit Skill (zLanqing/codex-claude-academic-skills, 4.6k stars), Data Scientist (majiayu000/claude-skill-registry, 666 stars) and Power Analysis (gaasher/Agent-Loop-Skills, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Science?

travisjneuman (a GitHub user) maintains it in travisjneuman/.claude, which has 101 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 6, 2026.

Source: travisjneuman/.claude on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.