Data Scientist
magnus919/hermes-profiles
PhD-level expertise in data science, statistics, and machine learning.
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
$ npx skills add magnus919/agent-skills --skill data-scientist -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/agent-skills data-scientist --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-scientist .claude/skills/data-scientist && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-scientist" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-scientist into .claude/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/agent-skills/tree/main/data-scientistType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/agent-skills --skill data-scientist -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/agent-skills data-scientist --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/data-scientist .agents/skills/data-scientist && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-scientist" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-scientist into .agents/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill data-scientist -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/agent-skills data-scientist --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/data-scientist .cursor/skills/data-scientist && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-scientist" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-scientist into .cursor/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/agent-skills.git --path data-scientist--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/agent-skills --skill data-scientist -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/agent-skills data-scientist --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/data-scientist .gemini/skills/data-scientist && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-scientist" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-scientist into .gemini/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/agent-skills data-scientistInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/agent-skills --skill data-scientist -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/data-scientist .github/skills/data-scientist && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-scientist" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-scientist into .github/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill data-scientist -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/agent-skills data-scientist --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/data-scientist .opencode/skills/data-scientist && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-scientist" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-scientist into .opencode/skills/data-scientist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scientist", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-scientistA skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
Data Scientist is an agent skill from magnus919/agent-skills. Use for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research methodology, or data science project leadership. Load when the user asks about statistical methods, experimental design, model selection, A/B testing, hypothesis testing, power analysis, regression, causality, Bayesian analysis, or research methodology. For insurance, actuarial, claims, reserving, solvency, credibility, tail-risk, or…
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 34 other files, including scripts, reference files and assets (for example `README.md`, `assets/experimental-plan-template.md` and `assets/report-template.md`). Compatibility notes: Python 3.10+ with scipy, statsmodels, scikit-learn, pandas, numpy. PyTorch and sklearn are the primary ML frameworks. Hardware-aware via detect-compute.py…
It sits in Research & Science, covering Experimental design, Financial modeling and Statistics. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c545c2b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Python 3.10+ with scipy, statsmodels, scikit-learn, pandas, numpy. PyTorch and sklearn are the primary ML frameworks. Hardware-aware via detect-compute.py. Optional R engine via rpy2. Deep learning assumes NVIDIA GPU with CUDA or Apple MPS.
From compatibility in the SKILL.md frontmatter.
Data Scientist loads about 4.1k tokens when it runs, and up to ~43k if it reads all its reference files. Until then it costs about 190 tokens; SKILL.md has 1,500 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from magnus919/agent-skills at commit c545c2b, republished under its MIT licence (© magnus919). 1,500 words, ~4,068 tokens.
.claude/skills/data-scientist/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.This skill owns general statistical and machine-learning methodology. Route to
actuarial-risk-modeling when the primary context is insurance, claims, reserving,
solvency, credibility, risk classification, tail risk, or financial-risk statistical
modeling, because those tasks require domain-specific exposure, development, calibration,
and governance checks. Route to financial-modeling for deterministic operating models,
unit economics, SaaS metrics, pricing scenarios, fundraising, and cash-flow analysis.
Remain here when those contexts are incidental and the core question is general inference,
causal design, experimentation, or model methodology.
actuarial-risk-modeling.financial-modeling.A PhD-level data scientist masters eight competency domains. This skill encodes all of them. When loaded, the agent operates within this scope:
| # | Competency | What It Enables |
|---|---|---|
| 1 | Mathematical & Statistical Foundations | Probability theory, statistical inference, linear algebra, optimization, asymptotic theory — the language in which all methods are expressed |
| 2 | Research Design & Methodology | Formulating testable questions, study design (observational vs experimental), power analysis, bias identification, preregistration |
| 3 | Statistical Modeling & Inference | Parametric and nonparametric methods, regression (linear, GLM, mixed, GAM, nonparametric), Bayesian inference, time series, survival analysis, multivariate methods |
| 4 | Machine Learning & Computational Methods | Supervised/unsupervised/deep/reinforcement learning, learning theory, model selection, regularization, ensembles, transformers, probabilistic ML |
| 5 | Causal Inference & Experimentation | DAGs, potential outcomes, identification strategies (IV, RDD, DID, matching, synthetic control), A/B testing, sensitivity analysis |
| 6 | Reproducibility & MLOps | Version control, environment management, pipeline orchestration, experiment tracking, model deployment, monitoring |
| 7 | Communication & Impact | Scientific writing, visualization, uncertainty communication, stakeholder translation, peer review, grant writing |
| 8 | Research Leadership | Identifying novel research questions, literature synthesis, mentoring, cross-disciplinary collaboration, ethical conduct |
Important: This skill does not make the agent a domain expert in specific application fields (medicine, economics, biology, etc.). It provides the statistical and methodological expertise to collaborate with domain experts.
Before answering any data science question, classify it into one of these types. The classification determines the response structure and rigor required.
User asks a data question.
│
├─ "What model/technique should I use?"
│ → TYPE: ADVICE
│ → Respond with: options + tradeoffs + recommendation + what I'd need to know
│ → Mode: consultative, conditional recommendations
│
├─ "Is this result significant? / Analyze this data."
│ → TYPE: ANALYSIS
│ → Respond with: assumptions check → appropriate test → effect size → uncertainty → interpretation
│ → Mode: rigorous protocol, every step documented
│
├─ "Does X cause Y? / What drives Z?"
│ → TYPE: RESEARCH
│ → Respond with: causal framework → identification strategy → sensitivity → limitations
│ → Mode: causal language, no correlation claims without identification
│
├─ "How should I set up this experiment / study?"
│ → TYPE: DESIGN
│ → Respond with: design taxonomy → power analysis → blocking → randomization → analysis plan
│ → Mode: prescriptive, pre-registration-style
│
├─ "Review this analysis / paper / result."
│ → TYPE: REVIEW
│ → Respond with: methodology check → assumption audit → robustness → reproducibility → summary
│ → Mode: critical, constructive, specific
│
├─ "Compare these methods / Justify an approach."
│ → TYPE: METHODOLOGY
│ → Respond with: criteria → comparison table → recommendation with rationale
│ → Mode: structured, multi-dimensional evaluation
│
├─ "Run a research campaign / I need to find the best approach"
│ → TYPE: CAMPAIGN
│ → Respond with: load references/experimental-campaign-protocol.md
│ → Mode: pipeline orchestration, iterative, multi-experiment
│
├─ Unclear / exploratory
│ → TYPE: CLARIFY
│ → Respond with: ask about data type, question structure, available data, decision context
│ → Mode: investigative| Type | Must Include | Must Not Do |
|---|---|---|
| ADVICE | Tradeoffs, assumptions, when NOT to use | Give single answer without caveats |
| ANALYSIS | Assumption checks, effect sizes, CIs, diagnostics | Stop at p-value |
| RESEARCH | Identification strategy, sensitivity, causal framework | Claim causality from observational data without caveats |
| DESIGN | Power analysis, randomization scheme, sample size justification | Promise significance |
| REVIEW | Specific issues with evidence, reproducibility check | Vague criticism |
| METHODOLOGY | Criteria-based comparison, explicit rationale | Personal preference |
The most important question is never "which test do I use?" but "what am I willing to assume about how these data were generated?" Every statistical method is a set of assumptions expressed as mathematics. Violate the assumptions and the method produces nonsense with high confidence.
Sequence: Data generating process → assumptions → method selection → diagnostics → sensitivity → conclusion
| Use Frequentist When | Use Bayesian When |
|---|---|
| Well-established standard in your field | Prior information exists and should be used explicitly |
| P-values are expected by your audience | You need probabilistic statements about parameters |
| You need a clear decision boundary | Small sample sizes with strong domain knowledge |
| The analysis must be fully specified upfront | Complex hierarchical models |
| Speed / simplicity matters | You want posterior uncertainty quantification |
Never present only p-values. Report effect sizes with confidence intervals (frequentist) or credible intervals (Bayesian) in every case.
Assume your analysis will be audited by someone with your dataset and your code. What would they need to get the same results? If there's a researcher degrees-of-freedom choice (how to handle outliers, which covariates to include, which test to run), document the decision and justify it.
When the user presents an ambiguous data science request, translate it through these steps before touching any method:
Then map to a method using the framework above.
Example:
Assumptions precede methods. Never apply a method without checking whether its assumptions hold for your data. Every reference file in this skill includes assumption-checking guidance.
Effect sizes over p-values. Statistical significance tells you about sample size, not importance. Always report magnitude and precision (CI/CrI).
Causal questions need causal methods. If the question involves "effect of X on Y," you need identification strategy, not just regression. See references/causal-inference-framework.md.
Diagnose before trust. Every fitted model gets assumption diagnostics before interpretation. See scripts/assumption-diagnostics.py.
Uncertainty is not optional. Every estimate comes with uncertainty quantification. If you can't quantify uncertainty, say so and explain why.
Design before data. If you can influence data collection, do power analysis and randomization planning first. See references/experimental-design.md and scripts/power-analysis.py.
Reproducibility is non-negotiable. Code, data, environment, and random seeds must be documented. See assets/experimental-plan-template.md.
The simplest defensible model wins. Favor interpretability until complexity demonstrably improves predictions or inference. Justify complexity with evidence (cross-validation, model comparison, sensitivity analysis).
Know your compute. Before running any experiment, detect available hardware. The model architecture, batch size, and techniques you can use depend on available VRAM, CUDA, and RAM. See scripts/detect-compute.py. See references/docker-experiment-isolation.md for safe execution.
Before recommending or running any experiment, detect your compute environment. Run:
python3 scripts/detect-compute.py --minimalThis returns a JSON object that self-constrains what approaches are feasible:
model_size_tier: "cpu_only" — no deep learning; use sklearn/xgboost/lightgbmmodel_size_tier: "7B-13B" — full fine-tuning or LoRA feasible on available VRAMmodel_size_tier: "up_to_3B" — QLoRA recommended, full FT for tiny models onlyThe agent should detect compute before selecting methods, not after failing. Integrate this check at the start of any CAMPAIGN task or before Phase 4 (Moonshot Experiments) in the campaign protocol.
This skill ships with supporting reference files and scripts:
references/statistical-methodology.md — test selection decision tree, assumptions, diagnosticsreferences/experimental-design.md — design taxonomy, power analysis, A/B testingreferences/causal-inference-framework.md — DAGs, potential outcomes, identification strategiesreferences/regression-modeling.md — model hierarchy, assumption checks, interpretationreferences/bayesian-workflow.md — prior elicitation, MCMC diagnostics, model comparisonreferences/interpretability-workflow.md — explanation target, method selection, stability, slices, and causal limitsreferences/interpretability-sources.md — primary papers and reporting guidancetemplates/interpretability-report.md — versioned explanation and limitation recordscripts/power-analysis.py — compute sample size or minimum detectable effectscripts/assumption-diagnostics.py — run diagnostics on fitted modelsscripts/model-comparison.py — compare models with AIC, BIC, CV, WAICscripts/effect-size-calculator.py — compute effect sizes with confidence intervalsscripts/experimental-design.py — generate experimental designsscripts/detect-compute.py — probe hardware and constrain recommendations (Phase 1)references/experimental-campaign-protocol.md — multi-experiment campaign workflow (Phase 2)references/pytorch-integration.md — training loops, device management, transfer learning, distillationreferences/sklearn-integration.md — pipelines, model selection, preprocessing, ensemblesreferences/data-science-coding-workflow.md — project structure, experiment logging, reproducibilityreferences/subagent-experiment-supervision.md — self-healing experiment pattern with auto-repairreferences/docker-experiment-isolation.md — safe containerized execution with resource limitsLoad this skill when the user's request contains signals from any of these categories:
Statistical methods: hypothesis test, t-test, chi-square, ANOVA, regression, p-value, confidence interval, Bayesian, prior, posterior, MCMC, bootstrap, permutation
Research design: experiment, A/B test, clinical trial, observational study, cohort, case-control, randomization, confounding, bias, power analysis, sample size
Causal: causality, causal inference, effect of, impact, treatment effect, DAG, directed acyclic graph, instrumental variable, DID, difference-in-differences, RDD, regression discontinuity
Modeling: machine learning, predict, classification, clustering, feature selection, overfitting, cross-validation, regularization, ensemble, gradient boosting, neural network, deep learning
Interpretability and fairness: explainability, interpretability, feature attribution, SHAP, LIME, saliency, counterfactual explanation, model card, fairness slice, subgroup performance, bias diagnosis
General: data analysis, statistical analysis, analyze this data, methodology, what model should I use, review my analysis
© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 31 other files (scripts, references, assets) in data-scientist of magnus919/agent-skills.
Open the folder on GitHubat commit c545c2b
Data Scientist next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Scientist this skillmagnus919/agent-skills | 116 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Data Scientistmagnus919/hermes-profiles | 282 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Statistical Analystalirezarezvani/claude-skills | 28k | — | ~2.5k | Automated safety check: Pass | MIT | |
| Power Analysisgaasher/Agent-Loop-Skills | 174 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Experimentation Analyticsrampstackco/claude-skills | 941 | — | ~8.9k | Automated safety check: Pass | MIT | |
| Review Experiment Resultsharness/harness-skills | 115 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 |
magnus919/hermes-profiles
PhD-level expertise in data science, statistics, and machine learning.
alirezarezvani/claude-skills
Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.
gaasher/Agent-Loop-Skills
A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…
rampstackco/claude-skills
How to read experiment results without fooling yourself. An agent skill from rampstackco/claude-skills.
harness/harness-skills
Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data).
K-Dense-AI/scientific-agent-skills
Calculates sample sizes and statistical power for study planning.
magnus919/agent-skills
Organize durable agent research outputs as summaries, analysis, and evidence dossiers.
magnus919/agent-skills
Build portable, first-person colored ASCII city engines and small GIS-derived city packs.
magnus919/agent-skills
Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.
magnus919/agent-skills
Use Docker Compose to define, run, debug, and harden multi-container applications.
magnus919/agent-skills
Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.
magnus919/agent-skills
Build multi-agent AI systems with LangGraph — the low-level orchestration framework for stateful, graph-based agent workflows.
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…. Data Scientist is an agent skill from magnus919/agent-skills. Use for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research methodology, or data science project leadership.
Data Scientist fits situations like: phD-level expertise in data science; machine learning: rigorous statistical analysis; experimental design; causal inference.
Run `npx skills add magnus919/agent-skills --skill data-scientist -a claude-code`. Or copy the skill folder (data-scientist in magnus919/agent-skills) into .claude/skills/data-scientist in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/agent-skills --skill data-scientist -a codex`. Or copy the skill folder (data-scientist in magnus919/agent-skills) into .agents/skills/data-scientist in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill data-scientist -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-scientist, .gemini/skills/data-scientist, .github/skills/data-scientist and .opencode/skills/data-scientist in your project.
Going by SKILL.md and its folder, Data Scientist needs the command-line tools its instructions call (python3). Our summary lists: Python 3; Docker. Compatibility (from SKILL.md): Python 3.10+ with scipy, statsmodels, scikit-learn, pandas, numpy. PyTorch and sklearn are the primary ML frameworks. Hardware-aware via detect-compute.py. Optional R engine via rpy2. Deep learning assumes NVIDIA GPU with CUDA or Apple MPS..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Data Scientist is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 39k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Scientist: Data Scientist (magnus919/hermes-profiles, 282 stars), Statistical Analyst (alirezarezvani/claude-skills, 28k stars), Power Analysis (gaasher/Agent-Loop-Skills, 174 stars) and Experimentation Analytics (rampstackco/claude-skills, 941 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 116 GitHub stars. The repository holds 130 skills in this directory. The repository was last updated on October 8, 2026.
Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.