Agent skill

Statistical Data Analysis

by lingzhi227 in lingzhi227/agent-research-skills

Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

No licenceAuto-check passedData & Analytics

Install Statistical Data Analysis

skills CLI
$ npx skills add lingzhi227/agent-research-skills --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lingzhi227/agent-research-skills data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lingzhi227/agent-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
384
Token cost
~886 tokens
SKILL.md length
299 words
Files
4 (incl. scripts, references)
Skills in repo
31
Repo updated
First seen
Licence
None found

At a glance

Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

  • Works in 3 steps: Generate Analysis Code → 4-Round Code Review → Produce Results
  • Analyzing experiment results for a research paper
  • SKILL.md covers Input, References, Scripts and Workflow, plus 4 more sections
  • Runs Python scripts from its folder; calls python

What it does

You give it a data source (CSV, JSON, pickle or experiment logs) and a research goal or hypothesis. The agent writes analysis code in a fixed section order, from imports and data loading through dataset preparation, descriptive statistics, preprocessing, the statistical tests and extra results, using pandas, numpy, scipy, statsmodels and sklearn.

That code goes through four review rounds: code flaws, data handling, per-table checks and cross-table consistency. Two bundled Python scripts help. stat_summary.py detects data types, recommends and runs comparisons and reports effect sizes with significance stars, and needs numpy and scipy. format_pvalue.py formats p-values as stars, LaTeX or plain text using only the standard library. A table maps data types to tests such as the t-test, Mann-Whitney U, Wilcoxon, ANOVA and Kruskal-Wallis.

Reported results must carry an uncertainty measure such as a confidence interval, standard deviation or p-value, the chosen tests have to suit the data type, and every number has to come from the actual data rather than being made up.

When your agent uses it

  • Analyzing experiment results for a research paper
  • Choosing the right statistical test for two or more groups
  • Formatting p-values as stars or LaTeX for a results table
  • Reviewing analysis code for statistical or data-handling mistakes

Example prompts

  • “Compare accuracy across the methods in results.csv and report effect sizes with significance.”
  • “Write the analysis script for survey.json to test whether the treatment group scored higher.”
  • “Format the p-values 0.001, 0.05 and 0.23 as LaTeX.”

Requirements

  • Python with numpy and scipy for stat_summary.py
  • pandas, statsmodels and scikit-learn for the generated analysis code

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Generate Analysis Code
  2. 4-Round Code Review
  3. Produce Results

What it can do on your machine

Read from SKILL.md and the folder at commit 9e6c085. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Statistical Data Analysis loads about 886 tokens when it runs, and up to ~1.9k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 299 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~886
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 299 words (~886 tokens).

“Generate rigorous statistical analysis code with multi-round review.”

— opening of SKILL.md by lingzhi227
name
data-analysis
argument-hint
data-source

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (scripts, references) in skills/data-analysis of lingzhi227/agent-research-skills.

  • SKILL.md
  • references/review-prompts.md
  • scripts/format_pvalue.py
  • scripts/stat_summary.py

Open the folder on GitHubat commit 9e6c085

Compare with similar skills

Statistical Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Statistical Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Statistical Data Analysis this skilllingzhi227/agent-research-skills384—~886Automated safety check: PassNone
Q-EDA Exploratory AnalysisTyrealQ/q-skills108—~1.1kAutomated safety check: PassMIT
Tooluniverse Epigenomicswu-yc/LabClaw1.1k2 repos~14kAutomated safety check: PassNone
PyMC Bayesian Modelingdavila7/claude-code-templates32k12 repos~3.9kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.

    108 GitHub stars~1.1k tokensUpdated 14 days ago
    Data & AnalyticsAuto-check passed
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • PyMC Bayesian Modeling

    davila7/claude-code-templates

    Builds, fits, checks and compares Bayesian models in PyMC, from priors and NUTS sampling to variational inference, LOO and WAIC comparison, and diagnostics.

    32k GitHub starsUsed in 12 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Trains scikit-learn models with walk-forward validation on features from OHLCV data to predict return direction and turn the predictions into trading signals.

    35k GitHub stars~3.2k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from lingzhi227/agent-research-skills

All 31 skills in this repo
  • Backward Traceability

    lingzhi227/agent-research-skills

    Makes each number in a LaTeX paper link back to the code line that produced it, using hypertarget and hyperlink tags and compile-time `\num` formulas.

    384 GitHub stars~802 tokensUpdated 7 mo ago
    Auto-check passed
  • Excalidraw Canvas Toolkit

    lingzhi227/agent-research-skills

    Draws and refines Excalidraw diagrams on a live canvas through MCP tools or a REST API, with screenshots, file import and export, snapshots and Mermaid conversion.

    384 GitHub stars~3.8k tokensUpdated 7 mo ago
    Auto-check passed
  • Research Experiment Designer

    lingzhi227/agent-research-skills

    Plans research experiments in four progressive stages, from a first working implementation through baseline tuning and creative research to ablation studies.

    384 GitHub stars~752 tokensUpdated 7 mo ago
    Auto-check passed
  • Scientific Figure Generation

    lingzhi227/agent-research-skills

    Generates publication-quality scientific figures with matplotlib or seaborn through query expansion, a run-and-retry coding loop and a visual check of the rendered PNG.

    384 GitHub stars~809 tokensUpdated 7 mo ago
    Auto-check passed
  • Research Idea Generation

    lingzhi227/agent-research-skills

    Generates and iteratively refines research ideas for a given area, checking each one's novelty against Semantic Scholar and arXiv, and scoring it on interestingness, feasibility and novelty.

    384 GitHub stars~747 tokensUpdated 7 mo ago
    Auto-check passed
  • Academic LaTeX Formatter

    lingzhi227/agent-research-skills

    Sets up conference-specific LaTeX paper templates, checks a draft for formatting and submission issues, and auto-fixes common problems for venues like ICML, ICLR, NeurIPS, AAAI and ACL.

    384 GitHub stars~603 tokensUpdated 7 mo ago
    Auto-check passed

Questions about Statistical Data Analysis

What does Statistical Data Analysis do?

Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals. You give it a data source (CSV, JSON, pickle or experiment logs) and a research goal or hypothesis. The agent writes analysis code in a fixed section order, from imports and data loading through dataset preparation, descriptive statistics, preprocessing, the statistical tests and extra results, using pandas, numpy, scipy, statsmodels and sklearn.

When should I use Statistical Data Analysis?

Statistical Data Analysis fits situations like: analyzing experiment results for a research paper; choosing the right statistical test for two or more groups; formatting p-values as stars or LaTeX for a results table; reviewing analysis code for statistical or data-handling mistakes.

How do I install Statistical Data Analysis in Claude Code?

Run `npx skills add lingzhi227/agent-research-skills --skill data-analysis -a claude-code`. Or copy the skill folder (skills/data-analysis in lingzhi227/agent-research-skills) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Statistical Data Analysis in Codex?

Run `npx skills add lingzhi227/agent-research-skills --skill data-analysis -a codex`. Or copy the skill folder (skills/data-analysis in lingzhi227/agent-research-skills) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Statistical Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lingzhi227/agent-research-skills --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Statistical Data Analysis need to run?

Going by SKILL.md and its folder, Statistical Data Analysis needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python with numpy and scipy for stat_summary.py; pandas, statsmodels and scikit-learn for the generated analysis code.

Does Statistical Data Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Statistical Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Statistical Data Analysis use?

No licence was found for Statistical Data Analysis or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Statistical Data Analysis use?

About 886 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1k tokens, read only when the agent opens those files.

What are the alternatives to Statistical Data Analysis?

Skills that share tags, products or a category with Statistical Data Analysis: Q-EDA Exploratory Analysis (TyrealQ/q-skills, 108 stars), Tooluniverse Epigenomics (wu-yc/LabClaw, 1.1k stars), PyMC Bayesian Modeling (davila7/claude-code-templates, 32k stars) and Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Statistical Data Analysis?

lingzhi227 (a GitHub user) maintains it in lingzhi227/agent-research-skills, which has 384 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on February 27, 2026.

Source: lingzhi227/agent-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.