Agent skill

Data Analysis

by flonat in flonat/flonat-research

Deliver an end-to-end analysis pipeline: EDA, estimation, or publication output.

MITAuto-check passedData & Analytics

Install Data Analysis

skills CLI
$ npx skills add flonat/flonat-research --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
145
Token cost
~2.1k tokens
SKILL.md length
783 words
Files
10 (incl. references)
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Deliver an end-to-end analysis pipeline: EDA, estimation, or publication output.

  • Works in 5 steps: Setup → Exploratory Data Analysis → Estimation → …
  • The user requests an end-to-end analysis pipeline: EDA
  • SKILL.md covers Modes, When to Use, When NOT to Use and Shared References, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Analysis is an agent skill from flonat/flonat-research. Deliver an end-to-end analysis pipeline: EDA, estimation, or publication output. Use when the user requests an end-to-end analysis pipeline: EDA, estimation, or publication output.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `references/econ-visualisation.md`, `references/estimation-recipes.md` and `references/language-conventions.md`).

It sits in Data & Analytics, covering Data analysis and End-to-end testing. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • The user requests an end-to-end analysis pipeline: EDA
  • Publication output

Example prompts

  • “/data-analysis”

Requirements

  • Pre-approved tools (allowed-tools): Bash(uv*, Rscript*, R*, stata*, julia*, mkdir*, ls*, cp*), Read, Write, Edit, Glob, Grep, AskUserQuestion, Skill

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Setup
  2. Exploratory Data Analysis
  3. Estimation
  4. Publication Output
  5. Save & Review

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(uv*
    • Rscript*
    • R*
    • stata*
    • julia*
    • mkdir*
    • ls*
    • cp*)
    • Read
    • Write

    …and 5 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Analysis loads about 2.1k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 783 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 783 words, ~2,058 tokens.

Download SKILL.mdSave it as .claude/skills/data-analysis/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
data-analysis
description
Deliver an end-to-end analysis pipeline: EDA, estimation, or publication output. Use when the user requests an end-to-end analysis pipeline: EDA, estimation, or publication output.
allowed-tools
Bash(uv*, Rscript*, R*, stata*, julia*, mkdir*, ls*, cp*), Read, Write, Edit, Glob, Grep, AskUserQuestion, Skill
argument-hint
[data-path or project-path] [--mode eda|estimation|full]
agent-dependencies
code-review
skill-dependencies
experiment-design

Data Analysis Pipeline

Generate, execute, and verify analysis scripts across R, Python, Stata, and Julia.

Modes

ModeWhat it doesPhases
EDAExploratory data analysis only1–2
EstimationEstimation + publication output (requires locked design)1, 3–4
FullComplete pipeline1–5

Default: Full. Detect mode from user request or ask if ambiguous.

When to Use

  • "Analyse this data" / "Run EDA on this CSV" / "Estimate the model"
  • "Generate results tables" / "Create publication figures"
  • Any task requiring data → script → output pipeline

When NOT to Use

  • Experimental design or power analysis → experiment-design
  • Generating synthetic data for testing → synthetic-data
  • Auditing identification strategy → causal-design
  • Proofreading or compiling the paper → proofread, latex

Shared References

  • Method probing questions: shared/method-probing-questions.md — ask before running any analysis
  • Validation tiers: shared/validation-tiers.md — declare tier before examining results
  • Escalation protocol: shared/escalation-protocol.md — escalate when methodology answers are vague
  • Distribution diagnostics: shared/distribution-diagnostics.md — mandatory DV checks before model selection
  • Engagement-stratified sampling: shared/engagement-stratified-sampling.md — stratify by engagement tiers for social media data
  • Inter-coder reliability: shared/intercoder-reliability.md — per-category reliability for content analysis and LLM annotation

Workflow

Phase 1: Setup
  1. Detect project structure: Read CLAUDE.md, check for data/, code/, paper/ directories.
  2. Detect language: Check existing scripts, user preference, or ask. Read shared/multi-language-conventions.md for the chosen language's conventions.
  3. Locate data: Find datasets in data/raw/ or data/processed/. Never modify data/raw/ (per data-sensitivity rule). For social-media datasets, follow shared/engagement-stratified-sampling.md when constructing analysis samples.
  4. Confirm validation tier per shared/validation-tiers.md. Tier dictates claim-strength language allowed in Phase 4 outputs and how strict the locked-design gate (step 5) is enforced.
  5. Check for locked design: Look for analysis plan in log/plans/, .context/project-recap.md, or MEMORY.md estimand registry. If running Estimation or Full mode and no design exists, stop and warn: "No locked research design found. Run experiment-design or causal-design first, or confirm the specification before proceeding." Use shared/method-probing-questions.md to probe gaps if the user pushes back on the gate.
Phase 2: Exploratory Data Analysis

Generate and execute an EDA script that produces:

  1. Data overview: dimensions, types, missingness summary
  2. Univariate distributions: histograms/density for continuous, bar charts for categorical
  3. Bivariate relationships: correlation matrix, key scatterplots, cross-tabulations
  4. Outlier detection: box plots, IQR-based flags
  5. Distribution diagnostics (mandatory): run distribution_diagnostics() from shared/distribution-diagnostics.md on every DV and key IVs. Report skewness, zero proportion, overdispersion, and model recommendation. Flag if OLS is inappropriate.
  6. Balance tables (if treatment variable identified): pre-treatment covariate balance

Output routing:

  • Exploratory figures → output/figures/ (not paper/figures/)
  • Summary statistics → output/tables/ as .csv
  • EDA script → code/01_eda.R (or .py/.do/.jl)

EDA mode stops here.

Phase 3: Estimation

Gate check: Verify the research design is locked before proceeding. The specification (estimand, identifying assumptions, main model) must be documented. This enforces the design-before-results rule. If the user resists the gate, follow shared/escalation-protocol.md — escalate rather than accommodate. For analyses involving human or LLM coding (content analysis, annotation), require per-category reliability per shared/intercoder-reliability.md before estimation.

Generate estimation script(s) covering:

  1. Main specification — as defined in the locked design
  2. Robustness checks — pre-committed alternatives (different SEs, controls, subsamples)
  3. Diagnostics — specification-appropriate tests (first-stage F for IV, parallel trends for DiD, bandwidth sensitivity for RDD)

Read references/estimation-recipes.md for language-specific estimation patterns.

Output:

  • Estimation script → code/02_estimation.R (or equivalent)
  • Coefficient estimates → output/results/ as .rds/.pkl/.dta for downstream table generation
Show full SKILL.md (260 more words)Show less
Phase 4: Publication Output

Generate publication-ready tables and figures. Read shared/publication-output.md for format standards and references/table-formatting.md for language-specific recipes.

  1. Main results table — booktabs three-line format, exported as .tex to paper/tables/
  2. Robustness tables — same format, appendix naming convention
  3. Publication figures — coefficient plots, event study plots, mechanism figures → paper/figures/ as PDF
  4. Inline statistics — export \newcommand definitions for key numbers referenced in text

Critical rule: All numbers in .tex files must come from generated files via \input{}. Never hard-code results (per no-hardcoded-results rule). Scripts in code/, outputs in paper/ (per overleaf-separation rule).

Output script → code/03_tables_figures.R (or equivalent)

Phase 5: Save & Review
  1. Verify outputs exist: Check all expected files in paper/tables/ and paper/figures/
  2. Run the code-review agent on all generated scripts (auto-invoke via skill-routing mechanism)
  3. Log the analysis: Record what was done, which scripts were created, which outputs were generated
  4. Suggest next steps: compilation with latex, or additional analyses

Script Structure

Every generated script follows this header template:

# ============================================================
# Script: [filename]
# Purpose: [one-line description]
# Inputs: [list of input files]
# Outputs: [list of output files]
# Dependencies: [packages used]
# Author: [from git config]
# Date: [today]
# ============================================================

Cross-References

ResourceWhen read
shared/multi-language-conventions.mdPhase 1 (language setup)
shared/publication-output.mdPhase 4 (table/figure format)
references/estimation-recipes.mdPhase 3 (estimation code patterns)
references/econ-visualisation.mdPhase 2 & 4 (economics figure/table conventions)
references/table-formatting.mdPhase 4 (language-specific table export)
references/language-conventions.mdPhase 1 (additional language notes)
design-before-results rulePhase 3 gate check
data-sensitivity rulePhase 1 (data access)
no-hardcoded-results rulePhase 4 (output routing)
overleaf-separation rulePhase 4 (file placement)
the code-review agentPhase 5 (auto-invoked)
experiment-design skillSuggested if no design exists
causal-design skillSuggested if no design exists
econ-plots skillEconomics-specific figures
r-econometrics skillR-specific estimation
econ-data skillData download from public APIs

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in skills/data-analysis of flonat/flonat-research.

  • SKILL.md
  • references/econ-visualisation.md
  • references/estimation-recipes.md
  • references/language-conventions.md
  • references/matplotlib.md
  • references/plotly.md
  • references/seaborn.md
  • references/statsmodels.md
  • references/sympy.md
  • references/table-formatting.md

Open the folder on GitHubat commit da27600

Compare with similar skills

Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Analysis this skillflonat/flonat-research145—~2.1kAutomated safety check: PassMIT
Exploratory Data Analysisspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2063 repos~3.4kAutomated safety check: NotesMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT

Similar skills

  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    206 GitHub starsUsed in 3 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Yichen Wecom Local Vault

    mcncarl/yichen-skills

    Read, decrypt, query, search, and export local WeCom/企业微信 5.x desktop databases on macOS into a private read-only vault.

    4.3k GitHub stars~1.3k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    145 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    145 GitHub stars~4.4k tokensUpdated 9 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    145 GitHub stars~1.2k tokensUpdated 9 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    145 GitHub stars~488 tokensUpdated 9 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    145 GitHub stars~1.6k tokensUpdated 9 days ago
    Auto-check passed
  • Pre Commit Audit

    flonat/flonat-research

    Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

    145 GitHub stars~2.8k tokensUpdated 9 days ago
    Auto-check: notes

Questions about Data Analysis

What does Data Analysis do?

Deliver an end-to-end analysis pipeline: EDA, estimation, or publication output. Data Analysis is an agent skill from flonat/flonat-research. Deliver an end-to-end analysis pipeline: EDA, estimation, or publication output.

When should I use Data Analysis?

Data Analysis fits situations like: the user requests an end-to-end analysis pipeline: EDA; publication output.

How do I install Data Analysis in Claude Code?

Run `npx skills add flonat/flonat-research --skill data-analysis -a claude-code`. Or copy the skill folder (skills/data-analysis in flonat/flonat-research) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Data Analysis in Codex?

Run `npx skills add flonat/flonat-research --skill data-analysis -a codex`. Or copy the skill folder (skills/data-analysis in flonat/flonat-research) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Data Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Analysis is instructions for the agent only. Its frontmatter pre-approves these tools: Bash(uv*, Rscript*, R*, stata*, julia*, mkdir*, ls*, cp*), Read, Write, Edit, Glob, Grep, AskUserQuestion, Skill.

Does Data Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Analysis use?

Data Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Analysis use?

About 2.1k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.

What are the alternatives to Data Analysis?

Skills that share tags, products or a category with Data Analysis: Exploratory Data Analysis (spacering-net/codeg, 3.8k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 206 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Analysis?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.