Agent skill

Data Analysis

by pedrohcgs in pedrohcgs/claude-code-my-workflow

End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures.

MITAuto-check: notesData & Analytics

Install Data Analysis

skills CLI
$ npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pedrohcgs/claude-code-my-workflow data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
1.6k
Token cost
~2.3k tokens
SKILL.md length
798 words
Files
1
Skills in repo
59
Repo updated
First seen
Licence
MIT

At a glance

End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures.

  • Works in 6 steps: Pre-Flight Report → Setup and Data Loading → Exploratory Data Analysis → …
  • User says analyze this dataset
  • SKILL.md covers Constraints, Workflow Phases, Script Structure and Important, plus 1 more section
  • Calls git

What it does

Data Analysis is an agent skill from pedrohcgs/claude-code-my-workflow. End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures. Use when user says "analyze this dataset", "run a regression on X", "explore this CSV", "full analysis workflow", "get me summary stats and a regression", or points at a .csv/.rds/.dta and asks for empirical results. Produces numbered R scripts in scripts/R/ and outputs to output/.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data analysis and CSV and tabular files. The repository describes itself as: A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols. The licence is MIT.

When your agent uses it

  • User says analyze this dataset
  • Run a regression on X
  • Explore this CSV
  • Full analysis workflow

Example prompts

  • “analyze this dataset”
  • “run a regression on X”
  • “explore this CSV”
  • “/data-analysis”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Write, Edit, Bash, Agent, Task, Monitor

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Pre-Flight Report
  2. Setup and Data Loading
  3. Exploratory Data Analysis
  4. Main Analysis
  5. Publication-Ready Output
  6. Save and Review

What it can do on your machine

Read from SKILL.md and the folder at commit ae72617. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Write
    • Edit
    • Bash
    • Agent
    • Task
    • Monitor

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • psantanna.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Analysis loads about 2.3k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 798 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Write, Edit, Bash, Agent, Task, Monitor

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pedrohcgs/claude-code-my-workflow at commit ae72617, republished under its MIT licence (© pedrohcgs). 798 words, ~2,263 tokens.

Download SKILL.mdSave it as .claude/skills/data-analysis/SKILL.md (or your agent's skills folder).
name
data-analysis
description
End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures. Use when user says "analyze this dataset", "run a regression on X", "explore this CSV", "full analysis workflow", "get me summary stats and a regression", or points at a `.csv`/`.rds`/`.dta` and asks for empirical results. Produces numbered R scripts in `scripts/R/` and outputs to `output/`.
allowed-tools
Read, Grep, Glob, Write, Edit, Bash, Agent, Task, Monitor
argument-hint
[dataset path or description of analysis goal]

Data Analysis Workflow

Run an end-to-end data analysis in R: load, explore, analyze, and produce publication-ready output.

Input: $ARGUMENTS — a dataset path (e.g., data/county_panel.csv) or a description of the analysis goal (e.g., "regress wages on education with state fixed effects using CPS data").


Constraints

  • Follow R code conventions in .claude/rules/r-code-conventions.md
  • Save all scripts to scripts/R/ with descriptive names
  • Save all outputs (figures, tables, RDS) to output/
  • Use saveRDS() for every computed object — Quarto slides may need them
  • Use project theme for all figures (check for custom theme in .claude/rules/)
  • Run r-reviewer on the generated script before presenting results

Workflow Phases

Phase 0: Pre-Flight Report

Before writing any analysis code, produce a Pre-Flight Report showing you read the inputs. This prevents the common failure mode where the agent hallucinates variable names or skips project conventions.

Output block (in your response to the user, before Phase 1):

markdown
## Pre-Flight Report

**Dataset:** [path]
- Variables found: [list from head()/names()]
- Rows: [count]
- Key types: [e.g., "outcome=numeric, treatment=binary, state=factor"]
- Missing-data summary: [% missing per key var]

**Project conventions read:**
- `.claude/rules/r-code-conventions.md` — [one-line summary of most relevant rule]
- `.claude/rules/content-invariants.md` — [INV-9, INV-10, INV-11, INV-12 applicable]

**Task interpretation:** [one sentence restating what the user asked for]

**Plan:** [3-5 bullet outline of the R script structure]

If any input cannot be read (missing file, unreadable format), stop and ask the user before proceeding.

Phase 1: Setup and Data Loading
  1. Create R script with proper header (title, author, purpose, inputs, outputs)
  2. Load required packages at top (library(), never require())
  3. Set seed once at top in YYYYMMDD format (per r-code-conventions.md), e.g. set.seed(20260415) (INV-9)
  4. Load and inspect the dataset
Phase 2: Exploratory Data Analysis

Generate diagnostic outputs:

  • Summary statistics: summary(), missingness rates, variable types
  • Distributions: Histograms for key continuous variables
  • Relationships: Scatter plots, correlation matrices
  • Time patterns: If panel data, plot trends over time
  • Group comparisons: If treatment/control, compare pre-treatment means

Save all diagnostic figures to output/diagnostics/.

Phase 3: Main Analysis

Based on the research question:

  • Regression analysis: Use fixest for panel data, lm/glm for cross-section
  • Standard errors: Cluster at the appropriate level (document why)
  • Multiple specifications: Start simple, progressively add controls
  • Effect sizes: Report standardized effects alongside raw coefficients
Specification ledger — every run, kept or not

Every specification estimated in this phase, including the ones dropped and the ones that failed, gets one row appended to quality_reports/spec-ledger.md as it is run. It is the record a referee's "what else did you try?" deserves: it keeps the search visible rather than preventing it, and it records specifications without advising which to run.

  • Append only. Never edit or delete a row; a correction is a new row. A commit that changes a committed row fails the repo-hygiene gate.
  • Status is kept, dropped or failed; Why says why for anything not kept.
  • Estimate is optional. On restricted data, leave it empty until the number has cleared disclosure review (confidential-data.md).
  • Escape a | inside a formula as \|, or the table breaks. A specification containing a backtick (a Stata local macro such as `x') goes in a double-backtick span: `` reg y `controls' ``.
bash
LEDGER=quality_reports/spec-ledger.md
[ -f "$LEDGER" ] || printf '%s\n' "# Specification ledger" "" \
  "Every specification estimated, kept or not, one row each, appended as it is run. Never edit a past row; a correction is a new row." "" \
  "| Date | Commit | Script:line | Outcome | Specification | Sample | Status | Why | Estimate |" \
  "|---|---|---|---|---|---|---|---|---|" > "$LEDGER"
if REV=$(git rev-parse --short HEAD 2>/dev/null); then       # -dirty = anything uncommitted outside quality_reports/, untracked files included
  [ -n "$(git status --porcelain -- . ':(exclude)quality_reports' 2>/dev/null)" ] && REV="$REV-dirty"
else
  REV="no-commit"
fi
printf '| %s | %s ' "$(date +%F)" "$REV" >> "$LEDGER"          # one printf + one heredoc per specification
cat >> "$LEDGER" <<'EOF'
| scripts/R/03_analyze.R:42 | log_wage | `feols(log_wage ~ treat + age \| id + year, cluster = ~id)` | panel 2010-2019 | kept | main specification | 0.082 (0.021) |
EOF
Show full SKILL.md (351 more words)Show less
Phase 4: Publication-Ready Output

Tables:

  • Use modelsummary for regression tables (preferred) or stargazer
  • Include all standard elements: coefficients, SEs, significance stars, N, R-squared
  • Export as .tex for LaTeX inclusion and .html for quick viewing

Figures:

  • Use ggplot2 with project theme
  • Set bg = "transparent" for Beamer compatibility
  • Include proper axis labels (sentence case, units)
  • Export with explicit dimensions: ggsave(width = X, height = Y)
  • Save as both .pdf and .png
Phase 5: Save and Review
  1. saveRDS() for all key objects (regression results, summary tables, processed data)
  2. Create output/ subdirectories as needed with dir.create(..., recursive = TRUE)
  3. Run the r-reviewer agent on the generated script:
Delegate to the r-reviewer agent:
"Review the script at scripts/R/[script_name].R"
  1. Address any Critical or High issues from the review.

Script Structure

Follow this template:

r
# ============================================================
# [Descriptive Title]
# Author: [from project context]
# Purpose: [What this script does]
# Inputs: [Data files]
# Outputs: [Figures, tables, RDS files]
# ============================================================

# 0. Setup ----
library(tidyverse)
library(fixest)
library(modelsummary)

set.seed(20260415)  # YYYYMMDD per r-code-conventions.md (INV-9)

dir.create("output/analysis", recursive = TRUE, showWarnings = FALSE)

# 1. Data Loading ----
# [Load and clean data]

# 2. Exploratory Analysis ----
# [Summary stats, diagnostic plots]

# 3. Main Analysis ----
# [Regressions, estimation]

# 4. Tables and Figures ----
# [Publication-ready output]

# 5. Export ----
# [saveRDS for all objects, ggsave for all figures]

Important

  • Reproduce, don't guess. If the user specifies a regression, run exactly that.
  • Show your work. Print summary statistics before jumping to regression.
  • Check for issues. Look for multicollinearity, outliers, perfect prediction.
  • Use relative paths. All paths relative to repository root.
  • No hardcoded values. Use variables for sample restrictions, date ranges, etc.

Long-running fits: use the Monitor tool (Apr 2026)

For regressions, simulations, or bootstrap loops that take more than a couple of minutes, launch via Bash with run_in_background: true and then use Anthropic's Monitor tool to stream R stdout into the conversation in real time. Pattern:

  1. Background-launch with Bash run_in_background: true, sending all output to a log: mkdir -p output && Rscript scripts/R/03_analyze.R > output/03_analyze.log 2>&1. The background job notifies you by itself when the process exits.
  2. Start Monitor with a command that follows that log and filters for milestones and failures, e.g. tail -f output/03_analyze.log | grep --line-buffered -E "Coefficients table written|Error|Execution halted". Monitor has no job-id parameter: the stdout of its own command is the event stream. tail -f never exits, so set timeout_ms above the expected runtime (or persistent: true) and stop the monitor with TaskStop once the job finishes.
  3. Continue or course-correct based on what the stream reveals.

This avoids the polling-loop anti-pattern (sleep 30; check; sleep 30; check) and avoids burning cache on idle waits. Especially useful when paired with the Cost-Conscious Parallelism section of the guide.

© pedrohcgs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/data-analysis of pedrohcgs/claude-code-my-workflow.

Open the folder on GitHubat commit ae72617

Compare with similar skills

Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Analysis this skillpedrohcgs/claude-code-my-workflow1.6k—~2.3kAutomated safety check: NotesMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2052 repos~3.4kAutomated safety check: NotesMIT
Eqtl Catalogue Region FetchClawBio/ClawBio1.2k1 repos~4.3kAutomated safety check: PassMIT
CSV Data Analysis5zjk5/prompt-engineering127—~2.6kAutomated safety check: PassNone
Gwas Catalog Region FetchClawBio/ClawBio1.2k1 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    205 GitHub starsUsed in 2 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Data & AnalyticsAuto-check passed
  • CSV Data Analysis

    5zjk5/prompt-engineering

    This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.

    127 GitHub stars~2.6k tokensUpdated 22 days ago
    Data & AnalyticsAuto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from pedrohcgs/claude-code-my-workflow

All 59 skills in this repo
  • Devils Advocate

    pedrohcgs/claude-code-my-workflow

    Adversarial 5-7 question challenge to a deck's pedagogical choices — ordering, prerequisites, cognitive load, motivation.

    1.6k GitHub starsUsed in 2 repos~641 tokens
    Auto-check passed
  • Vaccinate

    pedrohcgs/claude-code-my-workflow

    Qualify a check before it is allowed to clear anything — prove it can detect the failure it is meant to catch.

    1.6k GitHub stars~2.1k tokensUpdated 10 days ago
    Auto-check: notes
  • Compile Latex

    pedrohcgs/claude-code-my-workflow

    Compile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex).

    1.6k GitHub starsUsed in 1 repo~492 tokens
    Auto-check: notes
  • Context Status

    pedrohcgs/claude-code-my-workflow

    Show current context status and session health. An agent skill from pedrohcgs/claude-code-my-workflow.

    1.6k GitHub starsUsed in 1 repo~613 tokens
    Auto-check: notes
  • Capture Environment

    pedrohcgs/claude-code-my-workflow

    Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt /…

    1.6k GitHub stars~2.8k tokensUpdated 10 days ago
    Auto-check: notes
  • Checkpoint

    pedrohcgs/claude-code-my-workflow

    Save a structured state snapshot before stopping or handing off.

    1.6k GitHub stars~2.8k tokensUpdated 10 days ago
    Auto-check: notes

Questions about Data Analysis

What does Data Analysis do?

End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures. Data Analysis is an agent skill from pedrohcgs/claude-code-my-workflow. End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures.

When should I use Data Analysis?

Data Analysis fits situations like: user says analyze this dataset; run a regression on X; explore this CSV; full analysis workflow.

How do I install Data Analysis in Claude Code?

Run `npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a claude-code`. Or copy the skill folder (.claude/skills/data-analysis in pedrohcgs/claude-code-my-workflow) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Data Analysis in Codex?

Run `npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a codex`. Or copy the skill folder (.claude/skills/data-analysis in pedrohcgs/claude-code-my-workflow) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Data Analysis need to run?

Going by SKILL.md and its folder, Data Analysis needs the command-line tools its instructions call (git). Its frontmatter pre-approves these tools: Read, Grep, Glob, Write, Edit, Bash, Agent, Task, Monitor.

Does Data Analysis access the network?

SKILL.md names 1 domain. As links in the text: psantanna.com. This is read from the text; nothing was executed.

Is Data Analysis safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Data Analysis use?

Data Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Analysis use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Analysis?

Skills that share tags, products or a category with Data Analysis: Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 205 stars), Eqtl Catalogue Region Fetch (ClawBio/ClawBio, 1.2k stars) and CSV Data Analysis (5zjk5/prompt-engineering, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Analysis?

pedrohcgs (a GitHub user) maintains it in pedrohcgs/claude-code-my-workflow, which has 1,639 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on September 27, 2026.

Source: pedrohcgs/claude-code-my-workflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.