Excel and CSV Data Analysis
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow data-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/data-analysis .claude/skills/data-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-analysis" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/data-analysis into .claude/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/data-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow data-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/data-analysis .agents/skills/data-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-analysis" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/data-analysis into .agents/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow data-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/data-analysis .cursor/skills/data-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-analysis" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/data-analysis into .cursor/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pedrohcgs/claude-code-my-workflow.git --path .claude/skills/data-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow data-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/data-analysis .gemini/skills/data-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-analysis" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/data-analysis into .gemini/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pedrohcgs/claude-code-my-workflow data-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/data-analysis .github/skills/data-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-analysis" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/data-analysis into .github/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow data-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/data-analysis .opencode/skills/data-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-analysis" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/data-analysis into .opencode/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-analysisEnd-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures.
Data Analysis is an agent skill from pedrohcgs/claude-code-my-workflow. End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures. Use when user says "analyze this dataset", "run a regression on X", "explore this CSV", "full analysis workflow", "get me summary stats and a regression", or points at a .csv/.rds/.dta and asks for empirical results. Produces numbered R scripts in scripts/R/ and outputs to output/.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data analysis and CSV and tabular files. The repository describes itself as: A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ae72617. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGrepGlobWriteEditBashAgentTaskMonitorFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
psantanna.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Analysis loads about 2.3k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 798 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Grep, Glob, Write, Edit, Bash, Agent, Task, MonitorAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from pedrohcgs/claude-code-my-workflow at commit ae72617, republished under its MIT licence (© pedrohcgs). 798 words, ~2,263 tokens.
.claude/skills/data-analysis/SKILL.md (or your agent's skills folder).Run an end-to-end data analysis in R: load, explore, analyze, and produce publication-ready output.
Input: $ARGUMENTS — a dataset path (e.g., data/county_panel.csv) or a description of the analysis goal (e.g., "regress wages on education with state fixed effects using CPS data").
.claude/rules/r-code-conventions.mdscripts/R/ with descriptive namesoutput/saveRDS() for every computed object — Quarto slides may need them.claude/rules/)Before writing any analysis code, produce a Pre-Flight Report showing you read the inputs. This prevents the common failure mode where the agent hallucinates variable names or skips project conventions.
Output block (in your response to the user, before Phase 1):
## Pre-Flight Report
**Dataset:** [path]
- Variables found: [list from head()/names()]
- Rows: [count]
- Key types: [e.g., "outcome=numeric, treatment=binary, state=factor"]
- Missing-data summary: [% missing per key var]
**Project conventions read:**
- `.claude/rules/r-code-conventions.md` — [one-line summary of most relevant rule]
- `.claude/rules/content-invariants.md` — [INV-9, INV-10, INV-11, INV-12 applicable]
**Task interpretation:** [one sentence restating what the user asked for]
**Plan:** [3-5 bullet outline of the R script structure]If any input cannot be read (missing file, unreadable format), stop and ask the user before proceeding.
library(), never require())r-code-conventions.md), e.g. set.seed(20260415) (INV-9)Generate diagnostic outputs:
summary(), missingness rates, variable typesSave all diagnostic figures to output/diagnostics/.
Based on the research question:
fixest for panel data, lm/glm for cross-sectionEvery specification estimated in this phase, including the ones dropped and the ones that failed, gets one row appended to quality_reports/spec-ledger.md as it is run. It is the record a referee's "what else did you try?" deserves: it keeps the search visible rather than preventing it, and it records specifications without advising which to run.
kept, dropped or failed; Why says why for anything not kept.confidential-data.md).| inside a formula as \|, or the table breaks. A specification containing a backtick (a Stata local macro such as `x') goes in a double-backtick span: `` reg y `controls' ``.LEDGER=quality_reports/spec-ledger.md
[ -f "$LEDGER" ] || printf '%s\n' "# Specification ledger" "" \
"Every specification estimated, kept or not, one row each, appended as it is run. Never edit a past row; a correction is a new row." "" \
"| Date | Commit | Script:line | Outcome | Specification | Sample | Status | Why | Estimate |" \
"|---|---|---|---|---|---|---|---|---|" > "$LEDGER"
if REV=$(git rev-parse --short HEAD 2>/dev/null); then # -dirty = anything uncommitted outside quality_reports/, untracked files included
[ -n "$(git status --porcelain -- . ':(exclude)quality_reports' 2>/dev/null)" ] && REV="$REV-dirty"
else
REV="no-commit"
fi
printf '| %s | %s ' "$(date +%F)" "$REV" >> "$LEDGER" # one printf + one heredoc per specification
cat >> "$LEDGER" <<'EOF'
| scripts/R/03_analyze.R:42 | log_wage | `feols(log_wage ~ treat + age \| id + year, cluster = ~id)` | panel 2010-2019 | kept | main specification | 0.082 (0.021) |
EOFTables:
modelsummary for regression tables (preferred) or stargazer.tex for LaTeX inclusion and .html for quick viewingFigures:
ggplot2 with project themebg = "transparent" for Beamer compatibilityggsave(width = X, height = Y).pdf and .pngsaveRDS() for all key objects (regression results, summary tables, processed data)output/ subdirectories as needed with dir.create(..., recursive = TRUE)Delegate to the r-reviewer agent:
"Review the script at scripts/R/[script_name].R"Follow this template:
# ============================================================
# [Descriptive Title]
# Author: [from project context]
# Purpose: [What this script does]
# Inputs: [Data files]
# Outputs: [Figures, tables, RDS files]
# ============================================================
# 0. Setup ----
library(tidyverse)
library(fixest)
library(modelsummary)
set.seed(20260415) # YYYYMMDD per r-code-conventions.md (INV-9)
dir.create("output/analysis", recursive = TRUE, showWarnings = FALSE)
# 1. Data Loading ----
# [Load and clean data]
# 2. Exploratory Analysis ----
# [Summary stats, diagnostic plots]
# 3. Main Analysis ----
# [Regressions, estimation]
# 4. Tables and Figures ----
# [Publication-ready output]
# 5. Export ----
# [saveRDS for all objects, ggsave for all figures]For regressions, simulations, or bootstrap loops that take more than a couple of minutes, launch via Bash with run_in_background: true and then use Anthropic's Monitor tool to stream R stdout into the conversation in real time. Pattern:
run_in_background: true, sending all output to a log: mkdir -p output && Rscript scripts/R/03_analyze.R > output/03_analyze.log 2>&1. The background job notifies you by itself when the process exits.tail -f output/03_analyze.log | grep --line-buffered -E "Coefficients table written|Error|Execution halted". Monitor has no job-id parameter: the stdout of its own command is the event stream. tail -f never exits, so set timeout_ms above the expected runtime (or persistent: true) and stop the monitor with TaskStop once the job finishes.This avoids the polling-loop anti-pattern (sleep 30; check; sleep 30; check) and avoids burning cache on idle waits. Especially useful when paired with the Cost-Conscious Parallelism section of the guide.
© pedrohcgs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/data-analysis of pedrohcgs/claude-code-my-workflow.
Open the folder on GitHubat commit ae72617
Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Analysis this skillpedrohcgs/claude-code-my-workflow | 1.6k | — | ~2.3k | Automated safety check: Notes | MIT | |
| Excel and CSV Data Analysisbytedance/deer-flow | 83k | 4 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Exploratory Data AnalysisOleafly/Oleafly | 205 | 2 repos | ~3.4k | Automated safety check: Notes | MIT | |
| Eqtl Catalogue Region FetchClawBio/ClawBio | 1.2k | 1 repos | ~4.3k | Automated safety check: Pass | MIT | |
| CSV Data Analysis5zjk5/prompt-engineering | 127 | — | ~2.6k | Automated safety check: Pass | None | |
| Gwas Catalog Region FetchClawBio/ClawBio | 1.2k | 1 repos | ~3.5k | Automated safety check: Pass | MIT |
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
Oleafly/Oleafly
Perform bounded, local exploratory analysis of explicitly supported scientific files.
ClawBio/ClawBio
Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.
5zjk5/prompt-engineering
This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.
ClawBio/ClawBio
Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.
pipeshub-ai/pipeshub-ai
Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.
pedrohcgs/claude-code-my-workflow
Adversarial 5-7 question challenge to a deck's pedagogical choices — ordering, prerequisites, cognitive load, motivation.
pedrohcgs/claude-code-my-workflow
Qualify a check before it is allowed to clear anything — prove it can detect the failure it is meant to catch.
pedrohcgs/claude-code-my-workflow
Compile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex).
pedrohcgs/claude-code-my-workflow
Show current context status and session health. An agent skill from pedrohcgs/claude-code-my-workflow.
pedrohcgs/claude-code-my-workflow
Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt /…
pedrohcgs/claude-code-my-workflow
Save a structured state snapshot before stopping or handing off.
Categories
End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures. Data Analysis is an agent skill from pedrohcgs/claude-code-my-workflow. End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures.
Data Analysis fits situations like: user says analyze this dataset; run a regression on X; explore this CSV; full analysis workflow.
Run `npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a claude-code`. Or copy the skill folder (.claude/skills/data-analysis in pedrohcgs/claude-code-my-workflow) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a codex`. Or copy the skill folder (.claude/skills/data-analysis in pedrohcgs/claude-code-my-workflow) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pedrohcgs/claude-code-my-workflow --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.
Going by SKILL.md and its folder, Data Analysis needs the command-line tools its instructions call (git). Its frontmatter pre-approves these tools: Read, Grep, Glob, Write, Edit, Bash, Agent, Task, Monitor.
SKILL.md names 1 domain. As links in the text: psantanna.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Data Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Data Analysis: Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 205 stars), Eqtl Catalogue Region Fetch (ClawBio/ClawBio, 1.2k stars) and CSV Data Analysis (5zjk5/prompt-engineering, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pedrohcgs (a GitHub user) maintains it in pedrohcgs/claude-code-my-workflow, which has 1,639 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on September 27, 2026.
Source: pedrohcgs/claude-code-my-workflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.