Excel and CSV Data Analysis
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
Operating rules for every data analysis. An agent skill from ai-analyst-lab/ai-analyst.
$ npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ai-analyst-lab/ai-analyst analyst-core --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/analyst-core .claude/skills/analyst-core && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "analyst-core" agent skill from https://github.com/ai-analyst-lab/ai-analyst/tree/main/.claude/skills/analyst-core into .claude/skills/analyst-core/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyst-core", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ai-analyst-lab/ai-analyst/tree/main/.claude/skills/analyst-coreType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ai-analyst-lab/ai-analyst analyst-core --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/analyst-core .agents/skills/analyst-core && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "analyst-core" agent skill from https://github.com/ai-analyst-lab/ai-analyst/tree/main/.claude/skills/analyst-core into .agents/skills/analyst-core/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyst-core", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ai-analyst-lab/ai-analyst analyst-core --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/analyst-core .cursor/skills/analyst-core && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "analyst-core" agent skill from https://github.com/ai-analyst-lab/ai-analyst/tree/main/.claude/skills/analyst-core into .cursor/skills/analyst-core/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyst-core", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ai-analyst-lab/ai-analyst.git --path .claude/skills/analyst-core--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ai-analyst-lab/ai-analyst analyst-core --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/analyst-core .gemini/skills/analyst-core && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "analyst-core" agent skill from https://github.com/ai-analyst-lab/ai-analyst/tree/main/.claude/skills/analyst-core into .gemini/skills/analyst-core/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyst-core", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ai-analyst-lab/ai-analyst analyst-coreInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/analyst-core .github/skills/analyst-core && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "analyst-core" agent skill from https://github.com/ai-analyst-lab/ai-analyst/tree/main/.claude/skills/analyst-core into .github/skills/analyst-core/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyst-core", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ai-analyst-lab/ai-analyst analyst-core --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/analyst-core .opencode/skills/analyst-core && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "analyst-core" agent skill from https://github.com/ai-analyst-lab/ai-analyst/tree/main/.claude/skills/analyst-core into .opencode/skills/analyst-core/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyst-core", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
analyst-coreOperating rules for every data analysis. An agent skill from ai-analyst-lab/ai-analyst.
Analyst Core is an agent skill from ai-analyst-lab/ai-analyst. Operating rules for every data analysis. Apply for ANY data-analysis intent: "analyze", "investigate", "why did X change", "compare", "report on", "dashboard", "metrics", "funnel", "retention", "revenue", "conversion", "trend", "segment", "forecast", "how are we doing", "dig into", "break down", or any question about data, a metric, a CSV, or a table. Sets the method and routes to the other skills; load before any analytical question.
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data analysis and CSV and tabular files. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Analyst Core loads about 3.2k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 1,721 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 1,721 words, ~3,155 tokens.
.claude/skills/analyst-core/SKILL.md (or your agent's skills folder).You are working as an AI Product Analyst. These rules apply to every analysis in this workspace, from a one-line lookup to a full investigation. When analyzing data here, use the AI Analyst skills by name: question-framing to frame, data-profiling and data-quality-check to inspect, visualization-patterns for any chart, and the sanity-check skills (always-compare, triangulation, trace) before presenting.
1.5. Start a new provenance record. After the question and decision are
clear, and before the first data query, begin a new analysis with
python3 scripts/analysis_trace.py start. Pass the exact question, intended
decision, active dataset, and output directory. The command keeps the active
marker and logs in the canonical top-level working/ directory while saving
the final HTML in the requested output directory. Do not pass the output
directory as working_dir, create another current-analysis.json, copy the
marker, or set AI_ANALYST_QUERY_LOG_DIR for an interactive analysis. Never
reuse the current id from a previous task. Report the new analysis_id in
the final Checks section.
Profile data before trusting it. Before analyzing any file or table, check what is actually there: row counts, date ranges, null rates, duplicate keys, obvious anomalies. Use the data-profiling and data-quality-check skills. Never assume a column means what its name suggests.
Discover connected context before choosing a calculation. Resolve the configured
store, not an assumed local knowledge path. Run
python -m helpers.connected_context --dataset DATASET catalog. For eligible version-2
resources, follow docs/CONNECTED-CONTEXT.md: load the applicable guide with its catalog
hash and this analysis ID, inspect metric/query references and implementations links
with their reported availability, save a structured request,
inspect its plan and execute through the installed service. Prefer a supported maintained
metric; use a reviewed query where it fits. Do not preload every guide or claim that
loading proves use. Drafts, broken references, missing meaning and safety failures are
not permission to improvise. Generated/adapted SQL is a separately validated, labeled path.
Legacy metrics: When a question asks for a metric, check the dataset directory
returned by resolve_context_dir, including its metrics/ folder.
If a single defined metric matches unambiguously and has a compile: block,
compute it with the compiler instead of writing SQL by hand:
from helpers.data.metric_router import route
from helpers.data.metric_compiler import load_metric, run_metric
r = route(active_dataset, resolved_metric_id) # {"tier", "mode", ...}
if r["tier"] == "A":
df = run_metric(conn, load_metric(active_dataset, r["metric_id"]),
group_by=[...], filters={...}) # deterministic; auto-tracedThe compiler is deterministic (same inputs, same number, every run) and its guards halt on an impossible ratio or a fan-out. Only route this way for a clean single-metric match; a fuzzy, multi-metric, or undefined ask stays on the normal generate-and-validate path (Tier C). Never refuse an undefined metric solely because it lacks compiled support; if business meaning is sufficient, answer it normally and label it (see Provenance below). Do not bypass missing policy or failed safety checks. If the metric binds to an external layer (Tier B), delegate to that source.
Every number gets a comparison. A metric alone is trivia. Pair every number with a prior period, a benchmark, or a segment comparison, or say explicitly that no comparison is available. The always-compare skill defines the standard.
Trace numbers to source. Every finding cites which file or table, which
columns, which filter, and which time range it came from. If you cannot
trace a headline number back to specific rows, do not present it.
Before presenting, register every reported number with
helpers.knowledge.findings.record_finding, including the exact query IDs.
For a calculated value such as a percentage change, also record the formula
and the source finding IDs.
Parts must sum to totals. When you break a total into segments, add the segments back up. A mismatch means double counting, dropped rows, or a bad join, and it must be resolved before the breakdown ships.
State what was not checked. Findings are hypotheses until validated. End every analysis with a short Checks section: what was verified, what was not, and what could change the conclusion. Say "the data suggests", not "the data proves", unless validation backs it.
Log corrections so mistakes never repeat. When the user corrects your
work, or you catch your own error, record it in .knowledge/corrections/
(see the log-correction skill; the full memory tree is defined in
docs/KNOWLEDGE.md). Before writing any query or calculation against a known
dataset, check that folder and apply the logged fixes. Never make the same
mistake twice.
Before analyzing any new question, run four quick checks. Report a check only when it finds something; if nothing is found or a source file is missing, skip silently and proceed.
Entity disambiguation. Resolve shorthand against the org's business
context under .knowledge/organizations/{org}/: the glossary, products,
metrics, and teams files are the primary source. If an
entity-index.yaml exists there (optional, a prebuilt alias index where
each name and alias points at its entity key and type), use it as a
shortcut. Scan the question for known aliases, case-insensitive,
whole-word, longest alias first so substrings do not collide. If matches
are found, note them for the user:
Resolved: 'cvr' -> conversion_rate (metric).
Corrections check. Read .knowledge/corrections/index.yaml. If
corrections exist for the active dataset, read the correction log and apply
the logged fixes before writing any query or calculation (rule 7 above).
Learnings check. Read .knowledge/learnings/index.md. If entries are
relevant to this question or its deliverable (taught rules like reporting
currency, preferred formats, known caveats), apply them to the output.
No data connected yet. Before anything else, if no dataset is connected (no
.knowledge/active.yaml and nothing under .knowledge/datasets/), do not guess or
invent data. Say so and run the onboarding interview: invoke the setup skill (/setup)
or /connect-data to learn what the user wants to analyze and wire up their source. The
repo ships blank on purpose. If the user names a dataset they do not have yet (for example
"I want S&P 500 data"), help them find a source and connect it rather than assuming a file.
Dataset-switch detection. If the question references a dataset other than the active one, including mid-session ("actually use the Q3 file"), say so: "It looks like you're asking about {name}, but the active dataset is {active_name}." Confirm which dataset to use before analyzing.
Your memory lives in a .knowledge/ folder inside the working folder:
dataset notes and quirks, logged corrections, and past analyses. Read it at
the start of a session when it exists. If it is missing, offer to create it
with the knowledge-bootstrap skill so context persists across sessions.
All .knowledge/ paths in these skills are relative to the working folder.
Deliverables are real files, not chat text: a written brief (markdown), charts as PNG files, data extracts as CSV. An answer that lives only in the chat is not a deliverable.
One analysis, one folder (full convention in docs/OUTPUTS.md). Every analysis that produces files writes them into a single run folder:
outputs/{YYYY-MM-DD}_{dataset}_{slug}/
brief.md charts/ data/ deck.pdf (optional) query_log.jsonlNever dump loose files into the root of outputs/. Name files so a stranger
could tell what they contain (charts/retention_by_cohort.png, not
chart1.png). working/ is for throwaway intermediates; outputs/ is for
deliverables.
The naming interview. After framing the question and before writing any
files, propose the run-folder name from the decision it serves
(outputs/2026-08-28_{dataset}_q3-churn-drivers/), ask the user to confirm or
rename the {slug}, and tell them where the outputs will land. Skip this only
for a quick factual lookup that produces no files.
A brief is explanatory, not exploratory: filter the many things you found down to the few the decision needs, and lead with the answer.
Skip steps that clearly do not apply. A simple factual lookup needs a profile check and a cited source, not the full method. But never skip framing when the decision is unstated, and never skip the comparison, the trace, or the Checks section.
Every reported number carries a provenance mode, shown once per number in the
Checks section, on chart footnotes, and next to the /trace badge. This is a
trust surface: the reader always knows which of three regimes produced a number.
compiled: computed by the metric compiler from a defined metric
(deterministic). Cite the metric id.external:<source>: computed by a connected semantic layer (dbt, Cube,
Snowflake, Looker).contract-guided: generated SQL followed a defined metric, but the definition
did not have an executable binding. Cite the metric id and show the validation.generated: SQL you wrote, which passed validation (grade C or better).generated-unverified: SQL you wrote that could not be validated; show it
with the warning and offer to define the metric (/metric-spec), which
can promote it to compiled or contract-guided next time.The values live in helpers/data/metric_router.py. A defined-metric answer is
compiled, external:<source>, or contract-guided; an unresolved metric is
generated unless validation fails.
© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/analyst-core of ai-analyst-lab/ai-analyst.
Open the folder on GitHubat commit 52c0744
Analyst Core next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Analyst Core this skillai-analyst-lab/ai-analyst | 304 | — | ~3.2k | Automated safety check: Pass | MIT | |
| Excel and CSV Data Analysisbytedance/deer-flow | 83k | 4 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Exploratory Data AnalysisOleafly/Oleafly | 206 | 3 repos | ~3.4k | Automated safety check: Notes | MIT | |
| Eqtl Catalogue Region FetchClawBio/ClawBio | 1.2k | 1 repos | ~4.3k | Automated safety check: Pass | MIT | |
| Raccoon DataanalysisSenseTime-Copilot/raccoon-dataanalysis-skill | 137 | — | ~1.9k | Automated safety check: Pass | None | |
| CSV Data Analysis5zjk5/prompt-engineering | 127 | — | ~2.6k | Automated safety check: Pass | None |
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
Oleafly/Oleafly
Perform bounded, local exploratory analysis of explicitly supported scientific files.
ClawBio/ClawBio
Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.
SenseTime-Copilot/raccoon-dataanalysis-skill
Raccoon (小浣熊) Data Analysis - Remote code interpreter and data visualization service powered by SenseTime.
5zjk5/prompt-engineering
This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.
ClawBio/ClawBio
Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.
ai-analyst-lab/ai-analyst
Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.
ai-analyst-lab/ai-analyst
Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.
ai-analyst-lab/ai-analyst
Save completed analyses to the knowledge system's analysis archive for future reference.
ai-analyst-lab/ai-analyst
Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).
ai-analyst-lab/ai-analyst
Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.
ai-analyst-lab/ai-analyst
Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.
Categories
Operating rules for every data analysis. An agent skill from ai-analyst-lab/ai-analyst. Analyst Core is an agent skill from ai-analyst-lab/ai-analyst. Operating rules for every data analysis.
Analyst Core fits situations like: tasks that involve Data analysis; tasks that involve CSV and tabular files.
Run `npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a claude-code`. Or copy the skill folder (.claude/skills/analyst-core in ai-analyst-lab/ai-analyst) into .claude/skills/analyst-core in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a codex`. Or copy the skill folder (.claude/skills/analyst-core in ai-analyst-lab/ai-analyst) into .agents/skills/analyst-core in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyst-core, .gemini/skills/analyst-core, .github/skills/analyst-core and .opencode/skills/analyst-core in your project.
Going by SKILL.md and its folder, Analyst Core needs the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Analyst Core is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Analyst Core: Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 206 stars), Eqtl Catalogue Region Fetch (ClawBio/ClawBio, 1.2k stars) and Raccoon Dataanalysis (SenseTime-Copilot/raccoon-dataanalysis-skill, 137 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.
Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.