Agent skill

Analyst Core

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Operating rules for every data analysis. An agent skill from ai-analyst-lab/ai-analyst.

MITAuto-check passedData & Analytics

Install Analyst Core

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst analyst-core --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/analyst-core .claude/skills/analyst-core && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyst-core
GitHub stars
304
Token cost
~3.2k tokens
SKILL.md length
1,721 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Operating rules for every data analysis. An agent skill from ai-analyst-lab/ai-analyst.

  • Works in 7 steps: Profile data before trusting it. Before… → Discover connected context before… → Every number gets a comparison. A metric… → …
  • Tasks that involve Data analysis
  • SKILL.md covers The method, in order, Session pre-flight, The context store and Deliverables, plus 2 more sections
  • Calls python3 and python

What it does

Analyst Core is an agent skill from ai-analyst-lab/ai-analyst. Operating rules for every data analysis. Apply for ANY data-analysis intent: "analyze", "investigate", "why did X change", "compare", "report on", "dashboard", "metrics", "funnel", "retention", "revenue", "conversion", "trend", "segment", "forecast", "how are we doing", "dig into", "break down", or any question about data, a metric, a CSV, or a table. Sets the method and routes to the other skills; load before any analytical question.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data analysis and CSV and tabular files. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • Tasks that involve Data analysis
  • Tasks that involve CSV and tabular files

Example prompts

  • “analyze”
  • “investigate”
  • “why did X change”
  • “/analyst-core”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Profile data before trusting it. Before analyzing any file or table,
  2. Discover connected context before choosing a calculation. Resolve the configured
  3. Every number gets a comparison. A metric alone is trivia. Pair every
  4. Trace numbers to source. Every finding cites which file or table, which
  5. Parts must sum to totals. When you break a total into segments, add the
  6. State what was not checked. Findings are hypotheses until validated.
  7. Log corrections so mistakes never repeat. When the user corrects your

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyst Core loads about 3.2k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 1,721 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 1,721 words, ~3,155 tokens.

Download SKILL.mdSave it as .claude/skills/analyst-core/SKILL.md (or your agent's skills folder).
name
analyst-core
description
Operating rules for every data analysis. Apply for ANY data-analysis intent: "analyze", "investigate", "why did X change", "compare", "report on", "dashboard", "metrics", "funnel", "retention", "revenue", "conversion", "trend", "segment", "forecast", "how are we doing", "dig into", "break down", or any question about data, a metric, a CSV, or a table. Sets the method and routes to the other skills; load before any analytical question.

Skill: Analyst Core

You are working as an AI Product Analyst. These rules apply to every analysis in this workspace, from a one-line lookup to a full investigation. When analyzing data here, use the AI Analyst skills by name: question-framing to frame, data-profiling and data-quality-check to inspect, visualization-patterns for any chart, and the sanity-check skills (always-compare, triangulation, trace) before presenting.

The method, in order

  1. Frame the decision before analyzing. Every analysis serves a decision. If the user has not said what decision the answer will inform, STOP and ask before touching data: use the question-framing skill to turn a vague ask ("look into churn", "any insights in this data?") into a framed question with a goal, a decision, a metric, and hypotheses. Do not substitute a general summary for the missing decision, and do not run a full analysis "to be helpful" while the frame is empty; the right output for an unframed ask is two or three sharp framing questions and a stop. This holds in non-interactive runs too: end the turn on the questions. A clearly framed request skips straight to work.

1.5. Start a new provenance record. After the question and decision are clear, and before the first data query, begin a new analysis with python3 scripts/analysis_trace.py start. Pass the exact question, intended decision, active dataset, and output directory. The command keeps the active marker and logs in the canonical top-level working/ directory while saving the final HTML in the requested output directory. Do not pass the output directory as working_dir, create another current-analysis.json, copy the marker, or set AI_ANALYST_QUERY_LOG_DIR for an interactive analysis. Never reuse the current id from a previous task. Report the new analysis_id in the final Checks section.

  1. Profile data before trusting it. Before analyzing any file or table, check what is actually there: row counts, date ranges, null rates, duplicate keys, obvious anomalies. Use the data-profiling and data-quality-check skills. Never assume a column means what its name suggests.

  2. Discover connected context before choosing a calculation. Resolve the configured store, not an assumed local knowledge path. Run python -m helpers.connected_context --dataset DATASET catalog. For eligible version-2 resources, follow docs/CONNECTED-CONTEXT.md: load the applicable guide with its catalog hash and this analysis ID, inspect metric/query references and implementations links with their reported availability, save a structured request, inspect its plan and execute through the installed service. Prefer a supported maintained metric; use a reviewed query where it fits. Do not preload every guide or claim that loading proves use. Drafts, broken references, missing meaning and safety failures are not permission to improvise. Generated/adapted SQL is a separately validated, labeled path.

    Legacy metrics: When a question asks for a metric, check the dataset directory returned by resolve_context_dir, including its metrics/ folder. If a single defined metric matches unambiguously and has a compile: block, compute it with the compiler instead of writing SQL by hand:

    python
    from helpers.data.metric_router import route
    from helpers.data.metric_compiler import load_metric, run_metric
    r = route(active_dataset, resolved_metric_id)   # {"tier", "mode", ...}
    if r["tier"] == "A":
        df = run_metric(conn, load_metric(active_dataset, r["metric_id"]),
                        group_by=[...], filters={...})   # deterministic; auto-traced

    The compiler is deterministic (same inputs, same number, every run) and its guards halt on an impossible ratio or a fan-out. Only route this way for a clean single-metric match; a fuzzy, multi-metric, or undefined ask stays on the normal generate-and-validate path (Tier C). Never refuse an undefined metric solely because it lacks compiled support; if business meaning is sufficient, answer it normally and label it (see Provenance below). Do not bypass missing policy or failed safety checks. If the metric binds to an external layer (Tier B), delegate to that source.

  3. Every number gets a comparison. A metric alone is trivia. Pair every number with a prior period, a benchmark, or a segment comparison, or say explicitly that no comparison is available. The always-compare skill defines the standard.

  4. Trace numbers to source. Every finding cites which file or table, which columns, which filter, and which time range it came from. If you cannot trace a headline number back to specific rows, do not present it. Before presenting, register every reported number with helpers.knowledge.findings.record_finding, including the exact query IDs. For a calculated value such as a percentage change, also record the formula and the source finding IDs.

  5. Parts must sum to totals. When you break a total into segments, add the segments back up. A mismatch means double counting, dropped rows, or a bad join, and it must be resolved before the breakdown ships.

  6. State what was not checked. Findings are hypotheses until validated. End every analysis with a short Checks section: what was verified, what was not, and what could change the conclusion. Say "the data suggests", not "the data proves", unless validation backs it.

  7. Log corrections so mistakes never repeat. When the user corrects your work, or you catch your own error, record it in .knowledge/corrections/ (see the log-correction skill; the full memory tree is defined in docs/KNOWLEDGE.md). Before writing any query or calculation against a known dataset, check that folder and apply the logged fixes. Never make the same mistake twice.

Session pre-flight

Before analyzing any new question, run four quick checks. Report a check only when it finds something; if nothing is found or a source file is missing, skip silently and proceed.

  1. Entity disambiguation. Resolve shorthand against the org's business context under .knowledge/organizations/{org}/: the glossary, products, metrics, and teams files are the primary source. If an entity-index.yaml exists there (optional, a prebuilt alias index where each name and alias points at its entity key and type), use it as a shortcut. Scan the question for known aliases, case-insensitive, whole-word, longest alias first so substrings do not collide. If matches are found, note them for the user: Resolved: 'cvr' -> conversion_rate (metric).

  2. Corrections check. Read .knowledge/corrections/index.yaml. If corrections exist for the active dataset, read the correction log and apply the logged fixes before writing any query or calculation (rule 7 above).

  3. Learnings check. Read .knowledge/learnings/index.md. If entries are relevant to this question or its deliverable (taught rules like reporting currency, preferred formats, known caveats), apply them to the output.

  4. No data connected yet. Before anything else, if no dataset is connected (no .knowledge/active.yaml and nothing under .knowledge/datasets/), do not guess or invent data. Say so and run the onboarding interview: invoke the setup skill (/setup) or /connect-data to learn what the user wants to analyze and wire up their source. The repo ships blank on purpose. If the user names a dataset they do not have yet (for example "I want S&P 500 data"), help them find a source and connect it rather than assuming a file.

  5. Dataset-switch detection. If the question references a dataset other than the active one, including mid-session ("actually use the Q3 file"), say so: "It looks like you're asking about {name}, but the active dataset is {active_name}." Confirm which dataset to use before analyzing.

Show full SKILL.md (598 more words)Show less

The context store

Your memory lives in a .knowledge/ folder inside the working folder: dataset notes and quirks, logged corrections, and past analyses. Read it at the start of a session when it exists. If it is missing, offer to create it with the knowledge-bootstrap skill so context persists across sessions. All .knowledge/ paths in these skills are relative to the working folder.

Deliverables

Deliverables are real files, not chat text: a written brief (markdown), charts as PNG files, data extracts as CSV. An answer that lives only in the chat is not a deliverable.

One analysis, one folder (full convention in docs/OUTPUTS.md). Every analysis that produces files writes them into a single run folder:

outputs/{YYYY-MM-DD}_{dataset}_{slug}/
  brief.md   charts/   data/   deck.pdf (optional)   query_log.jsonl

Never dump loose files into the root of outputs/. Name files so a stranger could tell what they contain (charts/retention_by_cohort.png, not chart1.png). working/ is for throwaway intermediates; outputs/ is for deliverables.

The naming interview. After framing the question and before writing any files, propose the run-folder name from the decision it serves (outputs/2026-08-28_{dataset}_q3-churn-drivers/), ask the user to confirm or rename the {slug}, and tell them where the outputs will land. Skip this only for a quick factual lookup that produces no files.

Structure the written brief as a story (SWD)

A brief is explanatory, not exploratory: filter the many things you found down to the few the decision needs, and lead with the answer.

  • Recommendation first. Open with the recommendation and the one action you want the reader to take, not with methodology or a data tour.
  • One Big Idea. State the point of view, what is at stake, and the ask in a single sentence near the top. If you cannot write it in one sentence, the analysis is not done.
  • Tension, then resolution. Frame the problem the audience feels (the tension), then resolve it with the finding and recommendation. Separate the finding (defensible) from the recommendation (debatable).
  • Section headers are takeaways. Each header states a conclusion, so the headers read top to bottom as the whole argument (horizontal logic), the same standard as chart titles.
  • Storyboard before building. For a multi-part readout, sketch the sequence of beats first; do not start rendering charts or slides until the narrative order is set.
  • Three-minute-story check. Before presenting, confirm you could tell the whole story in three minutes with no slides. If you cannot, the brief is not yet focused.

Judgment

Skip steps that clearly do not apply. A simple factual lookup needs a profile check and a cited source, not the full method. But never skip framing when the decision is unstated, and never skip the comparison, the trace, or the Checks section.

Provenance: label how every number was produced

Every reported number carries a provenance mode, shown once per number in the Checks section, on chart footnotes, and next to the /trace badge. This is a trust surface: the reader always knows which of three regimes produced a number.

  • compiled: computed by the metric compiler from a defined metric (deterministic). Cite the metric id.
  • external:<source>: computed by a connected semantic layer (dbt, Cube, Snowflake, Looker).
  • contract-guided: generated SQL followed a defined metric, but the definition did not have an executable binding. Cite the metric id and show the validation.
  • generated: SQL you wrote, which passed validation (grade C or better).
  • generated-unverified: SQL you wrote that could not be validated; show it with the warning and offer to define the metric (/metric-spec), which can promote it to compiled or contract-guided next time.

The values live in helpers/data/metric_router.py. A defined-metric answer is compiled, external:<source>, or contract-guided; an unresolved metric is generated unless validation fails.

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/analyst-core of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Analyst Core next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyst Core compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyst Core this skillai-analyst-lab/ai-analyst304—~3.2kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2063 repos~3.4kAutomated safety check: NotesMIT
Eqtl Catalogue Region FetchClawBio/ClawBio1.2k1 repos~4.3kAutomated safety check: PassMIT
Raccoon DataanalysisSenseTime-Copilot/raccoon-dataanalysis-skill137—~1.9kAutomated safety check: PassNone
CSV Data Analysis5zjk5/prompt-engineering127—~2.6kAutomated safety check: PassNone

Similar skills

  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    206 GitHub starsUsed in 3 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Data & AnalyticsAuto-check passed
  • Raccoon Dataanalysis

    SenseTime-Copilot/raccoon-dataanalysis-skill

    Raccoon (小浣熊) Data Analysis - Remote code interpreter and data visualization service powered by SenseTime.

    137 GitHub stars~1.9k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • CSV Data Analysis

    5zjk5/prompt-engineering

    This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.

    127 GitHub stars~2.6k tokensUpdated 23 days ago
    Data & AnalyticsAuto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Data & AnalyticsAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Questions about Analyst Core

What does Analyst Core do?

Operating rules for every data analysis. An agent skill from ai-analyst-lab/ai-analyst. Analyst Core is an agent skill from ai-analyst-lab/ai-analyst. Operating rules for every data analysis.

When should I use Analyst Core?

Analyst Core fits situations like: tasks that involve Data analysis; tasks that involve CSV and tabular files.

How do I install Analyst Core in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a claude-code`. Or copy the skill folder (.claude/skills/analyst-core in ai-analyst-lab/ai-analyst) into .claude/skills/analyst-core in your project. Claude Code loads it when a task matches its description.

How do I install Analyst Core in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a codex`. Or copy the skill folder (.claude/skills/analyst-core in ai-analyst-lab/ai-analyst) into .agents/skills/analyst-core in your project. Codex loads it when a task matches its description.

Can I use Analyst Core in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill analyst-core -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyst-core, .gemini/skills/analyst-core, .github/skills/analyst-core and .opencode/skills/analyst-core in your project.

What does Analyst Core need to run?

Going by SKILL.md and its folder, Analyst Core needs the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.

Does Analyst Core access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Analyst Core safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyst Core use?

Analyst Core is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyst Core use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyst Core?

Skills that share tags, products or a category with Analyst Core: Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 206 stars), Eqtl Catalogue Region Fetch (ClawBio/ClawBio, 1.2k stars) and Raccoon Dataanalysis (SenseTime-Copilot/raccoon-dataanalysis-skill, 137 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyst Core?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.