Agent skill

Programmatic Eda

by nimrodfisher in nimrodfisher/data-analytics-skills

Systematic exploratory data analysis. An agent skill from nimrodfisher/data-analytics-skills.

MITAuto-check passedData & Analytics

Install Programmatic Eda

skills CLI
$ npx skills add nimrodfisher/data-analytics-skills --skill programmatic-eda -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nimrodfisher/data-analytics-skills programmatic-eda --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nimrodfisher/data-analytics-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/01-data-quality-validation/programmatic-eda .claude/skills/programmatic-eda && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
programmatic-eda
GitHub stars
465
Token cost
~582 tokens
SKILL.md length
258 words
Files
11 (incl. scripts, references, assets)
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Systematic exploratory data analysis. An agent skill from nimrodfisher/data-analytics-skills.

  • Works in 7 steps: Load and overview — run… → Null profile — run… → Outlier detection — run… → …
  • Tasks that involve Data analysis
  • Runs Python scripts from its folder
  • Tasks that involve Data cleaning

What it does

Programmatic Eda is an agent skill from nimrodfisher/data-analytics-skills. Systematic exploratory data analysis. Activate when a dataset needs profiling — structure check, nulls, outliers, distributions, correlations — before deeper analysis begins.

Its SKILL.md is about 580 tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts, reference files and assets (for example `assets/eda_report_template.md`, `assets/findings_summary.md` and `references/eda_checklist.md`).

It sits in Data & Analytics, covering Data analysis, Data cleaning and DataFrames. The repository describes itself as: A comprehensive list of Claude & Codex skills for a wide range of data analytics tasks. The licence is MIT.

When your agent uses it

  • Tasks that involve Data analysis
  • Tasks that involve Data cleaning
  • Tasks that involve DataFrames

Example prompts

  • “/programmatic-eda”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Load and overview — run scripts/data_overview.py to get row count, dtypes, memory usage, and a sample. Confirm grain (what one row…
  2. Null profile — run scripts/null_profiler.py; compare output against thresholds in references/quality_thresholds.md and flag columns above…
  3. Outlier detection — run scripts/outlier_detector.py (IQR + z-score) on numeric columns; document flagged values and decide: real signal or…
  4. Distribution summary — run scripts/distribution_summary.py for descriptive stats and univariate histograms on each numeric column.
  5. Correlation exploration — run scripts/correlation_explorer.py; flag pairs with |r| > 0.8 as potential multicollinearity or redundancy.
  6. EDA checklist sign-off — work through references/eda_checklist.md and confirm each item before declaring the dataset profiled.
  7. Write findings — fill assets/eda_report_template.md with full profiling output; distil top issues into assets/findings_summary.md.

What it can do on your machine

Read from SKILL.md and the folder at commit 9449d36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Programmatic Eda loads about 582 tokens when it runs, and up to ~2.9k if it reads all its reference files. Until then it costs about 48 tokens; SKILL.md has 258 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~582
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from nimrodfisher/data-analytics-skills at commit 9449d36, republished under its MIT licence (© nimrodfisher). 258 words, ~582 tokens.

Download SKILL.mdSave it as .claude/skills/programmatic-eda/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
programmatic-eda
description
Systematic exploratory data analysis. Activate when a dataset needs profiling — structure check, nulls, outliers, distributions, correlations — before deeper analysis begins.

When to use

  • You receive a new dataset and need to understand its shape and quality before analysis
  • An analysis produces surprising numbers and you want to verify the underlying data first
  • A stakeholder asks "is this data reliable?" or "what's in this table?"
  • You're about to run a model or statistical test and need data-quality assurance

Process

  1. Load and overview — run scripts/data_overview.py to get row count, dtypes, memory usage, and a sample. Confirm grain (what one row represents).
  2. Null profile — run scripts/null_profiler.py; compare output against thresholds in references/quality_thresholds.md and flag columns above limits.
  3. Outlier detection — run scripts/outlier_detector.py (IQR + z-score) on numeric columns; document flagged values and decide: real signal or data error?
  4. Distribution summary — run scripts/distribution_summary.py for descriptive stats and univariate histograms on each numeric column.
  5. Correlation exploration — run scripts/correlation_explorer.py; flag pairs with |r| > 0.8 as potential multicollinearity or redundancy.
  6. EDA checklist sign-off — work through references/eda_checklist.md and confirm each item before declaring the dataset profiled.
  7. Write findings — fill assets/eda_report_template.md with full profiling output; distil top issues into assets/findings_summary.md.

For pattern recipes (e.g. polars vs pandas equivalents, chunked reads for large files), see references/pandas_polars_recipes.md.

Inputs the skill needs

  • Required: dataset path (CSV / Parquet / Excel) or a DataFrame already in scope
  • Required: business context — what does one row represent?
  • Optional: quality threshold overrides (defaults in references/quality_thresholds.md)
  • Optional: columns to skip (PII, binary blobs, high-cardinality IDs)

Output

  • assets/eda_report_template.md (filled) — full profiling report with per-column stats
  • assets/findings_summary.md (filled) — top 3–5 quality issues and recommended next steps
  • Console output / plots from scripts for interactive inspection

© nimrodfisher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references, assets) in 01-data-quality-validation/programmatic-eda of nimrodfisher/data-analytics-skills.

  • SKILL.md
  • assets/eda_report_template.md
  • assets/findings_summary.md
  • references/eda_checklist.md
  • references/pandas_polars_recipes.md
  • references/quality_thresholds.md
  • scripts/correlation_explorer.py
  • scripts/data_overview.py
  • scripts/distribution_summary.py
  • scripts/null_profiler.py
  • scripts/outlier_detector.py

Open the folder on GitHubat commit 9449d36

Compare with similar skills

Programmatic Eda next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Programmatic Eda compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Programmatic Eda this skillnimrodfisher/data-analytics-skills465—~582Automated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent883—~2.5kAutomated safety check: PassApache-2.0
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Investigate Datawalkthru-earth/geocoding-playground153—~935Automated safety check: PassCC-BY-4.0

Similar skills

  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    883 GitHub stars~2.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Investigate Data

    walkthru-earth/geocoding-playground

    Investigates geocoder data quality issues by querying live S3 parquet files via MotherDuck MCP.

    153 GitHub stars~935 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Code Engineer

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

    148 GitHub stars~2.8k tokensUpdated 6 days ago
    Data & AnalyticsAuto-check passed

More from nimrodfisher/data-analytics-skills

All 31 skills in this repo
  • Ab Test Analysis

    nimrodfisher/data-analytics-skills

    Rigorous A/B test statistical analysis. An agent skill from nimrodfisher/data-analytics-skills.

    465 GitHub stars~708 tokensUpdated 12 days ago
    Auto-check passed
  • Analysis Assumptions Log

    nimrodfisher/data-analytics-skills

    Track and document analytical assumptions and decisions. An agent skill from nimrodfisher/data-analytics-skills.

    465 GitHub stars~578 tokensUpdated 12 days ago
    Auto-check passed
  • Analysis QA Checklist

    nimrodfisher/data-analytics-skills

    Pre-delivery quality assurance for analysis work. An agent skill from nimrodfisher/data-analytics-skills.

    465 GitHub stars~470 tokensUpdated 12 days ago
    Auto-check passed
  • Business Metrics Calculator

    nimrodfisher/data-analytics-skills

    Standard business metric calculation with industry benchmarks.

    465 GitHub stars~668 tokensUpdated 12 days ago
    Auto-check passed
  • Cohort Analysis

    nimrodfisher/data-analytics-skills

    Time-based cohort analysis with retention and behaviour tracking.

    465 GitHub stars~660 tokensUpdated 12 days ago
    Auto-check passed
  • Context Packager

    nimrodfisher/data-analytics-skills

    Efficiently package context for AI-assisted analysis. An agent skill from nimrodfisher/data-analytics-skills.

    465 GitHub stars~500 tokensUpdated 12 days ago
    Auto-check passed

Questions about Programmatic Eda

What does Programmatic Eda do?

Systematic exploratory data analysis. An agent skill from nimrodfisher/data-analytics-skills. Programmatic Eda is an agent skill from nimrodfisher/data-analytics-skills. Systematic exploratory data analysis.

When should I use Programmatic Eda?

Programmatic Eda fits situations like: tasks that involve Data analysis; tasks that involve Data cleaning; tasks that involve DataFrames.

How do I install Programmatic Eda in Claude Code?

Run `npx skills add nimrodfisher/data-analytics-skills --skill programmatic-eda -a claude-code`. Or copy the skill folder (01-data-quality-validation/programmatic-eda in nimrodfisher/data-analytics-skills) into .claude/skills/programmatic-eda in your project. Claude Code loads it when a task matches its description.

How do I install Programmatic Eda in Codex?

Run `npx skills add nimrodfisher/data-analytics-skills --skill programmatic-eda -a codex`. Or copy the skill folder (01-data-quality-validation/programmatic-eda in nimrodfisher/data-analytics-skills) into .agents/skills/programmatic-eda in your project. Codex loads it when a task matches its description.

Can I use Programmatic Eda in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nimrodfisher/data-analytics-skills --skill programmatic-eda -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/programmatic-eda, .gemini/skills/programmatic-eda, .github/skills/programmatic-eda and .opencode/skills/programmatic-eda in your project.

What does Programmatic Eda need to run?

Going by SKILL.md and its folder, Programmatic Eda needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Programmatic Eda access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Programmatic Eda safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Programmatic Eda use?

Programmatic Eda is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Programmatic Eda use?

About 582 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Programmatic Eda?

Skills that share tags, products or a category with Programmatic Eda: Pandas Pro (Jeffallan/claude-skills, 12k stars), Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars), Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 883 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Programmatic Eda?

nimrodfisher (a GitHub user) maintains it in nimrodfisher/data-analytics-skills, which has 465 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on September 25, 2026.

Source: nimrodfisher/data-analytics-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.