Agent skill

Data Exploration

by w95 in w95/awesome-claude-corporate-skills

Profile and explore datasets to understand their shape, quality, and patterns before analysis.

MITAuto-check passedData & Analytics

Install Data Exploration

skills CLI
$ npx skills add w95/awesome-claude-corporate-skills --skill data-exploration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install w95/awesome-claude-corporate-skills data-exploration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/w95/awesome-claude-corporate-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/10-data-analytics/data-exploration .claude/skills/data-exploration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-exploration
GitHub stars
235
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
678 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Profile and explore datasets to understand their shape, quality, and patterns before analysis.

  • Works in 3 steps: Structural Understanding → Column-Level Profiling → Relationship Discovery
  • Encountering a new dataset
  • SKILL.md covers Data Profiling Methodology, Quality Assessment Framework, Pattern Discovery Techniques and Schema Understanding and…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Exploration is an agent skill from w95/awesome-claude-corporate-skills. Profile and explore datasets to understand their shape, quality, and patterns before analysis. Use when encountering a new dataset, assessing data quality, discovering column distributions, identifying nulls and outliers, or deciding which dimensions to analyze.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data analysis and Data cleaning. The repository describes itself as: 166 production-ready Claude AI skills organized by corporate role — executive leadership, finance, HR, marketing, sales, legal, operations, engineering, product, data, customer…. The licence is MIT.

When your agent uses it

  • Encountering a new dataset
  • Assessing data quality
  • Discovering column distributions
  • Identifying nulls and outliers

Example prompts

  • “/data-exploration”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Structural Understanding
  2. Column-Level Profiling
  3. Relationship Discovery

What it can do on your machine

Read from SKILL.md and the folder at commit 78dbc7c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown and sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Exploration loads about 2k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 678 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from w95/awesome-claude-corporate-skills at commit 78dbc7c, republished under its MIT licence (© w95). 678 words, ~1,974 tokens.

Download SKILL.mdSave it as .claude/skills/data-exploration/SKILL.md (or your agent's skills folder).
name
data-exploration
description
Profile and explore datasets to understand their shape, quality, and patterns before analysis. Use when encountering a new dataset, assessing data quality, discovering column distributions, identifying nulls and outliers, or deciding which dimensions to analyze.

Data Exploration Skill

Systematic methodology for profiling datasets, assessing data quality, discovering patterns, and understanding schemas.

Data Profiling Methodology

Phase 1: Structural Understanding

Before analyzing any data, understand its structure:

Table-level questions:

  • How many rows and columns?
  • What is the grain (one row per what)?
  • What is the primary key? Is it unique?
  • When was the data last updated?
  • How far back does the data go?

Column classification: Categorize each column as one of:

  • Identifier: Unique keys, foreign keys, entity IDs
  • Dimension: Categorical attributes for grouping/filtering (status, type, region, category)
  • Metric: Quantitative values for measurement (revenue, count, duration, score)
  • Temporal: Dates and timestamps (created_at, updated_at, event_date)
  • Text: Free-form text fields (description, notes, name)
  • Boolean: True/false flags
  • Structural: JSON, arrays, nested structures
Phase 2: Column-Level Profiling

For each column, compute:

All columns:

  • Null count and null rate
  • Distinct count and cardinality ratio (distinct / total)
  • Most common values (top 5-10 with frequencies)
  • Least common values (bottom 5 to spot anomalies)

Numeric columns (metrics):

min, max, mean, median (p50)
standard deviation
percentiles: p1, p5, p25, p75, p95, p99
zero count
negative count (if unexpected)

String columns (dimensions, text):

min length, max length, avg length
empty string count
pattern analysis (do values follow a format?)
case consistency (all upper, all lower, mixed?)
leading/trailing whitespace count

Date/timestamp columns:

min date, max date
null dates
future dates (if unexpected)
distribution by month/week
gaps in time series

Boolean columns:

true count, false count, null count
true rate
Phase 3: Relationship Discovery

After profiling individual columns:

  • Foreign key candidates: ID columns that might link to other tables
  • Hierarchies: Columns that form natural drill-down paths (country > state > city)
  • Correlations: Numeric columns that move together
  • Derived columns: Columns that appear to be computed from others
  • Redundant columns: Columns with identical or near-identical information

Quality Assessment Framework

Completeness Score

Rate each column:

  • Complete (>99% non-null): Green
  • Mostly complete (95-99%): Yellow -- investigate the nulls
  • Incomplete (80-95%): Orange -- understand why and whether it matters
  • Sparse (<80%): Red -- may not be usable without imputation
Consistency Checks

Look for:

  • Value format inconsistency: Same concept represented differently ("USA", "US", "United States", "us")
  • Type inconsistency: Numbers stored as strings, dates in various formats
  • Referential integrity: Foreign keys that don't match any parent record
  • Business rule violations: Negative quantities, end dates before start dates, percentages > 100
  • Cross-column consistency: Status = "completed" but completed_at is null
Accuracy Indicators

Red flags that suggest accuracy issues:

  • Placeholder values: 0, -1, 999999, "N/A", "TBD", "test", "xxx"
  • Default values: Suspiciously high frequency of a single value
  • Stale data: Updated_at shows no recent changes in an active system
  • Impossible values: Ages > 150, dates in the far future, negative durations
  • Round number bias: All values ending in 0 or 5 (suggests estimation, not measurement)
Timeliness Assessment
  • When was the table last updated?
  • What is the expected update frequency?
  • Is there a lag between event time and load time?
  • Are there gaps in the time series?
Show full SKILL.md (270 more words)Show less

Pattern Discovery Techniques

Distribution Analysis

For numeric columns, characterize the distribution:

  • Normal: Mean and median are close, bell-shaped
  • Skewed right: Long tail of high values (common for revenue, session duration)
  • Skewed left: Long tail of low values (less common)
  • Bimodal: Two peaks (suggests two distinct populations)
  • Power law: Few very large values, many small ones (common for user activity)
  • Uniform: Roughly equal frequency across range (often synthetic or random)
Temporal Patterns

For time series data, look for:

  • Trend: Sustained upward or downward movement
  • Seasonality: Repeating patterns (weekly, monthly, quarterly, annual)
  • Day-of-week effects: Weekday vs. weekend differences
  • Holiday effects: Drops or spikes around known holidays
  • Change points: Sudden shifts in level or trend
  • Anomalies: Individual data points that break the pattern
Segmentation Discovery

Identify natural segments by:

  • Finding categorical columns with 3-20 distinct values
  • Comparing metric distributions across segment values
  • Looking for segments with significantly different behavior
  • Testing whether segments are homogeneous or contain sub-segments
Correlation Exploration

Between numeric columns:

  • Compute correlation matrix for all metric pairs
  • Flag strong correlations (|r| > 0.7) for investigation
  • Note: Correlation does not imply causation -- flag this explicitly
  • Check for non-linear relationships (e.g., quadratic, logarithmic)

Schema Understanding and Documentation

Schema Documentation Template

When documenting a dataset for team use:

markdown
## Table: [schema.table_name]

**Description**: [What this table represents]
**Grain**: [One row per...]
**Primary Key**: [column(s)]
**Row Count**: [approximate, with date]
**Update Frequency**: [real-time / hourly / daily / weekly]
**Owner**: [team or person responsible]

### Key Columns

| Column | Type | Description | Example Values | Notes |
|--------|------|-------------|----------------|-------|
| user_id | STRING | Unique user identifier | "usr_abc123" | FK to users.id |
| event_type | STRING | Type of event | "click", "view", "purchase" | 15 distinct values |
| revenue | DECIMAL | Transaction revenue in USD | 29.99, 149.00 | Null for non-purchase events |
| created_at | TIMESTAMP | When the event occurred | 2024-01-15 14:23:01 | Partitioned on this column |

### Relationships
- Joins to `users` on `user_id`
- Joins to `products` on `product_id`
- Parent of `event_details` (1:many on event_id)

### Known Issues
- [List any known data quality issues]
- [Note any gotchas for analysts]

### Common Query Patterns
- [Typical use cases for this table]
Schema Exploration Queries

When connected to a data warehouse, use these patterns to discover schema:

sql
-- List all tables in a schema (PostgreSQL)
SELECT table_name, table_type
FROM information_schema.tables
WHERE table_schema = 'public'
ORDER BY table_name;

-- Column details (PostgreSQL)
SELECT column_name, data_type, is_nullable, column_default
FROM information_schema.columns
WHERE table_name = 'my_table'
ORDER BY ordinal_position;

-- Table sizes (PostgreSQL)
SELECT relname, pg_size_pretty(pg_total_relation_size(relid))
FROM pg_catalog.pg_statio_user_tables
ORDER BY pg_total_relation_size(relid) DESC;

-- Row counts for all tables (general pattern)
-- Run per-table: SELECT COUNT(*) FROM table_name
Lineage and Dependencies

When exploring an unfamiliar data environment:

  1. Start with the "output" tables (what reports or dashboards consume)
  2. Trace upstream: What tables feed into them?
  3. Identify raw/staging/mart layers
  4. Map the transformation chain from raw data to analytical tables
  5. Note where data is enriched, filtered, or aggregated

© w95, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in 10-data-analytics/data-exploration of w95/awesome-claude-corporate-skills.

Open the folder on GitHubat commit 78dbc7c

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in w95/awesome-claude-corporate-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Data Exploration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Exploration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Exploration this skillw95/awesome-claude-corporate-skills2351 repos~2kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Code EngineeropenJiuwen-ai/sciencediscovery148—~2.8kAutomated safety check: PassApache-2.0
Data Analysisxiaoyuge886/aigc1981 repos~794Automated safety check: PassMIT
Math Modeling Data Cleaning and Chartsyushui2022/MathModel-Skill4521 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Code Engineer

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

    148 GitHub stars~2.8k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Data Analysis

    xiaoyuge886/aigc

    Perform data analysis tasks including data cleaning, statistical analysis, visualization, and insight generation.

    198 GitHub starsUsed in 1 repo~794 tokens
    Data & AnalyticsAuto-check passed
  • Math Modeling Data Cleaning and Charts

    yushui2022/MathModel-Skill

    Cleans raw or scraped competition data and produces exploratory charts and a figure plan as one stage of a mathematical modeling paper workflow.

    452 GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Data Analyst

    ailabs-393/ai-labs-claude-skills

    This skill should be used when analyzing CSV datasets, handling missing values through intelligent imputation, and creating interactive dashboards to visualize data trends.

    454 GitHub stars~2.7k tokensUpdated 11 mo ago
    Data & AnalyticsAuto-check passed

More from w95/awesome-claude-corporate-skills

All 43 skills in this repo
  • Data Context Extractor

    w95/awesome-claude-corporate-skills

    Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts.

    235 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Competitive Analysis

    w95/awesome-claude-corporate-skills

    Framework for competitive landscape analysis across any industry.

    235 GitHub starsUsed in 1 repo~4k tokens
    Auto-check passed
  • Account Research

    w95/awesome-claude-corporate-skills

    Research a company using Common Room data. An agent skill from w95/awesome-claude-corporate-skills.

    235 GitHub stars~1.5k tokensUpdated 7 mo ago
    Auto-check passed
  • Call Prep

    w95/awesome-claude-corporate-skills

    Prepare for a customer or prospect call using Common Room signals.

    235 GitHub stars~1.5k tokensUpdated 7 mo ago
    Auto-check passed
  • Compose Outreach

    w95/awesome-claude-corporate-skills

    Generate personalized outreach messages using Common Room signals.

    235 GitHub stars~1.4k tokensUpdated 7 mo ago
    Auto-check passed
  • SQL Queries

    w95/awesome-claude-corporate-skills

    Write correct, performant SQL across all major data warehouse dialects (Snowflake, BigQuery, Databricks, PostgreSQL, etc.).

    235 GitHub starsUsed in 3 repos~2.8k tokens
    Auto-check passed

Questions about Data Exploration

What does Data Exploration do?

Profile and explore datasets to understand their shape, quality, and patterns before analysis. Data Exploration is an agent skill from w95/awesome-claude-corporate-skills. Profile and explore datasets to understand their shape, quality, and patterns before analysis.

When should I use Data Exploration?

Data Exploration fits situations like: encountering a new dataset; assessing data quality; discovering column distributions; identifying nulls and outliers.

How do I install Data Exploration in Claude Code?

Run `npx skills add w95/awesome-claude-corporate-skills --skill data-exploration -a claude-code`. Or copy the skill folder (10-data-analytics/data-exploration in w95/awesome-claude-corporate-skills) into .claude/skills/data-exploration in your project. Claude Code loads it when a task matches its description.

How do I install Data Exploration in Codex?

Run `npx skills add w95/awesome-claude-corporate-skills --skill data-exploration -a codex`. Or copy the skill folder (10-data-analytics/data-exploration in w95/awesome-claude-corporate-skills) into .agents/skills/data-exploration in your project. Codex loads it when a task matches its description.

Can I use Data Exploration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add w95/awesome-claude-corporate-skills --skill data-exploration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-exploration, .gemini/skills/data-exploration, .github/skills/data-exploration and .opencode/skills/data-exploration in your project.

What does Data Exploration need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Exploration is instructions for the agent only.

Does Data Exploration access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Exploration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Exploration use?

Data Exploration is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Exploration use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Exploration?

Skills that share tags, products or a category with Data Exploration: Pandas Pro (Jeffallan/claude-skills, 12k stars), Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars), Code Engineer (openJiuwen-ai/sciencediscovery, 148 stars) and Data Analysis (xiaoyuge886/aigc, 198 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Exploration?

w95 (a GitHub user) maintains it in w95/awesome-claude-corporate-skills, which has 235 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on February 26, 2026.

Source: w95/awesome-claude-corporate-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.