Agent skill

Data Profiling

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Deep-profile the active dataset: distributions, temporal patterns, correlations, completeness gaps, anomalies.

MITAuto-check passedData & Analytics

Install Data Profiling

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill data-profiling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst data-profiling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/data-profiling .claude/skills/data-profiling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-profiling
GitHub stars
304
Token cost
~2.4k tokens
SKILL.md length
662 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Deep-profile the active dataset: distributions, temporal patterns, correlations, completeness gaps, anomalies.

  • Works in 4 steps: Connect and Profile Schema → Run Deep Profiling per Table → Correlation and Anomaly Analysis on Key… → …
  • Profile this data
  • SKILL.md covers Purpose, When to Use, Instructions and Output Format, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Profiling is an agent skill from ai-analyst-lab/ai-analyst. Deep-profile the active dataset: distributions, temporal patterns, correlations, completeness gaps, anomalies. Use after connecting a new dataset. Trigger on "profile this data", "deep-profile the dataset", "run a data profile", "check distributions", "find anomalies in the data", "how complete is this data". For a first-contact overview use data-map; for one column use distribution-profiler; for schema use data-inspect.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data analysis and Performance optimization. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • Profile this data
  • Deep-profile the dataset
  • Run a data profile
  • Check distributions

Example prompts

  • “profile this data”
  • “deep-profile the dataset”
  • “run a data profile”
  • “/data-profiling”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Connect and Profile Schema
  2. Run Deep Profiling per Table
  3. Correlation and Anomaly Analysis on Key Tables
  4. Generate Profile Report

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Profiling loads about 2.4k tokens when it runs. Until then it costs about 110 tokens; SKILL.md has 662 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 662 words, ~2,431 tokens.

Download SKILL.mdSave it as .claude/skills/data-profiling/SKILL.md (or your agent's skills folder).
name
data-profiling
description
Deep-profile the active dataset: distributions, temporal patterns, correlations, completeness gaps, anomalies. Use after connecting a new dataset. Trigger on "profile this data", "deep-profile the dataset", "run a data profile", "check distributions", "find anomalies in the data", "how complete is this data". For a first-contact overview use data-map; for one column use distribution-profiler; for schema use data-inspect.

Skill: Data Profiling

Purpose

Deep-profile the active dataset to understand schema structure, value distributions, temporal patterns, correlations, completeness gaps, and anomalies. Produces a comprehensive profile report that serves as the foundation for analysis planning and data quality assessment.

When to Use

  • After connecting a new dataset (post-bootstrap, pre-analysis)
  • Before the first analysis on any dataset
  • When explicitly invoked by the user
  • When the existing profile is stale (check last_profiled in manifest.yaml)

DISAMBIGUATION: this is the DEEP statistical profile (distributions, correlations, anomalies). For cross-table relationships/health and the first-contact "tell me about this data" overview, use data-map; for a single column's distribution, use distribution-profiler; for a plain schema listing, use /data (data-inspect).

Instructions

Step 1: Connect and Profile Schema
python
from helpers.data.data_helpers import get_connection_for_profiling
from helpers.data.schema_profiler import profile_source

# Get connection (auto-detects DuckDB vs CSV from active dataset)
conn_info = get_connection_for_profiling()

# Run full schema profile — introspects all tables: column names, types,
# nullability, row counts, sample values, basic statistics, date detection
schema = profile_source(conn_info)

Record the output. schema contains the full table inventory with column-level metadata. Use this to identify:

  • Which tables exist and their row counts
  • Which columns are date columns (for temporal analysis in Step 2)
  • Which columns are numeric (for distribution and correlation analysis)
  • Which columns have nulls (for completeness deep-dive in Step 2)
Step 2: Run Deep Profiling per Table

For each table in the schema, load the data and run the deep profiling functions. Prioritize tables with the most rows and the most date/numeric columns.

python
from helpers.data.data_helpers import read_table
from helpers.data.deep_profiler import (
    profile_distributions,
    profile_temporal_patterns,
    profile_completeness,
)

for table_info in schema["tables"]:
    table_name = table_info["name"]
    df = read_table(table_name)

    # Distribution analysis on all numeric columns
    distributions = profile_distributions(df)

    # Completeness assessment — null rates, zeros, empty strings, constant cols
    completeness = profile_completeness(df)

    # Temporal pattern analysis (only if the table has date columns)
    temporal = None
    if table_info.get("date_columns"):
        primary_date = table_info["date_columns"][0]
        temporal = profile_temporal_patterns(df, primary_date, freq="D")

Important: For large tables (>50K rows), profile_source() already samples. But read_table() loads the full CSV. If a table has >100K rows, sample before running deep profiling:

python
if len(df) > 100_000:
    df = df.sample(n=100_000, random_state=42)
Step 3: Correlation and Anomaly Analysis on Key Tables

Run correlation and anomaly detection on tables that contain key business metrics (revenue, counts, rates). Identify these tables by looking for columns with names like revenue, amount, total, count, rate, price, quantity.

python
from helpers.data.deep_profiler import profile_correlations, profile_anomalies

# Correlations — find relationships between numeric columns
correlations = profile_correlations(df, threshold=0.5)

# Anomaly detection — requires a date column and pre-aggregated data
# Aggregate to daily granularity first if the table has event-level rows
if table_info.get("date_columns"):
    primary_date = table_info["date_columns"][0]
    # Only run on tables with a clear date + metric pattern
    metric_cols = [c for c in df.select_dtypes(include="number").columns
                   if c not in ("id", table_name.rstrip("s") + "_id")]
    if metric_cols:
        # Aggregate to daily for anomaly detection
        daily = df.groupby(pd.to_datetime(df[primary_date]).dt.date)[metric_cols].sum().reset_index()
        daily.rename(columns={daily.columns[0]: primary_date}, inplace=True)
        anomalies = profile_anomalies(daily, date_col=primary_date,
                                       metric_cols=metric_cols, window=14)
Step 4: Generate Profile Report

Write the full profile report to .knowledge/datasets/{active}/last_profile.md. Use schema_to_markdown() for the schema portion, then append the deep profiling results.

python
from helpers.data.data_helpers import schema_to_markdown, detect_active_source

source = detect_active_source()
active_dataset = source["source"]

# Build the schema markdown section
schema_md = schema_to_markdown(schema)

Assemble the full report and write it to:

.knowledge/datasets/{active_dataset}/last_profile.md

Output Format

markdown
# Data Profile: {dataset_name}
**Profiled at:** {ISO timestamp}
**Source:** {connection type} ({path or schema prefix})
**Tables:** {count}  |  **Total rows:** {sum}

---

## Summary of Findings

| Severity | Count | Details |
|----------|-------|---------|
| BLOCKER  | X     | {brief list} |
| WARNING  | X     | {brief list} |
| INFO     | X     | {brief list} |

---

## Schema Overview

{output of schema_to_markdown()}

---

## Distribution Analysis

### {table_name}

| Column | Shape | Skewness | Outliers (IQR) | Recommended Transform |
|--------|-------|----------|----------------|----------------------|
| {col}  | {shape} | {skew} | {n_outliers}  | {transform or "none"} |

---

## Temporal Patterns

### {table_name} ({date_column})

- **Date range:** {min} to {max}
- **Coverage:** {actual}/{expected} periods ({pct}%)
- **Gaps:** {count} gaps found {list if any}
- **Trend:** {trend direction}
- **Seasonality:** {detected or not}
- **Day-of-week pattern:** {summary}

---

## Completeness

### {table_name}

| Column | Status | Null % | Zeros | Empty Strings | Constant? |
|--------|--------|--------|-------|---------------|-----------|
| {col}  | {status} | {pct} | {count} | {count}    | {yes/no}  |

---

## Correlations

### {table_name}

| Column A | Column B | Correlation | Strength | Direction |
|----------|----------|-------------|----------|-----------|
| {col_a}  | {col_b}  | {r}         | {strength} | {direction} |

---

## Anomalies

### {table_name}

{anomaly summary}

| Metric | Spikes | Drops | Details |
|--------|--------|-------|---------|
| {metric} | {count} | {count} | {top anomalies with dates} |

---

## Recommendations

- **BLOCKER items:** {must fix before analysis}
- **WARNING items:** {note as caveats}
- **Suggested analysis focus:** {tables/columns with most signal}
Severity Classification

Apply these rules consistently across all sections:

SeverityCondition
BLOCKER>50% nulls in a key metric column; entire date ranges missing (coverage <50%); constant columns that should have variance; very strong correlations (r>0.95) suggesting duplicate columns
WARNING5-50% nulls; heavy-tailed or bimodal distributions in metric columns; date coverage 50-90%; moderate anomalies detected; skewness >3 suggesting data quality issues
INFO<5% nulls; normal or mild skew distributions; full date coverage; no anomalies; expected correlations (e.g., quantity and revenue)
Show full SKILL.md (279 more words)Show less

Edge Cases

  1. No date columns in any table: Skip temporal analysis and anomaly detection entirely. Note in the report: "No temporal columns detected -- temporal analysis skipped."
  2. Single-column tables or lookup tables: Run completeness only. Skip distributions, correlations, and anomalies. Flag as "lookup table" in the report.
  3. All columns are non-numeric: Skip distribution and correlation analysis. Focus on completeness and categorical cardinality.
  4. Very wide tables (>50 columns): Profile all columns for completeness, but limit distribution analysis to the top 20 numeric columns by variance. Note which columns were skipped.
  5. Empty tables (0 rows): Log as BLOCKER. Do not attempt profiling -- report the table as empty and move on.
  6. DuckDB connection fails: Fall back to CSV via read_table(). The schema profiler handles this internally, but deep profiling should also use the CSV path.

Anti-Patterns

  1. Never skip profiling because "the data looks clean." Surprises hide in distributions and temporal patterns that summary stats miss.
  2. Never run anomaly detection on raw event rows. Always aggregate to daily or weekly granularity first. Running on raw rows will flag every row as an "anomaly" relative to rolling stats.
  3. Never profile in isolation from schema context. Always run profile_source() first (Step 1) so you know which columns are dates, which are numeric, and what the cardinality looks like before deep profiling.
  4. Never treat all WARNING items equally. A 10% null rate in a segmentation column is more impactful than 10% nulls in a free-text notes column. Contextualize severity by the column's role in analysis.
  5. Never skip the report write. Even if profiling runs smoothly, always write last_profile.md so future sessions can reference it without re-profiling.

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/data-profiling of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Data Profiling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Profiling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Profiling this skillai-analyst-lab/ai-analyst304—~2.4kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Exploratory Data Analysisspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Antv L7antvis/L74.1k—~1.4kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2063 repos~3.4kAutomated safety check: NotesMIT

Similar skills

  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Antv L7

    antvis/L7

    Comprehensive guide for AntV L7 geospatial visualization library.

    4.1k GitHub stars~1.4k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    206 GitHub starsUsed in 3 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Questions about Data Profiling

What does Data Profiling do?

Deep-profile the active dataset: distributions, temporal patterns, correlations, completeness gaps, anomalies. Data Profiling is an agent skill from ai-analyst-lab/ai-analyst. Deep-profile the active dataset: distributions, temporal patterns, correlations, completeness gaps, anomalies.

When should I use Data Profiling?

Data Profiling fits situations like: profile this data; deep-profile the dataset; run a data profile; check distributions.

How do I install Data Profiling in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill data-profiling -a claude-code`. Or copy the skill folder (.claude/skills/data-profiling in ai-analyst-lab/ai-analyst) into .claude/skills/data-profiling in your project. Claude Code loads it when a task matches its description.

How do I install Data Profiling in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill data-profiling -a codex`. Or copy the skill folder (.claude/skills/data-profiling in ai-analyst-lab/ai-analyst) into .agents/skills/data-profiling in your project. Codex loads it when a task matches its description.

Can I use Data Profiling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill data-profiling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-profiling, .gemini/skills/data-profiling, .github/skills/data-profiling and .opencode/skills/data-profiling in your project.

What does Data Profiling need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Profiling is instructions for the agent only. Our summary lists: Python 3.

Does Data Profiling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Profiling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Profiling use?

Data Profiling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Profiling use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Profiling?

Skills that share tags, products or a category with Data Profiling: Pandas Pro (Jeffallan/claude-skills, 12k stars), Exploratory Data Analysis (spacering-net/codeg, 3.8k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars) and Antv L7 (antvis/L7, 4.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Profiling?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.