Agent skill

Data Quality Check

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Validate data completeness, consistency, and coverage before any analysis, flagging issues with severity ratings.

MITAuto-check passedData & Analytics

Install Data Quality Check

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill data-quality-check -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst data-quality-check --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/data-quality-check .claude/skills/data-quality-check && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-quality-check
GitHub stars
304
Token cost
~3.3k tokens
SKILL.md length
761 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Validate data completeness, consistency, and coverage before any analysis, flagging issues with severity ratings.

  • Works in 6 steps: Completeness Checks → Consistency Checks → Coverage Checks → …
  • Check data quality
  • SKILL.md covers Purpose, When to Use, Instructions and Examples, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Quality Check is an agent skill from ai-analyst-lab/ai-analyst. Validate data completeness, consistency, and coverage before any analysis, flagging issues with severity ratings. Run at the start of every new analysis. Trigger on "check data quality", "is the data clean", "validate the data", "run a quality check", "what's the coverage", and on named-table questions: "tell me about the {table} table", "describe {table}", "what's in {table}"; pair schema answers with a minimum DQ probe.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • Check data quality
  • Is the data clean
  • Validate the data
  • Run a quality check

Example prompts

  • “check data quality”
  • “is the data clean”
  • “validate the data”
  • “/data-quality-check”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Completeness Checks
  2. Consistency Checks
  3. Coverage Checks
  4. Statistical Sanity Checks
  5. Time-Series Anomaly Scan
  6. Data Freshness Check

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, sql and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Quality Check loads about 3.3k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 761 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 761 words, ~3,289 tokens.

Download SKILL.mdSave it as .claude/skills/data-quality-check/SKILL.md (or your agent's skills folder).
name
data-quality-check
description
Validate data completeness, consistency, and coverage before any analysis, flagging issues with severity ratings. Run at the start of every new analysis. Trigger on "check data quality", "is the data clean", "validate the data", "run a quality check", "what's the coverage", and on named-table questions: "tell me about the {table} table", "describe {table}", "what's in {table}"; pair schema answers with a minimum DQ probe.

Skill: Data Quality Check

Purpose

Validate data completeness, consistency, and coverage before any analysis begins, flagging issues with severity ratings so the analyst knows what blocks analysis vs. what to note as a caveat.

When to Use

Apply this skill at the start of every new analysis, when connecting to a new data source, or when results look suspicious. Run quality checks BEFORE drawing conclusions from data.

Also fires on table-scoped questions. Any question that names a specific table ("tell me about {table}", "describe {table}", "what's in {table}", "show me {table}") triggers this skill. Schema-only answers are insufficient — pair the schema description with a minimum DQ probe:

  • Row count
  • Null rate per column (flag anything >5%)
  • Date range on the primary timestamp column
  • Duplicate check on the primary key
  • Surface anything from .knowledge/datasets/{active}/quirks.md for that table

If the table is large enough that probing is expensive (>100M rows or warehouse cost concerns), tell the user and ask before running the full probe — but always run at minimum row count + PK duplicate check.

Instructions

Primary method — run the named structural validators

Do not hand-roll the core checks as ad-hoc SQL. Query the rows once, then run the tested validators in helpers/validation/structural_validator.py, so the checks are identical every time and can't be skipped or mis-written. The validators operate on a DataFrame, so pull the row-level slice you're about to analyze with the repo connection first:

python
from helpers.data.connection_manager import ConnectionManager
from helpers.validation.structural_validator import run_structural_checks

cm = ConnectionManager(); cm.connect()
df = cm.query("select * from orders where order_date >= '2024-12-01'")   # the slice under analysis

result = run_structural_checks(df, {
    "primary_key": ["ORDER_ID"],                         # uniqueness + nulls
    "required_columns": ["TOTAL_AMOUNT", "STATUS"],      # completeness
    "completeness_threshold": 0.95,
    "date_column": "ORDER_DATE",                         # gap / range
    "value_domain": {"column": "STATUS",
                     "valid_values": ["completed", "cancelled", "returned"]},
    "min_rows": 1,
})
print(result["overall_ok"], result["checks_passed"], "/", result["checks_run"])
for name, d in result["details"].items():
    print(name, "->", "OK" if (d.get("ok") or d.get("valid")) else f"FAIL ({d.get('severity','')})")

run_structural_checks returns {overall_ok, checks_run, checks_passed, checks_failed, details}; each entry in details is a named validator's result carrying a severity — map it to the BLOCKER/WARNING/ INFO rules below. For one targeted check, call the validator directly, e.g. validate_primary_key(df, ["ORDER_ID"]). For referential integrity, pass parent_df + child_key + parent_key in the config.

Example: run_structural_checks(products_df, {"primary_key": ["product_id"], "min_rows": 1}) passes (PRODUCT_ID is a real PK); validate_primary_key(products_df, ["CATEGORY"]) flags it (6 duplicates). The check actually runs — it is not a description.

The SQL templates in the Check Sequence below show what each validator does under the hood and cover extras the validators don't (ad-hoc segment coverage, etc.). Use them to explain or extend the results, not to replace the named validators.

Check Sequence

Run these checks in order. Stop and report blockers immediately.

1. Completeness Checks
sql
-- Null rate per column
SELECT
    column_name,
    COUNT(*) AS total_rows,
    COUNT(*) - COUNT(column_name) AS null_count,
    ROUND(100.0 * (COUNT(*) - COUNT(column_name)) / COUNT(*), 1) AS null_pct
FROM table_name
GROUP BY column_name;

-- Missing date ranges (for time-series data)
WITH date_spine AS (
    SELECT generate_series(MIN(date_col), MAX(date_col), INTERVAL '1 day') AS expected_date
    FROM table_name
)
SELECT expected_date
FROM date_spine
LEFT JOIN table_name ON date_col = expected_date
WHERE table_name.date_col IS NULL;

-- Unexpected zeros in numeric columns
SELECT column_name, COUNT(*) AS zero_count
FROM table_name
WHERE numeric_column = 0
GROUP BY column_name;

Severity rules:

  • BLOCKER: Primary key has nulls, >50% nulls in a critical analysis column, entire date ranges missing
  • WARNING: 5-50% nulls in an analysis column, scattered missing dates, unexpected zeros in revenue/count columns
  • INFO: <5% nulls in non-critical columns, weekend gaps in business-day data
2. Consistency Checks
sql
-- Duplicate detection
SELECT id_column, COUNT(*) AS dupes
FROM table_name
GROUP BY id_column
HAVING COUNT(*) > 1;

-- Referential integrity
SELECT child.fk_column, COUNT(*)
FROM child_table child
LEFT JOIN parent_table parent ON child.fk_column = parent.pk_column
WHERE parent.pk_column IS NULL
GROUP BY child.fk_column;

-- Date format consistency
SELECT DISTINCT LENGTH(date_column), LEFT(date_column, 4)
FROM table_name
WHERE date_column IS NOT NULL;

Severity rules:

  • BLOCKER: Duplicate primary keys, broken referential integrity affecting >10% of rows
  • WARNING: Mixed date formats, inconsistent casing in categorical columns, orphan records <10%
  • INFO: Minor casing inconsistencies, trailing whitespace
Show full SKILL.md (320 more words)Show less
3. Coverage Checks

Use check_temporal_coverage() for time-series gap detection and check_value_domain() for categorical completeness:

python
from helpers.data.sql_helpers import check_temporal_coverage, check_value_domain

# Temporal coverage — detect missing days/weeks/months
coverage = check_temporal_coverage(df, "order_date", freq="D")
if coverage["status"] == "FAIL":
    print(f"BLOCKER: {coverage['message']}")

# Value domain — verify expected categories exist
domain = check_value_domain(df["device_type"], ["desktop", "mobile", "tablet"])
if domain["status"] == "FAIL":
    print(f"WARNING: {domain['message']}")

SQL checks for segment coverage:

sql
-- Expected segments present
SELECT segment_column, COUNT(*) AS row_count,
       MIN(date_col) AS earliest, MAX(date_col) AS latest
FROM table_name
GROUP BY segment_column
ORDER BY row_count DESC;

-- Missing cohorts
SELECT date_trunc('month', created_at) AS cohort_month, COUNT(DISTINCT user_id)
FROM users
GROUP BY 1
ORDER BY 1;

Severity rules:

  • BLOCKER: Key segments entirely missing, temporal coverage <80%
  • WARNING: Some segments have <10% of expected rows, coverage 80-95%, unexpected category values
  • INFO: Minor imbalances in segment sizes, coverage >95%
4. Statistical Sanity Checks

Use the helper functions for systematic outlier and null concentration checks:

python
from helpers.validation.data_quality_extras import check_null_concentration, check_outliers

# Null concentration — flags columns with high null rates
null_results = check_null_concentration(df)
for r in null_results:
    if r["status"] == "FAIL":
        print(f"BLOCKER: {r['column']} — {r['detail']}")
    elif r["status"] == "WARN":
        print(f"WARNING: {r['column']} — {r['detail']}")

# Outlier detection — IQR method (default) or z-score
for col in numeric_columns:
    iqr_result = check_outliers(df[col], method="iqr")
    zscore_result = check_outliers(df[col], method="zscore")
    # Use IQR as primary, z-score as cross-check
    if iqr_result["status"] in ("WARN", "FAIL"):
        print(f"WARNING: {col} — {iqr_result['detail']}")

For domain-specific sanity checks (impossible values, suspicious distributions):

python
from helpers.validation.data_quality_extras import sanity_check

stats, issues = sanity_check(df, "conversion_rate")   # issues: [(severity, message), ...]

Severity rules:

  • BLOCKER: Impossible values (negative revenue, conversion rate >100%, future dates), >95% nulls
  • WARNING: Extreme outliers (>3 IQR), >50% nulls, highly skewed distributions
  • INFO: Moderate outliers, slight skew, <5% nulls
5. Time-Series Anomaly Scan

For each date-indexed metric column in the dataset:

python
from helpers.validation.data_quality_extras import anomaly_scan

result = anomaly_scan(daily_df, "date", "orders", window=14, threshold=2.0)   # result["anomalies"], result["summary"]

Sequencing: Run after basic data profiling in the Data Explorer step, on data already aggregated to daily or weekly granularity (rolling bands on raw event rows are meaningless).

Severity rules:

  • WARNING: Any anomaly detected — present as starting point for investigation
  • INFO: No anomalies found — note that the metric appears stable

Output format:

Notable patterns detected:
  - [metric] spiked [X]% above normal on [date range]
  - [metric] dropped [X]% below normal on [date range]

These are observations, not conclusions — present as starting points for investigation.

6. Data Freshness Check

For each table with a date/timestamp column:

python
from helpers.validation.data_quality_extras import freshness_check

fresh = freshness_check(df, "order_date")   # fresh["max_date"], fresh["days_ago"], fresh["cadence"], fresh["status"], fresh["note"]

Output format:

Data freshness:
  - events: most recent = [date] ([N] days ago) [OK/WARNING]
  - orders: most recent = [date] ([N] days ago) [OK/WARNING]
  - users: most recent = [date] ([N] days ago) [OK/WARNING]

Severity rules:

  • WARNING: Data is stale relative to inferred cadence
  • INFO: Data is fresh or dataset is static/historical
Output Format
markdown
# Data Quality Report: [Dataset Name]
## Date: [YYYY-MM-DD]
## Analyst: AI Product Analyst

### Summary
| Severity | Count | Details |
|----------|-------|---------|
| BLOCKER  | X     | [Must fix before analysis] |
| WARNING  | X     | [Note as caveat in analysis] |
| INFO     | X     | [For awareness only] |

### BLOCKERS
[List each blocker with: what's wrong, which column/table, how many rows affected, suggested fix]

### WARNINGS
[List each warning with: what's wrong, potential impact on analysis, recommended handling]

### INFO
[List each info item briefly]

### Data Profile
| Table | Rows | Columns | Date Range | Key Columns |
|-------|------|---------|------------|-------------|
| ... | ... | ... | ... | ... |

### Recommendation
[Can analysis proceed? With what caveats?]
- PROCEED: No blockers, warnings noted
- PROCEED WITH CAUTION: No blockers, significant warnings — note in findings
- BLOCKED: Blockers found — fix data before analyzing

Examples

Example 1: Clean dataset
markdown
### Summary
| Severity | Count | Details |
|----------|-------|---------|
| BLOCKER  | 0     | — |
| WARNING  | 1     | 8% null in `referral_source` column |
| INFO     | 2     | Weekend gaps in daily data; minor casing inconsistency in `country` |

### Recommendation
PROCEED — the null referral_source values should be noted as "unknown" in any segmentation by acquisition channel. All other columns are complete and consistent.
Example 2: Problematic dataset
markdown
### Summary
| Severity | Count | Details |
|----------|-------|---------|
| BLOCKER  | 2     | Duplicate order IDs (1,247 rows); revenue column has negative values (-$45K total) |
| WARNING  | 3     | March 2025 data missing entirely; `device_type` has 12% nulls; conversion rates >1.0 for 89 rows |
| INFO     | 1     | `country` has mixed casing ("US" vs "us") |

### BLOCKERS
1. **Duplicate order_ids**: 1,247 rows have duplicate `order_id` values. This will inflate revenue calculations. Must deduplicate before analysis — keep earliest record per order_id.
2. **Negative revenue**: 342 rows have negative `revenue` values totaling -$45K. These may be refunds. Must classify and handle separately (exclude from revenue analysis or create separate refund analysis).

### Recommendation
BLOCKED — Fix duplicate order_ids and classify negative revenue before proceeding. Estimated fix time: 15 minutes with SQL dedup + refund classification.

Anti-Patterns

  1. Never skip quality checks because "the data looks fine" — surprises hide in the tails
  2. Never treat all nulls the same — 2% nulls in a non-critical column ≠ 50% nulls in a key metric
  3. Never fix data silently — always document what you changed and why in the quality report
  4. Never analyze data with known blockers — fix blockers first, or the entire analysis is unreliable
  5. Never assume dates are clean — check for future dates, time zone issues, and format inconsistencies
  6. Never ignore outliers — investigate whether they're real (whale users) or errors (test accounts, data bugs)

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/data-quality-check of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Data Quality Check next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Quality Check compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Quality Check this skillai-analyst-lab/ai-analyst304—~3.3kAutomated safety check: PassMIT
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 10 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Questions about Data Quality Check

What does Data Quality Check do?

Validate data completeness, consistency, and coverage before any analysis, flagging issues with severity ratings. Data Quality Check is an agent skill from ai-analyst-lab/ai-analyst. Validate data completeness, consistency, and coverage before any analysis, flagging issues with severity ratings.

When should I use Data Quality Check?

Data Quality Check fits situations like: check data quality; is the data clean; validate the data; run a quality check.

How do I install Data Quality Check in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill data-quality-check -a claude-code`. Or copy the skill folder (.claude/skills/data-quality-check in ai-analyst-lab/ai-analyst) into .claude/skills/data-quality-check in your project. Claude Code loads it when a task matches its description.

How do I install Data Quality Check in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill data-quality-check -a codex`. Or copy the skill folder (.claude/skills/data-quality-check in ai-analyst-lab/ai-analyst) into .agents/skills/data-quality-check in your project. Codex loads it when a task matches its description.

Can I use Data Quality Check in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill data-quality-check -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-quality-check, .gemini/skills/data-quality-check, .github/skills/data-quality-check and .opencode/skills/data-quality-check in your project.

What does Data Quality Check need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Quality Check is instructions for the agent only. Our summary lists: Python 3.

Does Data Quality Check access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Quality Check safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Quality Check use?

Data Quality Check is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Quality Check use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Quality Check?

Skills that share tags, products or a category with Data Quality Check: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Issues Deduplication (JetBrains/ideavim, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Quality Check?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.