Agent skill

Dataset Quality Audit

by zebbern in zebbern/claude-code-guide

Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…

MITAuto-check passedData & Analytics

Install Dataset Quality Audit

skills CLI
$ npx skills add zebbern/claude-code-guide --skill dataset-quality-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zebbern/claude-code-guide dataset-quality-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zebbern/claude-code-guide.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dataset-quality-audit .claude/skills/dataset-quality-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-quality-audit
GitHub stars
4.7k
Token cost
~996 tokens
SKILL.md length
280 words
Files
3 (incl. scripts)
Skills in repo
46
Repo updated
First seen
Licence
MIT

At a glance

Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…

  • Tasks that involve Data cleaning
  • SKILL.md covers Capabilities, Quick Start, Detailed Usage and Output Format (JSON), plus 2 more sections
  • Runs Python scripts from its folder; calls python3 and pip
  • Tasks that involve CSV and tabular files

What it does

Dataset Quality Audit is an agent skill from zebbern/claude-code-guide. Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and actionable suggestions. Triggered when users ask to check data quality, find missing or duplicate values, detect outliers, validate formats, profile data, or clean data.

Its SKILL.md is about 1000 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/data_quality_checker.py`).

It sits in Data & Analytics, covering Data cleaning, CSV and tabular files and Excel spreadsheets. It works with Microsoft Excel. The repository describes itself as: Claude Code Guide - Setup, Commands, workflows, agents, skills & tips-n-tricks from beginner to power user! The licence is MIT.

When your agent uses it

  • Tasks that involve Data cleaning
  • Tasks that involve CSV and tabular files
  • Tasks that involve Excel spreadsheets

Example prompts

  • “/dataset-quality-audit”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9cde898. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Quality Audit loads about 996 tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 280 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~996

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zebbern/claude-code-guide at commit 9cde898, republished under its MIT licence (© zebbern). 280 words, ~996 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-quality-audit/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
dataset-quality-audit
description
Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and actionable suggestions. Triggered when users ask to check data quality, find missing or duplicate values, detect outliers, validate formats, profile data, or clean data.
license
MIT

dataset-quality-audit

A data quality auditing tool that runs 12-dimension quality checks on tabular data, producing per-dimension scores (0–100), an overall grade, and actionable fix suggestions.

Capabilities

DimensionDescription
Missing ValuesCount and percentage of null/NaN values per column
Duplicate RowsNumber and percentage of fully duplicated rows
Type ConsistencyMixed types within a single column (e.g., numbers mixed with text)
Value Range / OutliersOutlier detection using the IQR method
Format ComplianceConsistency of date, email, phone number, and other formatted fields
Uniqueness ConstraintsWhether ID-type columns contain duplicates
Whitespace IssuesLeading/trailing spaces, empty strings, whitespace-only values
Constant ColumnsColumns with only a single unique value (zero information)
Distribution SkewnessWhether numeric columns have excessive skewness
Column NamingSpaces, special characters, or inconsistent casing in column names
Cardinality AnomaliesUnusually high or low number of unique values
Cross-Column ConsistencyLogical checks across columns (e.g., start date before end date)

Quick Start

bash
# Basic quality check
python3 scripts/data_quality_checker.py data.csv

# Save report as JSON
python3 scripts/data_quality_checker.py data.csv --output report.json

# Specify ID columns (for uniqueness checks)
python3 scripts/data_quality_checker.py users.csv --id-columns "user_id,email"

# Specify date columns (for format checks)
python3 scripts/data_quality_checker.py orders.csv --date-columns "created_at,updated_at"

Detailed Usage

Basic Invocation
bash
python3 scripts/data_quality_checker.py <data-file> [options]
Parameters
ParameterShortRequiredDefaultDescription
input—Yes—Path to input file (CSV/TSV/Excel/JSON)
--output-oNostdoutPath for the JSON report output
--id-columns-idNoAuto-detectComma-separated column names that should be unique
--date-columns-dcNoAuto-detectComma-separated column names containing dates
--sample-sNoAll rowsNumber of rows to sample (useful for large files)
--encoding-eNoutf-8File encoding

Output Format (JSON)

json
{
  "file": "data.csv",
  "rows": 10000,
  "columns": 15,
  "overall_score": 78.5,
  "grade": "B",
  "dimensions": {
    "missing_values": {
      "score": 85.0,
      "issues": [
        {"column": "age", "missing_count": 150, "missing_pct": 1.5, "suggestion": "Fill with median or mode"}
      ]
    },
    "duplicates": {
      "score": 95.0,
      "issues": [...]
    }
  },
  "top_suggestions": [
    "Column 'age' has 1.5% missing values — consider filling with the median",
    "Found 200 fully duplicated rows — consider deduplication"
  ]
}

Grading Scale

GradeScore RangeMeaning
A+95–100Excellent quality — ready for use as-is
A90–95Good quality — minor issues only
B80–90Moderate quality — recommended to fix before use
C60–80Poor quality — significant cleaning required
D40–60Very poor quality — many issues need attention
F0–40Essentially unusable — requires re-collection or major cleanup

Dependencies

  • Python 3.8+
  • pandas
  • numpy
bash
pip install pandas numpy

© zebbern, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/dataset-quality-audit of zebbern/claude-code-guide.

  • SKILL.md
  • LICENSE
  • scripts/data_quality_checker.py

Open the folder on GitHubat commit 9cde898

Compare with similar skills

Dataset Quality Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Quality Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Quality Audit this skillzebbern/claude-code-guide4.7k—~996Automated safety check: PassMIT
Visual Skillsnpc-live/clawfirm156—~7.4kAutomated safety check: PassNone
XLSX Spreadsheet ToolkitXiaomiMiMo/MiMo-Code14k—~2.9kAutomated safety check: PassApache-2.0
CSV and Excel MergerOneWave-AI/claude-skills328—~1.6kAutomated safety check: PassMIT
Excel Spreadsheet Creation and Editinganthropics/skills180k4 repos~2.1kAutomated safety check: PassProprietary
Excel and CSV Data Analysisbytedance/deer-flow84k4 repos~2.2kAutomated safety check: PassMIT

Similar skills

  • Visual Skills

    npc-live/clawfirm

    A skill your agent uses whenever the user provides data (CSV, JSON, table, pasted numbers, or any structured dataset) and expects a visual output — even if they don't say 'chart' or 'visualize'.

    156 GitHub stars~7.4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • XLSX Spreadsheet Toolkit

    XiaomiMiMo/MiMo-Code

    Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.

    14k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    328 GitHub stars~1.6k tokensUpdated 7 days ago
    Documents & OfficeAuto-check passed
  • Official

    Creates, edits and analyzes spreadsheets (.xlsx, .xlsm, .csv, .tsv) with openpyxl and pandas, writing live formulas and recalculating to confirm zero formula errors.

    180k GitHub starsUsed in 4 repos~2.1k tokens
    Documents & OfficeAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    84k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Raccoon Dataanalysis

    SenseTime-Copilot/raccoon-dataanalysis-skill

    Raccoon (小浣熊) Data Analysis - Remote code interpreter and data visualization service powered by SenseTime.

    137 GitHub stars~1.9k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed

More from zebbern/claude-code-guide

All 46 skills in this repo
  • Localization Toolkit

    zebbern/claude-code-guide

    This skill should be used when setting up, auditing, or enforcing internationalization/localization in UI codebases (React/TS, i18next or similar, JSON locales), including installing/configuring the…

    4.7k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Audit Flow

    zebbern/claude-code-guide

    Interactive system flow tracing across CODE, API, AUTH, DATA, NETWORK layers with SQLite persistence and Mermaid export.

    4.7k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Chart Image

    zebbern/claude-code-guide

    Generate publication-quality PNG chart images from data, supporting line, bar, area, candlestick, pie, and heatmap charts.

    4.7k GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Code To Diagram

    zebbern/claude-code-guide

    Analyze codebases and automatically generate architecture diagrams, flowcharts, and org charts.

    4.7k GitHub stars~972 tokensUpdated today
    Auto-check passed
  • Code Vuln Audit

    zebbern/claude-code-guide

    Scan code for security issues: dependency vulnerabilities (npm/pip audit), secret leaks (regex and entropy analysis), and OWASP anti-patterns like SQL injection, XSS, or command injection.

    4.7k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Idor Testing

    zebbern/claude-code-guide

    This skill should be used when the user asks to "test for insecure direct object references," "find IDOR vulnerabilities," "exploit broken access control," "enumerate user IDs or object references,"…

    4.7k GitHub starsUsed in 8 repos~3.1k tokens
    Auto-check passed

Works with

Questions about Dataset Quality Audit

What does Dataset Quality Audit do?

Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…. Dataset Quality Audit is an agent skill from zebbern/claude-code-guide. Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and actionable suggestions.

When should I use Dataset Quality Audit?

Dataset Quality Audit fits situations like: tasks that involve Data cleaning; tasks that involve CSV and tabular files; tasks that involve Excel spreadsheets.

How do I install Dataset Quality Audit in Claude Code?

Run `npx skills add zebbern/claude-code-guide --skill dataset-quality-audit -a claude-code`. Or copy the skill folder (skills/dataset-quality-audit in zebbern/claude-code-guide) into .claude/skills/dataset-quality-audit in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Quality Audit in Codex?

Run `npx skills add zebbern/claude-code-guide --skill dataset-quality-audit -a codex`. Or copy the skill folder (skills/dataset-quality-audit in zebbern/claude-code-guide) into .agents/skills/dataset-quality-audit in your project. Codex loads it when a task matches its description.

Can I use Dataset Quality Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zebbern/claude-code-guide --skill dataset-quality-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-quality-audit, .gemini/skills/dataset-quality-audit, .github/skills/dataset-quality-audit and .opencode/skills/dataset-quality-audit in your project.

What does Dataset Quality Audit need to run?

Going by SKILL.md and its folder, Dataset Quality Audit needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and pip). Our summary lists: Python 3.

Does Dataset Quality Audit access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Dataset Quality Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Dataset Quality Audit use?

Dataset Quality Audit is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Quality Audit use?

About 996 tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dataset Quality Audit?

Skills that share tags, products or a category with Dataset Quality Audit: Visual Skills (npc-live/clawfirm, 156 stars), XLSX Spreadsheet Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), CSV and Excel Merger (OneWave-AI/claude-skills, 328 stars) and Excel Spreadsheet Creation and Editing (anthropics/skills, 180k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Quality Audit?

zebbern (a GitHub user) maintains it in zebbern/claude-code-guide, which has 4,652 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.

Source: zebbern/claude-code-guide on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.