Agent skill

Data Quality Audit

by mohitagw15856 in mohitagw15856/pm-claude-skills

Audit a dataset for the quality problems that silently break analysis — missingness, duplicates, outliers, type and range errors, consistency, and freshness — and produce a prioritised fix list.

MITAuto-check passedData & Analytics

Install Data Quality Audit

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill data-quality-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills data-quality-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-quality-audit .claude/skills/data-quality-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-quality-audit
GitHub stars
1.4k
Token cost
~802 tokens
SKILL.md length
387 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Audit a dataset for the quality problems that silently break analysis — missingness, duplicates, outliers, type and range errors, consistency, and freshness — and produce a prioritised fix list.

  • Works in 5 steps: Summary → Quality scorecard → Specific issues → …
  • Asked to assess data quality
  • SKILL.md covers Working from a brief, Required Inputs, Output Format and Quality Checks, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Quality Audit is an agent skill from mohitagw15856/pm-claude-skills. Audit a dataset for the quality problems that silently break analysis — missingness, duplicates, outliers, type and range errors, consistency, and freshness — and produce a prioritised fix list. Use when asked to assess data quality, audit a dataset, check data before analysis, or explain why numbers look off. Produces a structured quality report across the standard dimensions, the specific issues found (with the checks to run), severity, and how to fix each.

Its SKILL.md is about 800 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to assess data quality
  • Audit a dataset
  • Check data before analysis
  • Explain why numbers look off

Example prompts

  • “/data-quality-audit”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Summary
  2. Quality scorecard
  3. Specific issues
  4. Fix plan (prioritised)
  5. Guardrails

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Quality Audit loads about 802 tokens when it runs. Until then it costs about 121 tokens; SKILL.md has 387 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~121
When it runs · the whole SKILL.md, loaded when a task matches
~802

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 387 words, ~802 tokens.

Download SKILL.mdSave it as .claude/skills/data-quality-audit/SKILL.md (or your agent's skills folder).
name
data-quality-audit
description
Audit a dataset for the quality problems that silently break analysis — missingness, duplicates, outliers, type and range errors, consistency, and freshness — and produce a prioritised fix list. Use when asked to assess data quality, audit a dataset, check data before analysis, or explain why numbers look off. Produces a structured quality report across the standard dimensions, the specific issues found (with the checks to run), severity, and how to fix each.

Data Quality Audit Skill

Bad analysis usually starts with bad data nobody checked. This skill audits a dataset across the dimensions that matter, names the specific issues (and the exact check to confirm each), and prioritises fixes by how much they distort the answer.

Working from a brief

Given a dataset description, sample rows, or a schema, produce the full audit anyway — infer the likely issues for that kind of data and give the concrete check (SQL/pandas-style) to verify each. If given actual data, ground the findings in it. Never just say "check for errors"; specify them.

Required Inputs

Ask for (if not already provided):

  • The dataset — schema, a sample, or a description (what each column is, the grain)
  • What it'll be used for (the analysis/decision it feeds — focuses the audit)
  • Source & freshness (where it comes from, how often it updates)
  • Known issues the user already suspects

Output Format

1. Summary

Overall read (🟢 usable / 🟡 fix-first / 🔴 don't trust yet) and the one issue most likely to mislead.

2. Quality scorecard
DimensionCheckFindingSeverity
Completenessnulls / missing per key column
Uniquenessduplicate rows / keys
Validitytype, format, range, allowed values
Consistencycross-field & cross-table agreement
Accuracysanity vs known totals / reality
Timelinessfreshness, gaps in the time series
3. Specific issues

For each real issue: what it is, the check to confirm it (a concrete query/snippet), why it matters for the intended use, and severity.

Show full SKILL.md (154 more words)Show less
4. Fix plan (prioritised)

Ordered by impact-on-the-decision: what to fix first, how (drop / impute / dedupe / cast / clamp / re-source), and what to flag rather than fix.

5. Guardrails

2–3 automated checks to add so these issues get caught next time (e.g. a not-null assertion, a row-count delta alarm, an allowed-values test).

Quality Checks

  • Covers all six dimensions, not just missing values
  • Each issue comes with a concrete check to confirm it, not just a label
  • Severity is judged against the intended use of the data
  • Fix plan is prioritised by impact and says fix-vs-flag
  • Recommends guardrails to prevent recurrence

Anti-Patterns

  • Only checking for nulls and calling it done
  • "Clean your data" with no specific issues or checks
  • Treating all issues as equally severe regardless of the decision
  • Fixing data silently with no record of what was changed

Example Trigger Phrases

  • "Assess data quality."
  • "Audit a dataset."
  • "Check data before analysis."
  • "Explain why numbers look off."

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/data-quality-audit of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Data Quality Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Quality Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Quality Audit this skillmohitagw15856/pm-claude-skills1.4k—~802Automated safety check: PassMIT
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 11 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Data Quality Audit

What does Data Quality Audit do?

Audit a dataset for the quality problems that silently break analysis — missingness, duplicates, outliers, type and range errors, consistency, and freshness — and produce a prioritised fix list. Data Quality Audit is an agent skill from mohitagw15856/pm-claude-skills. Audit a dataset for the quality problems that silently break analysis — missingness, duplicates, outliers, type and range errors, consistency, and freshness — and produce a prioritised fix list.

When should I use Data Quality Audit?

Data Quality Audit fits situations like: asked to assess data quality; audit a dataset; check data before analysis; explain why numbers look off.

How do I install Data Quality Audit in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill data-quality-audit -a claude-code`. Or copy the skill folder (skills/data-quality-audit in mohitagw15856/pm-claude-skills) into .claude/skills/data-quality-audit in your project. Claude Code loads it when a task matches its description.

How do I install Data Quality Audit in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill data-quality-audit -a codex`. Or copy the skill folder (skills/data-quality-audit in mohitagw15856/pm-claude-skills) into .agents/skills/data-quality-audit in your project. Codex loads it when a task matches its description.

Can I use Data Quality Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill data-quality-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-quality-audit, .gemini/skills/data-quality-audit, .github/skills/data-quality-audit and .opencode/skills/data-quality-audit in your project.

What does Data Quality Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Quality Audit is instructions for the agent only.

Does Data Quality Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Quality Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Quality Audit use?

Data Quality Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Quality Audit use?

About 802 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Quality Audit?

Skills that share tags, products or a category with Data Quality Audit: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Pandas Pro (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Quality Audit?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,433 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 8, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.