Agent skill

Data Quality Auditor

by borghei in borghei/Claude-Skills

Audit data quality across pipelines, warehouses, and stores.

MITAuto-check passedData & Analytics

Install Data Quality Auditor

skills CLI
$ npx skills add borghei/Claude-Skills --skill data-quality-auditor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills data-quality-auditor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/data-quality-auditor .claude/skills/data-quality-auditor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-quality-auditor
GitHub stars
874
Token cost
~2.1k tokens
SKILL.md length
889 words
Files
10 (incl. scripts, references)
Skills in repo
364
Repo updated
First seen
Licence
MIT

At a glance

Audit data quality across pipelines, warehouses, and stores.

  • Designing a DQ program
  • SKILL.md covers Core Capabilities, When to Use, Clarify First and The six DQ dimensions, plus 5 more sections
  • Runs Python scripts from its folder; calls python
  • Defining DQ dimensions

What it does

Data Quality Auditor is an agent skill from borghei/Claude-Skills. Audit data quality across pipelines, warehouses, and stores. Use when designing a DQ program, defining DQ dimensions, building rule-based checks, detecting schema drift, monitoring freshness SLAs, or responding to a DQ incident.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `references/data-quality-dimensions.md`, `references/dq-anti-patterns.md` and `references/dq-check-catalog.md`).

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Designing a DQ program
  • Defining DQ dimensions
  • Building rule-based checks
  • Detecting schema drift

Example prompts

  • “/data-quality-auditor”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Quality Auditor loads about 2.1k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 889 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 889 words, ~2,089 tokens.

Download SKILL.mdSave it as .claude/skills/data-quality-auditor/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
data-quality-auditor
description
Audit data quality across pipelines, warehouses, and stores. Use when designing a DQ program, defining DQ dimensions, building rule-based checks, detecting schema drift, monitoring freshness SLAs, or responding to a DQ incident.
license
MIT + Commons Clause
metadata.version
1.1.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
engineering
metadata.updated
2026-06-17
metadata.tags
data-quality, dq, freshness, schema-drift, great-expectations, dbt-tests, soda, data-observability, data-engineering

Data Quality Auditor

End-to-end data quality (DQ) practice: define DQ dimensions, write rule-based checks, detect schema drift, monitor freshness SLAs, respond to DQ incidents, build a maturity-graded program. Tool-agnostic — works whether you use Great Expectations, dbt tests, Soda Core, Monte Carlo, custom SQL, or hand-rolled scripts.

This skill is audit-focused, not pipeline-focused. For pipeline design, ETL, Spark/dbt, see engineering/senior-data-engineer.

Core Capabilities

  • Six DQ dimensions — completeness, accuracy, consistency, timeliness/freshness, validity, uniqueness (plus integrity, conformity, reasonableness); at least one check per dimension at production stage.
  • DQ check catalog — five categories (volume, freshness, schema, values, distribution) of ~50 specific check patterns applied per dataset.
  • Schema drift detection — snapshot baseline schemas and diff added/removed/changed columns, types, and ordinals with severity.
  • Freshness SLA monitoring — per-table max-age budgets with alerting-ready output.
  • Incident response — severity classification (Sev1-4) and a 7-step playbook (acknowledge → quarantine → triage → contain → fix-forward → notify → post-incident) with recovery patterns.
  • Maturity & governance — a five-level maturity model and anti-pattern catalog to grade and improve a DQ program.

When to Use

SituationSkill applies
Setting up DQ from scratch on a new pipelineYes — start with DQ dimensions + check catalog
Auditing existing pipelines for missing DQYes — dq_check_runner.py
Detecting schema drift in upstream sourcesYes — schema_drift_detector.py
Monitoring freshness / SLA on data assetsYes — freshness_monitor.py
Responding to a DQ incident (bad data in prod)Yes — incident response playbook
Designing a DQ governance modelYes — DQ maturity model
Compliance evidence (SOC 2 PI1, GDPR, ISO 27001)Yes — checks produce auditable artifacts
Building data pipelines for the first timeUse engineering/senior-data-engineer first

Clarify First

Before running the audit, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Dataset & target — which table, pipeline, or store to audit (the data the scripts read via --data)
  • Which check — DQ rule checks, schema drift, or freshness SLA (selects dq_check_runner.py vs schema_drift_detector.py vs freshness_monitor.py)
  • Thresholds & SLAs — per-dimension pass thresholds and the freshness max-age budget (sets the pass/fail line and --max-age-min)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

The six DQ dimensions

Industry-standard taxonomy. Every dataset should have at least one check per dimension when at production stage.

DimensionQuestionExample check
CompletenessAre required fields populated?users.email IS NOT NULL — fail if > 0.1% nulls
AccuracyDo values match reality?Reconciliation against source-of-truth system; sample-based human review
ConsistencyDo values agree across systems / time?users.email in DB matches Salesforce; row count today within 5% of yesterday
Timeliness / FreshnessIs data current to expectation?events_table.max(event_time) is < 1h old; pipeline runs SLA
ValidityDo values conform to format / schema / business rules?Email regex matches; country code in ISO 3166-1; status in known enum
UniquenessAre entities not duplicated?users.user_id is unique; no two rows with same (user_id, day)

Some teams add: Integrity (referential — FKs resolve), Conformity (matches a published standard), Reasonableness (passes basic sanity checks beyond strict validity).

The DQ check catalog

Checks group into five categories applied per dataset — Volume, Freshness, Schema, Values, and Distribution. See the category summary and the full ~50-pattern catalog in references/dq-check-catalog.md.

Show full SKILL.md (375 more words)Show less

Tools

ToolPurposeCommand
dq_check_runner.pyRun/profile DQ checks against tabular data; per-table pass/fail/warning with value vs thresholdpython scripts/dq_check_runner.py --data t.json --checks checks.json --format json
schema_drift_detector.pyDiff a current schema against a baseline snapshot (added/removed/changed columns, types, ordinals)python scripts/schema_drift_detector.py --baseline base.json --current cur.json
freshness_monitor.pyCheck a freshness SLA: current age vs max-age budget, alerting-ready outputpython scripts/freshness_monitor.py --data t.json --column updated_at --max-age-min 60

All scripts: stdlib only, argparse CLI, JSON or human-readable output (see Scope re: live DB integration).

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

  • references/data-quality-dimensions.md — the 6 dimensions in depth: how to measure, threshold guidance, alerting strategy, and common pitfalls per dimension. Read when defining checks for a specific dimension.
  • references/dq-check-catalog.md — the full catalog of ~50 specific check patterns with detection heuristics and tool snippets (Great Expectations / dbt / Soda / SQL). Read when writing concrete checks.
  • references/dq-incident-response.md — the full incident playbook: severity classification, the 7-step response loop, recovery patterns (backfill, quarantine, DLQ, idempotent reprocessing, rollback), and post-incident writeup/notification templates. Read when responding to a DQ incident.
  • references/dq-maturity-model.md — the five-level DQ maturity model (L0 reactive → L4 data-as-product) and what to invest in at each level. Read when grading or planning a DQ program.
  • references/dq-workflows.md — the five end-to-end workflows (new dataset, schema drift, freshness SLA, auditing existing pipelines, post-incident improvement) with the exact script commands. Read when executing a concrete DQ task.
  • references/dq-anti-patterns.md — the catalog of DQ anti-patterns to avoid. Read when reviewing an existing DQ program for smells.

Scope & Limitations

This skill covers:

  • Auditing data quality across pipelines, warehouses, and stores against the six DQ dimensions.
  • Rule-based check design, schema drift detection, and freshness SLA monitoring (tool-agnostic).
  • DQ incident response and a maturity-graded governance program.
  • Producing auditable DQ artifacts for compliance evidence (SOC 2 PI1, GDPR accuracy, ISO 27001).

This skill does NOT cover:

  • Pipeline design, ETL, dbt/Spark implementation — use engineering/senior-data-engineer.
  • Live database querying; scripts are stdlib-only and read JSON inputs or simulate. For production, integrate your DB driver of choice (psycopg / mysqlclient / google-cloud-bigquery / etc.).
  • engineering/senior-data-engineer — pipeline design, ETL, dbt, Spark
  • engineering/observability-designer — observability for data infrastructure (adjacent to DQ)
  • engineering/chaos-engineering — DQ checks benefit from chaos testing
  • ra-qm-team/gdpr-dsgvo-expert — DQ underpins GDPR Art. 5(1)(d) "accuracy"
  • ra-qm-team/soc2-compliance-expert — SOC 2 PI1 (Processing Integrity) requires DQ controls

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references) in engineering/data-quality-auditor of borghei/Claude-Skills.

  • SKILL.md
  • references/data-quality-dimensions.md
  • references/dq-anti-patterns.md
  • references/dq-check-catalog.md
  • references/dq-incident-response.md
  • references/dq-maturity-model.md
  • references/dq-workflows.md
  • scripts/dq_check_runner.py
  • scripts/freshness_monitor.py
  • scripts/schema_drift_detector.py

Open the folder on GitHubat commit c9a1487

Compare with similar skills

Data Quality Auditor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Quality Auditor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Quality Auditor this skillborghei/Claude-Skills874—~2.1kAutomated safety check: PassMIT
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 9 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from borghei/Claude-Skills

All 364 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    874 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    874 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    874 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    874 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Data Quality Auditor

What does Data Quality Auditor do?

Audit data quality across pipelines, warehouses, and stores. Data Quality Auditor is an agent skill from borghei/Claude-Skills. Audit data quality across pipelines, warehouses, and stores.

When should I use Data Quality Auditor?

Data Quality Auditor fits situations like: designing a DQ program; defining DQ dimensions; building rule-based checks; detecting schema drift.

How do I install Data Quality Auditor in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill data-quality-auditor -a claude-code`. Or copy the skill folder (engineering/data-quality-auditor in borghei/Claude-Skills) into .claude/skills/data-quality-auditor in your project. Claude Code loads it when a task matches its description.

How do I install Data Quality Auditor in Codex?

Run `npx skills add borghei/Claude-Skills --skill data-quality-auditor -a codex`. Or copy the skill folder (engineering/data-quality-auditor in borghei/Claude-Skills) into .agents/skills/data-quality-auditor in your project. Codex loads it when a task matches its description.

Can I use Data Quality Auditor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill data-quality-auditor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-quality-auditor, .gemini/skills/data-quality-auditor, .github/skills/data-quality-auditor and .opencode/skills/data-quality-auditor in your project.

What does Data Quality Auditor need to run?

Going by SKILL.md and its folder, Data Quality Auditor needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Data Quality Auditor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Quality Auditor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Quality Auditor use?

Data Quality Auditor is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Quality Auditor use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Data Quality Auditor?

Skills that share tags, products or a category with Data Quality Auditor: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Issues Deduplication (JetBrains/ideavim, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Quality Auditor?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.