Agent skill

Data Validation

by majiayu000 in majiayu000/claude-skill-registry

QA an analysis before sharing with stakeholders — methodology checks, accuracy verification, and bias detection.

MITAuto-check passedData & Analytics

Install Data Validation

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill data-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry data-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/data-validation-yongjianwan-agentskill .claude/skills/data-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-validation
GitHub stars
666
Used in
2 other repos
Token cost
~2.4k tokens
SKILL.md length
1,005 words
Files
2
Skills in repo
1,273
Repo updated
First seen
Licence
MIT

At a glance

QA an analysis before sharing with stakeholders — methodology checks, accuracy verification, and bias detection.

  • Works in 5 steps: Calculate the same metric two different… → Spot-check individual records -- pick a… → Compare to known benchmarks -- match… → …
  • Reviewing an analysis for errors
  • SKILL.md covers Pre-Delivery QA Checklist, Common Data Analysis Pitfalls, Result Sanity Checking and Documentation Standards for…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Validation is an agent skill from majiayu000/claude-skill-registry. QA an analysis before sharing with stakeholders — methodology checks, accuracy verification, and bias detection. Use when reviewing an analysis for errors, checking for survivorship bias, validating aggregation logic, or preparing documentation for reproducibility.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in Data & Analytics, covering Data cleaning and Reproducible research. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Reviewing an analysis for errors
  • Checking for survivorship bias
  • Validating aggregation logic
  • Preparing documentation for reproducibility

Example prompts

  • “/data-validation”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Calculate the same metric two different ways and verify they match
  2. Spot-check individual records -- pick a few specific entities and trace their data manually
  3. Compare to known benchmarks -- match against published dashboards, finance reports, or prior analyses
  4. Reverse engineer -- if total revenue is X, does per-user revenue times user count approximately equal X?
  5. Boundary checks -- what happens when you filter to a single day, a single user, or a single category? Are those micro-results sensible?

What it can do on your machine

Read from SKILL.md and the folder at commit 2d14a69. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql, markdown and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Validation loads about 2.4k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 1,005 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 2d14a69, republished under its MIT licence (© majiayu000). 1,005 words, ~2,383 tokens.

Download SKILL.mdSave it as .claude/skills/data-validation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
data-validation
description
QA an analysis before sharing with stakeholders — methodology checks, accuracy verification, and bias detection. Use when reviewing an analysis for errors, checking for survivorship bias, validating aggregation logic, or preparing documentation for reproducibility.

Data Validation Skill

Pre-delivery QA checklist, common data analysis pitfalls, result sanity checking, and documentation standards for reproducibility.

Pre-Delivery QA Checklist

Run through this checklist before sharing any analysis with stakeholders.

Data Quality Checks
  • Source verification: Confirmed which tables/data sources were used. Are they the right ones for this question?
  • Freshness: Data is current enough for the analysis. Noted the "as of" date.
  • Completeness: No unexpected gaps in time series or missing segments.
  • Null handling: Checked null rates in key columns. Nulls are handled appropriately (excluded, imputed, or flagged).
  • Deduplication: Confirmed no double-counting from bad joins or duplicate source records.
  • Filter verification: All WHERE clauses and filters are correct. No unintended exclusions.
Calculation Checks
  • Aggregation logic: GROUP BY includes all non-aggregated columns. Aggregation level matches the analysis grain.
  • Denominator correctness: Rate and percentage calculations use the right denominator. Denominators are non-zero.
  • Date alignment: Comparisons use the same time period length. Partial periods are excluded or noted.
  • Join correctness: JOIN types are appropriate (INNER vs LEFT). Many-to-many joins haven't inflated counts.
  • Metric definitions: Metrics match how stakeholders define them. Any deviations are noted.
  • Subtotals sum: Parts add up to the whole where expected. If they don't, explain why (e.g., overlap).
Reasonableness Checks
  • Magnitude: Numbers are in a plausible range. Revenue isn't negative. Percentages are between 0-100%.
  • Trend continuity: No unexplained jumps or drops in time series.
  • Cross-reference: Key numbers match other known sources (dashboards, previous reports, finance data).
  • Order of magnitude: Total revenue is in the right ballpark. User counts match known figures.
  • Edge cases: What happens at the boundaries? Empty segments, zero-activity periods, new entities.
Presentation Checks
  • Chart accuracy: Bar charts start at zero. Axes are labeled. Scales are consistent across panels.
  • Number formatting: Appropriate precision. Consistent currency/percentage formatting. Thousands separators where needed.
  • Title clarity: Titles state the insight, not just the metric. Date ranges are specified.
  • Caveat transparency: Known limitations and assumptions are stated explicitly.
  • Reproducibility: Someone else could recreate this analysis from the documentation provided.

Common Data Analysis Pitfalls

Join Explosion

The problem: A many-to-many join silently multiplies rows, inflating counts and sums.

How to detect:

sql
-- Check row count before and after join
SELECT COUNT(*) FROM table_a;  -- 1,000
SELECT COUNT(*) FROM table_a a JOIN table_b b ON a.id = b.a_id;  -- 3,500 (uh oh)

How to prevent:

  • Always check row counts after joins
  • If counts increase, investigate the join relationship (is it really 1:1 or 1:many?)
  • Use COUNT(DISTINCT a.id) instead of COUNT(*) when counting entities through joins
Survivorship Bias

The problem: Analyzing only entities that exist today, ignoring those that were deleted, churned, or failed.

Examples:

  • Analyzing user behavior of "current users" misses churned users
  • Looking at "companies using our product" ignores those who evaluated and left
  • Studying properties of "successful" outcomes without "unsuccessful" ones

How to prevent: Ask "who is NOT in this dataset?" before drawing conclusions.

Incomplete Period Comparison

The problem: Comparing a partial period to a full period.

Examples:

  • "January revenue is $500K vs. December's $800K" -- but January isn't over yet
  • "This week's signups are down" -- checked on Wednesday, comparing to a full prior week

How to prevent: Always filter to complete periods, or compare same-day-of-month / same-number-of-days.

Denominator Shifting

The problem: The denominator changes between periods, making rates incomparable.

Examples:

  • Conversion rate improves because you changed how you count "eligible" users
  • Churn rate changes because the definition of "active" was updated

How to prevent: Use consistent definitions across all compared periods. Note any definition changes.

Average of Averages

The problem: Averaging pre-computed averages gives wrong results when group sizes differ.

Example:

  • Group A: 100 users, average revenue $50
  • Group B: 10 users, average revenue $200
  • Wrong: Average of averages = ($50 + $200) / 2 = $125
  • Right: Weighted average = (100*$50 + 10*$200) / 110 = $63.64

How to prevent: Always aggregate from raw data. Never average pre-aggregated averages.

Show full SKILL.md (406 more words)Show less
Timezone Mismatches

The problem: Different data sources use different timezones, causing misalignment.

Examples:

  • Event timestamps in UTC vs. user-facing dates in local time
  • Daily rollups that use different cutoff times

How to prevent: Standardize all timestamps to a single timezone (UTC recommended) before analysis. Document the timezone used.

Selection Bias in Segmentation

The problem: Segments are defined by the outcome you're measuring, creating circular logic.

Examples:

  • "Users who completed onboarding have higher retention" -- obviously, they self-selected
  • "Power users generate more revenue" -- they became power users BY generating revenue

How to prevent: Define segments based on pre-treatment characteristics, not outcomes.

Result Sanity Checking

Magnitude Checks

For any key number in your analysis, verify it passes the "smell test":

Metric TypeSanity Check
User countsDoes this match known MAU/DAU figures?
RevenueIs this in the right order of magnitude vs. known ARR?
Conversion ratesIs this between 0% and 100%? Does it match dashboard figures?
Growth ratesIs 50%+ MoM growth realistic, or is there a data issue?
AveragesIs the average reasonable given what you know about the distribution?
PercentagesDo segment percentages sum to ~100%?
Cross-Validation Techniques
  1. Calculate the same metric two different ways and verify they match
  2. Spot-check individual records -- pick a few specific entities and trace their data manually
  3. Compare to known benchmarks -- match against published dashboards, finance reports, or prior analyses
  4. Reverse engineer -- if total revenue is X, does per-user revenue times user count approximately equal X?
  5. Boundary checks -- what happens when you filter to a single day, a single user, or a single category? Are those micro-results sensible?
Red Flags That Warrant Investigation
  • Any metric that changed by more than 50% period-over-period without an obvious cause
  • Counts or sums that are exact round numbers (suggests a filter or default value issue)
  • Rates exactly at 0% or 100% (may indicate incomplete data)
  • Results that perfectly confirm the hypothesis (reality is usually messier)
  • Identical values across time periods or segments (suggests the query is ignoring a dimension)

Documentation Standards for Reproducibility

Analysis Documentation Template

Every non-trivial analysis should include:

markdown
## Analysis: [Title]

### Question
[The specific question being answered]

### Data Sources
- Table: [schema.table_name] (as of [date])
- Table: [schema.other_table] (as of [date])
- File: [filename] (source: [where it came from])

### Definitions
- [Metric A]: [Exactly how it's calculated]
- [Segment X]: [Exactly how membership is determined]
- [Time period]: [Start date] to [end date], [timezone]

### Methodology
1. [Step 1 of the analysis approach]
2. [Step 2]
3. [Step 3]

### Assumptions and Limitations
- [Assumption 1 and why it's reasonable]
- [Limitation 1 and its potential impact on conclusions]

### Key Findings
1. [Finding 1 with supporting evidence]
2. [Finding 2 with supporting evidence]

### SQL Queries
[All queries used, with comments]

### Caveats
- [Things the reader should know before acting on this]
Code Documentation

For any code (SQL, Python) that may be reused:

python
"""
Analysis: Monthly Cohort Retention
Author: [Name]
Date: [Date]
Data Source: events table, users table
Last Validated: [Date] -- results matched dashboard within 2%

Purpose:
    Calculate monthly user retention cohorts based on first activity date.

Assumptions:
    - "Active" means at least one event in the month
    - Excludes test/internal accounts (user_type != 'internal')
    - Uses UTC dates throughout

Output:
    Cohort retention matrix with cohort_month rows and months_since_signup columns.
    Values are retention rates (0-100%).
"""
Version Control for Analyses
  • Save queries and code in version control (git) or a shared docs system
  • Note the date of the data snapshot used
  • If an analysis is re-run with updated data, document what changed and why
  • Link to prior versions of recurring analyses for trend comparison

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/analysis/data-validation-yongjianwan-agentskill of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 2d14a69

Used in 2 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Data Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Validation this skillmajiayu000/claude-skill-registry6662 repos~2.4kAutomated safety check: PassMIT
Datalineage Summarygoogle/skills21k—~1.7kAutomated safety check: PassApache-2.0
Data Lineage TrackerDrchronx/ai-agent-research-starter-kit135—~324Automated safety check: PassCustom licence
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Bio Outlier Splicing DetectionGPTomics/bioSkills1.2k2 repos~5.1kAutomated safety check: PassMIT
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT

Similar skills

  • Datalineage Summary

    google/skills

    Official

    Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS.

    21k GitHub stars~1.7k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Data Lineage Tracker

    Drchronx/ai-agent-research-starter-kit

    Track data lineage and reproducibility for academic projects.

    135 GitHub stars~324 tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Detects aberrant splicing in single rare-disease patients vs a control panel using FRASER 2.0 (Bioconductor; Beta-binomial autoencoder on Intron Jaccard Index, default delta cutoff 0.1, q…

    1.2k GitHub starsUsed in 2 repos~5.1k tokens
    Research & ScienceAuto-check passed
  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 10 days ago
    Data & AnalyticsAuto-check: notes

More from majiayu000/claude-skill-registry

All 1,273 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Pyzotero

    majiayu000/claude-skill-registry

    Interact with Zotero reference management libraries using the pyzotero Python client.

    666 GitHub starsUsed in 5 repos~1.6k tokens
    Auto-check: notes
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed

Questions about Data Validation

What does Data Validation do?

QA an analysis before sharing with stakeholders — methodology checks, accuracy verification, and bias detection. Data Validation is an agent skill from majiayu000/claude-skill-registry. QA an analysis before sharing with stakeholders — methodology checks, accuracy verification, and bias detection.

When should I use Data Validation?

Data Validation fits situations like: reviewing an analysis for errors; checking for survivorship bias; validating aggregation logic; preparing documentation for reproducibility.

How do I install Data Validation in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill data-validation -a claude-code`. Or copy the skill folder (skills/analysis/data-validation-yongjianwan-agentskill in majiayu000/claude-skill-registry) into .claude/skills/data-validation in your project. Claude Code loads it when a task matches its description.

How do I install Data Validation in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill data-validation -a codex`. Or copy the skill folder (skills/analysis/data-validation-yongjianwan-agentskill in majiayu000/claude-skill-registry) into .agents/skills/data-validation in your project. Codex loads it when a task matches its description.

Can I use Data Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill data-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-validation, .gemini/skills/data-validation, .github/skills/data-validation and .opencode/skills/data-validation in your project.

What does Data Validation need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Validation is instructions for the agent only. Our summary lists: Python 3.

Does Data Validation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Validation use?

Data Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Validation use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Validation?

Skills that share tags, products or a category with Data Validation: Datalineage Summary (google/skills, 21k stars), Data Lineage Tracker (Drchronx/ai-agent-research-starter-kit, 135 stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars) and Bio Outlier Splicing Detection (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Validation?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 1,273 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.