Agent skill

Data Validation Patterns

by revfactory in revfactory/harness-100

Migration data validation patterns: row count comparison, checksums, sampling validation, FK integrity, and business rule validation query design guide.

Apache-2.0Auto-check passedProduct & Project Management

Install Data Validation Patterns

skills CLI
$ npx skills add revfactory/harness-100 --skill data-validation-patterns -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install revfactory/harness-100 data-validation-patterns --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .claude/skills && cp -r skills-src/en/34-data-migration/.claude/skills/data-validation-patterns .claude/skills/data-validation-patterns && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-validation-patterns
GitHub stars
1.3k
Token cost
~1.7k tokens
SKILL.md length
58 words
Files
1
Skills in repo
464
Repo updated
First seen
Licence
Apache-2.0

At a glance

Migration data validation patterns: row count comparison, checksums, sampling validation, FK integrity, and business rule validation query design guide.

  • Requests involving data validation
  • SKILL.md covers Validation Layer Model, Level 1: Count Validation, Level 2: Schema Validation and Level 3: Data Value Validation, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Migration validation

What it does

Data Validation Patterns is an agent skill from revfactory/harness-100. Migration data validation patterns: row count comparison, checksums, sampling validation, FK integrity, and business rule validation query design guide. Use this skill for requests involving 'data validation', 'migration validation', 'checksum', 'row count comparison', 'integrity validation', 'regression testing', 'Go/No-Go checklist', etc. Enhances validation-engineer's validation design capabilities. Note: schema mapping and rollback planning are outside the scope of this skill.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Product & Project Management, covering Feature launches and release readiness and QA and bug reports. The licence is Apache-2.0.

When your agent uses it

  • Requests involving data validation
  • Migration validation
  • Row count comparison
  • Integrity validation

Example prompts

  • “data validation”
  • “migration validation”
  • “checksum”
  • “/data-validation-patterns”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 8e8d35c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql, python and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Validation Patterns loads about 1.7k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 58 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from revfactory/harness-100 at commit 8e8d35c, republished under its Apache-2.0 licence (© revfactory). 58 words, ~1,707 tokens.

Download SKILL.mdSave it as .claude/skills/data-validation-patterns/SKILL.md (or your agent's skills folder).
name
data-validation-patterns
description
Migration data validation patterns: row count comparison, checksums, sampling validation, FK integrity, and business rule validation query design guide. Use this skill for requests involving 'data validation', 'migration validation', 'checksum', 'row count comparison', 'integrity validation', 'regression testing', 'Go/No-Go checklist', etc. Enhances validation-engineer's validation design capabilities. Note: schema mapping and rollback planning are outside the scope of this skill.

Data Validation Patterns — Migration Validation Patterns Guide

A systematic collection of patterns and queries for verifying data integrity before and after migration.

Validation Layer Model

Level 5: Business Rule Validation  <- Domain-specific rules
Level 4: Cross-Reference Validation <- Inter-table relationships
Level 3: Data Value Validation      <- Transformation accuracy
Level 2: Schema Validation          <- Structural match
Level 1: Count Validation           <- Row/column count match

Level 1: Count Validation

sql
-- Source
SELECT 'orders' AS table_name, COUNT(*) AS row_count FROM source.orders
UNION ALL
SELECT 'customers', COUNT(*) FROM source.customers
UNION ALL
SELECT 'products', COUNT(*) FROM source.products;

-- Target (same query)
SELECT 'orders' AS table_name, COUNT(*) AS row_count FROM target.orders
UNION ALL ...;

-- Difference comparison
SELECT s.table_name,
       s.row_count AS source_count,
       t.row_count AS target_count,
       s.row_count - t.row_count AS diff,
       CASE WHEN s.row_count = t.row_count THEN 'PASS' ELSE 'FAIL' END AS status
FROM source_counts s JOIN target_counts t ON s.table_name = t.table_name;

Level 2: Schema Validation

sql
-- Column count comparison
SELECT s.table_name,
       s.col_count AS source_cols,
       t.col_count AS target_cols,
       CASE WHEN s.col_count = t.col_count THEN 'PASS' ELSE 'CHECK' END
FROM (SELECT table_name, COUNT(*) col_count
      FROM source_information_schema.columns GROUP BY table_name) s
JOIN (SELECT table_name, COUNT(*) col_count
      FROM target_information_schema.columns GROUP BY table_name) t
ON s.table_name = t.table_name;

-- NULL constraint comparison
-- PK/FK constraint comparison
-- Index comparison

Level 3: Data Value Validation

Checksum Comparison
sql
-- Row-level checksum (PostgreSQL)
SELECT id, md5(ROW(order_id, customer_id, total_amount, created_at)::text) AS row_hash
FROM orders;

-- Full table checksum
SELECT md5(string_agg(row_hash, '' ORDER BY id)) AS table_hash
FROM (
    SELECT id, md5(ROW(*)::text) AS row_hash FROM orders
) t;

-- MySQL checksum
CHECKSUM TABLE orders;
Sampling Validation
python
def sample_validation(source_conn, target_conn, table, pk_col, sample_size=1000):
    """Compare N randomly sampled rows one-by-one"""
    # 1. Random PK extraction
    pks = source_conn.execute(
        f"SELECT {pk_col} FROM {table} ORDER BY RANDOM() LIMIT {sample_size}"
    ).fetchall()

    mismatches = []
    for pk in pks:
        source_row = source_conn.execute(
            f"SELECT * FROM {table} WHERE {pk_col} = %s", (pk,)
        ).fetchone()
        target_row = target_conn.execute(
            f"SELECT * FROM {table} WHERE {pk_col} = %s", (pk,)
        ).fetchone()

        if not rows_equal(source_row, target_row):
            mismatches.append({
                'pk': pk, 'source': source_row, 'target': target_row,
                'diff_columns': find_diff_columns(source_row, target_row)
            })

    return {
        'table': table,
        'sample_size': sample_size,
        'mismatches': len(mismatches),
        'match_rate': (sample_size - len(mismatches)) / sample_size,
        'details': mismatches[:10]  # Top 10 only
    }
Aggregate Comparison
sql
-- Numeric column aggregate comparison
SELECT
    COUNT(*) AS cnt,
    SUM(total_amount) AS sum_amount,
    AVG(total_amount) AS avg_amount,
    MIN(total_amount) AS min_amount,
    MAX(total_amount) AS max_amount,
    COUNT(DISTINCT customer_id) AS unique_customers
FROM orders
WHERE created_at BETWEEN '2024-01-01' AND '2024-12-31';

Level 4: Cross-Reference Validation

sql
-- FK integrity: Does orders.customer_id exist in customers.id?
SELECT o.order_id, o.customer_id
FROM target.orders o
LEFT JOIN target.customers c ON o.customer_id = c.id
WHERE c.id IS NULL;
-- Result must be 0 rows to pass

-- Reverse: Do all customers with orders exist?
SELECT DISTINCT o.customer_id
FROM source.orders o
WHERE o.customer_id NOT IN (SELECT id FROM target.customers);

Level 5: Business Rule Validation

sql
-- Rule 1: Order total = sum of order items
SELECT o.order_id, o.total_amount, SUM(oi.price * oi.quantity) AS calc_total,
       ABS(o.total_amount - SUM(oi.price * oi.quantity)) AS diff
FROM target.orders o
JOIN target.order_items oi ON o.id = oi.order_id
GROUP BY o.order_id, o.total_amount
HAVING ABS(o.total_amount - SUM(oi.price * oi.quantity)) > 0.01;

-- Rule 2: Status transition validity
SELECT * FROM target.orders
WHERE status = 'SHIPPED' AND paid_at IS NULL;
-- Shipping without payment is invalid -> must be 0 rows

-- Rule 3: Date ordering
SELECT * FROM target.orders
WHERE created_at > paid_at OR paid_at > shipped_at;
-- Created > Paid > Shipped order violation -> must be 0 rows

Go/No-Go Checklist

markdown
## Migration Go/No-Go Decision

### Must Pass (all must be PASS for Go)
- [ ] L1: 100% row count match across all tables
- [ ] L2: Schema structure match (column count, types, constraints)
- [ ] L3: 100% checksum match (or 99.99% sample match)
- [ ] L4: 0 FK integrity violations
- [ ] L5: 0 critical business rule violations

### Warnings Allowed (document and Go)
- [ ] Date/time millisecond differences (from timezone conversion)
- [ ] String trailing whitespace differences
- [ ] Decimal trailing digit rounding differences

### Automatic Decision
All PASS -> Go
1+ mandatory FAIL -> No-Go
Warnings only -> Conditional Go (approval required)

Validation Automation Framework

python
class MigrationValidator:
    def __init__(self, source, target, tables):
        self.source = source
        self.target = target
        self.tables = tables
        self.results = []

    def run_all(self):
        for table in self.tables:
            self.results.append({
                'table': table,
                'row_count': self.check_row_count(table),
                'checksum': self.check_checksum(table),
                'fk_integrity': self.check_fk(table),
                'business_rules': self.check_rules(table),
            })
        return self.generate_report()

    def verdict(self):
        failed = [r for r in self.results if not r['all_pass']]
        return "GO" if not failed else "NO-GO"

© revfactory, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in en/34-data-migration/.claude/skills/data-validation-patterns of revfactory/harness-100.

Open the folder on GitHubat commit 8e8d35c

Compare with similar skills

Data Validation Patterns next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Validation Patterns compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Validation Patterns this skillrevfactory/harness-1001.3k—~1.7kAutomated safety check: PassApache-2.0
QA Releasejitpass/jit162—~1.1kAutomated safety check: PassCustom licence
Releasecodewhale-hq/Codewhale41k—~189Automated safety check: PassMIT
Release Readinesspetrkindlmann/qa-skills163—~6.7kAutomated safety check: PassMIT
.NET MAUI Release Readinessdotnet/maui23k—~15kAutomated safety check: PassMIT
Schematicblader/schematic239—~2.2kAutomated safety check: PassMIT

Similar skills

  • QA Release

    jitpass/jit

    Run jit's pre-release QA — a team of QA-engineer subagents (functionality, integrations, UX, bug-hunting, code review) that exercise a release candidate on this real Mac and hand back a consolidated…

    162 GitHub stars~1.1k tokensUpdated 2 days ago
    Product & Project ManagementAuto-check passed
  • Release

    codewhale-hq/Codewhale

    Prepare a named version: preflight, version consistency, build/package, smoke test, checksums/notes, and release readiness.

    41k GitHub stars~189 tokensUpdated today
    Testing & QAAuto-check passed
  • Release Readiness

    petrkindlmann/qa-skills

    Validate release readiness with evidence-based go/no-go decisions.

    163 GitHub stars~6.7k tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • Official

    Produces evidence-backed ship-readiness verdicts for .NET MAUI Servicing Releases and Previews, and drafts public-safe release handoff pages from the result.

    23k GitHub stars~15k tokensUpdated today
    Product & Project ManagementAuto-check passed
  • Schematic

    blader/schematic

    Reverse engineer a detailed product and technical specification document from a git branch's implementation.

    239 GitHub stars~2.2k tokensUpdated 6 mo ago
    Product & Project ManagementAuto-check passed
  • Onboarding Validation

    open-edge-platform/edge-ai-suites

    Validate the get-started experience of Open Edge Platform (OEP) software components from the perspective of a first-time user.

    140 GitHub stars~3.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from revfactory/harness-100

All 464 skills in this repo
  • Anti Bot Analyzer

    revfactory/harness-100

    A skill for analyzing website anti-bot defense mechanisms and developing legitimate evasion strategies.

    1.3k GitHub stars~1.1k tokensUpdated 6 mo ago
    Auto-check passed
  • API Error Design Patterns

    revfactory/harness-100

    Reference for designing how an API reports failures: structured error codes, response shapes, client-friendly messages, an error catalog and retry or fallback advice.

    1.3k GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • API Security Checklist

    revfactory/harness-100

    Walks a backend-dev agent through OWASP API Top 10 checks, authentication and authorization patterns, and defense code during API design.

    1.3k GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Arg Parser Generator

    revfactory/harness-100

    Methodology for systematically designing and generating CLI tool argument parser structures.

    1.3k GitHub stars~1.2k tokensUpdated 6 mo ago
    Auto-check passed
  • Audience Segmentation

    revfactory/harness-100

    Audience segmentation skill used by the analyst and curator agents.

    1.3k GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • Audio Storytelling

    revfactory/harness-100

    Audio storytelling skill used by the podcast scriptwriter and show note editor.

    1.3k GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Data Validation Patterns

What does Data Validation Patterns do?

Migration data validation patterns: row count comparison, checksums, sampling validation, FK integrity, and business rule validation query design guide. Data Validation Patterns is an agent skill from revfactory/harness-100. Migration data validation patterns: row count comparison, checksums, sampling validation, FK integrity, and business rule validation query design guide.

When should I use Data Validation Patterns?

Data Validation Patterns fits situations like: requests involving data validation; migration validation; row count comparison; integrity validation.

How do I install Data Validation Patterns in Claude Code?

Run `npx skills add revfactory/harness-100 --skill data-validation-patterns -a claude-code`. Or copy the skill folder (en/34-data-migration/.claude/skills/data-validation-patterns in revfactory/harness-100) into .claude/skills/data-validation-patterns in your project. Claude Code loads it when a task matches its description.

How do I install Data Validation Patterns in Codex?

Run `npx skills add revfactory/harness-100 --skill data-validation-patterns -a codex`. Or copy the skill folder (en/34-data-migration/.claude/skills/data-validation-patterns in revfactory/harness-100) into .agents/skills/data-validation-patterns in your project. Codex loads it when a task matches its description.

Can I use Data Validation Patterns in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add revfactory/harness-100 --skill data-validation-patterns -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-validation-patterns, .gemini/skills/data-validation-patterns, .github/skills/data-validation-patterns and .opencode/skills/data-validation-patterns in your project.

What does Data Validation Patterns need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Validation Patterns is instructions for the agent only. Our summary lists: Python 3.

Does Data Validation Patterns access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Validation Patterns safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Validation Patterns use?

Data Validation Patterns is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Validation Patterns use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Validation Patterns?

Skills that share tags, products or a category with Data Validation Patterns: QA Release (jitpass/jit, 162 stars), Release (codewhale-hq/Codewhale, 41k stars), Release Readiness (petrkindlmann/qa-skills, 163 stars) and .NET MAUI Release Readiness (dotnet/maui, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Validation Patterns?

revfactory (a GitHub user) maintains it in revfactory/harness-100, which has 1,290 GitHub stars. The repository holds 464 skills in this directory. The repository was last updated on March 22, 2026.

Source: revfactory/harness-100 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.