Agent skill

Data Validator

by FerroxLabs in FerroxLabs/wayland

Data validation and quality expertise covering Great Expectations patterns, schema validation, statistical validation, referential integrity checks, data profiling, anomaly detection, data…

Apache-2.0Auto-check passedData & Analytics

Install Data Validator

skills CLI
$ npx skills add FerroxLabs/wayland --skill data-validator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FerroxLabs/wayland data-validator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-engineering/data-validator .claude/skills/data-validator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-validator
GitHub stars
608
Token cost
~3.3k tokens
SKILL.md length
325 words
Files
1
Skills in repo
1,194
Repo updated
First seen
Licence
Apache-2.0

At a glance

Data validation and quality expertise covering Great Expectations patterns, schema validation, statistical validation, referential integrity checks, data profiling, anomaly detection, data…

  • The user asks about data validator
  • SKILL.md covers Overview, Great Expectations Patterns, Schema Validation and Statistical Validation, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Data validator best practices

What it does

Data Validator is an agent skill from FerroxLabs/wayland. Data validation and quality expertise covering Great Expectations patterns, schema validation, statistical validation, referential integrity checks, data profiling, anomaly detection, data contracts, quality scoring, and automated testing strategies for ensuring data reliability throughout the pipeline. Use when the user asks about data validator, data validator best practices, or needs guidance on data validator implementation. Do NOT use when the user needs a different specialized skill or is asking about an…

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Forms and validation, Anomaly detection and Test strategy. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.

When your agent uses it

  • The user asks about data validator
  • Data validator best practices
  • Needs guidance on data validator implementation
  • The user needs a different specialized skill

Example prompts

  • “/data-validator”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, yaml, sql and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Validator loads about 3.3k tokens when it runs. Until then it costs about 140 tokens; SKILL.md has 325 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~140
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 325 words, ~3,251 tokens.

Download SKILL.mdSave it as .claude/skills/data-validator/SKILL.md (or your agent's skills folder).
name
data-validator
description
Data validation and quality expertise covering Great Expectations patterns, schema validation, statistical validation, referential integrity checks, data profiling, anomaly detection, data contracts, quality scoring, and automated testing strategies for ensuring data reliability throughout the pipeline. Use when the user asks about data validator, data validator best practices, or needs guidance on data validator implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.
license
Apache-2.0
metadata.author
foundry-skills
metadata.version
1.0.0
metadata.tags
data-science sql testing
metadata.category
data-engineering
metadata.subcategory
pipelines-etl
metadata.disclaimer
none
metadata.difficulty
intermediate

Data Validator

Overview

Data quality is the foundation upon which all downstream analytics, ML models, and business decisions depend. This skill covers tools, techniques, and patterns for validating data at every stage of the pipeline.

Great Expectations Patterns

Setup and Expectation Suites
python
import great_expectations as gx

context = gx.get_context()
datasource = context.sources.add_pandas("pandas_datasource")

# Build expectation suite
suite = context.add_expectation_suite("customer_quality_suite")

# Table-level
suite.add_expectation(gx.expectations.ExpectTableRowCountToBeBetween(min_value=10000, max_value=10000000))
suite.add_expectation(gx.expectations.ExpectTableColumnCountToEqual(value=15))

# Column-level: customer_id
suite.add_expectation(gx.expectations.ExpectColumnValuesToNotBeNull(column="customer_id"))
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeUnique(column="customer_id"))
suite.add_expectation(gx.expectations.ExpectColumnValuesToMatchRegex(
    column="customer_id", regex=r"^CUS-[A-Z0-9]{8}$"
))

# Column-level: email (allow 1% non-matching for legacy data)
suite.add_expectation(gx.expectations.ExpectColumnValuesToMatchRegex(
    column="email",
    regex=r"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$",
    mostly=0.99
))

# Numeric column: revenue
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeBetween(
    column="lifetime_revenue", min_value=0, max_value=10000000, mostly=0.999
))

# Categorical column
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeInSet(
    column="status", value_set=["active", "inactive", "suspended", "pending"]
))
Running Validations in Pipelines
python
def validate_data(**context):
    ge_context = gx.get_context()
    checkpoint = ge_context.add_or_update_checkpoint(
        name="customer_checkpoint",
        validations=[{
            "batch_request": {"datasource_name": "warehouse", "data_asset_name": "dim_customer"},
            "expectation_suite_name": "customer_quality_suite",
        }],
        action_list=[
            {"name": "store_validation_result", "action": {"class_name": "StoreValidationResultAction"}},
            {"name": "update_data_docs", "action": {"class_name": "UpdateDataDocsAction"}},
        ]
    )
    result = checkpoint.run()
    if not result.success:
        raise ValueError("Data quality check failed")

Schema Validation

Schema Contract (YAML)
yaml
name: dim_customer
version: "2.1"
owner: data-engineering
sla_hours: 6
allow_extra_columns: false
min_rows: 50000

columns:
  - name: customer_id
    type: string
    nullable: false
    unique: true
    pattern: "^CUS-[A-Z0-9]{8}$"
  - name: email
    type: string
    nullable: true
    max_length: 255
  - name: lifetime_revenue
    type: float
    nullable: false
    min_value: 0
  - name: status
    type: string
    nullable: false
    allowed_values: ["active", "inactive", "suspended", "pending"]
Schema Drift Detection
python
class SchemaDriftDetector:
    def __init__(self, schema_store):
        self.store = schema_store

    def detect_drift(self, table_name: str, current_df) -> dict:
        previous_schema = self.store.get_schema(table_name)
        current_schema = self._extract_schema(current_df)
        if previous_schema is None:
            self.store.save_schema(table_name, current_schema)
            return {'status': 'new_table', 'changes': []}

        changes = []
        prev_cols = {c['name']: c for c in previous_schema['columns']}
        curr_cols = {c['name']: c for c in current_schema['columns']}

        for name in set(curr_cols) - set(prev_cols):
            changes.append({'type': 'column_added', 'column': name})
        for name in set(prev_cols) - set(curr_cols):
            changes.append({'type': 'column_removed', 'column': name})
        for name in set(prev_cols) & set(curr_cols):
            if prev_cols[name]['dtype'] != curr_cols[name]['dtype']:
                changes.append({'type': 'type_changed', 'column': name,
                    'from': prev_cols[name]['dtype'], 'to': curr_cols[name]['dtype']})

        if changes:
            self.store.save_schema(table_name, current_schema)
        return {'status': 'drift_detected' if changes else 'no_change', 'changes': changes}

Statistical Validation

python
from scipy import stats
import numpy as np

class StatisticalValidator:
    @staticmethod
    def detect_outliers_iqr(series, multiplier=1.5):
        Q1, Q3 = series.quantile(0.25), series.quantile(0.75)
        IQR = Q3 - Q1
        lower, upper = Q1 - multiplier * IQR, Q3 + multiplier * IQR
        outliers = series[(series < lower) | (series > upper)]
        return {'count': len(outliers), 'percentage': len(outliers) / len(series) * 100,
                'lower_bound': lower, 'upper_bound': upper}

    @staticmethod
    def compare_distributions(current, historical, significance=0.05):
        stat, p_value = stats.ks_2samp(current.dropna(), historical.dropna())
        return {'test': 'kolmogorov_smirnov', 'statistic': stat, 'p_value': p_value,
                'significant_drift': p_value < significance}

    @staticmethod
    def validate_proportions(series, expected_proportions):
        observed = series.value_counts(normalize=True).sort_index()
        expected = pd.Series(expected_proportions).sort_index()
        all_cats = sorted(set(observed.index) | set(expected.index))
        observed = observed.reindex(all_cats, fill_value=0)
        expected = expected.reindex(all_cats, fill_value=0)
        stat, p_value = stats.chisquare(observed * len(series), expected * len(series))
        return {'test': 'chi_squared', 'p_value': p_value, 'significant_difference': p_value < 0.05}

Referential Integrity

python
class ReferentialIntegrityChecker:
    def __init__(self, engine):
        self.engine = engine

    def check_fk(self, child_table, child_col, parent_table, parent_col):
        query = f"""
            SELECT COUNT(*) AS orphan_count
            FROM {child_table} c
            LEFT JOIN {parent_table} p ON c.{child_col} = p.{parent_col}
            WHERE p.{parent_col} IS NULL AND c.{child_col} IS NOT NULL
        """
        result = pd.read_sql(query, self.engine).iloc[0]
        return {'child_table': child_table, 'parent_table': parent_table,
                'orphan_count': int(result['orphan_count']),
                'passed': result['orphan_count'] == 0}

Anomaly Detection

python
class AnomalyDetector:
    def detect_volume_anomaly(self, current_count, historical_counts, z_threshold=3.0):
        mean, std = np.mean(historical_counts), np.std(historical_counts)
        if std == 0:
            return {'is_anomaly': current_count != mean}
        z_score = (current_count - mean) / std
        return {'is_anomaly': abs(z_score) > z_threshold, 'z_score': z_score,
                'expected_range': (mean - z_threshold * std, mean + z_threshold * std)}

    def detect_freshness_anomaly(self, latest_timestamp, expected_frequency_hours):
        from datetime import datetime, timezone
        age_hours = (datetime.now(timezone.utc) - latest_timestamp).total_seconds() / 3600
        return {'is_stale': age_hours > expected_frequency_hours * 2, 'age_hours': age_hours}

Data Quality Scoring

python
class DataQualityScorer:
    def __init__(self):
        self.dimensions = {
            'completeness': 0.25, 'uniqueness': 0.15, 'validity': 0.25,
            'consistency': 0.15, 'timeliness': 0.10, 'accuracy': 0.10,
        }

    def score(self, validation_results: dict) -> dict:
        scores = {}
        weighted_total = 0
        for dimension, weight in self.dimensions.items():
            if dimension in validation_results:
                dim_score = validation_results[dimension]
                scores[dimension] = {'score': dim_score, 'weight': weight,
                    'grade': 'A' if dim_score >= 95 else 'B' if dim_score >= 85 else 'C' if dim_score >= 70 else 'D' if dim_score >= 50 else 'F'}
                weighted_total += dim_score * weight
        overall = weighted_total / sum(self.dimensions[d] for d in scores) if scores else 0
        return {'overall_score': round(overall, 2), 'dimensions': scores}

Data Contracts

yaml
contract:
  name: dim_customer
  version: "3.0"
  owner: data-engineering
  consumers: [analytics, marketing-ml, customer-success]

  schema:
    fields:
      - { name: customer_id, type: string, required: true, unique: true }
      - { name: email, type: string, required: false, pii: true }
      - { name: lifetime_revenue, type: "decimal(12,2)", required: true, min: 0 }

  quality:
    freshness: { max_age_hours: 24, field: updated_at }
    volume: { min_rows: 50000, max_row_change_pct: 20 }
    completeness: { email: 95, phone: 80 }

  sla:
    availability: 99.9
    update_frequency: daily

  breaking_changes:
    notification_days: 14
    channels: [{ slack: "#data-contracts" }]

Automated Testing: dbt Integration

sql
-- tests/generic/test_revenue_positive.sql
{% test positive_revenue(model, column_name) %}
SELECT * FROM {{ model }} WHERE {{ column_name }} < 0
{% endtest %}

-- tests/singular/assert_revenue_reconciliation.sql
WITH warehouse AS (
    SELECT SUM(net_amount) AS total FROM {{ ref('fct_orders') }}
    WHERE order_date = '{{ var("check_date") }}'
),
source AS (
    SELECT SUM(amount) AS total FROM {{ source('stripe', 'charges') }}
    WHERE DATE(created) = '{{ var("check_date") }}'
)
SELECT * FROM warehouse w CROSS JOIN source s
WHERE ABS(w.total - s.total) / GREATEST(s.total, 1) > 0.01

Validation Decision Framework

StageWhat to ValidateFailure Action
IngestionSchema, encoding, row countReject file, alert source team
StagingTypes, nulls, basic rangesQuarantine records, log to DLQ
TransformBusiness rules, referential integrityBlock downstream, alert owner
LoadRow counts match, no duplicatesRollback, retry with investigation
ServingFreshness, availability, SLAAlert consumers, serve stale with warning

When to Use

Use this skill when:

  • Designing or implementing data validator solutions
  • Reviewing or improving existing data validator approaches
  • Making architectural or implementation decisions about data validator
  • Learning data validator patterns and best practices
  • Troubleshooting data validator-related issues

Do NOT use this skill when:

  • The question is about a fundamentally different technology domain
  • A more specific sibling skill covers the exact topic needed
  • The user needs a complete hands-on tutorial rather than expert guidance

Output Format

markdown
# Data Validator Analysis

## Context Assessment
[Situation summary and constraints]

## Recommended Approach
[Primary recommendation with rationale]

## Implementation Steps
1. [Step with specific details]
2. [Step with specific details]
3. [Step with specific details]

## Trade-offs and Considerations
- [Key trade-off 1]
- [Key trade-off 2]

## Next Steps
- [Immediate action item]
- [Follow-up action item]

Example

Input: "Help me implement data validator for a medium-scale production application"

Output: A structured analysis covering current state assessment, recommended data validator approach with specific patterns, implementation roadmap with milestones, and risk mitigation strategies tailored to the application scale and constraints.

Edge Cases

  • Legacy system integration: When data validator must coexist with legacy approaches, provide a gradual migration path rather than a complete rewrite
  • Scale mismatch: When the solution complexity exceeds the project scale, recommend a simpler approach and note when to revisit
  • Team skill gaps: When the team lacks experience with the recommended approach, include learning resources and simpler alternatives
  • Conflicting requirements: When constraints conflict (e.g., performance vs. maintainability), explicitly state the trade-off and recommend based on stated priorities

© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/process/resources/skills-library/bodies/skills/data-engineering/data-validator of FerroxLabs/wayland.

Open the folder on GitHubat commit 4c030c7

Compare with similar skills

Data Validator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Validator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Validator this skillFerroxLabs/wayland608—~3.3kAutomated safety check: PassApache-2.0
Data Quality Frameworkswshobson/agents40k10 repos~1.1kAutomated safety check: PassMIT
Profiling Tablesastronomer/agents450—~964Automated safety check: PassApache-2.0
Data QualityJoelLewis/finance_skills205—~11kAutomated safety check: PassMIT
Glue DiagnosticsKilo-Org/kilo-marketplace189—~2kAutomated safety check: PassMIT
Stat Edaasgard-ai-platform/skills241—~954Automated safety check: PassMIT

Similar skills

  • Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

    40k GitHub starsUsed in 10 repos~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Profiling Tables

    astronomer/agents

    Deep-dive data profiling for a specific table. An agent skill from astronomer/agents.

    450 GitHub stars~964 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Data Quality

    JoelLewis/finance_skills

    Design and operate data quality programs for financial data — validation rules, pricing validation, data lineage, exception management, profiling, and governance.

    205 GitHub stars~11k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Glue Diagnostics

    Kilo-Org/kilo-marketplace

    A skill your agent uses to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following…

    189 GitHub stars~2k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • Stat Eda

    asgard-ai-platform/skills

    Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks.

    241 GitHub stars~954 tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Data Researcher

    majiayu000/claude-skill-registry

    Data discovery and analysis specialist focused on extracting actionable insights from complex datasets, identifying patterns and anomalies, and transforming raw data into strategic intelligence.

    666 GitHub starsUsed in 1 repo~4.6k tokens
    Data & AnalyticsAuto-check passed

More from FerroxLabs/wayland

All 1,194 skills in this repo
  • Star Office Helper

    FerroxLabs/wayland

    Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.

    608 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Openclaw Setup

    FerroxLabs/wayland

    OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.

    608 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Tvcontrol Setup

    FerroxLabs/wayland

    Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.

    608 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Ab Testing Specialist

    FerroxLabs/wayland

    End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

    608 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Academic Writer

    FerroxLabs/wayland

    Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…

    608 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Accessibility Auditor

    FerroxLabs/wayland

    Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…

    608 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed

Questions about Data Validator

What does Data Validator do?

Data validation and quality expertise covering Great Expectations patterns, schema validation, statistical validation, referential integrity checks, data profiling, anomaly detection, data…. Data Validator is an agent skill from FerroxLabs/wayland. Data validation and quality expertise covering Great Expectations patterns, schema validation, statistical validation, referential integrity checks, data profiling, anomaly detection, data contracts, quality scoring, and automated testing strategies for ensuring data reliability throughout the pipeline.

When should I use Data Validator?

Data Validator fits situations like: the user asks about data validator; data validator best practices; needs guidance on data validator implementation; the user needs a different specialized skill.

How do I install Data Validator in Claude Code?

Run `npx skills add FerroxLabs/wayland --skill data-validator -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-engineering/data-validator in FerroxLabs/wayland) into .claude/skills/data-validator in your project. Claude Code loads it when a task matches its description.

How do I install Data Validator in Codex?

Run `npx skills add FerroxLabs/wayland --skill data-validator -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-engineering/data-validator in FerroxLabs/wayland) into .agents/skills/data-validator in your project. Codex loads it when a task matches its description.

Can I use Data Validator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill data-validator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-validator, .gemini/skills/data-validator, .github/skills/data-validator and .opencode/skills/data-validator in your project.

What does Data Validator need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Validator is instructions for the agent only. Our summary lists: Python 3.

Does Data Validator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Validator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Validator use?

Data Validator is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Validator use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Validator?

Skills that share tags, products or a category with Data Validator: Data Quality Frameworks (wshobson/agents, 40k stars), Profiling Tables (astronomer/agents, 450 stars), Data Quality (JoelLewis/finance_skills, 205 stars) and Glue Diagnostics (Kilo-Org/kilo-marketplace, 189 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Validator?

FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.

Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.