Agent skill

Schema Review

by snowplow in snowplow/iglu-central

Reviews JSON Schema files in iglu-central for common issues found in PR reviews.

Apache-2.0Auto-check passedDevelopment

Install Schema Review

skills CLI
$ npx skills add snowplow/iglu-central --skill schema-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install snowplow/iglu-central schema-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/snowplow/iglu-central.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/schema-review .claude/skills/schema-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
schema-review
GitHub stars
123
Token cost
~3.8k tokens
SKILL.md length
1,681 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reviews JSON Schema files in iglu-central for common issues found in PR reviews.

  • Works in 5 steps: Identify files to review → Parse each schema file → Apply validation checklist → …
  • Reviewing schema files before PR submission
  • SKILL.md covers What This Skill Checks, Validation Checklist, Review Process and Output Format, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Schema Review is an agent skill from snowplow/iglu-central. Reviews JSON Schema files in iglu-central for common issues found in PR reviews. Checks field constraints, type definitions, warehouse compatibility, and consistency patterns. Use when reviewing schema files before PR submission or when validating changes to schemas in the schemas/ directory.

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Pull requests. The repository describes itself as: Contains all JSON Schemas, Avros and Thrifts for Iglu Central. The licence is Apache-2.0.

When your agent uses it

  • Reviewing schema files before PR submission
  • Validating changes to schemas in the schemas/ directory

Example prompts

  • “Use the schema-review skill to review JSON Schema files in iglu-central for common issues found in PR reviews”
  • “/schema-review”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Identify files to review
  2. Parse each schema file
  3. Apply validation checklist
  4. Check related schemas for awareness
  5. Categorize findings

What it can do on your machine

Read from SKILL.md and the folder at commit 8efc158. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Schema Review loads about 3.8k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 1,681 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from snowplow/iglu-central at commit 8efc158, republished under its Apache-2.0 licence (© snowplow). 1,681 words, ~3,807 tokens.

Download SKILL.mdSave it as .claude/skills/schema-review/SKILL.md (or your agent's skills folder).
name
schema-review
description
Reviews JSON Schema files in iglu-central for common issues found in PR reviews. Checks field constraints, type definitions, warehouse compatibility, and consistency patterns. Use when reviewing schema files before PR submission or when validating changes to schemas in the schemas/ directory.

JSON Schema Review

Reviews JSON Schema files for issues that commonly appear in PR feedback.

What This Skill Checks

When invoked, reviews schema files for:

  • Field constraints (maxLength, minimum, maximum)
  • Type definitions (nullable fields, integer vs number)
  • Warehouse compatibility (column types, query performance)
  • Required field validation
  • Schema evolution breaking changes (required → optional across patch versions)
  • Documentation completeness
  • Consistency with related schemas

Validation Checklist

String Field Constraints

Check: All string fields have maxLength

Strings without maxLength accept any length, which can cause problems when users (intentionally or accidentally) send very large strings into the warehouse. This can significantly slow down warehouse modeling queries.

Problematic:

json
"field_name": {
  "type": "string",
  "description": "User identifier"
}

Fixed:

json
"field_name": {
  "type": "string",
  "description": "User identifier",
  "maxLength": 128
}

Why: Protects warehouse models from unexpectedly large strings. Also acts as a contract with downstream consumers about what to expect from the data.

What to do: Consider whether this field could realistically receive very large values, and whether you want to protect against that. Comment in your PR on your reasoning either way.

Severity: Warning


Check: Consider trade-offs of minLength = maxLength

When minLength equals maxLength, Snowplow's Redshift loader creates CHAR columns as an optimization. CHAR columns cannot be altered later, so this blocks schema evolution if you need to change the length.

Problematic if flexibility needed:

json
"snowplow_id": {
  "type": "string",
  "minLength": 26,
  "maxLength": 26,
  "pattern": "^[0-9a-zA-Z]{26}$"
}

Fixed for flexibility:

json
"snowplow_id": {
  "type": "string",
  "maxLength": 26,
  "pattern": "^[0-9a-zA-Z]{26}$"
}

Why: Trade-off between optimization and flexibility. CHAR is more efficient for truly fixed-length fields (like three-character country codes), but prevents future schema evolution.

What to do: If you're certain this field will always be exactly this length, minLength = maxLength is fine. If you might want to increase maxLength in future versions, omit minLength. Comment on your reasoning in the PR.

Severity: Warning

Optional Field Types

Check: Consider allowing null for optional fields

Snowplow has an established pattern where optional fields (not in the required array) also allow null in their type definition. This isn't required by JSON Schema, but it provides better cross-platform compatibility.

Current:

json
"properties": {
  "optional_field": {
    "type": "string",
    "description": "Optional identifier"
  }
},
"required": ["some_other_field"]

With pattern applied:

json
"properties": {
  "optional_field": {
    "type": ["string", "null"],
    "description": "Optional identifier"
  }
},
"required": ["some_other_field"]

Why: Cross-platform compatibility. Some languages cannot easily omit map properties and must use null instead. This pattern makes Iglu more permissive about accepting either missing fields or explicit null values, preventing validation errors when payloads contain "optional_field": null. Doesn't break anything if omitted, but following this pattern improves developer experience across different SDK implementations.

Severity: Suggestion

Integer Fields

Check: Integers should have maximum constraint

Integers without maximum default to 64-bit (BIGINT) in warehouses. Setting maximum to 2147483647 (max 32-bit value) creates smaller INT32 columns instead, which is a nice storage optimization for many use cases.

Problematic if optimization wanted:

json
"count": {
  "type": "integer",
  "description": "Number of items"
}

Fixed for 32-bit storage:

json
"count": {
  "type": "integer",
  "description": "Number of items",
  "maximum": 2147483647
}

Why: Storage optimization. But note you cannot easily increase to 64-bit range in future schema versions if you set maximum to 32-bit range now.

What to do: Consider whether this value could ever exceed 32-bit range. If you're confident it won't, the optimization is worthwhile. Be aware of the evolution trade-off.

Severity: Suggestion


Check: Use integer not number for whole numbers

Number type creates double/float columns in warehouses. Use integer for counts, IDs, and other whole numbers.

Problematic:

json
"autonomous_system_number": {
  "type": "number",
  "description": "ASN number"
}

Fixed:

json
"autonomous_system_number": {
  "type": "integer",
  "minimum": 0,
  "maximum": 4294967295,
  "description": "ASN number"
}

Why: Number type creates double columns. Integer type creates int columns, which is more appropriate for whole number data.

What to do: Use number if your data includes decimal values. Use integer if values are always whole numbers. You know your data best - just be aware of the warehouse column type impact.

Severity: Warning

Money and Decimal Fields

Check: Money fields need multipleOf: 0.01

Number type for monetary values creates double columns, which lose precision for financial calculations.

Problematic:

json
"price": {
  "type": "number",
  "description": "Item price"
}

Fixed:

json
"price": {
  "type": "number",
  "multipleOf": 0.01,
  "description": "Item price"
}

Why: Creates DECIMAL columns instead of DOUBLE in warehouses. Preserves precision for monetary calculations.

Severity: Critical (for monetary fields)

Required Field Validation

Check: Field becoming nullable in a non-first schema version

When a schema shares a MODEL number with a previous version (e.g. 1-0-2 when 1-0-1 exists, or 1-1-0 when 1-0-0 exists), compare each field's effective nullability against the previous version. A field is nullable if either of these is true:

  • It is absent from the required array
  • Its type allows null (e.g. ["string", "null"])

Any field that was non-nullable in the previous version and is nullable in the new version is a breaking change for Snowplow's Redshift integration. This covers two distinct cases:

Case 1 — removed from required:

json
// 1-0-1
"required": ["email", "eventName", "dataFields"]

// 1-0-2 — removes email from required (BREAKING)
"required": ["eventName", "dataFields"]

Case 2 — type changed to allow null:

json
// 1-0-1
"email": { "type": "string", "maxLength": 254 }

// 1-0-2 — type now allows null (BREAKING)
"email": { "type": ["string", "null"], "maxLength": 254 }

Both cases mean that a NOT NULL column in Redshift would need to become nullable — a migration that Snowplow's RDB Loader does not support.

Why: Snowplow's RDB Loader treats this as a breaking schema evolution. Rather than attempting the unsupported DDL migration, the loader routes events validated against the new version into a recovery table (named <base_table>_<revision>_<addition>_recovered_<hash>) instead of the main table. This is not immediately visible — events pass validation and appear to load successfully, but land in the wrong table where consumers won't find them.

What to do: Any change that makes a previously non-nullable field nullable requires a new major version (e.g. 2-0-0). Options if a major bump is undesirable:

  • Keep the field required/non-nullable and accept that events omitting it will fail validation
  • Discuss with a maintainer — silently routing events to a recovery table is almost always worse than validation failures

Severity: Critical


Check: Required fields that accept null

Fields in the required array that also accept null is unusual - it means the property must exist but can be null.

Verify intentionality:

json
"properties": {
  "status": {
    "type": ["string", "null"]
  }
},
"required": ["status"]

Why: This pattern is unusual but not strictly wrong. Verify this is what you intend rather than the field being truly required with a value, or optional without being in the required array.

Severity: Warning (verify intent)


Check: Descriptions should match schema validation

If a description makes any claim about the rules of the schema, that claim should be consistent with what the JSON Schema actually validates.

Problematic:

json
"properties": {
  "item_type": {
    "type": "string",
    "description": "Item type (required)"
  }
},
"required": ["item_id"]

Fixed:

json
"properties": {
  "item_type": {
    "type": "string",
    "description": "Item type"
  }
},
"required": ["item_id", "item_type"]

Why: Documentation should accurately reflect what the schema enforces. If the description claims something is required, the schema should require it.

Severity: Warning

Documentation

Check: Properties should have descriptions

Missing descriptions reduce clarity in Console and code generation.

Problematic:

json
"item_id": {
  "type": "string",
  "maxLength": 128
}

Fixed:

json
"item_id": {
  "type": "string",
  "maxLength": 128,
  "description": "Unique identifier for the item"
}

Severity: Suggestion


Check: Description grammar and clarity

Check for typos, grammar issues, unclear wording.

Common issues:

  • "Event occurrs when" → "Event occurs when"
  • "Event occurs when a customer cancels for a subscription" → "Event occurs when a customer cancels a subscription"
  • "The total number items" → "The total number of items"

Severity: Suggestion

Show full SKILL.md (649 more words)Show less
Schema Versioning

Check: Avoid $supersedes and $supersededBy

These fields exist in the Iglu spec but are not used in Iglu Central.

json
"$supersedes": ["1-0-2"]

Why: There is no precedent for using these fields in Iglu Central. This was previously discussed and decided against.

What to do: If you believe you have a special case that requires these fields, flag for discussion with a maintainer before proceeding.

Severity: Warning (flag for review if present)

Consistency Checks

Check: Field definitions across related schemas

Field names used in multiple related schemas sometimes have different types and constraints.

Example variance:

  • Schema A: "location": {"type": "string", "maxLength": 128}
  • Schema B: "location": {"type": "string", "maxLength": 256}
  • Schema C: "location": {"type": ["string", "null"], "maxLength": 128}

Why: Schemas in Iglu Central come from many different sources, so consistency across vendors isn't always expected. However, it's useful to be aware of how similar fields are defined elsewhere.

What to do: Consider whether consistency is appropriate for your use case. Differences may be intentional and justified.

Severity: Suggestion


Check: Enum usage across related schemas

If a field uses enum in one schema, consider whether related schemas would benefit from similar constraints.

Why: Awareness of patterns in related schemas can inform your design decisions.

What to do: Consider whether enum usage would be appropriate for your field based on how it's used in similar contexts.

Severity: Suggestion

Review Process

When this skill is invoked:

  1. Identify files to review

    • If path provided: Review that file
    • If no path: Check git diff for changed schema files
    • Look for files matching: schemas/**/*.jsonschema/*
  2. Parse each schema file

    • Read the JSON schema
    • Extract properties, types, constraints, required fields
    • If the schema is not a first version (i.e. not *-0-0), also read the immediately preceding version to enable evolution checks. This applies to both revision bumps (e.g. 1-0-0 → 1-1-0) and addition bumps (e.g. 1-0-1 → 1-0-2)
  3. Apply validation checklist

    • Check each rule above against the schema
    • Note file path and property name for any issues
  4. Check related schemas for awareness

    • For modified schemas, find others in same vendor/category
    • Compare field definitions to surface potential patterns
  5. Categorize findings

    • Critical: Issues that cause precision loss or break financial calculations
    • Warning: Issues likely to cause problems or that need submitter consideration
    • Suggestion: Documentation, patterns, and awareness items

Output Format

Provide structured results:

✅ Passed ([X] checks)
  • List successful validations
❌ Critical Issues ([X] found)

For each issue:

File: [path] Property: [property_name] Issue: [brief description]

json
// Current
[problematic code]
json
// Suggested fix
[fixed code]

Why: [explanation of impact]

⚠️ Warnings ([X] found)

For each warning:

File: [path] Property: [property_name] Issue: [brief description]

json
// Current
[current code]
json
// Consider
[suggested alternative]

Why: [explanation of trade-offs] What to do: [guidance for submitter]

💡 Suggestions ([X] found)

[Same format as warnings]

📋 Summary
  • Ready for PR: Yes/No (No if critical issues exist)
  • Total issues: [X] critical, [Y] warnings, [Z] suggestions
  • Next steps: [specific actions to resolve issues]

Severity Guidelines

Critical - Causes data precision loss or silently misroutes events:

  • Money fields without multipleOf (precision loss)
  • Field becoming nullable in a non-first schema version — either by removal from required or by adding null to the type (Snowplow's RDB Loader silently routes events to a recovery table rather than the main table)

Warning - Requires submitter consideration and explanation:

  • String fields without maxLength (should comment on reasoning)
  • minLength = maxLength (should explain optimization vs flexibility choice)
  • Number type for whole numbers (should confirm data includes decimals)
  • Required but nullable fields (should confirm intentionality)
  • Description inconsistent with schema validation
  • $supersedes/$supersededBy present (should discuss with maintainer)

Suggestion - Documentation, patterns, and awareness:

  • Optional fields not allowing null (Snowplow pattern for better DX)
  • Integer without 32-bit optimization (storage trade-off)
  • Missing field descriptions
  • Grammar and clarity improvements
  • Consistency patterns with related schemas
  • Enum usage patterns

Repository Context

This is iglu-central, a JSON Schema registry for Snowplow. Schemas define:

  • Event structures
  • Entity contexts
  • Enrichment configurations

Schemas must be warehouse-compatible as they generate DDL for:

  • Redshift (column types, VARCHAR sizing)
  • BigQuery (struct definitions)
  • Databricks/Snowflake (table schemas)

Schema evolution is append-only. Breaking changes require new major versions.

© snowplow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/schema-review of snowplow/iglu-central.

Open the folder on GitHubat commit 8efc158

Compare with similar skills

Schema Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Schema Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Schema Review this skillsnowplow/iglu-central123—~3.8kAutomated safety check: PassApache-2.0
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Check PRonyx-dot-app/onyx32k2 repos~2.3kAutomated safety check: PassMIT
PR Design DocOpenHands/OpenHands90k—~2.4kAutomated safety check: PassMIT
WooCommerce Code Reviewwoocommerce/woocommerce11k3 repos~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Check PR

    onyx-dot-app/onyx

    Checks a GitHub, GitLab, or Perforce (p4) pull request (or merge request, or shelved changelist) for unresolved review comments, failing status checks, and incomplete PR descriptions.

    32k GitHub starsUsed in 2 repos~2.3k tokens
    DevelopmentAuto-check passed
  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • WooCommerce Code Review

    woocommerce/woocommerce

    Reviews WooCommerce code changes against the project's standards, flagging backend PHP architecture, naming, documentation, data integrity and testing violations.

    11k GitHub starsUsed in 3 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Record PR Demo

    payloadcms/payload

    A skill your agent uses when a Payload pull request needs a concise visual walkthrough for reviewers.

    45k GitHub stars~1k tokensUpdated yesterday
    DevelopmentAuto-check passed

Categories

Questions about Schema Review

What does Schema Review do?

Reviews JSON Schema files in iglu-central for common issues found in PR reviews. Schema Review is an agent skill from snowplow/iglu-central. Reviews JSON Schema files in iglu-central for common issues found in PR reviews.

When should I use Schema Review?

Schema Review fits situations like: reviewing schema files before PR submission; validating changes to schemas in the schemas/ directory.

How do I install Schema Review in Claude Code?

Run `npx skills add snowplow/iglu-central --skill schema-review -a claude-code`. Or copy the skill folder (.claude/skills/schema-review in snowplow/iglu-central) into .claude/skills/schema-review in your project. Claude Code loads it when a task matches its description.

How do I install Schema Review in Codex?

Run `npx skills add snowplow/iglu-central --skill schema-review -a codex`. Or copy the skill folder (.claude/skills/schema-review in snowplow/iglu-central) into .agents/skills/schema-review in your project. Codex loads it when a task matches its description.

Can I use Schema Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add snowplow/iglu-central --skill schema-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/schema-review, .gemini/skills/schema-review, .github/skills/schema-review and .opencode/skills/schema-review in your project.

What does Schema Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Schema Review is instructions for the agent only.

Does Schema Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Schema Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Schema Review use?

Schema Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Schema Review use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Schema Review?

Skills that share tags, products or a category with Schema Review: Finishing a Development Branch (obra/superpowers, 297k stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars), Check PR (onyx-dot-app/onyx, 32k stars) and PR Design Doc (OpenHands/OpenHands, 90k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Schema Review?

snowplow (a GitHub organization) maintains it in snowplow/iglu-central, which has 123 GitHub stars. The repository was last updated on October 1, 2026.

Source: snowplow/iglu-central on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.