Agent skill

Validation Methodology

by prime-radiant-inc in prime-radiant-inc/greenfield

Cross-cutting validation discipline. An agent skill from prime-radiant-inc/greenfield.

Apache-2.0Auto-check passedTesting & QA

Install Validation Methodology

skills CLI
$ npx skills add prime-radiant-inc/greenfield --skill validation-methodology -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install prime-radiant-inc/greenfield validation-methodology --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/prime-radiant-inc/greenfield.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/validation-methodology .claude/skills/validation-methodology && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
validation-methodology
GitHub stars
292
Token cost
~3k tokens
SKILL.md length
1,093 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
Apache-2.0

At a glance

Cross-cutting validation discipline. An agent skill from prime-radiant-inc/greenfield.

  • Tasks that involve Quality gates
  • SKILL.md covers Acceptance Criteria Format, Verification Methods, Quality Gate Criteria and Definition of Done Checklists, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve User stories

What it does

Validation Methodology is an agent skill from prime-radiant-inc/greenfield. Cross-cutting validation discipline. Acceptance criteria format, definition of done checklists, quality gate criteria, verification methods. Loaded by every analysis agent.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Quality gates and User stories. The repository describes itself as: A Claude Code plugin that reverse-engineers clean behavioral specs, test vectors, and acceptance criteria from any codebase, producing a provenance trail so a fresh team can… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Quality gates
  • Tasks that involve User stories

Example prompts

  • “/validation-methodology”

What it can do on your machine

Read from SKILL.md and the folder at commit 6e6d4b4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown and dot).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Validation Methodology loads about 3k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 1,093 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from prime-radiant-inc/greenfield at commit 6e6d4b4, republished under its Apache-2.0 licence (© prime-radiant-inc). 1,093 words, ~2,954 tokens.

Download SKILL.mdSave it as .claude/skills/validation-methodology/SKILL.md (or your agent's skills folder).
name
validation-methodology
description
Cross-cutting validation discipline. Acceptance criteria format, definition of done checklists, quality gate criteria, verification methods. Loaded by every analysis agent.

Validation Methodology

Specifications are claims. Claims require proof. This skill defines how the pipeline transforms claims into proof.

Acceptance Criteria Format

Every module spec receives formal acceptance criteria during Layer 4. Criteria use Given/When/Then format:

markdown
### AC-CONFIG-001: Config File Load
**Given** a valid config file exists at the documented path
**When** the application starts
**Then**
- The application reads every declared key from the file
- Unknown keys are logged as warnings, not errors
- Missing optional keys fall back to documented defaults

**Source:** official-docs, source-code, runtime-observation
**Confidence:** confirmed
**Priority:** P0
**Verification:** Automated test
ID Format: AC-{DOMAIN}-{NNN}
  • {DOMAIN}: Short uppercase token for the behavioral domain (matches the spec's domain tokens — e.g., CLI, CONFIG, API, AUDIO, PARSER)
  • {NNN}: Zero-padded three-digit sequence (001, 002, ...)
  • IDs are immutable — retired, never reused
Field Rules

Given/When/Then:

  • Use observable terms, not implementation details
  • "Given a valid config file exists" not "Given the ConfigLoader is initialized"
  • Each Then outcome must be verifiable without implementation internals
  • Describe observable state, not UI mechanisms — "system rejects malformed input" not "error banner is visible." Criteria must be testable across any frontend platform.

Priority:

  • P0: Must pass. Blocks pipeline. Implementation cannot ship without satisfying.
  • P1: Should pass. Reported as warnings. Divergence requires documented rationale.
  • P2: Nice to pass. Advisory.

Verification: One of the 5 methods below.

Output Location
workspace/raw/specs/validation/acceptance-criteria/
    _index.md                   # Summary with counts and links
    config-loading.md           # AC-CONFIG-001 through AC-CONFIG-NNN
    cli-flags.md                # AC-CLI-001 through ...
    ...

Verification Methods

Every acceptance criterion specifies how the implementer will verify it:

MethodIDWhen to Use
Automated Testautomated-testPrecondition, action, and outcome are all programmable
Manual Inspectionmanual-inspectionOutcome requires human judgment (UX quality, clarity)
Runtime Observationruntime-observationBehavior involves timing, async events, sustained observation
Protocol Captureprotocol-captureWire format, headers, API protocol compliance
Code Reviewcode-reviewArchitectural or quality criteria not directly testable at runtime

Quality Gate Criteria

dot
digraph gate_evaluation {
    rankdir=TB;

    "Agent produces output" [shape=ellipse];
    "Check DOD checklist for agent layer" [shape=box];
    "All DOD items pass?" [shape=diamond];
    "Proceed to gate evaluation" [shape=box];
    "Flag incomplete output for review" [shape=box];
    "Check gate criteria" [shape=box];
    "All gate criteria pass?" [shape=diamond];
    "Tag and proceed to next layer" [shape=box];
    "Remediate findings" [shape=box];
    "Attempts exhausted?" [shape=diamond];
    "STOP: Pipeline blocked" [shape=octagon, style=filled, fillcolor=red, fontcolor=white];

    "Agent produces output" -> "Check DOD checklist for agent layer";
    "Check DOD checklist for agent layer" -> "All DOD items pass?";
    "All DOD items pass?" -> "Proceed to gate evaluation" [label="yes"];
    "All DOD items pass?" -> "Flag incomplete output for review" [label="no"];
    "Flag incomplete output for review" -> "Proceed to gate evaluation";
    "Proceed to gate evaluation" -> "Check gate criteria";
    "Check gate criteria" -> "All gate criteria pass?";
    "All gate criteria pass?" -> "Tag and proceed to next layer" [label="yes"];
    "All gate criteria pass?" -> "Remediate findings" [label="no"];
    "Remediate findings" -> "Attempts exhausted?";
    "Attempts exhausted?" -> "STOP: Pipeline blocked" [label="yes"];
    "Attempts exhausted?" -> "Check gate criteria" [label="no"];
}
Gate 1 (spec-verifier, after Layer 3)
  • Zero contradictions between specs
  • All crypto claims and constants verified against source (not assumed)
  • Every behavioral claim has a provenance citation
  • Assumed claims are a small minority — most claims have direct evidence. A module dominated by assumptions needs more intelligence gathering before proceeding
  • All modules have specs
  • Zero uncited behavioral claims in output candidates
Gate 2 (spec-reviewer, after Layer 4)
  • Zero implementation leakage in specs
  • Zero P0 completeness gaps
  • All ACs have IDs and are linked to specs
  • P0 ACs have test vectors
  • All ACs are testable
  • All modules have ACs
  • Gap report reviewed
  • Zero contamination in output
Implementation Leakage Definition

"Zero implementation leakage" means every identifier, name, and reference in the output specs passes the reimplementor test: "Could the reimplementor reasonably redesign this?" If yes, it's an implementation detail that should have been abstracted. If no, it's an external contract that belongs in the spec.

Implementation details (must not appear in output specs):

CategoryWhat to look for
Internal namesVariable names, function signatures, class names, method names from the source — any language
Internal architectureModule boundaries, file organization, inheritance hierarchies described in prose
Framework-specific patternsState management internals, ORM patterns, framework lifecycle hooks by internal name
Build/deployment artifactsSource file paths, line numbers, chunk IDs, minified identifiers
Code structure language"calls X then Y", "inherits from Z", "implements interface W"
Internal feature gatesFeature flags, A/B test names, internal telemetry event names

External contracts (must be preserved):

CategoryWhat to look for
User-facing identifiersCLI flags, env vars, config keys, config file paths
Wire protocol fieldsAPI endpoints, request/response field names, header names
External system schemasDatabase tables/columns (shared), third-party API contracts, external CLI tool flags
Published constantsError message text, exit codes, timeout values
Standard namesProtocol names (OAuth, JWT), encoding names (UTF-8), algorithm names (SHA-256)

See the spec-sanitization skill for the full semantic classification (Implementation Detail vs. External Contract) and rewrite guidance.

Definition of Done Checklists

Layer 1 Agent DOD
  • All output files at correct workspace directory for source origin
  • All behavioral claims have <!-- cite: --> provenance citations
  • Output follows standard format for this agent's intelligence source
  • Session JSONL captured to workspace/provenance/sessions/
  • No raw source code excerpts longer than one line
Layer 2 Agent DOD
  • Feature inventory covers all sources consulted
  • Architecture document covers all identified components
  • API surface covers all discovered interfaces
  • Module map is complete (name, description, priority, dependencies)
  • No known gaps (or gaps explicitly documented with [GAP] marker)
  • Cross-references to Layer 1 artifacts are valid
  • Session JSONL captured
Show full SKILL.md (446 more words)Show less
Layer 3 Agent DOD
  • Spec covers: entry points, decision trees, state machines, error handling, edge cases
  • Zero implementation details (no function names, variable names, line numbers)
  • All behavioral claims have provenance citations
  • Assumed claims are a small minority; most claims have direct evidence
  • Self-assessment verification checklist completed at end of spec
  • Session JSONL captured
Layer 3 Spec Self-Assessment Checklist

Every Layer 3 spec file ends with:

markdown
## Verification Checklist
- [ ] All entry points documented
- [ ] All decision trees traced (every branch, including error paths)
- [ ] All state machines documented (states, transitions, triggers)
- [ ] All error conditions listed with observable symptoms
- [ ] All edge cases identified and documented
- [ ] Zero source code identifiers in this document
- [ ] All behavioral claims have provenance citations
- [ ] Assumed claims are a small minority; most claims have direct evidence
- [ ] No sections are empty or placeholder-only
- [ ] Cross-references to other specs are valid
Layer 4 Agent DOD
  • Test vectors have concrete input/output pairs (not abstract descriptions)
  • Runnable test harness scripts produced where feasible
  • ACs have AC-{MODULE}-{NNN} IDs and Given/When/Then structure
  • Every AC has a verification method assigned
  • Every AC has provenance links
  • Every P0 module has test vectors and ACs
  • Session JSONL captured
Gate Agent DOD
  • Report written to correct path
  • Clear PASS or FAIL status in summary
  • Every failing criterion individually documented
  • Quantitative summary (claims checked, contradictions, gaps, coverage %)
  • Session JSONL captured
Layer 5 Agent DOD
  • All raw specs have clean counterparts in workspace/output/
  • Zero source code identifiers in the output files
  • All 17 contamination detection patterns return 0 on workspace/output/ (see Implementation Leakage Definition)
  • All behavioral information preserved (feature count, AC count identical)
  • Provenance metadata preserved (confidence, source types present)
  • Validation artifacts sanitized and copied to the output
  • Session JSONL captured
Layer 6 Agent DOD (Second-Pass Review)
  • All 8 structural leakage checks executed
  • All 12 content contamination checks executed
  • All 9 behavioral completeness checks executed
  • Detailed audit reports written to workspace/raw/audit/
  • Clean audit summaries written to workspace/output/audit/ (outcomes only)
  • Overall PASS determination (or FAIL after 3 remediation attempts)
  • review-complete git tag applied
  • Session JSONL captured

Confidence Principles

The goal is specs you'd trust enough to implement from. Apply judgment, not arithmetic:

  • Every claim needs provenance. An uncited behavioral claim is an unverified claim. Zero uncited claims in output — this is a hard gate.
  • Assumed claims are a warning sign. A few assumptions are inevitable (indirect reasoning, naming conventions). But a module where most claims are assumed hasn't been analyzed — it's been guessed at. Gather more intelligence.
  • Confirmed claims build trust. Claims corroborated by multiple independent sources (source + runtime, source + docs) are the foundation. Critical modules should be dominated by confirmed claims.
  • P0 behaviors need test vectors. Every P0 acceptance criterion must have a concrete test vector. No exceptions.

When Confidence Is Low

  • Too many assumptions: Dispatch additional intelligence-gathering agents. Try a different mode (runtime observation often confirms what source analysis can only infer).
  • Claims lack corroboration: Check if other modes have evidence. A claim from source alone is inferred; the same claim confirmed by runtime observation becomes confirmed.
  • Uncited claims in output/: Hard failure — pipeline blocked until resolved. Trace each uncited claim back to its originating agent and require a citation.

© prime-radiant-inc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/validation-methodology of prime-radiant-inc/greenfield.

Open the folder on GitHubat commit 6e6d4b4

Compare with similar skills

Validation Methodology next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Validation Methodology compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Validation Methodology this skillprime-radiant-inc/greenfield292—~3kAutomated safety check: PassApache-2.0
QA Reviewdigipulse-engineering/GAAI-framework163—~3.4kAutomated safety check: PassCustom licence
Validation Firsthashgraph-online/awesome-codex-plugins1.3k—~4.3kAutomated safety check: PassApache-2.0
Build Doddanshapiro/kilroy222—~2.5kAutomated safety check: PassMIT
Build Scenario Teststamdogood/builder-essential-skills220—~1.7kAutomated safety check: PassMIT
Verification Gatesrohitg00/skillkit1.5k—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • QA Review

    digipulse-engineering/GAAI-framework

    Validate that implemented code fully satisfies Story acceptance criteria, respects rules, and introduces no regressions.

    163 GitHub stars~3.4k tokensUpdated 10 days ago
    Testing & QAAuto-check passed
  • Validation First

    hashgraph-online/awesome-codex-plugins

    Validation-first design for AI agent output — every spec requirement must be automatically verifiable.

    1.3k GitHub stars~4.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Build Dod

    danshapiro/kilroy

    A skill your agent uses when converting a spec, requirements document, or goal statement into a Definition of Done with acceptance criteria and integration test scenarios

    222 GitHub stars~2.5k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Build Scenario Tests

    tamdogood/builder-essential-skills

    Inspect an unfamiliar repository, turn a focused Markdown behavior scenario into a deterministic test in the repository's native test stack, run it, and preserve traceability between intent and code.

    220 GitHub stars~1.7k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Verification Gates

    rohitg00/skillkit

    Creates explicit validation checkpoints (verification gates) between project phases to catch errors early and ensure quality before proceeding.

    1.5k GitHub stars~1.7k tokensUpdated 4 mo ago
    Product & Project ManagementAuto-check passed
  • Test Scenarios

    phuryn/pm-skills

    Create comprehensive test scenarios from user stories with test objectives, starting conditions, user roles, step-by-step actions, and expected outcomes.

    27k GitHub stars~866 tokensUpdated 25 days ago
    Testing & QAAuto-check passed

More from prime-radiant-inc/greenfield

All 21 skills in this repo
  • Reverse Engineering Analysis Pipeline

    prime-radiant-inc/greenfield

    Master methodology for reverse-engineering a codebase into behavioral specs with cited evidence, reading every line across source, binaries, docs, runtime and git history.

    292 GitHub stars~3.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Community Intelligence Research

    prime-radiant-inc/greenfield

    Mines tutorials, forums, reviews, issues and changelogs for observed product behavior, using six search channels and consensus analysis.

    292 GitHub stars~4.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Containerized Target Execution

    prime-radiant-inc/greenfield

    Runs untrusted analysis targets inside Docker or Podman containers with memory, CPU and process limits, covering image builds, lifecycle, command execution and cleanup.

    292 GitHub stars~2.1k tokensUpdated 2 mo ago
    Auto-check passed
  • API Contract Detection

    prime-radiant-inc/greenfield

    Finds OpenAPI, GraphQL, Protobuf and JSON Schema files in a codebase and extracts behavioral claims from them as part of a reverse-engineering workflow.

    292 GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Documentation Research Methodology

    prime-radiant-inc/greenfield

    Method for extracting behavioral specifications from a product's public documentation: tiered search order, claim extraction rules, output structure, stop criteria and gap analysis.

    292 GitHub stars~4.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Ecosystem Analysis

    prime-radiant-inc/greenfield

    Layer 1 skill for SDK and ecosystem analysis. An agent skill from prime-radiant-inc/greenfield.

    292 GitHub stars~2.9k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Validation Methodology

What does Validation Methodology do?

Cross-cutting validation discipline. An agent skill from prime-radiant-inc/greenfield. Validation Methodology is an agent skill from prime-radiant-inc/greenfield. Cross-cutting validation discipline.

When should I use Validation Methodology?

Validation Methodology fits situations like: tasks that involve Quality gates; tasks that involve User stories.

How do I install Validation Methodology in Claude Code?

Run `npx skills add prime-radiant-inc/greenfield --skill validation-methodology -a claude-code`. Or copy the skill folder (skills/validation-methodology in prime-radiant-inc/greenfield) into .claude/skills/validation-methodology in your project. Claude Code loads it when a task matches its description.

How do I install Validation Methodology in Codex?

Run `npx skills add prime-radiant-inc/greenfield --skill validation-methodology -a codex`. Or copy the skill folder (skills/validation-methodology in prime-radiant-inc/greenfield) into .agents/skills/validation-methodology in your project. Codex loads it when a task matches its description.

Can I use Validation Methodology in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add prime-radiant-inc/greenfield --skill validation-methodology -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validation-methodology, .gemini/skills/validation-methodology, .github/skills/validation-methodology and .opencode/skills/validation-methodology in your project.

What does Validation Methodology need to run?

SKILL.md names no scripts, command-line tools or credentials: Validation Methodology is instructions for the agent only.

Does Validation Methodology access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Validation Methodology safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Validation Methodology use?

Validation Methodology is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Validation Methodology use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Validation Methodology?

Skills that share tags, products or a category with Validation Methodology: QA Review (digipulse-engineering/GAAI-framework, 163 stars), Validation First (hashgraph-online/awesome-codex-plugins, 1.3k stars), Build Dod (danshapiro/kilroy, 222 stars) and Build Scenario Tests (tamdogood/builder-essential-skills, 220 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Validation Methodology?

prime-radiant-inc (a GitHub organization) maintains it in prime-radiant-inc/greenfield, which has 292 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on August 6, 2026.

Source: prime-radiant-inc/greenfield on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.