Agent skill

Surrogate Verifier

by Mathews-Tom in Mathews-Tom/armory

Generates structured test assertions and failure diagnostics for skill packages from a definition and task prompt.

MITAuto-check passedDevelopment

Install Surrogate Verifier

skills CLI
$ npx skills add Mathews-Tom/armory --skill surrogate-verifier -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory surrogate-verifier --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/surrogate-verifier .claude/skills/surrogate-verifier && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
surrogate-verifier
GitHub stars
328
Token cost
~2.3k tokens
SKILL.md length
870 words
Files
4 (incl. references)
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

Generates structured test assertions and failure diagnostics for skill packages from a definition and task prompt.

  • Works in 7 steps: Skill Analysis → Assertion Design → Output → …
  • : verify this skill
  • SKILL.md covers Reference Files, Information Isolation Protocol, Workflow and Budget Parameters, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Surrogate Verifier is an agent skill from Mathews-Tom/armory. Generates structured test assertions and failure diagnostics for skill packages from a definition and task prompt. Triggers on: "verify this skill", "generate assertions", "surrogate verification", "diagnose skill failure". NOT for code review, use pr-review.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `evals/cases.yaml`, `references/assertion-patterns.md` and `references/diagnostic-templates.md`).

It sits in Development, covering Pull requests. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • : verify this skill
  • Generate assertions
  • Surrogate verification
  • Diagnose skill failure

Example prompts

  • “verify this skill”
  • “generate assertions”
  • “surrogate verification”
  • “/surrogate-verifier”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Skill Analysis
  2. Assertion Design
  3. Output
  4. Failure Classification
  5. Root-Cause Analysis
  6. Remediation Suggestions
  7. Diagnostic Output

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Surrogate Verifier loads about 2.3k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 870 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 870 words, ~2,306 tokens.

Download SKILL.mdSave it as .claude/skills/surrogate-verifier/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
surrogate-verifier
description
Generates structured test assertions and failure diagnostics for skill packages from a definition and task prompt. Triggers on: "verify this skill", "generate assertions", "surrogate verification", "diagnose skill failure". NOT for code review, use pr-review.
metadata.version
1.0.1
metadata.category
review
metadata.tags
verification, assertions, testing, co-evolution, diagnostics, eval
metadata.difficulty
advanced
metadata.phase
verify

Surrogate Verifier

Generate structured test assertions and failure diagnostics for skill packages through information-isolated verification. The verifier operates without access to the skill generator's reasoning — it sees only the skill definition, a task prompt, and the output artifacts. This isolation prevents confirmation bias and is the single largest contributor to skill quality in co-evolutionary generation (+30pp per EvoSkills).

Reference Files

FileContentsLoad When
references/assertion-patterns.mdAssertion catalog by skill category with weight guidanceAlways
references/diagnostic-templates.mdFailure diagnostic templates with root-cause categoriesWhen producing failure reports

Information Isolation Protocol

This is the most critical constraint. Violating isolation degrades verification quality.

The verifier MUST NOT access:

  • The generator's conversation history or reasoning chain
  • Prior evolution iterations or refinement context
  • The generator's internal notes or decision rationale
  • Any context beyond what is explicitly listed below

The verifier receives ONLY:

  1. The skill's SKILL.md content (the definition file)
  2. One or more task prompts representing intended use
  3. The skill's output (when diagnosing failures)
  4. The assertion results from scripts/eval_assertions.py (when diagnosing)

Implementation: When invoked by the test-engineer agent, this skill MUST be loaded into a separate Agent spawn using isolation: "worktree" or at minimum a fresh session with no shared context. The invoking agent passes artifacts as explicit text, not as conversation references.

Workflow

Mode 1: Assertion Generation

Generate assertions for a skill given its definition and task prompts.

Phase 1: Skill Analysis

Read the SKILL.md definition and extract:

  1. Stated capabilities — what the skill claims to do (from description + workflow sections)
  2. Output format — expected structure of the skill's output (markdown, JSON, tables, etc.)
  3. Error handling — documented failure modes and recovery paths
  4. Prerequisites — required tools, dependencies, or context
  5. Trigger boundaries — what the skill does NOT handle (negative scope)
Phase 2: Assertion Design

For each task prompt, generate 5-10 assertions covering these dimensions:

DimensionAssertion Types to UsePurpose
Output completenesscontains, matches_regexAll claimed sections/components present
Format complianceoutput_format, containsOutput matches declared structure
Factual signalscontains, not_containsKey domain terms present, hallmarks absent
Tool usagecalls_toolExpected tools were invoked
Negative constraintsnot_containsForbidden patterns absent

Weight assignment:

  • Output completeness assertions: weight 1.0 (must have)
  • Format compliance: weight 0.8 (structural correctness)
  • Factual signals: weight 0.6 (content quality)
  • Tool usage: weight 0.5 (method verification)
  • Negative constraints: weight 0.3 (absence checks are weaker signals)

See references/assertion-patterns.md for category-specific assertion catalogs.

Phase 3: Output

Produce assertions in the evals/cases.yaml schema format:

yaml
assertions:
  - type: contains
    target: "## Scalability"
    weight: 1.0
  - type: output_format
    target: markdown_table
    weight: 0.8
  - type: not_contains
    target: "TODO"
    weight: 0.3
  - type: calls_tool
    target: Read
    weight: 0.5

Context cap: Do not consume more than 70% of the available context window. If the skill definition is very long, focus assertion generation on the workflow phases and output format sections. Summarize rather than quote verbatim.

Mode 2: Failure Diagnostics

When an oracle returns fail, produce a structured diagnostic explaining why.

Input
  • The skill's SKILL.md (same as Mode 1)
  • The task prompt that was executed
  • The output that failed
  • The assertion results: which passed, which failed, with details
Phase 1: Failure Classification

Categorize each failed assertion into a root-cause category:

CategorySignalSeverity
Missing capabilitycontains assertion failed for a claimed featureHIGH
Format mismatchoutput_format assertion failedHIGH
Incomplete outputMultiple contains assertions failed in the same sectionMEDIUM
Hallucinated contentnot_contains assertion failed (forbidden pattern present)HIGH
Wrong tool usagecalls_tool assertion failedMEDIUM
Partial successSome assertions in a group pass, others failLOW
Show full SKILL.md (323 more words)Show less
Phase 2: Root-Cause Analysis

For each failed assertion:

  1. Identify the specific section of SKILL.md that promises the missing capability
  2. Compare what the skill definition instructs vs. what the output actually contains
  3. Hypothesize why the gap exists (missing workflow step, ambiguous instruction, wrong tool choice)
Phase 3: Remediation Suggestions

For each root cause, produce a concrete, actionable fix:

  • Missing capability: "Add a workflow step between Phase 2 and Phase 3 that explicitly generates [X]"
  • Format mismatch: "Change the output format instruction from 'produce a summary' to 'produce a markdown table with columns: [A, B, C]'"
  • Hallucinated content: "Add a negative constraint in the workflow: 'Do NOT include [X] unless [condition]'"
  • Wrong tool usage: "Replace 'use Bash to read the file' with 'use the Read tool for file contents'"
Phase 4: Diagnostic Output

Produce a structured diagnostic string:

DIAGNOSTIC: [skill-name] failed on [task-prompt-summary]

FAILED ASSERTIONS (N/M):
  1. [SEVERITY] type=contains target="..." — Missing capability: [explanation]
  2. [SEVERITY] type=output_format target="..." — Format mismatch: [explanation]

ROOT CAUSES:
  - [category]: [specific explanation with SKILL.md section reference]

REMEDIATION:
  1. [Concrete change to SKILL.md with exact section and wording]
  2. [Concrete change to workflow with step numbers]

See references/diagnostic-templates.md for worked examples per root-cause category.

Budget Parameters

Per EvoSkills Algorithm 1:

  • Context cap: 0.7 (70% of available context window)
  • Max surrogate retries: 15 per oracle round
  • Max oracle rounds: 5 (enforced by the orchestrating agent, not the verifier)

The verifier does not track its own budget — the test-engineer agent manages iteration limits.

Limitations

  • No execution capability: The verifier generates assertions but does not execute them. Execution is handled by scripts/run_evals.py (the oracle).
  • Text-only verification: Cannot verify visual outputs, interactive behaviors, or side effects. Assertions operate on the textual output only.
  • Single-turn scope: Each verification is independent. The verifier does not remember prior rounds (the orchestrating agent feeds context as needed).
  • Assertion granularity: The 5 assertion types cover common patterns but not all possible verification needs. Custom assertion types require extending scripts/eval_assertions.py.

Error Handling

ErrorResolution
Skill definition too largeSummarize to workflow phases + output format sections only
No assertions generatableReturn empty assertions list with warning; skill may be too vague
Ambiguous output formatDefault to contains assertions; avoid output_format checks
Context cap exceededTruncate diagnostic detail; preserve failed assertion list

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/surrogate-verifier of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml
  • references/assertion-patterns.md
  • references/diagnostic-templates.md

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Surrogate Verifier next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Surrogate Verifier compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Surrogate Verifier this skillMathews-Tom/armory328—~2.3kAutomated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Check PRonyx-dot-app/onyx32k2 repos~2.3kAutomated safety check: PassMIT
PR Design DocOpenHands/OpenHands90k—~2.4kAutomated safety check: PassMIT
WooCommerce Code Reviewwoocommerce/woocommerce11k3 repos~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Check PR

    onyx-dot-app/onyx

    Checks a GitHub, GitLab, or Perforce (p4) pull request (or merge request, or shelved changelist) for unresolved review comments, failing status checks, and incomplete PR descriptions.

    32k GitHub starsUsed in 2 repos~2.3k tokens
    DevelopmentAuto-check passed
  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • WooCommerce Code Review

    woocommerce/woocommerce

    Reviews WooCommerce code changes against the project's standards, flagging backend PHP architecture, naming, documentation, data integrity and testing violations.

    11k GitHub starsUsed in 3 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Record PR Demo

    payloadcms/payload

    A skill your agent uses when a Payload pull request needs a concise visual walkthrough for reviewers.

    45k GitHub stars~1k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    328 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    328 GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    328 GitHub stars~2.8k tokensUpdated 3 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    328 GitHub stars~3.1k tokensUpdated 3 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    328 GitHub stars~4.9k tokensUpdated 3 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    328 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about Surrogate Verifier

What does Surrogate Verifier do?

Generates structured test assertions and failure diagnostics for skill packages from a definition and task prompt. Surrogate Verifier is an agent skill from Mathews-Tom/armory. Generates structured test assertions and failure diagnostics for skill packages from a definition and task prompt.

When should I use Surrogate Verifier?

Surrogate Verifier fits situations like: : verify this skill; generate assertions; surrogate verification; diagnose skill failure.

How do I install Surrogate Verifier in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill surrogate-verifier -a claude-code`. Or copy the skill folder (skills/surrogate-verifier in Mathews-Tom/armory) into .claude/skills/surrogate-verifier in your project. Claude Code loads it when a task matches its description.

How do I install Surrogate Verifier in Codex?

Run `npx skills add Mathews-Tom/armory --skill surrogate-verifier -a codex`. Or copy the skill folder (skills/surrogate-verifier in Mathews-Tom/armory) into .agents/skills/surrogate-verifier in your project. Codex loads it when a task matches its description.

Can I use Surrogate Verifier in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill surrogate-verifier -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/surrogate-verifier, .gemini/skills/surrogate-verifier, .github/skills/surrogate-verifier and .opencode/skills/surrogate-verifier in your project.

What does Surrogate Verifier need to run?

SKILL.md names no scripts, command-line tools or credentials: Surrogate Verifier is instructions for the agent only.

Does Surrogate Verifier access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Surrogate Verifier safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Surrogate Verifier use?

Surrogate Verifier is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Surrogate Verifier use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Surrogate Verifier?

Skills that share tags, products or a category with Surrogate Verifier: Finishing a Development Branch (obra/superpowers, 297k stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars), Check PR (onyx-dot-app/onyx, 32k stars) and PR Design Doc (OpenHands/OpenHands, 90k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Surrogate Verifier?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 328 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.