Agent skill

Verification

by romiluz13 in romiluz13/cc10x

A skill your agent uses when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens.

MITAuto-check: notesAgent Workflows

Install Verification

skills CLI
$ npx skills add romiluz13/cc10x --skill verification -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install romiluz13/cc10x verification --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/cc10x/skills/verification .claude/skills/verification && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verification
GitHub stars
164
Token cost
~1.9k tokens
SKILL.md length
941 words
Files
6 (incl. references)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens.

  • Works in 3 steps: Completion: Did the agent finish its… → Truth: Is the work actually correct?… → Proof: Can you prove it with evidence?…
  • Judging whether a task reached its goal
  • SKILL.md covers Reference Files, The Gate Function, Self-Critique Gate (BEFORE… and Validation Levels, plus 5 more sections
  • Calls python3

What it does

Verification is an agent skill from romiluz13/cc10x. Use when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens. Task completion is not goal achievement.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `evals/README.md`, `evals/eval-01-should-pass-without-running.md` and `evals/eval-02-trust-agent-success-report.md`).

It sits in Agent Workflows. The repository describes itself as: The Loop Engine for Claude Code — engineer the loop, not the prompt. 1 router · 9 agents · 16 skills · 4 workflows. Fail-closed gates, test honesty, anti-anchored review. The licence is MIT.

When your agent uses it

  • Judging whether a task reached its goal
  • Not just finished: the gate function
  • Self-critique gate
  • Validation levels

Example prompts

  • “/verification”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Bash, Grep, Glob

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Completion: Did the agent finish its task? (contract says PASS)
  2. Truth: Is the work actually correct? (independent verification)
  3. Proof: Can you prove it with evidence? (exit codes, test output, screenshots)

What it can do on your machine

Read from SKILL.md and the folder at commit f346ebe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verification loads about 1.9k tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 941 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash, Grep, Glob

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from romiluz13/cc10x at commit f346ebe, republished under its MIT licence (© romiluz13). 941 words, ~1,875 tokens.

Download SKILL.mdSave it as .claude/skills/verification/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
verification
description
Use when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens. Task completion is not goal achievement.
allowed-tools
Read, Bash, Grep, Glob
user-invocable
false

Verification

Task completion is not goal achievement. Verify that the phase achieved its goal, not that prior agents said it did. If you cannot independently reproduce a claimed success, return FAIL.

Reference Files

  • references/live-production-testing.md — live/production verification strategy

The Gate Function

COMPLETION → TRUTH → PROOF
  1. Completion: Did the agent finish its task? (contract says PASS)
  2. Truth: Is the work actually correct? (independent verification)
  3. Proof: Can you prove it with evidence? (exit codes, test output, screenshots)

A PASS without proof is a claim. When a result surprises you in either direction — a pass that came too easily, or a failure you cannot explain — suspect the observation method before the system: confirm the check runs the right command against the current build and exercises the real code path before theorizing about the code.

<!-- Authoring rule (maintenance, not runtime): every gate in this skill encodes an observed failure mode; understand what a gate prevents before removing it. -->

Self-Critique Gate (BEFORE Verification Commands)

Before running any test, audit your own work:

Code Quality:

  • No TODO/FIXME/stubs in changed files
  • No commented-out code
  • No debug logging left in
  • Error handling covers the failure modes the change introduces
  • Types are correct (not as any or as unknown)

Implementation Completeness:

  • Every acceptance criterion from the plan has corresponding code
  • Every named scenario has a test
  • Every error path in the plan has handling
  • No silent failures (empty catches, discarded errors)
  • Build succeeds, type-check passes

Self-Critique Verdict: If any check fails, fix it BEFORE running verification. Don't run tests you know will fail on code quality issues — a known-failing run produces noise evidence that then has to be explained away.

Validation Levels

LevelWhatExit CodeWhen
DeterministicAutomated test0/1Always — the default
ProbabilisticTest with known flake rate0/1 (retry policy)When deterministic impossible (timing, network)
ManualHuman verifies with checklistn/aWhen automation not worth the cost
LiveProduction-like environment0/1When plan requires live proof

Every verification must state its validation level — unlabeled evidence lets weak proof masquerade as strong. If manual, state the checklist. Manual evidence is a validation level for verification, never a substitute for a TDD RED or GREEN exit code: with no test runner and no scripted check, require a runner or block. If deterministic, state the command + exit code. If live, state the harness command.

Production-Like Live Proof

When the plan includes ### Live Verification Strategy or a harness manifest:

  • Run python3 "${CLAUDE_PLUGIN_ROOT}/tools/live_harness_runner.py" --manifest <path> --mode proof
  • If stress required: also run --mode stress
  • Do NOT silently substitute replay fixtures or unit tests for required live proof

Flaky test handling: re-run once — one retry separates environment blips from real flake; more retries launder genuine failures. Pass on re-run → mark PASS with flaky: true. Fail both → FAIL. Never convert flaky pass into unconditional confidence.

Evidence Array Protocol (MANDATORY for PASS)

EVIDENCE:
  scenarios:
    - "[name] | Given [state] | When [action] | [command] → exit [code] | expected=[expected] | actual=[actual]"
  regressions:
    - "[test] → exit [code]: [result]"
  edge_cases:
    - "[case]: [command] → exit [code]: [result]"

Every scenario needs non-empty Expected and Actual. Every scenario maps to exactly one EVIDENCE entry. SCENARIOS_PASSED must equal EVIDENCE.scenarios with exit 0 + Result=PASS. Run each check against the revision being verified; if the code changes after a check, re-run the affected scenarios before citing them. Record the tested identity with the evidence: the commit SHA, plus a patch or digest when the tree is dirty; evidence that does not identify the revision it was produced from is unverified, not PASS.

Goal-Backward Lens

Walk backward from the goal to verify it was achieved:

  1. What was the goal? (re-read the plan's exit criteria)
  2. What would prove it? (name the specific evidence)
  3. Do I have that evidence? (run the verification)
  4. Does the evidence actually prove the goal? (not "tests pass" but "the tests test the right thing")
Show full SKILL.md (360 more words)Show less

Common Failures

FailureWhat happensFix
False greenTest passes without exercising the real code pathTest Honesty Gates (see integration-verifier)
Tautological checkExpected value recomputed the way the code computes it, e.g. expect(add(a, b)).toBe(a + b) — passes by constructionExpected values come from an independent source of truth: a known-good literal, a worked example, the spec
Scope skip"All tests pass" but untested scenarios existGoal-backward lens: name every scenario, verify each
Stale evidence"Tests pass" but you didn't run them this sessionRe-run. Evidence must be from THIS session.
Claim without proof"It works" with no command/exit codeEvidence array is mandatory for PASS
Pointer-as-proofPath cited as PASS without re-opening itRe-open the artifact; if it does not show the claim, downgrade or re-run. Never cite it as PASS
Untested claimA scenario or claim the return never addressesMark it untested with its reason; never assume it passed
Environment escapeTest fails with env signal (command not found, ECONNREFUSED)Classify as ENVIRONMENT not code. Mark BLOCKED.

Excuses and Tells

If you catch yourself in the left column before a PASS, do the right column first.

Excuse or tellReality
"Should work now" / "looks good" / "seems fine" / "probably"RUN the verification command. "Should" is not evidence.
"I'm confident it works"Confidence is not evidence. Exit code 0 is evidence.
"Builder reported success" / "the agent said it passed"Verify independently. Agents can be wrong or sycophantic.
"The tests cover this"Name the specific test — or it isn't covered.
"No regressions detected"List what was tested — or nothing was.
"Tests passed before" (another session)Stale evidence. Re-run in THIS session.
"It's just a typo fix" / "it's trivial"Small changes break things too. Run the tests.
"I'll verify after commit" / committing without fresh evidenceVerify BEFORE commit. A broken commit is harder to revert.
"The build succeeded so it works"Build success ≠ correct behavior. Run behavioral tests.
"I don't need to run the full suite"If your change touches shared code, you do.
Satisfaction before seeing exit code 0Evidence first, satisfaction after.
Weakening an assertion to make a test passThat launders the failure. Fix the code, not the test.

© romiluz13, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in plugins/cc10x/skills/verification of romiluz13/cc10x.

  • SKILL.md
  • evals/README.md
  • evals/eval-01-should-pass-without-running.md
  • evals/eval-02-trust-agent-success-report.md
  • evals/eval-03-partial-evidence-extrapolation.md
  • references/live-production-testing.md

Open the folder on GitHubat commit f346ebe

Compare with similar skills

Verification next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verification compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verification this skillromiluz13/cc10x164—~1.9kAutomated safety check: NotesMIT
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k11 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k35 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers296k2 repos~5.1kAutomated safety check: PassMIT
Claude Code Agent Developmentanthropics/claude-plugins-official38k8 repos~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 11 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 35 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    296k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 8 repos~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    795 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed

More from romiluz13/cc10x

All 22 skills in this repo
  • Building

    romiluz13/cc10x

    A skill your agent uses when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation…

    164 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Diff Driven Docs

    romiluz13/cc10x

    A skill your agent uses when a BUILD phase completes, a commit is staged, or a PR is about to be created, and the diff has not yet been reflected in documentation.

    164 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Planning

    romiluz13/cc10x

    A skill your agent uses when writing an execution plan or a decision RFC: task decomposition, context references, validation levels, risk-based testing, ADR format, plan completeness gate, and…

    164 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Agent Common

    romiluz13/cc10x

    A skill your agent uses when a cc10x agent starts a task: the shared preamble for the memory protocol, the contract format, and the output rules.

    164 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Cc10x Guide

    romiluz13/cc10x

    Answers questions about cc10x itself — what it is, how to install and configure it, how the router, workflows, memory, and hooks operate, and how to troubleshoot.

    164 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Cc10x Router

    romiluz13/cc10x

    Routes build, debug, review, plan, QA, and triage requests through the cc10x workflows (task graphs, workflow artifacts, gates); it is the single entry point for cc10x code work.

    164 GitHub stars~20k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Verification

What does Verification do?

A skill your agent uses when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens. Verification is an agent skill from romiluz13/cc10x. Use when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens.

When should I use Verification?

Verification fits situations like: judging whether a task reached its goal; not just finished: the gate function; self-critique gate; validation levels.

How do I install Verification in Claude Code?

Run `npx skills add romiluz13/cc10x --skill verification -a claude-code`. Or copy the skill folder (plugins/cc10x/skills/verification in romiluz13/cc10x) into .claude/skills/verification in your project. Claude Code loads it when a task matches its description.

How do I install Verification in Codex?

Run `npx skills add romiluz13/cc10x --skill verification -a codex`. Or copy the skill folder (plugins/cc10x/skills/verification in romiluz13/cc10x) into .agents/skills/verification in your project. Codex loads it when a task matches its description.

Can I use Verification in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add romiluz13/cc10x --skill verification -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verification, .gemini/skills/verification, .github/skills/verification and .opencode/skills/verification in your project.

What does Verification need to run?

Going by SKILL.md and its folder, Verification needs the command-line tools its instructions call (python3). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Bash, Grep, Glob.

Does Verification access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verification safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Verification use?

Verification is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verification use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 459 tokens, read only when the agent opens those files.

What are the alternatives to Verification?

Skills that share tags, products or a category with Verification: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 296k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verification?

romiluz13 (a GitHub user) maintains it in romiluz13/cc10x, which has 164 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 7, 2026.

Source: romiluz13/cc10x on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.