Agent skill

Verification Before Completion

by softspark in softspark/ai-toolkit

Forces verification commands before success claims. An agent skill from softspark/ai-toolkit.

Apache-2.0Auto-check passedAgent Workflows

Install Verification Before Completion

skills CLI
$ npx skills add softspark/ai-toolkit --skill verification-before-completion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install softspark/ai-toolkit verification-before-completion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/app/skills/verification-before-completion .claude/skills/verification-before-completion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verification-before-completion
GitHub stars
179
Token cost
~2.2k tokens
SKILL.md length
1,108 words
Files
1
Skills in repo
112
Repo updated
First seen
Licence
Apache-2.0

At a glance

Forces verification commands before success claims. An agent skill from softspark/ai-toolkit.

  • Works in 4 steps: Define the rubric BEFORE building — 3–6… → Launch the app and observe — actually… → Score with a fresh evaluator — have an… → …
  • Tasks that involve Verification before completion
  • SKILL.md covers The Iron Law, The Gate Function, When To Apply and Don't Assume It Exists, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Verification Before Completion is an agent skill from softspark/ai-toolkit. Forces verification commands before success claims. Evidence before assertions. Triggers: complete, fixed, passing, done, ready, verified.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Verification before completion. The repository describes itself as: Professional-grade AI coding toolkit: 94 skills, 44 agents, multi-platform (Claude, Cursor, Windsurf, Copilot, Gemini, Cline, Roo Code, Aider, Augment, Antigravity, Codex CLI… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Verification before completion

Example prompts

  • “Use the verification-before-completion skill to force verification commands before success claims. An agent skill from softspark/ai-toolkit”
  • “/verification-before-completion”

Requirements

  • Pre-approved tools (allowed-tools): Read

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Define the rubric BEFORE building — 3–6 criteria, each with a weight and an explicit pass bar. Example
  2. Launch the app and observe — actually run it and capture the behavior (screenshot, console, network). Do not infer from the source.
  3. Score with a fresh evaluator — have an independent agent grade the observed behavior against the rubric, not the implementer who wrote it…
  4. Gate on the threshold — below the bar (e.g. < 0.8) the claim is NOT verified: list the failing criteria as concrete defects and iterate…

What it can do on your machine

Read from SKILL.md and the folder at commit d64db2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verification Before Completion loads about 2.2k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 1,108 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from softspark/ai-toolkit at commit d64db2b, republished under its Apache-2.0 licence (© softspark). 1,108 words, ~2,248 tokens.

Download SKILL.mdSave it as .claude/skills/verification-before-completion/SKILL.md (or your agent's skills folder).
name
verification-before-completion
description
Forces verification commands before success claims. Evidence before assertions. Triggers: complete, fixed, passing, done, ready, verified.
allowed-tools
Read
user-invocable
false

Verification Before Completion

The Iron Law

NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE

If you haven't run the verification command in this message, you cannot claim it passes.

Claiming work is complete without verification is dishonesty, not efficiency.

The Gate Function

BEFORE claiming any status or expressing satisfaction:

1. IDENTIFY: What command proves this claim?
2. RUN: Execute the FULL command (fresh, complete)
3. READ: Full output, check exit code, count failures
4. VERIFY: Does output confirm the claim?
   - If NO: State actual status with evidence
   - If YES: State claim WITH evidence
5. ONLY THEN: Make the claim

Skip any step = lying, not verifying

When To Apply

ALWAYS before:

  • ANY variation of success/completion claims
  • ANY expression of satisfaction ("Great!", "Perfect!", "Done!")
  • ANY positive statement about work state
  • Committing, PR creation, task completion
  • Moving to next task
  • Delegating to agents

Don't Assume It Exists

A prompt — yours, the user's, or another agent's — naming a file, table, column, endpoint, env var, config key, or dependency is a hint, not a fact. The name existing in text is not the same as the thing existing in the repo.

  • Before you read from, write to, import, or quote any such resource, confirm it: Read the file, ls/grep the path, list the table, check the lockfile. One look beats a confident guess.
  • Never reconstruct a file's contents, a function's body, or an API's signature from its name alone. The name tells you almost nothing about the shape. If you haven't read it, you don't know it.
  • A plausible-sounding path or symbol is the easiest thing in the world to invent. Treat your own fluency here as a warning sign, not as evidence.
  • AUTHORIZED security work (CTF, sanctioned pentest, defensive review) still verifies — probing whether a path/endpoint/parameter actually exists is part of the job, not an exception to it.

Declare Ungrounded as a Real Outcome

When search, KB lookup, or tool calls come back with nothing relevant, the correct move is to say so and stop — not to backfill the hole from training memory and present it as established fact.

  • "Not found in <source>" is a complete, honest answer. State which source you checked and that it came up empty.
  • Do not promote a half-remembered detail to a definite claim just because the gap feels uncomfortable. An unverified recollection is a hypothesis; label it as one or leave it out.
  • If the missing piece blocks the task, surface the gap and ask — do not paper over it with a guess dressed as a finding.

Common Failures

ClaimRequiresNot Sufficient
Tests passTest command output: 0 failuresPrevious run, "should pass"
Linter cleanLinter output: 0 errorsPartial check, extrapolation
Build succeedsBuild command: exit 0Linter passing, logs look good
Bug fixedTest original symptom: passesCode changed, assumed fixed
Regression test worksRed-green cycle verifiedTest passes once
New gate works (lint rule, grep check, coverage/contract test)Seen it fail once on a planted violation, with the expected messageIt passes on the current code
Agent completedVCS diff shows changesAgent reports "success"
Requirements metLine-by-line checklistTests passing
No dead code (Art. VI.1)Grep for every removed/renamed symbol: 0 references"I cleaned up what I touched"
Behavior change covered (Art. VI.2)Integration test for the API surface + unit test + docs updatedUnit test on the helper only
Diff is clean (Art. VI.4)Re-read full diff: no orphaned imports, no stale docs, no skipped fixes"I only changed what I needed"
Resource exists (file/table/env var/endpoint/dep)Read/list/grep it and see itThe prompt mentioned it, or the name sounds real
File contents / API signatureOpen the file, read the actual linesInferring shape from the filename or symbol name
Answer is groundedA source you actually read returns itRecalling it from training memory

Red Flags — STOP

  • Using "should", "probably", "seems to"
  • Expressing satisfaction before verification
  • About to commit/push/PR without verification
  • Trusting agent success reports without independent check
  • Relying on partial verification
  • Thinking "just this once"

Rationalization Prevention

ExcuseReality
"Should work now"RUN the verification
"I'm confident"Confidence is not evidence
"Just this once"No exceptions
"Linter passed"Linter is not compiler
"Agent said success"Verify independently
"Partial check is enough"Partial proves nothing
"Different words so rule doesn't apply"Spirit over letter
Show full SKILL.md (468 more words)Show less

Pre-Completion Self-Audit

Before you type "done", run these fast yes/no gates over your own draft. Each one has a fix — if the answer is bad, do the fix, don't ship the draft.

Self-checkIf the honest answer is "no" / "yes, I did"
Did I actually run the verification command this message, or am I asserting success from memory?Run it now. Memory is not a test result.
Does every factual claim trace to output or a source I actually saw?Cite the evidence, or cut the claim.
Did I invent any path, filename, symbol, table, flag, or version number?Verify it exists, or remove it.
Did I quote file contents or an API signature I never opened?Open and confirm, or stop quoting.
Did a search/KB lookup come back empty that I then "filled in" anyway?Replace the fill-in with "not found in <source>".
Am I about to commit/push/PR on the strength of a claim I haven't proven?Prove it first.

Same ethos as the rest of this skill: evidence before assertions, applied to your own output one line at a time.

Key Patterns

Tests:

CORRECT:  [Run test command] [See: 34/34 pass] "All tests pass"
WRONG:    "Should pass now" / "Looks correct"

Regression tests (TDD Red-Green):

CORRECT:  Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)
WRONG:    "I've written a regression test" (without red-green verification)

New gates:

CORRECT:  Add gate → Run (pass) → Plant one violation in a throwaway copy or fixture → Run (MUST FAIL, expected message) → Remove the plant → Run (pass)
WRONG:    "Gate added, it's green" (a gate that has never been red may be checking nothing)

Requirements:

CORRECT:  Re-read plan → Create checklist → Verify each → Report gaps or completion
WRONG:    "Tests pass, phase complete"

Agent delegation:

CORRECT:  Agent reports success → Check VCS diff → Verify changes → Report actual state
WRONG:    Trust agent report at face value

Live-App Rubric Verification (optional)

When the success criterion is behavioral or visual (a UI flow, a generated app, a multi-step interaction), a pass/fail command is not enough — the proof is the running app, observed. Use a weighted rubric instead of a single assertion:

  1. Define the rubric BEFORE building — 3–6 criteria, each with a weight and an explicit pass bar. Example:

    CriterionWeightPass bar
    Core flow completes end-to-end0.40No error, reaches success state
    Empty / loading / error states render0.25All three visible
    Matches the requested layout0.20No major deviation
    No console errors0.15Console clean
  2. Launch the app and observe — actually run it and capture the behavior (screenshot, console, network). Do not infer from the source.

  3. Score with a fresh evaluator — have an independent agent grade the observed behavior against the rubric, not the implementer who wrote it (self-grading anchors high). Compute the weighted score.

  4. Gate on the threshold — below the bar (e.g. < 0.8) the claim is NOT verified: list the failing criteria as concrete defects and iterate. At or above the bar, state the score WITH the captured evidence.

The rubric is the verification command for work that has no green/red exit code. The same Iron Law applies: observed evidence before the claim, every time.

Constitutional Anchors

This skill enforces Constitution Art. VI.4 (Verify Before Claiming Done). The diff re-read is not optional: before any completion claim, confirm no orphaned references, no missing test coverage for changed paths, no stale docs. A task is not done while any of those exist.

The Bottom Line

Run the command. Read the output. THEN claim the result.

This is non-negotiable.

© softspark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in app/skills/verification-before-completion of softspark/ai-toolkit.

Open the folder on GitHubat commit d64db2b

Compare with similar skills

Verification Before Completion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verification Before Completion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verification Before Completion this skillsoftspark/ai-toolkit179—~2.2kAutomated safety check: PassApache-2.0
Show Me Your Work Decision Logcursor/plugins10k9 repos~1.6kAutomated safety check: PassNone
PUA Looptanweai/pua20k1 repos~1.1kAutomated safety check: PassMIT
Scope Creep Guardlennney/stop-that-shit2.5k1 repos~2kAutomated safety check: PassMIT
Verification Before Completionfarm-fe/farm5.6k46 repos~1kAutomated safety check: PassMIT
Incremental Implementationaddyosmani/agent-skills103k1 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 9 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • PUA Loop

    tanweai/pua

    Runs an unattended iterate-until-verified loop in which a user-set verify command, not the agent's own claim, decides when the task is finished.

    20k GitHub starsUsed in 1 repo~1.1k tokens
    Agent WorkflowsAuto-check passed
  • Scope Creep Guard

    lennney/stop-that-shit

    Keeps an agent focused on the requested work by applying a five-step ladder that checks for direct solutions, real gaps and speculative defenses before adding anything.

    2.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • A skill your agent uses when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any…

    5.6k GitHub starsUsed in 46 repos~1k tokens
    Agent WorkflowsAuto-check passed
  • Incremental Implementation

    addyosmani/agent-skills

    Delivers a change in thin vertical slices, each implemented, tested, verified and committed before the next, using vertical, contract-first or risk-first slicing.

    103k GitHub starsUsed in 1 repo~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Pushes an agent to keep verifying and changing approach after repeated failures, using a diagnosis line, evidence-based completion and confirmation before risky edits.

    20k GitHub stars~502 tokensUpdated 29 days ago
    Agent WorkflowsAuto-check passed

More from softspark/ai-toolkit

All 112 skills in this repo
  • Prepare Test Env

    softspark/ai-toolkit

    Prepare or verify a project QA environment with source identity, readiness, browser access, evidence paths and owned cleanup.

    179 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check: notes
  • A11y Validate

    softspark/ai-toolkit

    Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check: notes
  • Analyze

    softspark/ai-toolkit

    Analyzes code quality, complexity, patterns across codebase.

    179 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Autonomous Dev

    softspark/ai-toolkit

    Drives a brief, specification, issue or existing PR through implementation, review, tests and QA to a ready PR.

    179 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Brand Voice

    softspark/ai-toolkit

    Direct technical voice for docs, README, user-facing text. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • CI

    softspark/ai-toolkit

    Detect/generate/debug CI pipeline config (GitHub Actions, GitLab CI).

    179 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes

Categories

Questions about Verification Before Completion

What does Verification Before Completion do?

Forces verification commands before success claims. An agent skill from softspark/ai-toolkit. Verification Before Completion is an agent skill from softspark/ai-toolkit. Forces verification commands before success claims.

When should I use Verification Before Completion?

Verification Before Completion fits situations like: tasks that involve Verification before completion.

How do I install Verification Before Completion in Claude Code?

Run `npx skills add softspark/ai-toolkit --skill verification-before-completion -a claude-code`. Or copy the skill folder (app/skills/verification-before-completion in softspark/ai-toolkit) into .claude/skills/verification-before-completion in your project. Claude Code loads it when a task matches its description.

How do I install Verification Before Completion in Codex?

Run `npx skills add softspark/ai-toolkit --skill verification-before-completion -a codex`. Or copy the skill folder (app/skills/verification-before-completion in softspark/ai-toolkit) into .agents/skills/verification-before-completion in your project. Codex loads it when a task matches its description.

Can I use Verification Before Completion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softspark/ai-toolkit --skill verification-before-completion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verification-before-completion, .gemini/skills/verification-before-completion, .github/skills/verification-before-completion and .opencode/skills/verification-before-completion in your project.

What does Verification Before Completion need to run?

SKILL.md names no scripts, command-line tools or credentials: Verification Before Completion is instructions for the agent only. Its frontmatter pre-approves these tools: Read.

Does Verification Before Completion access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verification Before Completion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verification Before Completion use?

Verification Before Completion is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verification Before Completion use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verification Before Completion?

Skills that share tags, products or a category with Verification Before Completion: Show Me Your Work Decision Log (cursor/plugins, 10k stars), PUA Loop (tanweai/pua, 20k stars), Scope Creep Guard (lennney/stop-that-shit, 2.5k stars) and Verification Before Completion (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verification Before Completion?

softspark (a GitHub user) maintains it in softspark/ai-toolkit, which has 179 GitHub stars. The repository holds 112 skills in this directory. The repository was last updated on October 7, 2026.

Source: softspark/ai-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.