Agent skill

Verification Protocol

by PlamenTSV in PlamenTSV/plamen

How to prove a hypothesis is TRUE or FALSE using Move unit tests.

MITAuto-check passedTesting & QA

Install Verification Protocol

skills CLI
$ npx skills add PlamenTSV/plamen --skill verification-protocol -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PlamenTSV/plamen verification-protocol --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PlamenTSV/plamen.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agents/skills/aptos/verification-protocol .claude/skills/verification-protocol && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verification-protocol
GitHub stars
303
Token cost
~3.5k tokens
SKILL.md length
1,203 words
Files
3 (incl. references)
Skills in repo
87
Repo updated
First seen
Licence
MIT

At a glance

How to prove a hypothesis is TRUE or FALSE using Move unit tests.

  • Tasks that involve Unit testing
  • SKILL.md covers Evidence Source Tracking…, Pre-Verification Understanding, Pre-PoC Feasibility Gates… and Test File Template, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Verification Protocol is an agent skill from PlamenTSV/plamen. How to prove a hypothesis is TRUE or FALSE using Move unit tests.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/advanced.md` and `references/templates.md`).

It sits in Testing & QA, covering Unit testing. The repository describes itself as: Autonomous Web3 security audit agent for Claude Code. The licence is MIT.

When your agent uses it

  • Tasks that involve Unit testing

Example prompts

  • “/verification-protocol”

What it can do on your machine

Read from SKILL.md and the folder at commit 795962b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verification Protocol loads about 3.5k tokens when it runs, and up to ~7.9k if it reads all its reference files. Until then it costs about 22 tokens; SKILL.md has 1,203 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PlamenTSV/plamen at commit 795962b, republished under its MIT licence (© PlamenTSV). 1,203 words, ~3,465 tokens.

Download SKILL.mdSave it as .claude/skills/verification-protocol/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
verification-protocol
description
How to prove a hypothesis is TRUE or FALSE using Move unit tests.

Verification Protocol -- Aptos Move

How to prove a hypothesis is TRUE or FALSE using Move unit tests.


Evidence Source Tracking (MANDATORY)

CRITICAL: For EVERY piece of evidence used in verification, you MUST tag its source. Evidence from mocks or unverified external modules CANNOT support a REFUTED verdict.

Evidence Source Tags
TagMeaningValid for REFUTED?
[PROD-ONCHAIN]Production module verified on Aptos ExplorerYES
[PROD-SOURCE]Source code verified on-chain (Aptos Explorer source verification)YES
[CODE]Audited codebase (in-scope)YES
[MOCK]Mock/test moduleNO
[EXT-UNV]External module, unverified behaviorNO
[DOC]Documentation/spec onlyNO (needs verification)
Evidence Audit Table (REQUIRED in every verification output)

Before ANY verdict, fill this table:

markdown
### Evidence Audit
| Claim | Evidence Source | Tag | Valid for REFUTED? |
|-------|-----------------|-----|-------------------|
| "External module returns X" | Mock module | [MOCK] | NO |
| "State changes to Y" | protocol_module.move:123 | [CODE] | YES |
| "Coin transfer triggers Z" | Aptos Explorer source | [PROD-ONCHAIN] | YES |
Mock Rejection Rule

AUTOMATIC OVERRIDE: If ANY evidence supporting REFUTED has tag [MOCK] or [EXT-UNV]:

  • CANNOT return REFUTED
  • MUST return CONTESTED
  • Triggers production verification

Example:

markdown
## Verdict: REFUTED -> CONTESTED (mock evidence override)

### Evidence Audit
| Claim | Source | Tag | Valid? |
|-------|--------|-----|--------|
| "Staking returns shares" | test_staking.move:45 | [MOCK] | NO |

**Override reason**: REFUTED verdict relies on mock behavior at test_staking.move:45.
Production module behavior is UNVERIFIED. Must verify against on-chain source.

Pre-Verification Understanding

Before writing ANY test code, you MUST answer:

Question 1: What is the EXACT bug?
NOT: "Something is inconsistent"
NOT: "State is wrong"
NOT: "Capability leak possible"

YES: "[Variable/resource] is [read/written/moved] at [location] but should be
      [read/written/moved] at [other location] because [specific reason]"
Question 2: What OBSERVABLE difference proves it?
NOT: "Values are different"
NOT: "State changed"

YES: "Before operation: [resource/value] = [expected value]
      After operation: [resource/value] = [actual value]
      Expected: [what it should be]"
Question 3: What is the EXACT assertion?
NOT: assert!(bug_exists, 0)
NOT: assert!(!is_secure, 0)

YES: assert!(actual_value == expected_value, ERROR_CODE)
 OR: assert!(before != after, ERROR_CODE)  // "value changed when it shouldn't"
 OR: assert!(error > threshold, ERROR_CODE)  // "error exceeds acceptable threshold"

If you cannot answer all three -> ASK FOR CLARIFICATION


Pre-PoC Feasibility Gates (MANDATORY)

Before writing test code, verify these two gates. If either FAILS, adjust the hypothesis.

Gate F1: Reachability

Trace a call path from a permissionless entry point to the vulnerable code.

  • Entry point identified (public/external/entry function)
  • Call path traced through intermediary functions
  • All access checks on the path are passable by the attacker profile

If NO entry point reaches the vulnerable code → UNREACHABLE → FALSE_POSITIVE. If reachable only through a restricted path → document the restriction, adjust likelihood.

Gate F2: Math Bounds

Substitute real-world value domains into the expression that triggers the bug.

  • Parameter domains identified (token decimals, max supply, TVL range, fee range, time bounds)
  • Expression evaluated at worst-case feasible inputs
  • Result crosses the bug threshold

If the bug requires values outside feasible domains → INFEASIBLE → FALSE_POSITIVE. If feasible only at extreme but realistic parameters → document the threshold, proceed with adjusted severity.

Both gates PASS → proceed to PoC. Either gate FAILS → document and stop.


Test File Template

See templates.md in this directory for all Move test file templates and Move-specific test patterns.

Interpreting Results

Test PASSES -> Bug CONFIRMED

The assertion that "proves the bug" succeeded.

  • If assert!(after != before, 0) passes -> values ARE different (bug exists)
  • If assert!(error > threshold, 0) passes -> error IS above threshold (bug exists)
Test FAILS -> Check Why
FailureMeaningAction
Assertion failed (abort code)Bug doesn't exist as hypothesizedRe-examine hypothesis
Abort in setupModule initialization wrongFix setup (check init order, missing resources)
Abort in actionOperation blocked (access control, precondition)Check preconditions, signer requirements
ARITHMETIC_ERROR (0x20001)Overflow/underflow or division by zeroCheck calculations, validate inputs
RESOURCE_NOT_FOUNDMissing move_to in setupEnsure all required resources are initialized
ALREADY_EXISTSDuplicate resource creationCheck init called only once
Common Aptos-Specific Test Issues
IssueCauseFix
ENOT_FOUND on coin operationsAccount not registered for coin typeAdd coin::register<CoinType>(user) before operations
Timestamp not availabletimestamp module not initializedAdd timestamp::set_time_has_started_for_testing(aptos_framework)
Object not foundObject created at unexpected addressUse object::create_named_object with deterministic seed
Module not publishedTest module can't import protocol moduleCheck Move.toml dependencies and test address mapping
Signer mismatch@protocol_addr doesn't match expectedVerify #[test(...)] signer addresses match module publish address

Iteration Protocol

Attempt 1: Direct implementation of test strategy from hypothesis

Attempt 2: Adjust parameters

  • Different amounts (larger/smaller, boundary values)
  • Different timing (advance more/fewer seconds)
  • Different actors (swap attacker/victim roles)
  • Different resource initialization order

Attempt 3: Re-examine assumptions

  • Is setup correct? (all resources initialized, correct init order)
  • Are preconditions met? (correct signer, sufficient balance, required state)
  • Is the bug mechanism correctly understood?
  • Are module dependencies correctly configured in Move.toml?

After 5 attempts:

  • If still fails -> FALSE_POSITIVE with documented reasoning
  • Explain why the hypothesis was wrong

Severity Determination

CRITICAL
  • Direct fund theft possible (drain FungibleStore, mint unlimited tokens)
  • Protocol insolvency (assets < liabilities)
  • No special prerequisites needed (permissionless exploit)
  • Attacker profits significantly
  • Ref capability leak granting unrestricted mint/transfer/burn
HIGH
  • Fund loss with some setup (specific state required)
  • Broken core functionality (deposits, withdrawals, swaps non-functional)
  • Significant value at risk
  • Cumulative error compounds quickly
  • Ref capability leak with limited but significant blast radius
MEDIUM
  • Limited fund loss (bounded by rate limits, caps)
  • Requires specific conditions (timing, state, multi-step)
  • Edge cases with real impact
  • Moderate value at risk
  • Access control weakness that requires compromised friend module
LOW
  • Negligible direct impact
  • Extreme edge cases only
  • Admin/owner controlled risk with compensating controls
  • Informational with minor consequence

Show full SKILL.md (477 more words)Show less

Output Format

CONFIRMED
markdown
## Verdict: CONFIRMED

### Evidence Audit
| Claim | Evidence Source | Tag | Valid for REFUTED? |
|-------|-----------------|-----|-------------------|

### Bug Mechanism Verified
{Explain what the test proves in 2-3 sentences}

### Test File
`tests/audit/test_hypothesis_N.move`

### Test Output

{Paste relevant aptos move test output}


### Key Evidence
| Metric | Value |
|--------|-------|
| Before | {value} |
| After | {value} |
| Expected | {value} |
| Difference | {calculation} |

### Severity: {LEVEL}
{Justification in 1-2 sentences}

### RAG Evidence
- **Attack Vectors Consulted**: [list bug classes queried]
- **Similar Exploits Found**: [count and brief descriptions]
- **PoC Template Used**: [yes/no, which template]
- **Historical Precedent**: [describe any matching historical vulnerabilities]
FALSE_POSITIVE
markdown
## Verdict: FALSE_POSITIVE

### Evidence Audit
| Claim | Evidence Source | Tag | Valid for REFUTED? |
|-------|-----------------|-----|-------------------|

### Attempts Made

**Attempt 1:**
- Approach: {description}
- Result: {what happened}
- Learning: {insight}

**Attempt 2:**
- Approach: {description}
- Result: {what happened}
- Learning: {insight}

**Attempt 3:**
- Approach: {description}
- Result: {what happened}
- Learning: {insight}

### Why It's Not a Bug
{Explain the actual behavior and why hypothesis was wrong in 2-3 sentences}
CONTESTED (CRITICAL)
markdown
## Verdict: CONTESTED

### Evidence Audit
| Claim | Evidence Source | Tag | Valid for REFUTED? |
|-------|-----------------|-----|-------------------|

### Evidence Status
| Checkpoint | Status | Details |
|------------|--------|---------|
| External behavior verified against PRODUCTION | NO | Used mock behavior as evidence |
| All callers checked | YES | Checked A, B, C |
| Ref access paths fully traced | NO | Friend module re-export not analyzed |
| Profit calculated with attacker holding | NO | Only analyzed donation loss |

### Why This Cannot Be REFUTED
{Explain what evidence is missing to definitively rule out the bug}

### Escalation Required
- [ ] Fetch production module source from Aptos Explorer for {external dep}
- [ ] Re-analyze with attacker holding shares/tokens
- [ ] Check additional caller paths: {list}
- [ ] Trace Ref access through friend modules: {list}

### Current Assessment
Likely: {TRUE_POSITIVE / FALSE_POSITIVE / UNKNOWN}
Confidence: {LOW / MEDIUM}

Insufficient Evidence (HALT CONDITIONS) -- CRITICAL

MANDATORY: You MUST check ALL boxes before returning REFUTED. If ANY checkbox is NO -> Return CONTESTED, not REFUTED.

Before marking REFUTED, check:

  • External behavior verified against PRODUCTION (not mock)
    • Check Aptos Explorer for on-chain module source verification
    • If external module is marked 'UNVERIFIED' -> CANNOT use as evidence
    • If mock differs from production -> use PRODUCTION behavior
  • Attack path checked on ALL callers (not just main path)
    • Enumerate all public fun and public entry fun that reach the vulnerable code
    • Check public(friend) fun callers via friend module analysis
  • Ref capability paths fully traced
    • For Ref-related findings: trace every path from Ref creation to Ref usage
    • Check friend modules for transitive Ref access
    • Check if ExtendRef-derived signer enables unexpected access
  • Profit calculated with attacker HOLDING tokens (not just donating)
    • "Attacker loses by donating" is NOT sufficient evidence
    • Check: what if attacker holds X% of shares BEFORE donating?
  • Missing precondition documented
    • Document in structured format: precondition type + why it blocks
    • Types: STATE / ACCESS / TIMING / EXTERNAL / BALANCE
  • Searched other findings for matching postconditions
    • Read {scratchpad}/findings_inventory.md for CONFIRMED/PARTIAL findings
    • Check if ANY finding creates the postcondition that would enable this attack
    • If match found -> CONTESTED, not REFUTED (chain analysis will combine)
Evidence That Does NOT Count
  • "Mock shows X" -- mocks != production (CRITICAL: always verify against production)
  • "Standard Coin module" -- may have custom transfer hooks via fungible_asset dispatch
  • "Attacker loses by donating" -- may profit via shares held
  • "Function is private/friend" -- friend module may expose it publicly
  • "Requires admin signer" -- admin may be compromised or malicious
  • "Attacker cannot acquire X" -- another finding may CREATE this condition
  • "Ref is in private storage" -- friend module may provide access path
Anti-Downgrade Halt for VS/BLIND Findings (HARD RULE)

For findings from Validation Sweep ([VS-]) or Blind Spot Scanner ([BLIND-]): apply Rule 13's 5-question test BEFORE any downgrade. HALT: If test shows users harmed AND unavoidable AND undocumented -> you CANNOT return FALSE_POSITIVE. Minimum verdict: CONTESTED. Defense parity gaps (Module A has protection X, Module B lacks it for same action) are NEVER "by design" -> minimum severity: Medium, minimum verdict: CONTESTED. Violating this halt is a workflow error equivalent to using [MOCK] evidence for REFUTED.

Chain Analysis Integration

A finding is NEVER truly REFUTED until chain analysis completes.

If you mark a finding as REFUTED but document a missing precondition, the chain analyzer will search for other findings whose postconditions match your missing precondition. If found, the finding will be escalated to CONTESTED and combined into a chain hypothesis.

Example:

  • Your finding: "Drain attack blocked because attacker cannot get TransferRef"
  • Other finding: "Friend module exposes TransferRef via public function"
  • Chain: Other finding enables your finding -> Combined HIGH severity


Advanced Protocol Reference: See advanced.md for RAG queries before PoC, exchange rate finding severity, design flaw escalation, bidirectional role analysis, chain hypothesis, and Aptos-specific verification considerations.

© PlamenTSV, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in agents/skills/aptos/verification-protocol of PlamenTSV/plamen.

  • SKILL.md
  • references/advanced.md
  • references/templates.md

Open the folder on GitHubat commit 795962b

Compare with similar skills

Verification Protocol next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verification Protocol compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verification Protocol this skillPlamenTSV/plamen303—~3.5kAutomated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
Testing OpenLogi UIAprilNEA/OpenLogi23k—~1.1kAutomated safety check: PassApache-2.0
Go Testingcxuu/golang-skills1701 repos~1.3kAutomated safety check: PassApache-2.0
Contractssamchon/nestia2.2k—~1.3kAutomated safety check: PassMIT
Cohesion Over TestabilityEpicenterHQ/epicenter4.8k—~2kAutomated safety check: PassCustom licence

Similar skills

  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • Testing OpenLogi UI

    AprilNEA/OpenLogi

    Verifies OpenLogi's native GPUI interface with focused tests, the component gallery and a mock agent, choosing the evidence that fits each change.

    23k GitHub stars~1.1k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Go Testing

    cxuu/golang-skills

    A skill your agent uses when writing, reviewing, or improving Go test code — including table-driven tests, subtests, parallel tests, test helpers, test doubles, and assertions with cmp.Diff.

    170 GitHub starsUsed in 1 repo~1.3k tokens
    Testing & QAAuto-check passed
  • Contracts

    samchon/nestia

    Defines self-acknowledgments for production declarations and tests.

    2.2k GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Cohesion Over Testability

    EpicenterHQ/epicenter

    Collapse test-shaped production boundaries while preserving behavior and coverage.

    4.8k GitHub stars~2k tokensUpdated today
    Testing & QAAuto-check passed
  • JS-in-HTML Testing

    liaohch3/claude-tap

    Tests JavaScript embedded in an HTML file in two layers: pytest checks of the logic ported to Python, and Playwright runs in a real browser for the DOM.

    3.3k GitHub stars~924 tokensUpdated 15 days ago
    Testing & QAAuto-check passed

More from PlamenTSV/plamen

All 87 skills in this repo
  • Audit Prep

    PlamenTSV/plamen

    Prepare Solidity projects for a security audit — test coverage, test quality, NatSpec docs, code hygiene, dependency health, best-practice enforcement, deployment readiness, and project…

    303 GitHub stars~3.7k tokensUpdated 11 days ago
    Auto-check passed
  • Verification Protocol

    PlamenTSV/plamen

    Trigger Pattern Always (used by all verifier agents) - Inject Into security-verifier agents (Phase 5)

    303 GitHub stars~3.5k tokensUpdated 11 days ago
    Auto-check passed
  • Ability Analysis

    PlamenTSV/plamen

    Trigger Pattern Always (Aptos Move) - foundational security check - Inject Into Breadth agents, depth agents

    303 GitHub stars~3.3k tokensUpdated 11 days ago
    Auto-check passed
  • Ability Analysis

    PlamenTSV/plamen

    Trigger Pattern Always (Sui Move) -- foundational security check - Inject Into Breadth agents, depth agents

    303 GitHub stars~3.2k tokensUpdated 11 days ago
    Auto-check passed
  • Account Lifecycle

    PlamenTSV/plamen

    Trigger Pattern ACCOUNTCLOSING flag detected (close/CloseAccount usage) - Inject Into Breadth agents, depth agents

    303 GitHub stars~1.2k tokensUpdated 11 days ago
    Auto-check passed
  • Account Validation

    PlamenTSV/plamen

    Trigger Pattern Always required for Solana audits - Inject Into Breadth agents, depth agents

    303 GitHub stars~1.7k tokensUpdated 11 days ago
    Auto-check passed

Categories

Questions about Verification Protocol

What does Verification Protocol do?

How to prove a hypothesis is TRUE or FALSE using Move unit tests. Verification Protocol is an agent skill from PlamenTSV/plamen. How to prove a hypothesis is TRUE or FALSE using Move unit tests.

When should I use Verification Protocol?

Verification Protocol fits situations like: tasks that involve Unit testing.

How do I install Verification Protocol in Claude Code?

Run `npx skills add PlamenTSV/plamen --skill verification-protocol -a claude-code`. Or copy the skill folder (agents/skills/aptos/verification-protocol in PlamenTSV/plamen) into .claude/skills/verification-protocol in your project. Claude Code loads it when a task matches its description.

How do I install Verification Protocol in Codex?

Run `npx skills add PlamenTSV/plamen --skill verification-protocol -a codex`. Or copy the skill folder (agents/skills/aptos/verification-protocol in PlamenTSV/plamen) into .agents/skills/verification-protocol in your project. Codex loads it when a task matches its description.

Can I use Verification Protocol in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PlamenTSV/plamen --skill verification-protocol -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verification-protocol, .gemini/skills/verification-protocol, .github/skills/verification-protocol and .opencode/skills/verification-protocol in your project.

What does Verification Protocol need to run?

SKILL.md names no scripts, command-line tools or credentials: Verification Protocol is instructions for the agent only.

Does Verification Protocol access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verification Protocol safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verification Protocol use?

Verification Protocol is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verification Protocol use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Verification Protocol?

Skills that share tags, products or a category with Verification Protocol: TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Testing OpenLogi UI (AprilNEA/OpenLogi, 23k stars), Go Testing (cxuu/golang-skills, 170 stars) and Contracts (samchon/nestia, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verification Protocol?

PlamenTSV (a GitHub user) maintains it in PlamenTSV/plamen, which has 303 GitHub stars. The repository holds 87 skills in this directory. The repository was last updated on September 26, 2026.

Source: PlamenTSV/plamen on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.