Agent skill

Verification Protocol

by PlamenTSV in PlamenTSV/plamen

How to prove a hypothesis is TRUE or FALSE using Foundry tests.

MITAuto-check passedBackend & APIs

Install Verification Protocol

skills CLI
$ npx skills add PlamenTSV/plamen --skill verification-protocol -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PlamenTSV/plamen verification-protocol --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PlamenTSV/plamen.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agents/skills/evm/verification-protocol .claude/skills/verification-protocol && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verification-protocol
GitHub stars
303
Token cost
~3k tokens
SKILL.md length
926 words
Files
2 (incl. references)
Skills in repo
87
Repo updated
First seen
Licence
MIT

At a glance

How to prove a hypothesis is TRUE or FALSE using Foundry tests.

  • Tasks that involve Smart contracts
  • SKILL.md covers Evidence Source Tracking…, Pre-Verification Understanding, Pre-PoC Feasibility Gates… and Test File Template, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Verification Protocol is an agent skill from PlamenTSV/plamen. How to prove a hypothesis is TRUE or FALSE using Foundry tests.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/advanced.md`).

It sits in Backend & APIs, covering Smart contracts. The repository describes itself as: Autonomous Web3 security audit agent for Claude Code. The licence is MIT.

When your agent uses it

  • Tasks that involve Smart contracts

Example prompts

  • “/verification-protocol”

What it can do on your machine

Read from SKILL.md and the folder at commit 795962b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown and solidity).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verification Protocol loads about 3k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 21 tokens; SKILL.md has 926 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PlamenTSV/plamen at commit 795962b, republished under its MIT licence (© PlamenTSV). 926 words, ~3,033 tokens.

Download SKILL.mdSave it as .claude/skills/verification-protocol/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
verification-protocol
description
How to prove a hypothesis is TRUE or FALSE using Foundry tests.

Verification Protocol

How to prove a hypothesis is TRUE or FALSE using Foundry tests.


Evidence Source Tracking (MANDATORY)

CRITICAL: For EVERY piece of evidence used in verification, you MUST tag its source. Evidence from mocks or unverified external contracts CANNOT support a REFUTED verdict.

Evidence Source Tags
TagMeaningValid for REFUTED?
[PROD]Production contract (verified on-chain)YES
[MOCK]Mock/test contractNO
[CODE]Audited codebase (in-scope)YES
[EXT-UNV]External, unverified behaviorNO
[DOC]Documentation/spec onlyNO (needs verification)
Evidence Audit Table (REQUIRED in every verification output)

Before ANY verdict, fill this table:

markdown
### Evidence Audit
| Claim | Evidence Source | Tag | Valid for REFUTED? |
|-------|-----------------|-----|-------------------|
| "External returns X" | Mock contract | [MOCK] | NO |
| "State changes to Y" | Protocol.sol:123 | [CODE] | YES |
| "Transfer triggers Z" | Etherscan source | [PROD] | YES |
Mock Rejection Rule

AUTOMATIC OVERRIDE: If ANY evidence supporting REFUTED has tag [MOCK] or [EXT-UNV]:

  • CANNOT return REFUTED
  • MUST return CONTESTED
  • Triggers production verification (Step 4a.5)

Example:

markdown
## Verdict: REFUTED -> CONTESTED (mock evidence override)

### Evidence Audit
| Claim | Source | Tag | Valid? |
|-------|--------|-----|--------|
| "Staking returns shares" | StakingMock.sol:45 | [MOCK] | NO |

**Override reason**: REFUTED verdict relies on mock behavior at StakingMock.sol:45.
Production contract behavior is UNVERIFIED. Must fetch production source.

Pre-Verification Understanding

Before writing ANY test code, you MUST answer:

Question 1: What is the EXACT bug?
NOT: "Something is inconsistent"
NOT: "State is wrong"
NOT: "Reentrancy possible"

YES: "[Variable] is [read/written] at [location] but should be [read/written]
      at [other location] because [specific reason]"
Question 2: What OBSERVABLE difference proves it?
NOT: "Values are different"
NOT: "State changed"

YES: "Before operation: [variable] = [expected value]
      After operation: [variable] = [actual value]
      Expected: [what it should be]"
Question 3: What is the EXACT assertion?
NOT: assertTrue(bugExists)
NOT: assertFalse(isSecure)

YES: assertEq(actualValue, expectedValue, "description of what's wrong")
 OR: assertNotEq(before, after, "value changed when it shouldn't")
 OR: assertGt(error, threshold, "error exceeds acceptable threshold")

If you cannot answer all three -> ASK FOR CLARIFICATION


Pre-PoC Feasibility Gates (MANDATORY)

Before writing test code, verify these two gates. If either FAILS, adjust the hypothesis.

Gate F1: Reachability

Trace a call path from a permissionless entry point to the vulnerable code.

  • Entry point identified (public/external/entry function)
  • Call path traced through intermediary functions
  • All access checks on the path are passable by the attacker profile

If NO entry point reaches the vulnerable code → UNREACHABLE → FALSE_POSITIVE. If reachable only through a restricted path → document the restriction, adjust likelihood.

Gate F2: Math Bounds

Substitute real-world value domains into the expression that triggers the bug.

  • Parameter domains identified (token decimals, max supply, TVL range, fee range, time bounds)
  • Expression evaluated at worst-case feasible inputs
  • Result crosses the bug threshold

If the bug requires values outside feasible domains → INFEASIBLE → FALSE_POSITIVE. If feasible only at extreme but realistic parameters → document the threshold, proceed with adjusted severity.

Both gates PASS → proceed to PoC. Either gate FAILS → document and stop.


Test File Template

solidity
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.0;

import "forge-std/Test.sol";
import "forge-std/console.sol";

/**
 * @title Test_H{N}: {Title}
 *
 * BUG: {2 sentence description}
 * EXPECTED: {what should happen}
 * ACTUAL: {what does happen}
 */
contract Test_H{N} is Test {

    // === CONTRACTS ===
    // Declare target contract and any dependencies

    // === ACTORS ===
    address attacker = makeAddr("attacker");
    address victim = makeAddr("victim");
    address owner = makeAddr("owner");

    // === SETUP ===
    function setUp() public {
        // Deploy contracts
        // Set initial state
        // Fund actors if needed
    }

    // === TEST: Direct bug demonstration ===
    function test_H{N}_bug_demonstration() public {
        // 1. RECORD BEFORE
        console.log("=== BEFORE ===");
        uint256 valueBefore = target.criticalValue();
        console.log("Critical value:", valueBefore);

        // 2. ACTION
        console.log("=== ACTION ===");
        // Perform the operation that triggers the bug

        // 3. RECORD AFTER
        console.log("=== AFTER ===");
        uint256 valueAfter = target.criticalValue();
        console.log("Critical value:", valueAfter);

        // 4. PROVE BUG
        console.log("=== VERIFICATION ===");
        // THE ASSERTION THAT PROVES THE BUG
        // Design this so it PASSES when the bug EXISTS
    }

    // === TEST: Impact demonstration (optional) ===
    function test_H{N}_impact() public {
        // Show cumulative impact or attacker profit
    }
}

Interpreting Results

Test PASSES -> Bug CONFIRMED

The assertion that "proves the bug" succeeded.

  • If assertNotEq(after, before) passes -> values ARE different (bug exists)
  • If assertGt(error, threshold) passes -> error IS above threshold (bug exists)
Test FAILS -> Check Why
FailureMeaningAction
Assertion failed: values equalBug doesn't exist as hypothesizedRe-examine hypothesis
Revert in setupDeployment/config wrongFix setup
Revert in actionOperation blockedCheck preconditions
Arithmetic errorValues wrongCheck calculations

Iteration Protocol

Attempt 1: Direct implementation of test strategy from hypothesis

Attempt 2: Adjust parameters

  • Different amounts (larger/smaller)
  • Different timing (more/fewer blocks)
  • Different actors

Attempt 3: Re-examine assumptions

  • Is setup correct?
  • Are preconditions met?
  • Is the bug mechanism correctly understood?

After 5 attempts:

  • If still fails -> FALSE_POSITIVE with documented reasoning
  • Explain why the hypothesis was wrong

Severity Determination

CRITICAL
  • Direct fund theft possible
  • Protocol insolvency
  • No special prerequisites needed
  • Attacker profits significantly
HIGH
  • Fund loss with some setup
  • Broken core functionality
  • Significant value at risk
  • Cumulative error compounds quickly
MEDIUM
  • Limited fund loss
  • Requires specific conditions
  • Edge cases with real impact
  • Moderate value at risk
LOW
  • Negligible direct impact
  • Extreme edge cases only
  • Owner/admin controlled risk
  • Informational with minor consequence

Output Format

CONFIRMED
markdown
## Verdict: CONFIRMED

### Bug Mechanism Verified
{Explain what the test proves in 2-3 sentences}

### Test File
`test/audit/Test_H{N}.t.sol`

### Test Output

{Paste relevant forge test output}


### Key Evidence
| Metric | Value |
|--------|-------|
| Before | {value} |
| After | {value} |
| Expected | {value} |
| Difference | {calculation} |

### Severity: {LEVEL}
{Justification in 1-2 sentences}
FALSE_POSITIVE
markdown
## Verdict: FALSE_POSITIVE

### Attempts Made

**Attempt 1:**
- Approach: {description}
- Result: {what happened}
- Learning: {insight}

**Attempt 2:**
- Approach: {description}
- Result: {what happened}
- Learning: {insight}

**Attempt 3:**
- Approach: {description}
- Result: {what happened}
- Learning: {insight}

### Why It's Not a Bug
{Explain the actual behavior and why hypothesis was wrong in 2-3 sentences}
CONTESTED (NEW in v5 -- CRITICAL)
markdown
## Verdict: CONTESTED

### Evidence Status
| Checkpoint | Status | Details |
|------------|--------|---------|
| External behavior verified against PRODUCTION | NO | Used mock behavior as evidence |
| All callers checked | YES | Checked A, B, C |
| Profit calculated with attacker holding | NO | Only analyzed donation loss |

### Why This Cannot Be REFUTED
{Explain what evidence is missing to definitively rule out the bug}

### Escalation Required
- [ ] Fetch production contract source for {external dep}
- [ ] Re-analyze with attacker holding shares
- [ ] Check additional caller paths: {list}

### Current Assessment
Likely: {TRUE_POSITIVE / FALSE_POSITIVE / UNKNOWN}
Confidence: {LOW / MEDIUM}

Show full SKILL.md (405 more words)Show less

Insufficient Evidence (HALT CONDITIONS - CRITICAL)

MANDATORY: You MUST check ALL boxes before returning REFUTED. If ANY checkbox is NO -> Return CONTESTED, not REFUTED.

Before marking REFUTED, check:

  • External behavior verified against PRODUCTION (not mock)
    • Read {scratchpad}/external_production_behavior.md
    • If external dep is marked 'UNVERIFIED' -> CANNOT use as evidence
    • If mock differs from production -> use PRODUCTION behavior
  • Attack path checked on ALL callers (not just main path)
    • Use mcp__slither-analyzer__get_function_callers() to enumerate
  • Profit calculated with attacker HOLDING tokens (not just donating)
    • "Attacker loses by donating" is NOT sufficient evidence
    • Check: what if attacker holds X% of shares BEFORE donating?
  • Missing precondition documented
    • Document in structured format: precondition type + why it blocks
    • Types: STATE / ACCESS / TIMING / EXTERNAL / BALANCE
  • Searched other findings for matching postconditions
    • Read {scratchpad}/findings_inventory.md for CONFIRMED/PARTIAL findings
    • Check if ANY finding creates the postcondition that would enable this attack
    • If match found -> CONTESTED, not REFUTED (chain analysis will combine)
Evidence That Does NOT Count
  • "Mock shows X" -- mocks != production (CRITICAL: always verify against production)
  • "Standard ERC20" -- may have transfer hooks, side effects
  • "Attacker loses by donating" -- may profit via shares held
  • "Function is internal" -- may be called by public function
  • "Requires admin" -- admin may be compromised or malicious
  • "Attacker cannot acquire X" -- another finding may CREATE this condition
Anti-Downgrade Halt for VS/BLIND Findings (HARD RULE)

For findings from Validation Sweep ([VS-]) or Blind Spot Scanner ([BLIND-]): apply Rule 13's 5-question test BEFORE any downgrade. HALT: If test shows users harmed AND unavoidable AND undocumented -> you CANNOT return FALSE_POSITIVE. Minimum verdict: CONTESTED. Defense parity gaps (Contract A has protection X, Contract B lacks it for same action) are NEVER "by design" -> minimum severity: Medium, minimum verdict: CONTESTED. Violating this halt is a workflow error equivalent to using [MOCK] evidence for REFUTED.

Chain Analysis Integration

A finding is NEVER truly REFUTED until chain analysis completes.

If you mark a finding as REFUTED but document a missing precondition, the chain analyzer (Step 6b) will search for other findings whose postconditions match your missing precondition. If found, the finding will be escalated to CONTESTED and combined into a chain hypothesis.

Example:

  • Your finding: "Donation attack blocked because attacker cannot hold receipt tokens"
  • Other finding: "External protocol interaction returns transferable receipt tokens"
  • Chain: Other finding enables your finding -> Combined HIGH severity


Advanced Protocol Reference: See advanced.md for RAG queries before PoC, exchange rate finding severity, design flaw escalation, bidirectional role analysis, RAG confidence override, chain hypothesis protection, fork testing, and Foundry PoC methodology.

© PlamenTSV, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in agents/skills/evm/verification-protocol of PlamenTSV/plamen.

  • SKILL.md
  • references/advanced.md

Open the folder on GitHubat commit 795962b

Compare with similar skills

Verification Protocol next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verification Protocol compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verification Protocol this skillPlamenTSV/plamen303—~3kAutomated safety check: PassMIT
Fizz Convertpashov/skills1.2k2 repos~3.7kAutomated safety check: PassMIT
Solana Devsolana-foundation/solana-dev-skill571—~3.8kAutomated safety check: PassMIT
Feynman Auditor0xiehnnkta/nemesis-auditor2441 repos~11kAutomated safety check: PassMIT
Smart Contract Auditgreatpie/smart-contract-audit-skill101—~1.1kAutomated safety check: PassNone
RadarAuditware/radar154—~2.1kAutomated safety check: PassGPL-3.0

Similar skills

  • Fizz Convert

    pashov/skills

    Convert English-language properties in PROPERTIES.md (produced by the Fizz skill) into Solidity assertions inside the existing fuzz harness, then flip their checkboxes.

    1.2k GitHub starsUsed in 2 repos~3.7k tokens
    Backend & APIsAuto-check passed
  • Solana Dev

    solana-foundation/solana-dev-skill

    A skill your agent uses when user asks to "build a Solana dapp", "write an Anchor program", "create a token", "debug Solana errors", "set up wallet connection", "test my Solana program", "fuzz my…

    571 GitHub stars~3.8k tokensUpdated 4 days ago
    Backend & APIsAuto-check passed
  • Feynman Auditor

    0xiehnnkta/nemesis-auditor

    Deep business logic bug finder using the Feynman technique. An agent skill from 0xiehnnkta/nemesis-auditor.

    244 GitHub starsUsed in 1 repo~11k tokens
    Backend & APIsAuto-check passed
  • Smart Contract Audit

    greatpie/smart-contract-audit-skill

    Script-backed, out-of-box auditing workflow for Solidity/EVM repositories based on EVMbench detect/patch/exploit methodology.

    101 GitHub stars~1.1k tokensUpdated 7 mo ago
    Backend & APIsAuto-check passed
  • Radar

    Auditware/radar

    Use radar for smart contract security analysis, AST generation, and detection template development.

    154 GitHub stars~2.1k tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Solidity Auditor

    Gabson0x/bountyforge

    Security audit of Solidity code while you develop. An agent skill from Gabson0x/bountyforge.

    443 GitHub stars~3.7k tokensUpdated 20 days ago
    Backend & APIsAuto-check passed

More from PlamenTSV/plamen

All 87 skills in this repo
  • Audit Prep

    PlamenTSV/plamen

    Prepare Solidity projects for a security audit — test coverage, test quality, NatSpec docs, code hygiene, dependency health, best-practice enforcement, deployment readiness, and project…

    303 GitHub stars~3.7k tokensUpdated 11 days ago
    Auto-check passed
  • Verification Protocol

    PlamenTSV/plamen

    Trigger Pattern Always (used by all verifier agents) - Inject Into security-verifier agents (Phase 5)

    303 GitHub stars~3.5k tokensUpdated 11 days ago
    Auto-check passed
  • Ability Analysis

    PlamenTSV/plamen

    Trigger Pattern Always (Aptos Move) - foundational security check - Inject Into Breadth agents, depth agents

    303 GitHub stars~3.3k tokensUpdated 11 days ago
    Auto-check passed
  • Ability Analysis

    PlamenTSV/plamen

    Trigger Pattern Always (Sui Move) -- foundational security check - Inject Into Breadth agents, depth agents

    303 GitHub stars~3.2k tokensUpdated 11 days ago
    Auto-check passed
  • Account Lifecycle

    PlamenTSV/plamen

    Trigger Pattern ACCOUNTCLOSING flag detected (close/CloseAccount usage) - Inject Into Breadth agents, depth agents

    303 GitHub stars~1.2k tokensUpdated 11 days ago
    Auto-check passed
  • Account Validation

    PlamenTSV/plamen

    Trigger Pattern Always required for Solana audits - Inject Into Breadth agents, depth agents

    303 GitHub stars~1.7k tokensUpdated 11 days ago
    Auto-check passed

Categories

Questions about Verification Protocol

What does Verification Protocol do?

How to prove a hypothesis is TRUE or FALSE using Foundry tests. Verification Protocol is an agent skill from PlamenTSV/plamen. How to prove a hypothesis is TRUE or FALSE using Foundry tests.

When should I use Verification Protocol?

Verification Protocol fits situations like: tasks that involve Smart contracts.

How do I install Verification Protocol in Claude Code?

Run `npx skills add PlamenTSV/plamen --skill verification-protocol -a claude-code`. Or copy the skill folder (agents/skills/evm/verification-protocol in PlamenTSV/plamen) into .claude/skills/verification-protocol in your project. Claude Code loads it when a task matches its description.

How do I install Verification Protocol in Codex?

Run `npx skills add PlamenTSV/plamen --skill verification-protocol -a codex`. Or copy the skill folder (agents/skills/evm/verification-protocol in PlamenTSV/plamen) into .agents/skills/verification-protocol in your project. Codex loads it when a task matches its description.

Can I use Verification Protocol in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PlamenTSV/plamen --skill verification-protocol -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verification-protocol, .gemini/skills/verification-protocol, .github/skills/verification-protocol and .opencode/skills/verification-protocol in your project.

What does Verification Protocol need to run?

SKILL.md names no scripts, command-line tools or credentials: Verification Protocol is instructions for the agent only.

Does Verification Protocol access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verification Protocol safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verification Protocol use?

Verification Protocol is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verification Protocol use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.

What are the alternatives to Verification Protocol?

Skills that share tags, products or a category with Verification Protocol: Fizz Convert (pashov/skills, 1.2k stars), Solana Dev (solana-foundation/solana-dev-skill, 571 stars), Feynman Auditor (0xiehnnkta/nemesis-auditor, 244 stars) and Smart Contract Audit (greatpie/smart-contract-audit-skill, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verification Protocol?

PlamenTSV (a GitHub user) maintains it in PlamenTSV/plamen, which has 303 GitHub stars. The repository holds 87 skills in this directory. The repository was last updated on September 26, 2026.

Source: PlamenTSV/plamen on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.