Install the "prove-feature" agent skill from https://github.com/searlsco/prove_it/tree/main/.claude/skills/prove-feature into .claude/skills/prove-feature/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prove-feature", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add searlsco/prove_it --skill prove-feature -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "prove-feature" agent skill from https://github.com/searlsco/prove_it/tree/main/.claude/skills/prove-feature into .agents/skills/prove-feature/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prove-feature", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add searlsco/prove_it --skill prove-feature -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "prove-feature" agent skill from https://github.com/searlsco/prove_it/tree/main/.claude/skills/prove-feature into .cursor/skills/prove-feature/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prove-feature", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add searlsco/prove_it --skill prove-feature -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "prove-feature" agent skill from https://github.com/searlsco/prove_it/tree/main/.claude/skills/prove-feature into .gemini/skills/prove-feature/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prove-feature", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add searlsco/prove_it --skill prove-feature -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "prove-feature" agent skill from https://github.com/searlsco/prove_it/tree/main/.claude/skills/prove-feature into .github/skills/prove-feature/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prove-feature", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add searlsco/prove_it --skill prove-feature -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "prove-feature" agent skill from https://github.com/searlsco/prove_it/tree/main/.claude/skills/prove-feature into .opencode/skills/prove-feature/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prove-feature", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
prove-feature
GitHub stars
198
Token cost
~5.7k tokens
SKILL.md length
1,581 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT
At a glance
Create a temporary real project and prove a proveit feature works (or doesn't) end-to-end.
Works in 5 steps: Design the scenario BEFORE writing code → Write a self-contained Node.js test script → Implement the scenarios — failure first,… → …
You need to prove a feature
SKILL.md covers What "prove" means — read this…, Arguments, Method and Design principles, plus 2 more sections
Calls claude and node
What it does
Prove Feature is an agent skill from searlsco/prove_it. Create a temporary real project and prove a proveit feature works (or doesn't) end-to-end. Builds a disposable git repo, writes a focused config, runs real dispatches through the installed or local proveit, and produces a human-readable session transcript. Use when you need to prove a feature, reproduce a bug, or validate a fix against a real project — not just unit tests.
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Unit testing. It works with Git. The repository describes itself as: The verification harness that Claude Code should have shipped with. The licence is MIT.
When your agent uses it
You need to prove a feature
Reproduce a bug
Validate a fix against a real project — not just unit tests
Example prompts
“/prove-feature”
Requirements
Node.js
Workflow steps
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit eedd8da. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Shell commands in SKILL.md call:
claude
node
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Prove Feature loads about 5.7k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 1,581 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~98
When it runs· the whole SKILL.md, loaded when a task matches
~5.7k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/prove-feature/SKILL.md (or your agent's skills folder).
name
prove-feature
description
Create a temporary real project and prove a prove_it feature works (or doesn't) end-to-end. Builds a disposable git repo, writes a focused config, runs real dispatches through the installed or local prove_it, and produces a human-readable session transcript. Use when you need to prove a feature, reproduce a bug, or validate a fix against a real project — not just unit tests.
Prove a feature works (or doesn't)
Build a throwaway project and exercise a prove_it feature through the real
dispatcher pipeline. The output is a human-readable transcript the user can
read to confirm the system works end-to-end.
What "prove" means — read this first
Proving a feature means watching the feature do its actual job, not just
watching the dispatcher accept a config and return a decision.
If the feature is a reviewer that detects dead code, you must:
Create a project that contains dead code → run the reviewer → see it catch the dead code
Create a project that has no dead code → run the reviewer → see it pass clean
If the feature is a task that validates API design, you must:
Write an API file with real design violations → see the task reject it
Write a clean API file → see the task approve it
If the feature is a when-condition gate, you must:
Run with the condition unmet → see the task get skipped
Run with the condition met → see the task actually execute and produce its real output
The pattern is always the same: construct a realistic situation where the
feature's logic is exercised, then observe it both succeed and fail. A test
that only checks "did the dispatcher return allow/deny" without verifying the
feature itself inspected the right thing and made the right call is not a proof.
The critical question
Before writing any code, answer this: "What real-world situation does this
feature exist to handle, and how will I simulate that situation in a throwaway
project?"
If you can't answer that, you don't understand the feature well enough to prove
it yet. Stop and think harder.
Anti-patterns — do NOT do these
Plumbing-only tests. Testing that the dispatcher routes to a script and
returns the script's exit code is testing prove_it's plumbing, not the feature.
The feature is what the script does. You must create input that makes the
script's logic actually fire.
Trivial scripts. Writing #!/bin/bash\nexit 0 and #!/bin/bash\nexit 1
proves the dispatcher handles exit codes. It proves nothing about the feature.
The script must contain the feature's real logic (or a faithful stand-in), and
the test input must be realistic enough to exercise it.
Success-only tests. If you only prove the feature passes, you haven't
proved it works — you may have proved it does nothing. A reviewer that always
says "looks good" is broken. Always prove the feature can fail/reject/deny
before proving it can pass/approve/allow. Failure-first is how you know the
feature has teeth.
Config-exists tests. Writing a config, loading it, and checking it parsed
correctly is a config test, not a feature test. The feature is what happens
after the config is loaded.
Arguments
<feature description> — a short description of what to prove. Examples:
"the dead-code reviewer catches unused functions"
"the API design validator rejects endpoints without error schemas"
"linesChanged threshold triggers at the right count"
Method
Step 1: Design the scenario BEFORE writing code
This is the most important step. Think through:
What does the feature actually do? Not what config key enables it —
what real-world thing does it detect, enforce, or transform?
What does a realistic failing input look like? Build a project state
that the feature should catch. Be specific: if it's a code reviewer, write
actual bad code. If it's a design validator, write an actual bad design.
If it's a file-size gate, create an actual large file.
What does a realistic passing input look like? Build the "clean" version
of the same scenario. Same project structure, but with the problem fixed.
What side effects should you observe? Beyond pass/fail decisions, what
should the feature produce? Error messages with specific details? Log entries?
Modified files? The proof should verify these too — they're how you know the
feature understood the input, not just guessed.
Write down your scenario plan as comments in the script before implementing it.
Step 2: Write a self-contained Node.js test script
Create a single file at /tmp/prove_it_feature_<name>/prove.js that does
everything: creates the temp project, writes config, builds the realistic
scenario, runs dispatches, and prints results. This is not a node:test
test — it's a standalone script that prints a human-readable transcript.
Step 3: Implement the scenarios — failure first, then success
Fill in the // ... scenario implementation goes here ... section. Every
scenario must follow this structure:
Build the realistic project state. Write actual source files, configs,
scripts — whatever the feature needs to inspect. These must be realistic
enough that the feature's logic is meaningfully exercised. A dead-code
detector needs real code with real dead functions. A design validator needs
a real API spec with real violations.
Prove the feature catches the problem (failure/deny case first). Run the
dispatcher against the bad input. Assert not just the decision, but that
the reason or output shows the feature understood what was wrong. A deny
with a generic reason like "script exited 1" is not proof — the reason
should reference the specific problem (e.g., "unused function oldHandler
detected in src/routes.js").
Fix the problem in the project, then prove the feature approves. Modify
the project to resolve the issue (remove the dead code, add the missing
schema, etc.), then re-run. Assert that the feature now passes. This
confirms the feature is actually sensitive to the input, not just randomly
failing.
Check side effects. If the feature should produce logs, annotations,
modified files, or specific error messages, verify those exist and contain
the right content.
Follow these patterns for common feature types:
Show full SKILL.md (645 more words)Show less
Custom reviewer/validator tasks
This is the most common case. You're proving that a task (script, command, etc.)
correctly analyzes project state.
javascript
// ── Scenario: Dead code reviewer ──
header('Scenario 1: Dead code reviewer catches unused exports')
// Step 1: Write the reviewer script (this IS the feature)
writeFile('scripts/check-dead-code.sh', `#!/bin/bash
# Scan for exported functions that are never imported elsewhere
dead=$(grep -rn 'export function' src/ | while read line; do
fn=$(echo "$line" | sed 's/.*export function \\([a-zA-Z_]*\\).*/\\1/')
if ! grep -rq "$fn" src/ --include='*.js' -l | grep -v "$(echo "$line" | cut -d: -f1)" > /dev/null 2>&1; then
echo "$line"
fi
done)
if [ -n "$dead" ]; then
echo "Dead exports found:"
echo "$dead"
exit 1
fi
echo "No dead exports"
exit 0
`)
makeExecutable('scripts/check-dead-code.sh')
writeConfig({ enabled: true, hooks: [{ type: 'claude', event: 'PreToolUse', tasks: [{
name: 'dead-code-check',
type: 'script',
command: './scripts/check-dead-code.sh'
}] }] })
// Step 2: Create a project WITH dead code → feature should catch it
writeFile('src/utils.js', `
export function activeHelper() { return 'used' }
export function staleHelper() { return 'nobody calls me' }
`)
writeFile('src/main.js', `
import { activeHelper } from './utils.js'
console.log(activeHelper())
`)
gitIn(PROJECT_DIR, 'add', '.')
gitIn(PROJECT_DIR, 'commit', '-m', 'add code with dead export')
const r1 = invokePreToolUse('Bash', { command: 'echo editing src/main.js' })
check('Dead code present → reviewer denies', decision(r1) === 'deny')
check('Reason mentions the dead function', reason(r1).includes('staleHelper'),
`Expected reason to mention "staleHelper", got: ${reason(r1).slice(0, 200)}`)
printRaw('raw', r1)
// Step 3: Remove the dead code → feature should pass
writeFile('src/utils.js', `
export function activeHelper() { return 'used' }
`)
gitIn(PROJECT_DIR, 'add', '.')
gitIn(PROJECT_DIR, 'commit', '-m', 'remove dead export')
const r2 = invokePreToolUse('Bash', { command: 'echo editing src/main.js' })
check('Dead code removed → reviewer allows', decision(r2) === 'allow')
check('Reason confirms clean scan', reason(r2).includes('No dead exports'),
`Expected clean message, got: ${reason(r2).slice(0, 200)}`)
printRaw('raw', r2)
The key: the script contains real analysis logic, the project contains real
code, and we verify the feature's output references the specific problem.
When-condition gating
javascript
header('Scenario 2: When-condition gates task execution')
writeConfig({ enabled: true, hooks: [{ type: 'claude', event: 'PreToolUse', tasks: [{
name: 'gated-review', type: 'script', command: './scripts/check-dead-code.sh',
when: { envSet: 'RUN_DEAD_CODE_CHECK' }
}] }] })
// Without the env var → task should be skipped entirely
const r3 = invokePreToolUse('Bash', { command: 'echo test' })
check('No env var → task skipped (not denied)', decision(r3) !== 'deny')
printRaw('raw', r3)
// With the env var → task should actually run (and find the dead code if present)
const r4 = invokePreToolUse('Bash', { command: 'echo test' }, { RUN_DEAD_CODE_CHECK: '1' })
check('Env var set → task runs', decision(r4) !== '(silent)',
`Expected the task to run, got decision=${decision(r4)}`)
printRaw('raw', r4)
Stateful features (appeal, suspension, failure counting)
Run the same dispatch multiple times with the same session ID. Assert
that behavior changes across invocations (e.g., failure count increments,
backchannel appears, task gets suspended).
javascript
header('Scenario 3: Repeated failures trigger suspension')
// Use a script that always fails — simulating a reviewer that keeps catching problems
writeFile('scripts/always-fail.sh', '#!/bin/bash\necho "design violation: missing error schema"\nexit 1')
makeExecutable('scripts/always-fail.sh')
writeConfig({ enabled: true, hooks: [{ type: 'claude', event: 'PreToolUse', tasks: [{
name: 'strict-reviewer', type: 'script', command: './scripts/always-fail.sh'
}] }] })
// Run multiple times — behavior should change as failures accumulate
for (let i = 1; i <= 5; i++) {
const r = invokePreToolUse('Bash', { command: `echo attempt ${i}` })
info(`Attempt ${i}: decision=${decision(r)} reason=${reason(r).slice(0, 80)}`)
}
// (Assert on the specific suspension behavior expected by the feature)
Show the full terminal output to the user. The transcript IS the proof.
Design principles
The feature must do its actual job in the test. This is the #1 principle.
If you're proving a code reviewer, it must review real code and catch real
problems. If you're proving a linter gate, it must lint real files. If you're
proving a threshold trigger, you must cross the actual threshold. The
dispatcher plumbing (routing, config parsing, exit code handling) is assumed
to work — you're proving the feature, not the framework.
Failure first, then success. Always prove the feature can reject/deny/fail
before proving it can approve/pass/allow. A system that always says "yes" is
indistinguishable from a system that does nothing. The deny case is what proves
the feature has teeth. Only after seeing a legitimate deny should you fix the
input and verify the allow.
Verify the reason, not just the decision. A decision of "deny" could mean
anything. The reason is how you know the feature understood the problem. Check
that the reason references the specific issue (the dead function name, the
missing field, the threshold value). Generic reasons like "script failed" are
not proof.
One script, zero dependencies. The prove script must be a single file
that uses only Node.js stdlib. No test framework. No imports from the
prove_it repo (the dispatcher is invoked as a subprocess, not imported).
Isolated from real config. Always use a fake HOME and PROVE_IT_DIR.
Never touch ~/.claude/ or the user's real sessions.
FAKE_HOME is HOME. Several prove_it internals use process.env.HOME,
not CLAUDE_PROJECT_DIR. For example, findPlanFile() searches
HOME/.claude/plans/, not the project directory. When your scenario
involves plan files, place them under FAKE_HOME/.claude/plans/, not
PROJECT_DIR/.claude/plans/. If a task passes silently but has no
side effects, a wrong HOME-vs-project path split is the likely cause.
Agent tasks need skills installed in FAKE_HOME. When proving agent-type
tasks (type: 'agent' with promptType: 'skill'), the reviewer subprocess
(claude -p) looks for the skill at FAKE_HOME/.claude/skills/<name>/SKILL.md.
If you don't copy the skill file there, the agent will fail with "skill not
found" — which looks like a broken feature but is just a missing setup step.
Copy the skill from the repo before invoking the dispatcher:
Similarly, claude -p is available on PATH system-wide, so agent tasks can
and should fully execute their reviewer subprocess in the test environment.
Human-readable first. The output is for a human reading a terminal.
Use color, alignment, and section headers. Print the raw dispatcher output
for each scenario so the reader can verify without expanding tool calls.
Session transcript is mandatory. Always call printSessionLog() at the
end. The session .jsonl file is the ground truth for what the dispatcher
did. If it's empty, something is wrong.
Use the release binary by default. Set USE_LOCAL = false unless you're
testing a fix that isn't released yet. The whole point is to prove the
shipped system works.
Exit non-zero on failure. The script's exit code is the verdict.
Choosing local vs release
Scenario
USE_LOCAL
Why
Proving a shipped feature works
false
Tests what users actually run
Validating a fix before release
true
Tests the working tree
Reproducing a bug
false first
Confirm bug exists in release, then true to verify fix
Reporting
Present the full terminal output to the user. The output should look like:
── Environment ──
INFO prove_it: /opt/homebrew/bin/prove_it
INFO source: release (Homebrew)
INFO project: /tmp/prove_feature_abc123
INFO session: prove-1772134567890
── Scenario 1: Dead code reviewer catches unused exports ──
FAIL Dead code present → reviewer denies
PASS Reason mentions the dead function
raw: decision=deny reason=Dead exports found: src/utils.js:3 staleHelper
PASS Dead code removed → reviewer allows
PASS Reason confirms clean scan
raw: decision=allow reason=No dead exports
── Scenario 2: When-condition gates task execution ──
PASS No env var → task skipped (not denied)
PASS Env var set → task runs
── Session Transcript ──
TIME STATUS TASK REASON
──────── ─────── ─────────────────── ──────────────────────────────
19:25:35 DENY dead-code-check Dead exports found: staleHelper
19:25:35 PASS dead-code-check No dead exports
19:25:36 SKIP gated-review Skipped: $RUN_DEAD_CODE_CHECK not set
19:25:36 PASS gated-review No dead exports
── Summary ──
✓ All 6 checks passed
Prove Feature next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Run fast static checks that don't require generating downstreams, including git submodule verification, gofmt formatting, YAML linting, template validation (version-guard and unused-tmpl), mmv1 core…
Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need…
How to test and verify work in the dev-3.0 repo — which vitest config covers what, how to write a test that fits the house style, mocking Electrobun RPC and i18n providers, what coverage is actually…
Conventions for writing and reviewing unit and integration tests of Git::Repository facade methods in the ruby-git project, covering setup, cases, grouping and scope.
Create a temporary real project and prove a proveit feature works (or doesn't) end-to-end. Prove Feature is an agent skill from searlsco/prove_it. Create a temporary real project and prove a proveit feature works (or doesn't) end-to-end.
When should I use Prove Feature?
Prove Feature fits situations like: you need to prove a feature; reproduce a bug; validate a fix against a real project — not just unit tests.
How do I install Prove Feature in Claude Code?
Run `npx skills add searlsco/prove_it --skill prove-feature -a claude-code`. Or copy the skill folder (.claude/skills/prove-feature in searlsco/prove_it) into .claude/skills/prove-feature in your project. Claude Code loads it when a task matches its description.
How do I install Prove Feature in Codex?
Run `npx skills add searlsco/prove_it --skill prove-feature -a codex`. Or copy the skill folder (.claude/skills/prove-feature in searlsco/prove_it) into .agents/skills/prove-feature in your project. Codex loads it when a task matches its description.
Can I use Prove Feature in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add searlsco/prove_it --skill prove-feature -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prove-feature, .gemini/skills/prove-feature, .github/skills/prove-feature and .opencode/skills/prove-feature in your project.
What does Prove Feature need to run?
Going by SKILL.md and its folder, Prove Feature needs the command-line tools its instructions call (claude and node). Our summary lists: Node.js.
Does Prove Feature access the network?
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Is Prove Feature safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Prove Feature use?
Prove Feature is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Prove Feature use?
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to Prove Feature?
Skills that share tags, products or a category with Prove Feature: Run Pre Gen Checks (GoogleCloudPlatform/magic-modules, 974 stars), Test Pyramid (kubernetes-sigs/agent-sandbox, 4.2k stars), Walkthrough (smallnest/goal-workflow, 291 stars) and Verify Changes (h0x91b/dev-3.0, 309 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Prove Feature?
searlsco (a GitHub organization) maintains it in searlsco/prove_it, which has 198 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 22, 2026.
Source: searlsco/prove_it on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.