Agent skill

Reality Verification

by tzachbon in tzachbon/smart-ralph

This skill should be used when the user asks to "verify a fix", "reproduce failure", "diagnose issue", "check BEFORE/AFTER state", "VF task", "reality check", "check test quality", "mock-only…

MITAuto-check passedTesting & QA

Install Reality Verification

skills CLI
$ npx skills add tzachbon/smart-ralph --skill reality-verification -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tzachbon/smart-ralph reality-verification --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tzachbon/smart-ralph.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ralph-specum/skills/reality-verification .claude/skills/reality-verification && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
reality-verification
GitHub stars
558
Token cost
~877 tokens
SKILL.md length
272 words
Files
3 (incl. references)
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when the user asks to "verify a fix", "reproduce failure", "diagnose issue", "check BEFORE/AFTER state", "VF task", "reality check", "check test quality", "mock-only…

  • Asks to verify a fix
  • SKILL.md covers Goal Detection, Command Mapping, BEFORE/AFTER Documentation and VF Task Format, plus 2 more sections
  • Calls pnpm, gh and tsc
  • Reproduce failure

What it does

Reality Verification is an agent skill from tzachbon/smart-ralph. This skill should be used when the user asks to "verify a fix", "reproduce failure", "diagnose issue", "check BEFORE/AFTER state", "VF task", "reality check", "check test quality", "mock-only tests", or needs guidance on verifying fixes by reproducing failures before and after implementation, or detecting mock-heavy test anti-patterns.

Its SKILL.md is about 880 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/goal-detection-patterns.md` and `references/mock-quality-checks.md`).

It sits in Testing & QA. It works with pnpm. The repository describes itself as: Spec-driven development with smart compaction. Claude Code plugin combining Ralph Wiggum loop with structured specification workflow. The licence is MIT.

When your agent uses it

  • Asks to verify a fix
  • Reproduce failure
  • Check BEFORE/AFTER state
  • Check test quality

Example prompts

  • “verify a fix”
  • “reproduce failure”
  • “diagnose issue”
  • “/reality-verification”

What it can do on your machine

Read from SKILL.md and the folder at commit ac7251a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • gh
    • tsc

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reality Verification loads about 877 tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 272 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~877
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tzachbon/smart-ralph at commit ac7251a, republished under its MIT licence (© tzachbon). 272 words, ~877 tokens.

Download SKILL.mdSave it as .claude/skills/reality-verification/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
reality-verification
description
This skill should be used when the user asks to "verify a fix", "reproduce failure", "diagnose issue", "check BEFORE/AFTER state", "VF task", "reality check", "check test quality", "mock-only tests", or needs guidance on verifying fixes by reproducing failures before and after implementation, or detecting mock-heavy test anti-patterns.
version
0.2.0
user-invocable
false

Reality Verification

For fix goals: reproduce the failure BEFORE work, verify resolution AFTER.

Goal Detection

Classify user goals to determine if diagnosis is needed. See references/goal-detection-patterns.md for detailed patterns.

Quick reference:

  • Fix indicators: fix, repair, resolve, debug, patch, broken, failing, error, bug
  • Add indicators: add, create, build, implement, new
  • Conflict resolution: If both present, treat as Fix

Command Mapping

Goal KeywordsReproduction Command
CI, pipelinegh run view --log-failed
test, testsproject test command
type, typescriptpnpm check-types or tsc --noEmit
lintpnpm lint
buildpnpm build
E2E, UIPlaywright MCP browser tools
API, endpointWebFetch tool

For E2E/deployment verification, use MCP tools (Playwright MCP browser tools for UI, WebFetch tool for APIs).

BEFORE/AFTER Documentation

BEFORE State (Diagnosis)

Document in .progress.md under ## Reality Check (BEFORE):

markdown
## Reality Check (BEFORE)

**Goal type**: Fix
**Reproduction command**: `pnpm test`
**Failure observed**: Yes
**Output**:

FAIL src/auth.test.ts Expected: 200 Received: 401

**Timestamp**: 2026-01-16T10:30:00Z
AFTER State (Verification)

Document in .progress.md under ## Reality Check (AFTER):

markdown
## Reality Check (AFTER)

**Command**: `pnpm test`
**Result**: PASS
**Output**:

PASS src/auth.test.ts All tests passed

**Comparison**: BEFORE failed with 401, AFTER passes
**Verified**: Issue resolved

VF Task Format

Add as task 4.3 (after PR creation) for fix-type specs:

markdown
- [ ] 4.3 VF: Verify original issue resolved
  - **Do**:
    1. Read BEFORE state from .progress.md
    2. Re-run reproduction command: `<command>`
    3. Compare output with BEFORE state
    4. Document AFTER state in .progress.md
  - **Verify**: `grep -q "Verified: Issue resolved" ./specs/<name>/.progress.md`
  - **Done when**: AFTER shows issue resolved, documented in .progress.md
  - **Commit**: `chore(<name>): verify fix resolves original issue`

Test Quality Checks

When verifying test-related fixes, check for mock-only test anti-patterns. See references/mock-quality-checks.md for detailed patterns.

Quick reference red flags:

  • Mock declarations > 3x real assertions
  • Missing import of actual module under test
  • All assertions are mock interaction checks (toHaveBeenCalled)
  • No integration tests
  • Missing mock cleanup (afterEach)

Why This Matters

WithoutWith
"Fix CI" spec completes but CI still redCI verified green before merge
Tests "fixed" but original failure unknownBefore/after comparison proves fix
Silent regressionsExplicit failure reproduction
Manual verification requiredAutomated verification in workflow
Tests pass but only test mocksTests verify real behavior, not mock behavior
False sense of security from green testsConfidence that tests catch real bugs

© tzachbon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/ralph-specum/skills/reality-verification of tzachbon/smart-ralph.

  • SKILL.md
  • references/goal-detection-patterns.md
  • references/mock-quality-checks.md

Open the folder on GitHubat commit ac7251a

Compare with similar skills

Reality Verification next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reality Verification compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reality Verification this skilltzachbon/smart-ralph558—~877Automated safety check: PassMIT
DeerFlow Smoke Testbytedance/deer-flow84k—~2.5kAutomated safety check: NotesMIT
OpenWork Desktop CDP Driverdifferent-ai/openwork24k—~465Automated safety check: PassCustom licence
E2Egronxb/hot-updater1.8k—~1.6kAutomated safety check: PassCustom licence
Cloud Agents Starterscalar/scalar16k—~1.4kAutomated safety check: PassMIT
Desktop Dual Repo TestEynzof/Hermes-CN-Desktop1.8k—~1.8kAutomated safety check: PassCustom licence

Similar skills

  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    84k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • OpenWork Desktop CDP Driver

    different-ai/openwork

    Drives a running OpenWork desktop window over CDP from the shell to evaluate JS, take screenshots, start sessions and send prompts for hand checks.

    24k GitHub stars~465 tokensUpdated today
    Testing & QAAuto-check passed
  • E2E

    gronxb/hot-updater

    Run end-to-end OTA verification for examples/v0.85.0 with agent-device.

    1.8k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Minimal starter runbook for cloud agents to install dependencies, run packages, execute tests, and troubleshoot the Scalar monorepo quickly.

    16k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Desktop Dual Repo Test

    Eynzof/Hermes-CN-Desktop

    A skill your agent uses when starting or verifying Hermes Agent CN Desktop with the latest Hermes-CN-Desktop and Hermes-CN-Core branches — dev smoke test (pnpm tauri:dev), packaged beta/release…

    1.8k GitHub stars~1.8k tokensUpdated 19 days ago
    Testing & QAAuto-check passed
  • Local CI

    module-federation/core

    Run this repository's local CI parity commands and pnpm run ci:local jobs.

    2.7k GitHub stars~914 tokensUpdated yesterday
    Testing & QAAuto-check passed

More from tzachbon/smart-ralph

All 26 skills in this repo
  • Communication Style

    tzachbon/smart-ralph

    This skill should be used when generating spec artifacts (research.md, requirements.md, design.md, tasks.md), formatting agent output, structuring phase results, or when any Ralph agent needs…

    558 GitHub stars~551 tokensUpdated 23 days ago
    Auto-check passed
  • Interview Framework

    tzachbon/smart-ralph

    This skill should be used when a Ralph phase must identify critical user decisions, run a layered grill, persist partial answers, obtain explicit approval, or resume an interrupted phase interview…

    558 GitHub stars~2.2k tokensUpdated 23 days ago
    Auto-check passed
  • Spec Workflow

    tzachbon/smart-ralph

    This skill should be used when the user asks to "build a feature", "create a spec", "start spec-driven development", "run research phase", "generate requirements", "create design", "plan tasks"…

    558 GitHub stars~1.1k tokensUpdated 23 days ago
    Auto-check passed
  • Ralph Specum Design

    tzachbon/smart-ralph

    This skill should be used only when the user explicitly asks to use $ralph-specum-design, or explicitly asks Ralph Specum in Codex to run the design phase.

    558 GitHub stars~1.5k tokensUpdated 23 days ago
    Auto-check passed
  • Ralph Specum Tasks

    tzachbon/smart-ralph

    This skill should be used only when the user explicitly asks to use $ralph-specum-tasks, or explicitly asks Ralph Specum in Codex to run the tasks phase.

    558 GitHub stars~1.6k tokensUpdated 23 days ago
    Auto-check passed
  • Smart Ralph

    tzachbon/smart-ralph

    This skill should be used when the user asks about "ralph arguments", "quick mode", "commit spec", "max iterations", "ralph state file", "prototype overlay", "execution modes", "ralph loop"…

    558 GitHub stars~2k tokensUpdated 23 days ago
    Auto-check passed

Works with

Questions about Reality Verification

What does Reality Verification do?

This skill should be used when the user asks to "verify a fix", "reproduce failure", "diagnose issue", "check BEFORE/AFTER state", "VF task", "reality check", "check test quality", "mock-only…. Reality Verification is an agent skill from tzachbon/smart-ralph. This skill should be used when the user asks to "verify a fix", "reproduce failure", "diagnose issue", "check BEFORE/AFTER state", "VF task", "reality check", "check test quality", "mock-only tests", or needs guidance on verifying fixes by reproducing failures before and after implementation, or detecting mock-heavy test anti-patterns.

When should I use Reality Verification?

Reality Verification fits situations like: asks to verify a fix; reproduce failure; check BEFORE/AFTER state; check test quality.

How do I install Reality Verification in Claude Code?

Run `npx skills add tzachbon/smart-ralph --skill reality-verification -a claude-code`. Or copy the skill folder (plugins/ralph-specum/skills/reality-verification in tzachbon/smart-ralph) into .claude/skills/reality-verification in your project. Claude Code loads it when a task matches its description.

How do I install Reality Verification in Codex?

Run `npx skills add tzachbon/smart-ralph --skill reality-verification -a codex`. Or copy the skill folder (plugins/ralph-specum/skills/reality-verification in tzachbon/smart-ralph) into .agents/skills/reality-verification in your project. Codex loads it when a task matches its description.

Can I use Reality Verification in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tzachbon/smart-ralph --skill reality-verification -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/reality-verification, .gemini/skills/reality-verification, .github/skills/reality-verification and .opencode/skills/reality-verification in your project.

What does Reality Verification need to run?

Going by SKILL.md and its folder, Reality Verification needs the command-line tools its instructions call (pnpm, gh and tsc).

Does Reality Verification access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Reality Verification safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Reality Verification use?

Reality Verification is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reality Verification use?

About 877 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Reality Verification?

Skills that share tags, products or a category with Reality Verification: DeerFlow Smoke Test (bytedance/deer-flow, 84k stars), OpenWork Desktop CDP Driver (different-ai/openwork, 24k stars), E2E (gronxb/hot-updater, 1.8k stars) and Cloud Agents Starter (scalar/scalar, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reality Verification?

tzachbon (a GitHub user) maintains it in tzachbon/smart-ralph, which has 558 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on September 16, 2026.

Source: tzachbon/smart-ralph on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.