Agent skill

Verify CLI Run

by reticlehq in reticlehq/reticle

Prove that a command-line tool actually did what it said, instead of trusting its exit code and its output.

Apache-2.0Auto-check passedDevelopment

Install Verify CLI Run

skills CLI
$ npx skills add reticlehq/reticle --skill verify-cli-run -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install reticlehq/reticle verify-cli-run --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/verify-cli-run .claude/skills/verify-cli-run && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify-cli-run
GitHub stars
1.2k
Token cost
~2.6k tokens
SKILL.md length
1,330 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
Apache-2.0

At a glance

Prove that a command-line tool actually did what it said, instead of trusting its exit code and its output.

  • Works in 7 steps: Name the consequence BEFORE you run → Snapshot before → Check whether it is already true → …
  • Tasks that involve Linting and formatting
  • SKILL.md covers The rule that decides everything, 1. Name the consequence BEFORE…, 2. Snapshot before and 3. Check whether it is already…, plus 8 more sections
  • Calls git, claude and sh

What it does

Verify CLI Run is an agent skill from reticlehq/reticle. Prove that a command-line tool actually did what it said, instead of trusting its exit code and its output. Snapshots the filesystem before and after, names the expected consequence in advance, and returns one of four verdicts with what could not be seen. Use after running a build, a migration, a scaffolder, a formatter, a codegen step, or an AI coding CLI; when a command printed success and you are about to report "done"; when you are about to write "exit code 0, so it worked"; or when a tool claims it edited…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Linting and formatting and Project scaffolding. It works with Git. The repository describes itself as: AI agents can generate code, but still struggle to understand what they build. Reticle brings Jev-style machine-native runtime perception to web & desktop applications. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Linting and formatting
  • Tasks that involve Project scaffolding

Example prompts

  • “; when you are about to write”
  • “/verify-cli-run”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Name the consequence BEFORE you run
  2. Snapshot before
  3. Check whether it is already true
  4. Run it, and capture separately
  5. Diff, and read the diff as the evidence
  6. Reach a verdict, one of four
  7. Say what you did not see

What it can do on your machine

Read from SKILL.md and the folder at commit 15bddcf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • claude
    • sh
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • reticle.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify CLI Run loads about 2.6k tokens when it runs. Until then it costs about 141 tokens; SKILL.md has 1,330 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~141
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from reticlehq/reticle at commit 15bddcf, republished under its Apache-2.0 licence (© reticlehq). 1,330 words, ~2,591 tokens.

Download SKILL.mdSave it as .claude/skills/verify-cli-run/SKILL.md (or your agent's skills folder).
name
verify-cli-run
description
Prove that a command-line tool actually did what it said, instead of trusting its exit code and its output. Snapshots the filesystem before and after, names the expected consequence in advance, and returns one of four verdicts with what could not be seen. Use after running a build, a migration, a scaffolder, a formatter, a codegen step, or an AI coding CLI; when a command printed success and you are about to report "done"; when you are about to write "exit code 0, so it worked"; or when a tool claims it edited files. Needs nothing installed.
license
Apache-2.0
metadata.version
3.6.0
metadata.homepage
https://www.reticle.sh
metadata.repository
https://github.com/reticlehq/reticle

Prove the command did what it said

A command exited 0 and printed ✓ done. You have learned that the tool reached its own success branch. You have not learned that anything happened.

This is the whole problem with verifying a CLI: the two things everybody checks are both the tool describing itself. The exit code is chosen by the same code that did the work. The output is written by it. When a tool is wrong about what it did, it is wrong on both, in agreement. That is exactly why their agreement proves nothing.

This skill needs no tools installed. It is four rules and some git.

The rule that decides everything

Evidence for a consequence must not come from the thing that performed the action.

Grade every fact before you use it:

What you haveGradeCan it prove the command worked?
Files on disk, before vs afterconsequenceYes. The filesystem decided whether the write landed, not the tool
An API answering when you ask it afterwardsconsequenceYes. Another party, answering you rather than the tool
Exit codepresenceNo. The tool chose it
stdout / stderrcontextNo. The tool wrote it
The tool's own --verbose reportcontextNo

Everything in the bottom three rows is real information and none of it is proof. Use it to explain a verdict, never to reach one.

1. Name the consequence BEFORE you run

This is the method. An expectation written after you see the output can be talked into agreeing with whatever happened; one written in advance can only be met or missed.

Write it down in one line, in the transcript, before the command:

EXPECT: dist/index.js is rewritten, and no file outside dist/ changes.
EXPECT: src/utils.ts gains a function called parseConfig.
EXPECT: the migration creates 3 files under migrations/ and nothing else.

A consequence you cannot state in advance is one you cannot verify. If you cannot name one, say so and stop. That is an honest no-fault, not a pass.

2. Snapshot before

In a git repo. This is the good case, and it is one line:

bash
git status --porcelain > /tmp/before.txt

git status --porcelain is a near-perfect evidence channel: structured, cheap, and independent of the tool. Git's index records what landed on disk, not what the tool meant to do.

Not in a git repo, or the paths are outside it:

bash
find <declared-root> -type f -newermt '1970-01-01' -exec shasum -a 256 {} + | sort > /tmp/before.txt

Three rules for the roots you declare:

  • Exclude node_modules, .git/objects and large build caches unless they are the subject. Hashing a full node_modules is tens of thousands of files.
  • Declare them explicitly. A tool can write anywhere; you are watching a list you chose.
  • Anything outside that list is a blind spot, and you will report it in step 6 rather than pretend it did not happen.

3. Check whether it is already true

The commonest false green in CLI work: dist/index.js exists, so "the build produced dist/index.js" passes, on a no-op rebuild that did nothing at all.

Before running, check whether your expected consequence already holds. If it does, the command cannot prove it. Either pick a consequence the command changes, or delete the artifact first and say that you did.

4. Run it, and capture separately

bash
<the command> > /tmp/out.txt 2> /tmp/err.txt; echo "exit=$?"

Keep the exit code. You are not going to use it as proof. You are going to use it to explain the verdict, and to notice when it disagrees with the filesystem.

Two things worth knowing:

  • Exit codes wrap: sh -c 'exit 256' exits 0. A large code is not always what you think.
  • 130 is a SIGINT, 137 is usually a kill. Those are facts about how the run ended, not about whether it worked.
  • Some tools exit 0 when they declined to do anything. AI coding CLIs do this routinely.

5. Diff, and read the diff as the evidence

bash
git status --porcelain > /tmp/after.txt; diff /tmp/before.txt /tmp/after.txt
git diff -- <the paths you expected to change>

Now answer your step-1 expectation against this and nothing else.

Two checks worth making every time, because they cost nothing:

  • Nothing else changed. A build that also rewrote a lockfile, or a formatter that reformatted a file nobody asked about, is a finding.
  • Content, not just existence. A file that exists is weaker than a file that contains what you expected. But note the asymmetry: whether a file appeared is decided by the filesystem; what is inside it was written entirely by the tool. Existence is stronger evidence than content.

6. Reach a verdict, one of four

Not two. The two extra values are the point: they are the ones that stop ignorance being rounded towards good news.

VerdictWhen
yesThe consequence you named in step 1 is visible in the diff, and it was not already true
noThe diff contradicts it, or the tool claimed something the filesystem does not show
unknownYou could not see what you needed. Nothing was watching, the window was wrong, or the effect is somewhere you were not looking
no-faultEverything was watched, nothing was wrong, and nothing was declared to prove

unknown is not a failure and must never be reported as one. "I could not see" and "it is broken" send somebody in opposite directions: one says look again, the other says go and fix something.

And say what bought the yes. "dist/index.js changed on disk" is a verdict. "The build said it succeeded" is not, and if that is all you have, the honest answer is unknown.

Show full SKILL.md (469 more words)Show less

7. Say what you did not see

A verdict that cannot say what it missed is indistinguishable from one that saw everything. Name the gaps, every time. One line is enough:

Did not observe: the network (cannot see whether it called out); writes outside ./src and ./dist;
files created and deleted during the run; anything the tool's child processes did.

Four gaps are always present and worth naming by default:

  • The network. You cannot see that a command dialled an endpoint. If the claim is about a remote (a push, a deploy, a PR), the filesystem cannot answer it.
  • Writes outside your declared roots. ~/.config, global caches, /tmp. Most AI coding CLIs write to their own state directory on every run.
  • Transient files. Written and deleted inside the run. A before/after snapshot sees net effect only.
  • Child processes, and anything that outlives the command: a fork, a deferred flush, a background upload.

Two moves that cost nothing and catch a lot

Run it twice. A build, a formatter, a codegen step or a migration should produce an empty diff the second time. A tool that keeps changing things on repeat runs has a defect, and this needs no expected-value to check against.

bash
<command> && git status --porcelain > /tmp/a.txt && <command> && git status --porcelain > /tmp/b.txt && diff /tmp/a.txt /tmp/b.txt

Ask the far side. For a network tool, the filesystem cannot help, but the service can. gh pr view, git ls-remote, a curl of the deployed URL. That is another party answering you rather than the tool reporting on itself, so it is consequence-grade, and it is the only proof available for a remote claim.

Do not accept the tool's own local record as a substitute. git push writes .git/refs/remotes/origin/main, and that sha is git's transcription of what it believes the remote said. It is on disk, and it is still the tool describing itself. Ask the remote.

Worked example: an AI coding CLI

The strongest case, because these tools' product is file edits, and their transcript is the least reliable thing about them.

EXPECT: src/config.ts gains a function `parseConfig`. Nothing outside src/ changes.

$ git status --porcelain > /tmp/before.txt
$ grep -c "parseConfig" src/config.ts          # already-true check → 0
$ claude -p "add a parseConfig function to src/config.ts" ; echo "exit=$?"
$ git status --porcelain > /tmp/after.txt ; diff /tmp/before.txt /tmp/after.txt
 M src/config.ts
$ git diff src/config.ts | grep "^+.*parseConfig"
+export function parseConfig(raw: string): Config {

VERDICT: yes, bought by the working-tree diff, which is independent of the agent's report.
         It was not already true (grep returned 0 before).
Did not see: writes under ~/.claude (the tool records every session there); the network;
         whether a permission prompt declined an edit it also claimed.

Note what did not decide it: the exit code, and the paragraph in which the agent said it had added the function.

When to escalate past this skill

This method tops out in four places, and each has a real answer:

You need toUse
See that a command called out to a hosta proxy, or ask the far side afterwards
Drive an interactive panel or a TUIa pseudo-terminal; a shell cannot type into one
Verify a web app rather than a CLIReticle and the verify-ui-change skill
Produce a verdict somebody else can re-checka verification artifact, not a transcript

The one thing not to do

Do not weaken the expectation you wrote in step 1 because the diff did not match it. Rewriting the target after seeing the result is how a verification step becomes a rationalisation step, and it is invisible in a transcript afterwards, because the amended expectation reads exactly like the original one.

If the diff does not match, that is the finding. Report it.

© reticlehq, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/verify-cli-run of reticlehq/reticle.

Open the folder on GitHubat commit 15bddcf

Compare with similar skills

Verify CLI Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify CLI Run compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify CLI Run this skillreticlehq/reticle1.2k—~2.6kAutomated safety check: PassApache-2.0
Saleor Commit Workflowsaleor/saleor23k—~575Automated safety check: PassBSD-3-Clause
Qt Cpp ReviewSerial-Studio/Serial-Studio7.2k—~4.3kAutomated safety check: PassCustom licence
degit Project ScaffoldingRich-Harris/degit7.9k—~534Automated safety check: PassMIT
Qt Qml ReviewSerial-Studio/Serial-Studio7.2k—~3.7kAutomated safety check: PassCustom licence
Lint Commit PRTresjs/tres3.8k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Commits changes in the Saleor codebase and works through pre-commit hook failures from ruff, mypy, the GraphQL schema check and the migrations check.

    23k GitHub stars~575 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Qt Cpp Review

    Serial-Studio/Serial-Studio

    Qt6/C++ deep code review for Serial Studio. An agent skill from Serial-Studio/Serial-Studio.

    7.2k GitHub stars~4.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • degit Project Scaffolding

    Rich-Harris/degit

    Downloads a repository snapshot or template with degit into an empty folder, from GitHub, GitLab, Bitbucket, Sourcehut or a Gist, optionally at a branch, tag or commit.

    7.9k GitHub stars~534 tokensUpdated 25 days ago
    DevelopmentAuto-check passed
  • Qt Qml Review

    Serial-Studio/Serial-Studio

    Qt6/QML deep code review for Serial Studio. An agent skill from Serial-Studio/Serial-Studio.

    7.2k GitHub stars~3.7k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Lint Commit PR

    Tresjs/tres

    Lint local changes, auto-fix, conventional commit, and optionally create PR

    3.8k GitHub stars~1.1k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Qt6 QML Code Reviewer

    x-tools-author/x-tools

    Runs a 47-rule deterministic QML linter, then six parallel deep-analysis passes over bindings, layout, loaders, delegates, states, and performance.

    1.1k GitHub starsUsed in 1 repo~3.6k tokens
    DevelopmentAuto-check passed

More from reticlehq/reticle

All 18 skills in this repo
  • Verify Login Logout

    reticlehq/reticle

    Prove sign-in lands the user where they should be, sign-out really ends the session so a protected page sends them back to sign in, and an expired session asks to sign in again instead of breaking.

    1.2k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Verify Optimistic Update

    reticlehq/reticle

    Prove a UI that updates before the server answers puts things back and says so when the request fails, and keeps the change when it succeeds.

    1.2k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Verify Pagination

    reticlehq/reticle

    Prove the next page of a list, or the next batch of an infinite scroll, loads NEW rows from the server, repeats none, and that the end of the list is handled.

    1.2k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Agentic TDD

    reticlehq/reticle

    Applies red-green TDD to behavior unit tests cannot reach, by stating the expected outcome against the running app with Reticle before writing the feature.

    1.2k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Whole-App Health Sweep

    reticlehq/reticle

    Sweeps a running web app by clicking every reachable control, then reports dead buttons, console errors, failed requests and mismatches between API data and the screen.

    1.2k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Broken UI Debugger

    reticlehq/reticle

    Finds why a running web app misbehaves when the console is empty and the code looks fine, by reading the click, request, store and console together.

    1.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Verify CLI Run

What does Verify CLI Run do?

Prove that a command-line tool actually did what it said, instead of trusting its exit code and its output. Verify CLI Run is an agent skill from reticlehq/reticle. Prove that a command-line tool actually did what it said, instead of trusting its exit code and its output.

When should I use Verify CLI Run?

Verify CLI Run fits situations like: tasks that involve Linting and formatting; tasks that involve Project scaffolding.

How do I install Verify CLI Run in Claude Code?

Run `npx skills add reticlehq/reticle --skill verify-cli-run -a claude-code`. Or copy the skill folder (skills/verify-cli-run in reticlehq/reticle) into .claude/skills/verify-cli-run in your project. Claude Code loads it when a task matches its description.

How do I install Verify CLI Run in Codex?

Run `npx skills add reticlehq/reticle --skill verify-cli-run -a codex`. Or copy the skill folder (skills/verify-cli-run in reticlehq/reticle) into .agents/skills/verify-cli-run in your project. Codex loads it when a task matches its description.

Can I use Verify CLI Run in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add reticlehq/reticle --skill verify-cli-run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify-cli-run, .gemini/skills/verify-cli-run, .github/skills/verify-cli-run and .opencode/skills/verify-cli-run in your project.

What does Verify CLI Run need to run?

Going by SKILL.md and its folder, Verify CLI Run needs the command-line tools its instructions call (git, claude, sh and gh).

Does Verify CLI Run access the network?

SKILL.md names 1 domain. As links in the text: reticle.sh. This is read from the text; nothing was executed.

Is Verify CLI Run safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verify CLI Run use?

Verify CLI Run is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify CLI Run use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verify CLI Run?

Skills that share tags, products or a category with Verify CLI Run: Saleor Commit Workflow (saleor/saleor, 23k stars), Qt Cpp Review (Serial-Studio/Serial-Studio, 7.2k stars), degit Project Scaffolding (Rich-Harris/degit, 7.9k stars) and Qt Qml Review (Serial-Studio/Serial-Studio, 7.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify CLI Run?

reticlehq (a GitHub organization) maintains it in reticlehq/reticle, which has 1,190 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 8, 2026.

Source: reticlehq/reticle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.