Agent skill

Reticle Runtime Verification

by reticlehq in reticlehq/reticle

Installs Reticle's dev-only SDK in a running web app and verifies user-facing changes by driving a real flow, returning a verdict with the file and line to fix.

Apache-2.0Auto-check: notesTesting & QA

Install Reticle Runtime Verification

skills CLI
$ npx skills add reticlehq/reticle --skill reticle -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install reticlehq/reticle reticle --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin .claude/skills/reticle && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
reticle
GitHub stars
1.2k
Token cost
~2.8k tokens
SKILL.md length
1,546 words
Files
5
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Installs Reticle's dev-only SDK in a running web app and verifies user-facing changes by driving a real flow, returning a verdict with the file and line to fix.

  • Works in 2 steps: No recognisable dev script in… → Your host asks the human to approve a…
  • Installing and setting up Reticle in a web project
  • SKILL.md covers Finish the setup steps, Run this branch first, What YOU decide, and pass in and Then read what it gives you back, plus 3 more sections
  • Runs JavaScript scripts from its folder; calls curl and npx; reaches docs.reticle.sh; needs RETICLE_LICENSE_KEY

What it does

Reticle embeds a dev-only SDK inside the running web app and exposes it to the agent as reticle_* MCP tools, so the agent can look, act, observe and assert against the real DOM, network, routing, console and framework state instead of relying on screenshots. The plugin has already registered the MCP server, so no client restart is needed.

The skill first checks for a .reticle.json file. If it is missing, the project is onboarded with npx @reticlehq/server@latest init, which detects the framework and package manager, wires the build config, installs the SDK, registers the MCP server, starts the dev server, opens the app and waits for a session to connect; it exits non-zero if nothing connects. Either way the work is finished only when a verdict exists from reticle_act_and_wait, reticle_assert, or reticle_act with a step that declares an expectation.

The first run is a separate reticle_verify call with the explore action and a persona. It drives the app with a model inside the daemon and records the flow, so later checks replay it with no model involved. The agent asks you only when a step needs a decision, such as no recognizable dev script in package.json or a host prompt to approve a command, which it must never bypass.

When your agent uses it

  • Installing and setting up Reticle in a web project
  • Proving a user-facing change works before calling it done
  • Investigating a UI that is broken even though the tests pass
  • Checking DOM, network, routing and console state of a running app without screenshots

Example prompts

  • “Set up Reticle in this web app and run a first verification of the checkout flow.”
  • “I changed the login form validation. Prove it works in the running app before we merge.”
  • “The tests pass but the settings page renders blank. Find out why using Reticle.”
  • “/reticle verify the signup flow end to end and tell me which file to fix.”

Requirements

  • A web app with a recognizable dev script in package.json
  • npx, to run the Reticle server init command
  • The Reticle plugin with its MCP server registered

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. No recognisable dev script in package.json. Say so; do not invent one.
  2. Your host asks the human to approve a command. That prompt belongs to the host. Never bypass or suppress it, and take a refusal as the…

What it can do on your machine

Read from SKILL.md and the folder at commit 178e5c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • curl
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • docs.reticle.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • RETICLE_LICENSE_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reticle Runtime Verification loads about 2.8k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 1,546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:67
    one: it writes `RETICLE_LICENSE_KEY` to `.env` and keeps `.env` out of git.
  • NoteMentions a .env fileSKILL.md:134
    CLE_LICENSE_KEY=<key>` to the project's `.env` (create it if absent). The CLI loads a project-local `.env` on every invo
  • NoteMentions a .env fileSKILL.md:135
    2. Make sure `.env` is in `.gitignore`. A license key in git is a leaked credential, so check before you write and add t

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from reticlehq/reticle at commit 178e5c0, republished under its Apache-2.0 licence (© reticlehq). 1,546 words, ~2,829 tokens.

Download SKILL.mdSave it as .claude/skills/reticle/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
reticle
description
Install, instrument and verify this running web app from the inside (DOM, network, routing, console and framework state) instead of screenshots or guessing. Drives one real flow end to end and returns a verdict with the file:line to fix. Use when the user asks to set up or install Reticle, when a user-facing change needs proving before you call it done, when a test passes but the UI is broken, or when the user types /reticle.
license
Apache-2.0
metadata.version
3.7.0
metadata.homepage
https://www.reticle.sh
metadata.repository
https://github.com/reticlehq/reticle

Reticle

Reticle embeds a dev-only SDK in the running app and exposes it to you as reticle_* MCP tools. You look, act, observe, and assert against the real app. No screenshots.

Finish the setup steps

The plugin already registered the MCP server, so this file runs with no client restart. Finish the setup steps; ask the user only when a step needs their decision. Running init, fixing wiring it could not, starting the dev server and opening the browser are setup steps, not decisions.

The repo already answers which framework, package manager and port, so work those out rather than asking. Say what you did in one line.

Two places always need the user:

  1. No recognisable dev script in package.json. Say so; do not invent one.
  2. Your host asks the human to approve a command. That prompt belongs to the host. Never bypass or suppress it, and take a refusal as the answer. init writing a pre-approval rule for the reticle server is not that: it is a scoped, announced config change the human asked for by running the command, and it covers only Reticle's own tools.

Run this branch first

bash
cat .reticle.json 2>/dev/null || echo NOT_FOUND
  • NOT_FOUND → ONBOARD, then VERIFY.
  • File exists → VERIFY.

Either way you are finished only when a verdict exists: from reticle_act_and_wait, reticle_assert, or reticle_act { steps } where a step declares expect. Config files are not an install, and a listed session is not a result.


ONBOARD

One command wires the project. A second one proves a flow.

bash
npx @reticlehq/server@latest init

It detects the framework and package manager, wires the build config, installs the SDK, registers the MCP server, starts the dev server, opens the app, and waits for a session to connect from inside it. That connection IS the proof onboarding worked: the SDK is in the page and the tools have something to talk to. It exits non-zero if nothing connected, and prints exactly what is left to do.

Then prove a flow. That is the FIRST RUN, and it is a separate call:

reticle_act_and_wait { ref, action, until }

Drive the journey that matters and put the verdict on its LAST step: until names the end state before the action fires. What you drive is saved as a flow, so later runs replay it with no model. On a linked project (reticle connect; every plan, Free included, has monthly Harness credits), reticle_verify { action: "explore", persona: "<who does what>" } has the Reticle Harness drive the whole journey for you instead. Before driving anything, replay what is already saved: reticle_verify { action: "flows" } costs no model at all.

What YOU decide, and pass in

The command reads the repository. It cannot read the request, and these live only there.

flagwhat only you know
persona: "<what>" (on the FIRST RUN, not on init)which journey proves the thing the user asked for. Code can list the buttons; it cannot know checkout matters and the theme toggle does not.
--env KEY=VALUEwhat the app needs to reach a usable state: the key from .env.example, the mock backend, the variable that skips an auth wall. Repeatable.
--app <dir>which app in a monorepo. It can list the servable ones; only the request says which is being worked on.

Add --license <key> if the user gave you one: it writes RETICLE_LICENSE_KEY to .env and keeps .env out of git.

Framework, package manager, port, editor and MCP client are answerable from the repo you are sitting in, so work them out rather than asking.

Then read what it gives you back

A non-zero exit is a to-do list, not a failed install. The command names the cause and prints the REMAINING steps from wherever it stopped; it will not tell you to redo a phase that already worked. Do those and re-run, which is safe.

It is not finished until a verdict exists. Writing files is not an install, and neither is a connected session.

If that command could not run it

Do not choose this path. It is not the thorough version of the one above; it is what you fall back to when the command physically could not do the work. Use it only when init exited without ever printing starting: or ▸ WATCH (an older CLI that stops after writing files), or when it stopped in the same place twice after you did what it asked. A ⚠ in the report is not a reason: re-run the command, which is idempotent and names what is still outstanding.

bash
curl https://docs.reticle.sh/install-manual.md      # register the MCP, wire the SDK, prove it
curl https://docs.reticle.sh/troubleshooting.md     # nothing connected, click did nothing, verdict unknown

VERIFY

Verdicts come from reticle_act_and_wait, reticle_assert, and reticle_act { steps } when a step declares expect. Everything else (a bare act, look, navigate, observe) moves or reads the app and proves nothing, however many tools it used.

verified: "unknown" is not a pass. It means Reticle drove the app and could not tell what happened, so report it as unknown. verified: "no-fault" is not a pass either: the page settled and no channel complained, but nothing was declared to prove. Never weaken a check to make it pass. That converts a real signal into a false one, which is the failure this product exists to prevent.

Take the cheapest path that answers the question

Stop at the first row that fits.

The questionThe callCalls
"Did my edit break anything?"reticle_verify({ action: "change", files: ["src/App.tsx"] })1
"Does this known journey still work?"reticle_run({ tool: "reticle_flow_replay", args: { flowName: "login" } })1
"Does this new behaviour work?"reticle_act { steps: [...] } to the last page, then reticle_act_and_wait on the step that ENDS the journey2
No MCP reachable at allnpx @reticlehq/server verify <url> in the shell1, no MCP

reticle_flow_replay is not on the advertised tool list. It is reached through reticle_run exactly as written, which is the supported call shape and why you have to be told it exists. reticle_verify {action:"change"} answers unknown when no saved flow covers the files you changed: nothing ran, so nothing was proved. That is the honest answer and the signal to record one, never a pass.

Show full SKILL.md (564 more words)Show less

Driving by hand

Four calls for a login, not fourteen. Every call is a full model turn, and in a client that approves each one it is also a click.

  1. reticle_look({ action: "page", mode: "interactive" }) once, for the whole flow. Elements are addressable by role and name, so you do not need to add data-testid anywhere.
  2. reticle_act { steps: [...] } for the setup: every fill and every intermediate click in ONE call.
  3. reticle_act_and_wait({ ref, action, until }) for the step that ENDS the journey (the confirmation, the saved record, the last page), not the first click that looks like success. until names that end state before the action fires, which is the difference between a check and a rationalisation. Pass durable: true when the change should survive a reload. A step on the way with a consequence of its own (a save, a submit) gets its own expect in steps, so a journey that breaks at step three fails at step three.
  4. reticle_look({ action: "state" }) once at the end. If it comes back empty or hasCapabilities is false, the capabilities file init generated registered nothing, which is what its ℹ AGENT: finish the capabilities file line was telling you. Finish it yourself before you drive: open src/reticle-dev.*, register the app's store if it has one, and list the testids the flow you picked actually touches. That is a setup step rather than a decision: a few lines in a file init already wrote for you. Never report a clean install over an empty state read.

Assert what the app did, not what it shows. A UI that renders the value it just sent, rather than the value the server returned, passes every DOM-level check ever written:

until: { kind: "allOf", predicates: [
  { kind: "net", method: "POST", urlContains: "/api/refund", status: 200, bodyContains: "\"refunded\":1187.01" },
  { kind: "element", query: { testid: "receipt" } },
  { kind: "console", level: "error", absent: true },
]}

Then report what you drove, what it produced, and the file:line for anything broken.


License key

Enterprise features are unlocked by a signed key, verified locally with Ed25519 and no network call. If the user hands you one, at any point:

  1. Append RETICLE_LICENSE_KEY=<key> to the project's .env (create it if absent). The CLI loads a project-local .env on every invocation, so nothing else has to change.
  2. Make sure .env is in .gitignore. A license key in git is a leaked credential, so check before you write and add the line if it is missing.
  3. Confirm: npx @reticlehq/server license prints active, eval, missing or expired.

Never echo the key back in your reply, and never put it in a commit, a comment, or a feedback report. Details: curl https://docs.reticle.sh/enterprise.md.


Feedback is first-class

Reticle is built for agents. You are its user, and your experience is the only signal that decides what gets built or fixed, and it disappears when your context does. The moment Reticle breaks, lies, cannot see something you needed, costs too many calls, or gives you trouble during setup:

reticle_session { action: "feedback" } { kind, text }

kind is one of bug, gap, ambiguity, feature_request, improvement. One call, then carry on. If the tools are unreachable, file the same report from the shell:

bash
npx @reticlehq/server feedback --agent --kind <bug|gap|ambiguity|feature_request|improvement> "what happened"

Nothing is too minor. Report defects in Reticle; a bug you find in the app under test is Reticle working, and belongs in your answer to the user. Never put app source, secrets, or user data in a report.

Going deeper

Fetch the one page that answers the question rather than re-reading this file. Appending .md to any docs URL returns its source with no site chrome.

bash
curl https://docs.reticle.sh/llms.txt              # every page title and URL; read this first
curl https://docs.reticle.sh/frameworks.md         # per-framework SDK wiring
curl https://docs.reticle.sh/troubleshooting.md    # nothing connected, click did nothing, verdict unknown
curl https://docs.reticle.sh/agent-cheatsheet.md   # the verify loop on one screen
curl https://docs.reticle.sh/predicates.md         # every `until` predicate

© reticlehq, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in plugin of reticlehq/reticle.

  • SKILL.md
  • .claude-plugin/plugin.json
  • hooks/hooks.json
  • hooks/init-nudge.mjs
  • icon.svg

Open the folder on GitHubat commit 178e5c0

Compare with similar skills

Reticle Runtime Verification next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reticle Runtime Verification compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reticle Runtime Verification this skillreticlehq/reticle1.2k—~2.8kAutomated safety check: NotesApache-2.0
Playwright E2E Testsonyx-dot-app/onyx32k1 repos~2.8kAutomated safety check: NotesCustom licence
Diff-Driven Smoke TestsSkyvern-AI/skyvern23k—~5.2kAutomated safety check: PassAGPL-3.0
E2E VerificationChorus-AIDLC/Chorus1.2k—~1.5kAutomated safety check: NotesAGPL-3.0
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone
Frontend Playwright E2Eansible/ansible-ui113—~2.5kAutomated safety check: NotesApache-2.0

Similar skills

  • Playwright E2E Tests

    onyx-dot-app/onyx

    Write and maintain Playwright end-to-end tests for the Onyx application.

    32k GitHub starsUsed in 1 repo~2.8k tokens
    Testing & QAAuto-check: notes
  • Diff-Driven Smoke Tests

    Skyvern-AI/skyvern

    Reads your git diff, writes a handful of happy-path browser smoke tests, runs them with Skyvern or Chrome DevTools MCP and posts screenshot evidence to the PR.

    23k GitHub stars~5.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • E2E Verification

    Chorus-AIDLC/Chorus

    A skill your agent uses when manually verifying a Chorus frontend change in a real browser — finding local login credentials, driving the running dev server with the Playwright MCP, logging in…

    1.2k GitHub stars~1.5k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Frontend Playwright E2E

    ansible/ansible-ui

    Write, run, and debug Playwright E2E / integration / live tests.

    113 GitHub stars~2.5k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • LangBot Testing

    langbot-app/LangBot

    Tests LangBot's WebUI and core flows through an automated browser and backend logs, with a routing table to reference guides per feature area.

    18k GitHub stars~1k tokensUpdated today
    Testing & QAAuto-check: notes

More from reticlehq/reticle

All 19 skills in this repo
  • Agentic TDD

    reticlehq/reticle

    Applies red-green TDD to behavior unit tests cannot reach, by stating the expected outcome against the running app with Reticle before writing the feature.

    1.2k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Whole-App Health Sweep

    reticlehq/reticle

    Sweeps a running web app by clicking every reachable control, then reports dead buttons, console errors, failed requests and mismatches between API data and the screen.

    1.2k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Broken UI Debugger

    reticlehq/reticle

    Finds why a running web app misbehaves when the console is empty and the code looks fine, by reading the click, request, store and console together.

    1.2k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Drives and verifies Electron or Tauri desktop apps through Reticle, which sees the renderer and the IPC calls that a browser-based testing tool cannot observe.

    1.2k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Finds out why a passing test suite sits on top of a broken app by comparing what the running app does with what the tests claim, using Reticle.

    1.2k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Fix What I Pointed At

    reticlehq/reticle

    Picks up bugs a person flagged by pointing at elements in the running app, each mark carrying the element, their note and the source file and line, then fixes and verifies them.

    1.2k GitHub stars~878 tokensUpdated yesterday
    Auto-check passed

Questions about Reticle Runtime Verification

What does Reticle Runtime Verification do?

Installs Reticle's dev-only SDK in a running web app and verifies user-facing changes by driving a real flow, returning a verdict with the file and line to fix. Reticle embeds a dev-only SDK inside the running web app and exposes it to the agent as reticle_* MCP tools, so the agent can look, act, observe and assert against the real DOM, network, routing, console and framework state instead of relying on screenshots. The plugin has already registered the MCP server, so no client restart is needed.

When should I use Reticle Runtime Verification?

Reticle Runtime Verification fits situations like: installing and setting up Reticle in a web project; proving a user-facing change works before calling it done; investigating a UI that is broken even though the tests pass; checking DOM, network, routing and console state of a running app without screenshots.

How do I install Reticle Runtime Verification in Claude Code?

Run `npx skills add reticlehq/reticle --skill reticle -a claude-code`. Or copy the skill folder (plugin in reticlehq/reticle) into .claude/skills/reticle in your project. Claude Code loads it when a task matches its description.

How do I install Reticle Runtime Verification in Codex?

Run `npx skills add reticlehq/reticle --skill reticle -a codex`. Or copy the skill folder (plugin in reticlehq/reticle) into .agents/skills/reticle in your project. Codex loads it when a task matches its description.

Can I use Reticle Runtime Verification in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add reticlehq/reticle --skill reticle -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/reticle, .gemini/skills/reticle, .github/skills/reticle and .opencode/skills/reticle in your project.

What does Reticle Runtime Verification need to run?

Going by SKILL.md and its folder, Reticle Runtime Verification needs JavaScript for the scripts in its folder, the command-line tools its instructions call (curl and npx) and credentials named RETICLE_LICENSE_KEY. Our summary lists: A web app with a recognizable dev script in package.json; npx, to run the Reticle server init command; The Reticle plugin with its MCP server registered.

Does Reticle Runtime Verification access the network?

SKILL.md names 1 domain. In commands or code: docs.reticle.sh; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Reticle Runtime Verification safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Reticle Runtime Verification use?

Reticle Runtime Verification is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reticle Runtime Verification use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Reticle Runtime Verification?

Skills that share tags, products or a category with Reticle Runtime Verification: Playwright E2E Tests (onyx-dot-app/onyx, 32k stars), Diff-Driven Smoke Tests (Skyvern-AI/skyvern, 23k stars), E2E Verification (Chorus-AIDLC/Chorus, 1.2k stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reticle Runtime Verification?

reticlehq (a GitHub organization) maintains it in reticlehq/reticle, which has 1,199 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.

Source: reticlehq/reticle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.