Playwright E2E Tests
onyx-dot-app/onyx
Write and maintain Playwright end-to-end tests for the Onyx application.
Applies red-green TDD to behavior unit tests cannot reach, by stating the expected outcome against the running app with Reticle before writing the feature.
$ npx skills add reticlehq/reticle --skill agentic-tdd -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install reticlehq/reticle agentic-tdd --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentic-tdd .claude/skills/agentic-tdd && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentic-tdd" agent skill from https://github.com/reticlehq/reticle/tree/main/skills/agentic-tdd into .claude/skills/agentic-tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentic-tdd", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/reticlehq/reticle/tree/main/skills/agentic-tddType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add reticlehq/reticle --skill agentic-tdd -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install reticlehq/reticle agentic-tdd --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentic-tdd .agents/skills/agentic-tdd && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentic-tdd" agent skill from https://github.com/reticlehq/reticle/tree/main/skills/agentic-tdd into .agents/skills/agentic-tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentic-tdd", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add reticlehq/reticle --skill agentic-tdd -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install reticlehq/reticle agentic-tdd --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentic-tdd .cursor/skills/agentic-tdd && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentic-tdd" agent skill from https://github.com/reticlehq/reticle/tree/main/skills/agentic-tdd into .cursor/skills/agentic-tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentic-tdd", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/reticlehq/reticle.git --path skills/agentic-tdd--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add reticlehq/reticle --skill agentic-tdd -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install reticlehq/reticle agentic-tdd --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentic-tdd .gemini/skills/agentic-tdd && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentic-tdd" agent skill from https://github.com/reticlehq/reticle/tree/main/skills/agentic-tdd into .gemini/skills/agentic-tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentic-tdd", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install reticlehq/reticle agentic-tddInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add reticlehq/reticle --skill agentic-tdd -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentic-tdd .github/skills/agentic-tdd && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentic-tdd" agent skill from https://github.com/reticlehq/reticle/tree/main/skills/agentic-tdd into .github/skills/agentic-tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentic-tdd", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add reticlehq/reticle --skill agentic-tdd -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install reticlehq/reticle agentic-tdd --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentic-tdd .opencode/skills/agentic-tdd && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentic-tdd" agent skill from https://github.com/reticlehq/reticle/tree/main/skills/agentic-tdd into .opencode/skills/agentic-tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentic-tdd", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentic-tddApplies red-green TDD to behavior unit tests cannot reach, by stating the expected outcome against the running app with Reticle before writing the feature.
Unit tests cannot say that clicking a Deploy button posts to an endpoint, updates the store and shows a banner, because that outcome exists only when the whole app runs. This skill runs the test-first discipline one level up, using Reticle to drive the real app. The call `reticle_act_and_wait` takes an `until` argument, so the expected consequence is named before the action runs, which stops the agent from fitting the expectation to whatever the code happened to do.
In the red step the agent writes the expectation and must see a `verified: no` result; `unknown` is not red because it gives no signal, and an immediate yes means the expectation is too loose. In the green step it implements the smallest change and re-runs the same call unchanged, explaining aloud if the assertion truly had to change. Reticle is installed with an npx init command followed by an install-and-verify skill.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 178e5c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlnpxFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
docs.reticle.shFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentic TDD loads about 1.2k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 583 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from reticlehq/reticle at commit 178e5c0, republished under its Apache-2.0 licence (© reticlehq). 583 words, ~1,241 tokens.
.claude/skills/agentic-tdd/SKILL.md (or your agent's skills folder).Unit tests drive the units. They cannot say "clicking Deploy posts to /api/deploy, moves the store to deploying, and shows the banner": that outcome only exists when the whole app runs. So the loop stalls exactly where the interesting bugs are, and the agent falls back to writing code and hoping.
This runs the same discipline one level up, using Reticle to drive the real app. Not installed? RETICLE_INSTALL_SOURCE=npx_skill npx @reticlehq/server@latest init, then the install-and-verify skill.
The whole value of test-first is that the oracle is written while you still do not know the answer. An expectation written after seeing the result can always be adjusted into agreeing with whatever happened, and an agent is especially good at that adjustment. It will find a reading of the output under which the code it just wrote is correct.
reticle_act_and_wait({ ref, action, until }) enforces the order structurally: until is an argument to the action, so the consequence is named before the action runs. That is the red-green loop, made unfakeable.
Before you write the feature, state what the app must do:
reticle_act_and_wait({ sessionId, ref, action: "click", until: { kind: "allOf", predicates: [
{ kind: "net", method: "POST", urlContains: "/api/deploy", status: 200 },
{ kind: "signal", name: "deploy:started" },
{ kind: "element", query: { testid: "deploy-banner" } },
{ kind: "console", level: "error", absent: true },
]}})You want verified: "no" here. A red you did not see is a test you cannot trust. If this comes back yes before you have written anything, the expectation is not specific enough to the change. Tighten it until it fails for the right reason.
verified: "unknown" is not a red. It means Reticle could not tell, so the loop has no signal at all. Fix that before writing code, usually by naming a consequence the app can actually produce.
Write the smallest change that makes it hold, then re-run the same call, unchanged. That last word is the discipline: editing the predicate to match what you built converts TDD into narration. If the assertion has to change, say out loud why the original expectation was wrong.
Prefer re-asserting over re-driving when the verdict was unknown / unsettled: reticle_assert({ predicate, since, timeout_ms: 8000 }). Re-driving repeats a side effect that already happened.
Now change the implementation freely and re-run. Predicates are bound to behaviour (a request, a signal, a state path), not to markup, so a refactor that preserves behaviour stays green while a DOM-shaped test would go red for no reason.
A journey worth writing test-first is a journey worth re-running forever. Save it once:
reticle_run({ tool: "reticle_flow_save", sessionId, args: { flowName: "deploy" } })
reticle_run({ tool: "reticle_verify", sessionId, args: { action: "flows" } }) // every saved flow, no model per flowThat turns the red-green loop into a regression suite you never hand-wrote.
{ kind: "signal" }): the app declaring success in its own vocabulary. Strongest available.reticle_look { action: "state" }): what the app believes. Catches a UI that moved while the store did not.Never weaken a check to turn a verdict green. In this loop that is not a small sin: it is the loop running backwards, and it produces a green suite over a feature that does not work.
Predicate reference: curl https://docs.reticle.sh/predicates.md. Everything else: curl https://docs.reticle.sh/llms.txt.
© reticlehq, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/agentic-tdd of reticlehq/reticle.
Open the folder on GitHubat commit 178e5c0
Agentic TDD next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentic TDD this skillreticlehq/reticle | 1.2k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Playwright E2E Testsonyx-dot-app/onyx | 32k | 1 repos | ~2.8k | Automated safety check: Notes | Custom licence | |
| Diff-Driven Smoke TestsSkyvern-AI/skyvern | 23k | — | ~5.2k | Automated safety check: Pass | AGPL-3.0 | |
| E2E VerificationChorus-AIDLC/Chorus | 1.2k | — | ~1.5k | Automated safety check: Notes | AGPL-3.0 | |
| Playwright Testingchongdashu/vibejam-starter-pack | 149 | — | ~2.2k | Automated safety check: Pass | None | |
| Frontend Playwright E2Eansible/ansible-ui | 113 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 |
onyx-dot-app/onyx
Write and maintain Playwright end-to-end tests for the Onyx application.
Skyvern-AI/skyvern
Reads your git diff, writes a handful of happy-path browser smoke tests, runs them with Skyvern or Chrome DevTools MCP and posts screenshot evidence to the PR.
Chorus-AIDLC/Chorus
A skill your agent uses when manually verifying a Chorus frontend change in a real browser — finding local login credentials, driving the running dev server with the Playwright MCP, logging in…
chongdashu/vibejam-starter-pack
Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.
ansible/ansible-ui
Write, run, and debug Playwright E2E / integration / live tests.
langbot-app/LangBot
Tests LangBot's WebUI and core flows through an automated browser and backend logs, with a routing table to reference guides per feature area.
reticlehq/reticle
Sweeps a running web app by clicking every reachable control, then reports dead buttons, console errors, failed requests and mismatches between API data and the screen.
reticlehq/reticle
Finds why a running web app misbehaves when the console is empty and the code looks fine, by reading the click, request, store and console together.
reticlehq/reticle
Drives and verifies Electron or Tauri desktop apps through Reticle, which sees the renderer and the IPC calls that a browser-based testing tool cannot observe.
reticlehq/reticle
Finds out why a passing test suite sits on top of a broken app by comparing what the running app does with what the tests claim, using Reticle.
reticlehq/reticle
Picks up bugs a person flagged by pointing at elements in the running app, each mark carrying the element, their note and the source file and line, then fixes and verifies them.
reticlehq/reticle
Saves a user journey driven through the app as a deterministic regression check that replays with no model and no test code, using Reticle.
Works with
Categories
Applies red-green TDD to behavior unit tests cannot reach, by stating the expected outcome against the running app with Reticle before writing the feature. Unit tests cannot say that clicking a Deploy button posts to an endpoint, updates the store and shows a banner, because that outcome exists only when the whole app runs. This skill runs the test-first discipline one level up, using Reticle to drive the real app.
Agentic TDD fits situations like: building a user-facing feature test-first against the real running app; doing TDD on UI or full-stack work where a unit test cannot express the outcome; replacing mocks with a red-green loop that exercises the real app.
Run `npx skills add reticlehq/reticle --skill agentic-tdd -a claude-code`. Or copy the skill folder (skills/agentic-tdd in reticlehq/reticle) into .claude/skills/agentic-tdd in your project. Claude Code loads it when a task matches its description.
Run `npx skills add reticlehq/reticle --skill agentic-tdd -a codex`. Or copy the skill folder (skills/agentic-tdd in reticlehq/reticle) into .agents/skills/agentic-tdd in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add reticlehq/reticle --skill agentic-tdd -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentic-tdd, .gemini/skills/agentic-tdd, .github/skills/agentic-tdd and .opencode/skills/agentic-tdd in your project.
Going by SKILL.md and its folder, Agentic TDD needs the command-line tools its instructions call (curl and npx). Our summary lists: Reticle installed in the project; A running instance of the app under development.
SKILL.md names 1 domain. In commands or code: docs.reticle.sh; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentic TDD is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.2k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agentic TDD: Playwright E2E Tests (onyx-dot-app/onyx, 32k stars), Diff-Driven Smoke Tests (Skyvern-AI/skyvern, 23k stars), E2E Verification (Chorus-AIDLC/Chorus, 1.2k stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
reticlehq (a GitHub organization) maintains it in reticlehq/reticle, which has 1,197 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.
Source: reticlehq/reticle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.