Agent skill

Agentic TDD

by reticlehq in reticlehq/reticle

Applies red-green TDD to behavior unit tests cannot reach, by stating the expected outcome against the running app with Reticle before writing the feature.

Apache-2.0Auto-check passedTesting & QA

Install Agentic TDD

skills CLI
$ npx skills add reticlehq/reticle --skill agentic-tdd -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install reticlehq/reticle agentic-tdd --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/reticlehq/reticle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentic-tdd .claude/skills/agentic-tdd && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentic-tdd
GitHub stars
1.2k
Token cost
~1.2k tokens
SKILL.md length
583 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Applies red-green TDD to behavior unit tests cannot reach, by stating the expected outcome against the running app with Reticle before writing the feature.

  • Works in 4 steps: RED: write the expectation, watch it fail → GREEN: implement until the same call… → REFACTOR: the expectation is the safety… → …
  • Building a user-facing feature test-first against the real running app
  • SKILL.md covers Why this is TDD and not just…, 1. RED: write the expectation,…, 2. GREEN: implement until the… and 3. REFACTOR: the expectation…, plus 3 more sections
  • Calls curl and npx; reaches docs.reticle.sh

What it does

Unit tests cannot say that clicking a Deploy button posts to an endpoint, updates the store and shows a banner, because that outcome exists only when the whole app runs. This skill runs the test-first discipline one level up, using Reticle to drive the real app. The call `reticle_act_and_wait` takes an `until` argument, so the expected consequence is named before the action runs, which stops the agent from fitting the expectation to whatever the code happened to do.

In the red step the agent writes the expectation and must see a `verified: no` result; `unknown` is not red because it gives no signal, and an immediate yes means the expectation is too loose. In the green step it implements the smallest change and re-runs the same call unchanged, explaining aloud if the assertion truly had to change. Reticle is installed with an npx init command followed by an install-and-verify skill.

When your agent uses it

  • Building a user-facing feature test-first against the real running app
  • Doing TDD on UI or full-stack work where a unit test cannot express the outcome
  • Replacing mocks with a red-green loop that exercises the real app

Example prompts

  • “Build the Deploy button flow test-first; the click should post to the deploy endpoint and show a banner.”
  • “Write the failing expectation for the new checkout confirmation before implementing it.”
  • “Use Reticle to run a red-green loop on the settings form instead of mocks.”

Requirements

  • Reticle installed in the project
  • A running instance of the app under development

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. RED: write the expectation, watch it fail
  2. GREEN: implement until the same call passes
  3. REFACTOR: the expectation is the safety net
  4. Keep the loop for the next change

What it can do on your machine

Read from SKILL.md and the folder at commit 178e5c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • docs.reticle.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentic TDD loads about 1.2k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 583 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from reticlehq/reticle at commit 178e5c0, republished under its Apache-2.0 licence (© reticlehq). 583 words, ~1,241 tokens.

Download SKILL.mdSave it as .claude/skills/agentic-tdd/SKILL.md (or your agent's skills folder).
name
agentic-tdd
description
Test-driven development for behaviour a unit test cannot reach, by writing the expectation against the running app before writing the code. Declare the consequence first, watch it fail, implement, watch it pass. Use when building a user-facing feature, when the user asks for TDD on UI or full-stack work, when a unit test cannot express the outcome that matters, or when you want a red-green loop that runs against the real app instead of mocks.
license
Apache-2.0
metadata.version
3.7.0
metadata.homepage
https://www.reticle.sh
metadata.repository
https://github.com/reticlehq/reticle

Red, green, refactor: against the running app

Unit tests drive the units. They cannot say "clicking Deploy posts to /api/deploy, moves the store to deploying, and shows the banner": that outcome only exists when the whole app runs. So the loop stalls exactly where the interesting bugs are, and the agent falls back to writing code and hoping.

This runs the same discipline one level up, using Reticle to drive the real app. Not installed? RETICLE_INSTALL_SOURCE=npx_skill npx @reticlehq/server@latest init, then the install-and-verify skill.

Why this is TDD and not just testing afterwards

The whole value of test-first is that the oracle is written while you still do not know the answer. An expectation written after seeing the result can always be adjusted into agreeing with whatever happened, and an agent is especially good at that adjustment. It will find a reading of the output under which the code it just wrote is correct.

reticle_act_and_wait({ ref, action, until }) enforces the order structurally: until is an argument to the action, so the consequence is named before the action runs. That is the red-green loop, made unfakeable.

1. RED: write the expectation, watch it fail

Before you write the feature, state what the app must do:

reticle_act_and_wait({ sessionId, ref, action: "click", until: { kind: "allOf", predicates: [
  { kind: "net",     method: "POST", urlContains: "/api/deploy", status: 200 },
  { kind: "signal",  name: "deploy:started" },
  { kind: "element", query: { testid: "deploy-banner" } },
  { kind: "console", level: "error", absent: true },
]}})

You want verified: "no" here. A red you did not see is a test you cannot trust. If this comes back yes before you have written anything, the expectation is not specific enough to the change. Tighten it until it fails for the right reason.

verified: "unknown" is not a red. It means Reticle could not tell, so the loop has no signal at all. Fix that before writing code, usually by naming a consequence the app can actually produce.

2. GREEN: implement until the same call passes

Write the smallest change that makes it hold, then re-run the same call, unchanged. That last word is the discipline: editing the predicate to match what you built converts TDD into narration. If the assertion has to change, say out loud why the original expectation was wrong.

Prefer re-asserting over re-driving when the verdict was unknown / unsettled: reticle_assert({ predicate, since, timeout_ms: 8000 }). Re-driving repeats a side effect that already happened.

Show full SKILL.md (222 more words)Show less

3. REFACTOR: the expectation is the safety net

Now change the implementation freely and re-run. Predicates are bound to behaviour (a request, a signal, a state path), not to markup, so a refactor that preserves behaviour stays green while a DOM-shaped test would go red for no reason.

4. Keep the loop for the next change

A journey worth writing test-first is a journey worth re-running forever. Save it once:

reticle_run({ tool: "reticle_flow_save", sessionId, args: { flowName: "deploy" } })
reticle_run({ tool: "reticle_verify", sessionId, args: { action: "flows" } })   // every saved flow, no model per flow

That turns the red-green loop into a regression suite you never hand-wrote.

What to assert on, in order of strength

  1. A signal the app fires itself ({ kind: "signal" }): the app declaring success in its own vocabulary. Strongest available.
  2. State (reticle_look { action: "state" }): what the app believes. Catches a UI that moved while the store did not.
  3. Network: the request, method and status. Catches a mock standing in for the real thing.
  4. An element appearing: necessary, never sufficient. Anything can render.
  5. Absence of console errors: always include it, never rely on it alone. Absence-only predicates pass on a control wired to nothing.

Honesty

Never weaken a check to turn a verdict green. In this loop that is not a small sin: it is the loop running backwards, and it produces a green suite over a feature that does not work.


Predicate reference: curl https://docs.reticle.sh/predicates.md. Everything else: curl https://docs.reticle.sh/llms.txt.

© reticlehq, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agentic-tdd of reticlehq/reticle.

Open the folder on GitHubat commit 178e5c0

Compare with similar skills

Agentic TDD next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentic TDD compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentic TDD this skillreticlehq/reticle1.2k—~1.2kAutomated safety check: PassApache-2.0
Playwright E2E Testsonyx-dot-app/onyx32k1 repos~2.8kAutomated safety check: NotesCustom licence
Diff-Driven Smoke TestsSkyvern-AI/skyvern23k—~5.2kAutomated safety check: PassAGPL-3.0
E2E VerificationChorus-AIDLC/Chorus1.2k—~1.5kAutomated safety check: NotesAGPL-3.0
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone
Frontend Playwright E2Eansible/ansible-ui113—~2.5kAutomated safety check: NotesApache-2.0

Similar skills

  • Playwright E2E Tests

    onyx-dot-app/onyx

    Write and maintain Playwright end-to-end tests for the Onyx application.

    32k GitHub starsUsed in 1 repo~2.8k tokens
    Testing & QAAuto-check: notes
  • Diff-Driven Smoke Tests

    Skyvern-AI/skyvern

    Reads your git diff, writes a handful of happy-path browser smoke tests, runs them with Skyvern or Chrome DevTools MCP and posts screenshot evidence to the PR.

    23k GitHub stars~5.2k tokensUpdated today
    Testing & QAAuto-check passed
  • E2E Verification

    Chorus-AIDLC/Chorus

    A skill your agent uses when manually verifying a Chorus frontend change in a real browser — finding local login credentials, driving the running dev server with the Playwright MCP, logging in…

    1.2k GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Frontend Playwright E2E

    ansible/ansible-ui

    Write, run, and debug Playwright E2E / integration / live tests.

    113 GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • LangBot Testing

    langbot-app/LangBot

    Tests LangBot's WebUI and core flows through an automated browser and backend logs, with a routing table to reference guides per feature area.

    18k GitHub stars~1k tokensUpdated yesterday
    Testing & QAAuto-check: notes

More from reticlehq/reticle

All 19 skills in this repo
  • Whole-App Health Sweep

    reticlehq/reticle

    Sweeps a running web app by clicking every reachable control, then reports dead buttons, console errors, failed requests and mismatches between API data and the screen.

    1.2k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Broken UI Debugger

    reticlehq/reticle

    Finds why a running web app misbehaves when the console is empty and the code looks fine, by reading the click, request, store and console together.

    1.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Drives and verifies Electron or Tauri desktop apps through Reticle, which sees the renderer and the IPC calls that a browser-based testing tool cannot observe.

    1.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Finds out why a passing test suite sits on top of a broken app by comparing what the running app does with what the tests claim, using Reticle.

    1.2k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Fix What I Pointed At

    reticlehq/reticle

    Picks up bugs a person flagged by pointing at elements in the running app, each mark carrying the element, their note and the source file and line, then fixes and verifies them.

    1.2k GitHub stars~878 tokensUpdated today
    Auto-check passed
  • Reticle Flow Replay

    reticlehq/reticle

    Saves a user journey driven through the app as a deterministic regression check that replays with no model and no test code, using Reticle.

    1.2k GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Questions about Agentic TDD

What does Agentic TDD do?

Applies red-green TDD to behavior unit tests cannot reach, by stating the expected outcome against the running app with Reticle before writing the feature. Unit tests cannot say that clicking a Deploy button posts to an endpoint, updates the store and shows a banner, because that outcome exists only when the whole app runs. This skill runs the test-first discipline one level up, using Reticle to drive the real app.

When should I use Agentic TDD?

Agentic TDD fits situations like: building a user-facing feature test-first against the real running app; doing TDD on UI or full-stack work where a unit test cannot express the outcome; replacing mocks with a red-green loop that exercises the real app.

How do I install Agentic TDD in Claude Code?

Run `npx skills add reticlehq/reticle --skill agentic-tdd -a claude-code`. Or copy the skill folder (skills/agentic-tdd in reticlehq/reticle) into .claude/skills/agentic-tdd in your project. Claude Code loads it when a task matches its description.

How do I install Agentic TDD in Codex?

Run `npx skills add reticlehq/reticle --skill agentic-tdd -a codex`. Or copy the skill folder (skills/agentic-tdd in reticlehq/reticle) into .agents/skills/agentic-tdd in your project. Codex loads it when a task matches its description.

Can I use Agentic TDD in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add reticlehq/reticle --skill agentic-tdd -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentic-tdd, .gemini/skills/agentic-tdd, .github/skills/agentic-tdd and .opencode/skills/agentic-tdd in your project.

What does Agentic TDD need to run?

Going by SKILL.md and its folder, Agentic TDD needs the command-line tools its instructions call (curl and npx). Our summary lists: Reticle installed in the project; A running instance of the app under development.

Does Agentic TDD access the network?

SKILL.md names 1 domain. In commands or code: docs.reticle.sh; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Agentic TDD safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentic TDD use?

Agentic TDD is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentic TDD use?

About 1.2k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agentic TDD?

Skills that share tags, products or a category with Agentic TDD: Playwright E2E Tests (onyx-dot-app/onyx, 32k stars), Diff-Driven Smoke Tests (Skyvern-AI/skyvern, 23k stars), E2E Verification (Chorus-AIDLC/Chorus, 1.2k stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentic TDD?

reticlehq (a GitHub organization) maintains it in reticlehq/reticle, which has 1,197 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.

Source: reticlehq/reticle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.