Agent skill

Plan Tests

by genkovich in genkovich/sdd

A skill your agent uses to turn a feature's acceptance criteria into a test plan before any test is written — a table that maps every spec.md §5 acceptance criterion to at least one test, names the…

MITAuto-check passedTesting & QA

Install Plan Tests

skills CLI
$ npx skills add genkovich/sdd --skill plan-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install genkovich/sdd plan-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/genkovich/sdd.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/plan-tests .claude/skills/plan-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
plan-tests
GitHub stars
171
Token cost
~3.1k tokens
SKILL.md length
1,585 words
Files
2
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses to turn a feature's acceptance criteria into a test plan before any test is written — a table that maps every spec.md §5 acceptance criterion to at least one test, names the…

  • Works in 11 steps: Gate. test -f docs/features//spec.md →… → Pick the output target. Per the size… → Map levels — generic only. Name test… → …
  • Names the test levels (unit / integration / e2e / contract / load) without binding to a language
  • SKILL.md covers Owner, Inputs, Protocol and Definition of Done, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Plan Tests is an agent skill from genkovich/sdd. Use to turn a feature's acceptance criteria into a test plan before any test is written — a table that maps every spec.md §5 acceptance criterion to at least one test, names the test levels (unit / integration / e2e / contract / load) without binding to a language or framework, and fixes the integration and data strategy. Triggers on "plan tests for {slug}", "test plan for {slug}", "how do we test {slug}", "test strategy for {slug}", "/sdd:plan-tests {slug}", "план тестів для {slug}", "як тестувати {slug}"…

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `templates/test-plan.md`).

It sits in Testing & QA, covering Test generation, User stories and Test strategy. The repository describes itself as: Spec-Driven Development for Claude Code: 12 atomic Socratic skills + a TDD implement engine (agent-team & dynamic-workflow modes). The licence is MIT.

When your agent uses it

  • Names the test levels (unit / integration / e2e / contract / load) without binding to a language
  • Fixes the integration and data strategy
  • Plan tests for {slug}
  • Test plan for {slug}

Example prompts

  • “plan tests for {slug}”
  • “test plan for {slug}”
  • “how do we test {slug}”
  • “/plan-tests”

Workflow steps

11 steps, taken from the first numbered list in SKILL.md.

  1. Gate. test -f docs/features//spec.md → fail = refuse with the pointer above. Then read §5 (acceptance criteria — the rows of the coverage…
  2. Pick the output target. Per the size matrix: XS/S → write the plan inline in spec.md as a short ## Test plan section (a coverage table is…
  3. Map levels — generic only. Name test levels from a fixed vocabulary, never a tool or language: unit (pure logic — a rule, a calculation, a…
  4. Core mapping (the contract of this skill) — user chooses the level per AC. Build the AC→test table: every acceptance criterion in §5 maps…
  5. Edge cases & error paths. Every error/authorization acceptance criterion gets its own dedicated test row — never folded into the happy…
  6. Integration strategy — real, ephemeral dependency. For integration tests, the default is an ephemeral real dependency, e.g. a throwaway DB…
  7. NFR → load. For each §6 NFR that carries a number, write one concrete load scenario (target rate, duration, the metric and its threshold)…
  8. CI placement. Note which suites run where: fast suites (unit, contract) on every PR; the heavier ones (e2e, load) on a schedule or…
  9. Socratic walk + write. Walk the coverage table and the strategy choices with the 4-state actions from ../_shared/ask-style.md (Accept /…
  10. Structural self-check — per ../_shared/self-check.md: re-read the written plan from disk and verify 6 items: (1) every spec.md §5 AC id…
  11. Propose commit + handoff. test-plan: . Then emit the stage-handoff block per ../_shared/handoff.md — What I did (incl. «self-check: 6/6…

What it can do on your machine

Read from SKILL.md and the folder at commit 4403913. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Plan Tests loads about 3.1k tokens when it runs. Until then it costs about 179 tokens; SKILL.md has 1,585 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~179
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from genkovich/sdd at commit 4403913, republished under its MIT licence (© genkovich). 1,585 words, ~3,116 tokens.

Download SKILL.mdSave it as .claude/skills/plan-tests/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
plan-tests
description
Use to turn a feature's acceptance criteria into a test plan before any test is written — a table that maps every spec.md §5 acceptance criterion to at least one test, names the test levels (unit / integration / e2e / contract / load) without binding to a language or framework, and fixes the integration and data strategy. Triggers on "plan tests for {slug}", "test plan for {slug}", "how do we test {slug}", "test strategy for {slug}", "/sdd:plan-tests {slug}", "план тестів для {slug}", "як тестувати {slug}", "тест-план". Output: docs/features/{slug}/test-plan.md (separate file for M+), or inline in spec.md for XS/S per the size matrix. Hard-refuse if spec.md is missing → run `specify {slug}` first.
model
inherit
effort
medium

Skill: plan-tests

Turns an already-specified feature into a test plan: a table that ties every acceptance criterion in spec.md §5 to at least one named test, the levels those tests live at (unit / integration / e2e / contract / load), the integration strategy (a real dependency, spun up throwaway), and the test-data + cleanup approach. The plan is written before a single test exists — the next stage, implement, reads this map and writes the red tests against it, not "however it seems". This file is the spine; the output scaffold lives in templates/test-plan.md.

This skill keeps only its own machinery. Question phrasing is shared → ../_shared/ask-style.md. Depth (inline in the spec vs a separate file) follows the size matrix → ../_shared/size-matrix.md. It names test levels, never test tools — the concrete commands are detected by implement against the repo, not hard-coded here.

Plan prose follows artifact_language — the ## Test plan heading (parsed downstream), test names and level tokens stay English → ../_shared/artifact-language.md.

Owner

QA + the engineer who will implement the feature (co-authors). QA drives the level breakdown and the edge/error cases; the implementing engineer confirms each acceptance criterion has a reachable test and that the integration strategy fits the repo. The Tech Lead signs off that no acceptance criterion is left uncovered.

Inputs

  • <slug> — the same feature slug every earlier stage used.
  • Gate (hard-refuse if missing): docs/features/<slug>/spec.md. Its §5 acceptance criteria are the entire reason this plan exists — each one must map to a test. If spec.md is absent → STOP and point: «run specify <slug> first — the test plan maps its §5 acceptance criteria to tests».
  • (Optional) docs/features/<slug>/data-model.md — the entity shapes tell you what test data to build and what to seed/clean per suite. Read it if present.
  • (Optional) docs/features/<slug>/sad.md §6 sequence diagrams — each drawn flow is an e2e candidate; each cross-participant boundary is a contract-test candidate.
  • (Optional) docs/features/<slug>/.size — depth hint. Absent → default to M (separate test-plan.md file) and say so loudly in the handoff — «size M (default — no .size; run /sdd:classify-size <slug>)».

Protocol

  1. Gate. test -f docs/features/<slug>/spec.md → fail = refuse with the pointer above. Then read §5 (acceptance criteria — the rows of the coverage table) and §6 (NFRs — which drive load tests). Read data-model.md / sad.md §6 if present.
  2. Pick the output target. Per the size matrix: XS/S → write the plan inline in spec.md as a short ## Test plan section (a coverage table is enough — no separate file); M+ → a separate docs/features/<slug>/test-plan.md from the template. Confirm the target with one AskUserQuestion (phrasing per ../_shared/ask-style.md) when .size is absent.
  3. Map levels — generic only. Name test levels from a fixed vocabulary, never a tool or language: unit (pure logic — a rule, a calculation, a validator, no I/O), integration (the module against a real dependency it owns — DB, cache, queue), e2e (a full flow end to end, one per critical user story), contract (a boundary between two participants — an API shape or an event schema agreed by both sides), load (only when an NFR carries a number — throughput, p95 latency). When sad.md frontmatter target_surfaces declares a UI surface (web-frontend / mobile-app / desktop-app), add the frontend tiers — component (a UI component exercised in isolation), visual-regression (web — the rendered UI diffed against a baseline), e2e-through-UI (the flow driven through the real UI, not just the API). These are the "testing trophy", the dominant frontend testing vocabulary (web.dev / Kent C. Dodds) — a vocabulary, not a mandate (→ ../_shared/surfaces.md). When the design pipeline ran, source the UI rows from its artifacts: the e2e-through-UI path comes from ux-flows.md (the flow is the test's script) and the states a component test must cover come from screens.md (the manifest's state rows). Do not write tool names (no specific test runner, broker, visual-regression, or load tool) — implement detects what the repo already uses (e.g. Playwright / Storybook / a visual-diff tool).
  4. Core mapping (the contract of this skill) — user chooses the level per AC. Build the AC→test table: every acceptance criterion in §5 maps to ≥1 test. For each AC, propose a default level from a heuristic (pure logic/rule/validator → unit; behaviour against a real dependency the module owns → integration; a full user-story flow → e2e; a cross-participant API/event shape → contract; and — when a UI surface is declared — a UI piece → component, a user-facing flow → e2e-through-UI), then confirm the level(s) with the user via one AskUserQuestion (multiSelect — an AC may fan out to several levels, e.g. unit for the rule + e2e for the flow), phrased per ../_shared/ask-style.md. The user's choice is authoritative and is recorded in the table's Level column; implement reads it to write the test at the right level (it does not re-decide). A criterion with zero tests is the cardinal anti-pattern. Name each test descriptively from the criterion's intent (e.g. over-quota request is rejected), not from any framework convention.
  5. Edge cases & error paths. Every error/authorization acceptance criterion gets its own dedicated test row — never folded into the happy path. List the boundary and failure cases the spec implies (missing identifier, malformed input, dependency unavailable → the spec's fallback behaviour) as explicit rows with their expected outcome named in plain words (no status numbers, no error-code strings).
  6. Integration strategy — real, ephemeral dependency. For integration tests, the default is an ephemeral real dependency, e.g. a throwaway DB container spun up for the suite and torn down after (testcontainers-style). Mocking the datastore is an anti-pattern — a passing mock is not a passing production. State the seed strategy (factories/fixtures for the data shape) and the cleanup boundary (per-test vs per-suite); without cleanup the suite goes flaky and blocks CI.
  7. NFR → load. For each §6 NFR that carries a number, write one concrete load scenario (target rate, duration, the metric and its threshold) and name the tool generically: the load tool already in your repo, or e.g. k6 or Locust. If no NFR carries a number, mark the load section <!-- N/A: no numeric NFR --> — do not invent a load test.
  8. CI placement. Note which suites run where: fast suites (unit, contract) on every PR; the heavier ones (e2e, load) on a schedule or pre-release. The split is advice, not a pipeline config — implement and the repo's CI own the actual wiring.
  9. Socratic walk + write. Walk the coverage table and the strategy choices with the 4-state actions from ../_shared/ask-style.md (Accept / Fix / Save-as-OQ / Drop); on Fix, regenerate that one row (one round, second answer final). Maintain the edits-log per ../_shared/socratic-loop.md. On pass, write the plan to its target (separate file for M+, inline ## Test plan for XS/S).
  10. Structural self-check — per ../_shared/self-check.md: re-read the written plan from disk and verify 6 items: (1) every spec.md §5 AC id appears in the coverage table; (2) every error/authorization AC has its own dedicated row (not folded into a happy path); (3) every Level value ∈ the fixed vocabulary {unit, integration, e2e, contract, load, component, visual-regression, e2e-through-UI}; (4) zero tool names in the plan (no runner / broker / visual-diff / load-tool name outside the single "e.g. k6 / Locust" allowance); (5) the load section carries numbers (rate + duration + metric + threshold) or the literal <!-- N/A: no numeric NFR -->; (6) the plan sits at its size-correct target (inline ## Test plan for XS/S or route quick, separate test-plan.md for M+). Fix + re-check ≤2 cycles; surface anything unresolved.
  11. Propose commit + handoff. test-plan: <slug>. Then emit the stage-handoff block per ../_shared/handoff.md — What I did (incl. «self-check: 6/6 pass») + Review (test-plan.md, or spec.md ## Test plan for XS/S) + Run next (/clear, then /sdd:implement <slug>, which consumes this map to write the red tests).
Show full SKILL.md (368 more words)Show less

Definition of Done

  • The plan exists at its size-correct target: a separate docs/features/<slug>/test-plan.md for M+, or an inline ## Test plan section in spec.md for XS/S.
  • Every acceptance criterion in spec.md §5 maps to ≥1 named test — zero uncovered criteria.
  • Each error / authorization criterion has its own dedicated test row, not folded into a happy path.
  • Test levels are generic (unit / integration / e2e / contract / load; + component / visual-regression / e2e-through-UI when a UI surface is declared in target_surfaces) — no test-runner, broker, visual-regression, or load-tool name is hard-coded (the load tool is named only as "the one in your repo, or e.g. k6 / Locust"; UI tools are detected by implement).
  • Integration tests use an ephemeral real dependency (throwaway container), with the seed and cleanup boundary stated; no mocked datastore.
  • Every numeric §6 NFR has a load scenario (rate + duration + metric + threshold), or the load section is explicitly <!-- N/A -->.

Anti-patterns

  • An acceptance criterion with no test. The whole point of the map is that §5 is verifiable; an uncovered criterion means it isn't.
  • Naming a concrete tool or language — a specific runner, broker, or load tool. The legacy plan hard-coded k6; here load is "the tool already in your repo, or e.g. k6 / Locust", and the rest stay generic levels. implement detects the real commands.
  • Mocking the datastore. A passing mock is not a passing production — use a throwaway real dependency for integration.
  • e2e without a cleanup boundary. Leftover state makes the suite flaky and every flaky run blocks CI.
  • "100 % coverage" as the goal. The target is critical paths + happy + error paths mapped to acceptance criteria, not a line-count number.
  • A wishlist plan — "would be nice to add". A test plan is a commitment the next stage executes, not a backlog.
  • Inventing a load test with no numeric NFR. No number → <!-- N/A -->, not a fabricated throughput target.

References & template

  • ../_shared/ask-style.md — canonical question/option phrasing for steps 2 and 9.
  • ../_shared/self-check.md — the structural self-check contract step 10 runs.
  • ../_shared/size-matrix.md — inline-in-spec (XS/S) vs separate file (M+) depth.
  • ../_shared/surfaces.md — a declared UI surface adds the component / visual-regression / e2e-through-UI tiers (testing-trophy vocabulary); read from sad.md target_surfaces.
  • ./templates/test-plan.md — output scaffold: AC→test mapping table, generic test levels, ephemeral-dependency integration strategy, stack-agnostic load section. Its <!-- … --> comments are the per-section contract.

© genkovich, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/plan-tests of genkovich/sdd.

  • SKILL.md
  • templates/test-plan.md

Open the folder on GitHubat commit 4403913

Compare with similar skills

Plan Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Plan Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Plan Tests this skillgenkovich/sdd171—~3.1kAutomated safety check: PassMIT
Test Scenariosphuryn/pm-skills27k—~866Automated safety check: PassMIT
Test Automationrevfactory/harness-1001.3k—~1.7kAutomated safety check: PassApache-2.0
Senior QAnicepkg/auto-company1923 repos~1.1kAutomated safety check: NotesNone
Argent QA Flowsbbplayer-app/BBPlayer1.1k—~3.2kAutomated safety check: PassMIT
Senior QAalirezarezvani/claude-skills28k1 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Test Scenarios

    phuryn/pm-skills

    Create comprehensive test scenarios from user stories with test objectives, starting conditions, user roles, step-by-step actions, and expected outcomes.

    27k GitHub stars~866 tokensUpdated 23 days ago
    Testing & QAAuto-check passed
  • Test Automation

    revfactory/harness-100

    A full test automation pipeline. An agent skill from revfactory/harness-100.

    1.3k GitHub stars~1.7k tokensUpdated 6 mo ago
    Testing & QAAuto-check passed
  • Senior QA

    nicepkg/auto-company

    Comprehensive QA and testing skill for quality assurance, test automation, and testing strategies for ReactJS, NextJS, NodeJS applications.

    192 GitHub starsUsed in 3 repos~1.1k tokens
    Testing & QAAuto-check: notes
  • Argent QA Flows

    bbplayer-app/BBPlayer

    Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria.

    1.1k GitHub stars~3.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Senior QA

    alirezarezvani/claude-skills

    Generates unit tests, integration tests, and E2E tests for React/Next.js applications.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check passed
  • Testing Strategies

    ancoleman/ai-design-components

    Strategic guidance for choosing and implementing testing approaches across the test pyramid.

    526 GitHub stars~3.8k tokensUpdated 10 mo ago
    Testing & QAAuto-check passed

More from genkovich/sdd

All 21 skills in this repo
  • Fix

    genkovich/sdd

    A skill your agent uses to fix a reported bug spec-first: reproduce it, trace the symptom to the owning feature's acceptance criteria, pin it with a failing (RED) test, apply the minimal GREEN fix…

    171 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Implement

    genkovich/sdd

    A skill your agent uses to implement a feature from its tasks.json with test-driven development — writes a failing test first, makes it pass, refactors, gates, and commits per task.

    171 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Interview

    genkovich/sdd

    Use BEFORE roadmap or specify to get the idea OUT OF YOUR HEAD and onto disk — a Socratic interview that surfaces hidden assumptions, names tradeoffs, exposes imprecisions and proposes fresh angles…

    171 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Classify Size

    genkovich/sdd

    A skill your agent uses to classify a feature into XS/S/M/L/XL and write docs/features/{slug}/.size plus the pipeline route docs/features/{slug}/.route (quick|standard|full) so later skills know how…

    171 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Decide Adr

    genkovich/sdd

    A skill your agent uses to record a post-hoc or asynchronous architecture decision as a MADR ADR when it was NOT captured during the synchronous design pass — a choice made in code, in a chat, on a…

    171 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Tasks

    genkovich/sdd

    A skill your agent uses to break a designed feature into atomic, ≤1-day tasks with a dependency graph, a per-task Definition of Done, and a machine-readable tasks.json that the implement engine…

    171 GitHub stars~4.8k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Plan Tests

What does Plan Tests do?

A skill your agent uses to turn a feature's acceptance criteria into a test plan before any test is written — a table that maps every spec.md §5 acceptance criterion to at least one test, names the…. Plan Tests is an agent skill from genkovich/sdd.md §5 acceptance criterion to at least one test, names the test levels (unit / integration / e2e / contract / load) without binding to a language or framework, and fixes the integration and data strategy.

When should I use Plan Tests?

Plan Tests fits situations like: names the test levels (unit / integration / e2e / contract / load) without binding to a language; fixes the integration and data strategy; plan tests for {slug}; test plan for {slug}.

How do I install Plan Tests in Claude Code?

Run `npx skills add genkovich/sdd --skill plan-tests -a claude-code`. Or copy the skill folder (skills/plan-tests in genkovich/sdd) into .claude/skills/plan-tests in your project. Claude Code loads it when a task matches its description.

How do I install Plan Tests in Codex?

Run `npx skills add genkovich/sdd --skill plan-tests -a codex`. Or copy the skill folder (skills/plan-tests in genkovich/sdd) into .agents/skills/plan-tests in your project. Codex loads it when a task matches its description.

Can I use Plan Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add genkovich/sdd --skill plan-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/plan-tests, .gemini/skills/plan-tests, .github/skills/plan-tests and .opencode/skills/plan-tests in your project.

What does Plan Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Plan Tests is instructions for the agent only.

Does Plan Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Plan Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Plan Tests use?

Plan Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Plan Tests use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Plan Tests?

Skills that share tags, products or a category with Plan Tests: Test Scenarios (phuryn/pm-skills, 27k stars), Test Automation (revfactory/harness-100, 1.3k stars), Senior QA (nicepkg/auto-company, 192 stars) and Argent QA Flows (bbplayer-app/BBPlayer, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Plan Tests?

genkovich (a GitHub user) maintains it in genkovich/sdd, which has 171 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on September 5, 2026.

Source: genkovich/sdd on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.