Agent skill

Automated Test Planning

by testdouble in testdouble/han

Produce a standalone test plan by analyzing code for test coverage gaps and edge cases.

MITAuto-check passedTesting & QA

Install Automated Test Planning

skills CLI
$ npx skills add testdouble/han --skill automated-test-planning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install testdouble/han automated-test-planning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/testdouble/han.git skills-src && mkdir -p .claude/skills && cp -r skills-src/han-coding/skills/automated-test-planning .claude/skills/automated-test-planning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
automated-test-planning
GitHub stars
279
Token cost
~6.7k tokens
SKILL.md length
3,809 words
Files
4 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

Produce a standalone test plan by analyzing code for test coverage gaps and edge cases.

  • Works in 6 steps: Determine Scope → 5: Classify Size and Mode → Dispatch Testing Agents → …
  • You need to create
  • SKILL.md covers Operating Principles, Project Context, Step 1: Determine Scope and Step 1.5: Classify Size and Mode, plus 4 more sections
  • Runs Shell scripts from its folder; calls git and bash

What it does

Automated Test Planning is an agent skill from testdouble/han. Produce a standalone test plan by analyzing code for test coverage gaps and edge cases. Use when you need to create, generate, or draft a test plan for a branch, need to analyze test coverage, or need to identify what tests to write for specific files or directories. Does not produce a plain-language plan for a person to run tests by hand — use manual-test-planning for that. Does not write test code — use tdd to implement behavior test-first. Does not refine existing plans — use iterative-plan-review. Does not…

Its SKILL.md is about 6.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/template.md` and `scripts/detect-test-context.sh`).

It sits in Testing & QA, covering Test strategy, Test generation and Test-driven development. The repository describes itself as: Han: AI skills and agents for "Solo" product engineers and small teams. The licence is MIT.

When your agent uses it

  • You need to create
  • Draft a test plan for a branch
  • Need to analyze test coverage
  • Need to identify what tests to write for specific files

Example prompts

  • “/automated-test-planning”

Requirements

  • A Bash shell
  • Pre-approved tools (allowed-tools): Bash(git *), Bash(find *), Read, Grep, Glob, Agent, Bash(bash "${CLAUDE_PLUGIN_ROOT}/scripts/han-config-dir.sh")

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Determine Scope
  2. 5: Classify Size and Mode
  3. Dispatch Testing Agents
  4. Merge and Prioritize
  5. Generate Output
  6. Review the Output

What it can do on your machine

Read from SKILL.md and the folder at commit abba73a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(git *)
    • Bash(find *)
    • Read
    • Grep
    • Glob
    • Agent
    • Bash(bash "${CLAUDE_PLUGIN_ROOT}/scripts/han-config-dir.sh")

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Automated Test Planning loads about 6.7k tokens when it runs, and up to ~8.5k if it reads all its reference files. Until then it costs about 186 tokens; SKILL.md has 3,809 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~186
When it runs · the whole SKILL.md, loaded when a task matches
~6.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from testdouble/han at commit abba73a, republished under its MIT licence (© testdouble). 3,809 words, ~6,695 tokens.

Download SKILL.mdSave it as .claude/skills/automated-test-planning/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
automated-test-planning
description
Produce a standalone test plan by analyzing code for test coverage gaps and edge cases. Use when you need to create, generate, or draft a test plan for a branch, need to analyze test coverage, or need to identify what tests to write for specific files or directories. Does not produce a plain-language plan for a person to run tests by hand — use manual-test-planning for that. Does not write test code — use tdd to implement behavior test-first. Does not refine existing plans — use iterative-plan-review. Does not review code quality, security, or style — use code-review for full code review. Does not evaluate architectural testability or structural coupling — use architectural-analysis for architectural assessment.
allowed-tools
Bash(git *), Bash(find *), Read, Grep, Glob, Agent, Bash(bash "${CLAUDE_PLUGIN_ROOT}/scripts/han-config-dir.sh")
arguments
size
argument-hint
[size: small | medium | large | dynamic] [optional: file paths, directories, or description of what to test]

Operating Principles

  • Test behavior through the public API, never internals. Every recommended test verifies observable behavior at a public seam: the inputs a real caller supplies, the outputs and side effects they observe, and the interactions the unit has with the other objects and services it collaborates with. Do not recommend tests that reach into private methods, internal state, or implementation structure — those tests pin the how and break on every refactor. When two designs would produce the same observable behavior, the test must pass for both. If a behavior can only be observed by inspecting internals, that is a signal the behavior belongs at a different seam, not a license to test the internal; note it and move the recommendation to the public boundary that exposes it. This is the depth ceiling: cover the critical behaviors a caller depends on, and stop. Do not specify tests for every branch, every private helper, or every intermediate value.

  • YAGNI is a first-class operating principle for tests. Apply the evidence-based YAGNI rule from ../../references/yagni-rule.md. A test is worth recommending only when (a) the code under review commits to a behavior the test verifies and (b) the failure mode the test would catch is realistic for this codebase. Tests for code paths that don't exist yet, hypothetical adversaries the code doesn't face, hypothetical scaling problems the workload doesn't have, "completeness" with existing tests, or symmetry ("we have a test for create, so we should have one for delete") are YAGNI candidates and go to the Deferred Tests section with the trigger that would justify writing them. When many speculative low-level tests can be replaced by one durable behavioral test that catches the same realistic failure modes, recommend the single test instead. Every test is ongoing maintenance and a brittleness surface.

  • Dispatch in proportion to the question. The size band chosen in Step 1.5 caps the roster, and a narrow question gets one agent and a prose answer rather than a team and a full document. Never dispatch the full roster for a question a single agent can settle BECAUSE the dispatch overhead exceeds the work, and a reader who asked one question pays for a document they did not want.

Project Context

  • git installed: !which git 2>/dev/null || echo "not installed"
  • CLAUDE.md: !find . -maxdepth 1 -name "CLAUDE.md" -type f
  • project-discovery.md: !find . -maxdepth 3 -name "project-discovery.md" -type f
  • personal config directory: !bash "${CLAUDE_PLUGIN_ROOT}/scripts/han-config-dir.sh" 2>/dev/null || echo "$HOME/.claude"
  • project .han/config.md: !cat .han/config.md 2>/dev/null || echo ""

As your first action, use the Read tool on .han/config.md inside the personal config directory path above. A read that returns no file is no personal configuration: continue silently. When that file or the project .han/config.md probe supplies content, apply it per config-rule.md, which governs precedence between the two files, relative-path resolution, and what to do with a file that reads but cannot be used.

Step 1: Determine Scope

Resolve project config: read CLAUDE.md's ## Project Discovery section for test command (under ### Commands and Tests, not ### Frameworks and Tooling), language, and framework; fall back to project-discovery.md. Store found values for use in later steps.

Scope determination: Check git installed from Project Context. If empty or not installed, skip to Mode C below.

Run ${CLAUDE_SKILL_DIR}/scripts/detect-test-context.sh and parse its output. If git-available: false, skip to Mode C below.

Mode A: Full git context — git-available: true and the output contains a changed-files-start block with content.

  • If the user provided file paths, directories, or a description: use those as scope (do not go searching for plan files or try to locate plans)
  • Otherwise: use the changed files list from the script output as scope

Mode B: Uncommitted changes — git-available: true but output contains changed-files: none.

  • If the user provided scope: use it as-is
  • Otherwise: run git diff (unstaged changes), git diff --cached (staged changes), and git status --short (untracked files) to identify changed files; if any files are found, use those as scope
  • If no files found in any of those commands, fall through to Mode C

Mode C: No git / no changes found — git missing, not in a repo, or no changes detected in any state.

  • If the user provided file paths, directories, or a description of what to test: use those as-is
  • Otherwise: use Glob to discover source files in the current directory, excluding node_modules/, .git/, vendor/, dist/, build/, __pycache__/, lock files; present the discovered files and ask the user to confirm scope

Build a list of source files to analyze: expand directories to find source files; identify relevant source files from branch changes or project structure for descriptions.

Step 1.5: Classify Size and Mode

Default to small. Start the classification at small and only escalate to medium or large when the signals below clearly require it. When a signal is borderline, stay at the smaller band.

Classify from the user's own request first, and from Step 1's file list only when the request settles nothing. A narrow question asked on a branch with many changed files is still small BECAUSE Step 1 falls back to the whole changed-files list whenever the user named no scope, and that list describes the branch rather than the question.

  • Small (default) — the request names specific existing or proposed tests and asks whether they are worth writing, or asks one question a single agent can settle, or the scope from Step 1 is 1-3 files in a single subsystem. Defaults to focused mode.
  • Medium — 4-10 files, or one cross-cutting concern such as a single API contract, a schema migration, or a new permission check. Defaults to full mode.
  • Large — more than 10 files, multiple subsystems, architectural changes, or security or data implications. Defaults to full mode.

Size override. If $size is non-empty (the user passed small, medium, large, or dynamic as the first argument), use it: a band value is the size and skips the signal-based classification above, while dynamic forces the signal-based classification even when the config sets a default band. If $size is empty and either .han/config.md supplies a band via default-swarm-size (per ../../references/config-rule.md), use that band and skip the signal-based classification. Anything that is not one of the four accepted values is trailing context, not a size.

Bind $focus. Read the user's free-form argument string from the invocation (everything after the optional $size positional). If non-empty, bind $focus to that string verbatim; if empty, bind it to none provided. This binding is the description Step 2 passes to each agent prompt.

State the chosen band and mode in one line with the justification before dispatching anything (for example, Small: the request names two proposed contexts on one method, so focused mode, Medium: passed via $size, or Medium: from the project .han/config.md default-swarm-size, naming whichever file supplied it). Accept the user's override of the band, the mode, or both.

In focused mode, dispatch only han-core:test-engineer in Step 2, skip the conditional specialists, answer in prose per Step 4's focused-mode rules, and skip the two reviewers in Step 5. Step 3 runs in full either way. Its behavioral, prerequisite, and YAGNI sweeps decide what may be recommended at all, so skipping them would let a focused answer recommend a test that cannot be written without an out-of-scope production change. In full mode, run every step as written.

Step 2: Dispatch Testing Agents

In focused mode, dispatch item 1 alone and skip the conditional-dispatch section entirely.

Launch the testing agents in parallel using the Agent tool with run_in_background: true. Pass each agent the file list from Step 1. In Mode A or Mode B, include on branch {branch} in agent prompts if a branch name was detected by the script; in Mode C or when no branch was detected, omit the branch reference entirely. When $focus is anything other than none provided, include it in every agent prompt so they can focus their analysis; it is what the {any additional context from user arguments} placeholder below resolves to.

Always dispatch
  1. Launch han-core:test-engineer agent — prompt: "Analyze test coverage for the following files{on branch {branch} if applicable}: {file list}. Recommend tests only at the public API: the inputs a real caller supplies, the outputs and side effects they observe, and the interactions the unit has with the objects and services it collaborates with. Do not recommend tests that reach into private methods, internal state, or implementation structure — if two implementations would produce the same observable behavior, the test must pass for both. Cover the critical behaviors a caller depends on and stop; do not specify a test for every branch, private helper, or intermediate value. Apply the YAGNI rule from ../../references/yagni-rule.md — recommend a test only when the code commits to a behavior the test verifies AND the failure mode is realistic for this codebase. Symmetry, completeness, and hypothetical scaling are YAGNI; defer those to the Deferred Tests section with the trigger that would justify writing them. {any additional context from user arguments}"

  2. Launch han-core:edge-case-explorer agent — prompt: "Explore edge cases for the following files{on branch {branch} if applicable}: {file list}. Focus on inputs, integration points, and error paths observable through the public API and through interactions with collaborating objects and services — not internal state or private implementation. Apply the YAGNI rule from ../../references/yagni-rule.md — raise an edge case only when a real caller produces the input, the failure mode has plausible production trigger, or the case is critical-path correctness regardless of caller. Hypothetical adversaries the code doesn't face and symmetry-driven boundaries go to Dropped Edge Cases with the trigger that would justify revisiting. {any additional context from user arguments}"

Conditional dispatch

Skip this whole section in focused mode. Otherwise inspect the file list before launching, and skip any that do not apply.

  1. Launch han-core:concurrency-analyst agent — only if the file list touches threads, async/await, goroutines, actors, shared mutable state across requests, timers, locks, or message queues. Prompt: "Identify concurrency test gaps for the following files{on branch {branch} if applicable}: {file list}. Focus on race conditions, lock ordering, shared-resource contention, deadlock potential, and async error handling that should be covered by tests. {any additional context from user arguments}"

  2. Launch han-core:adversarial-security-analyst agent — only if the file list touches authentication, authorization, input validation, data isolation, session handling, crypto, file uploads, external API calls with secrets, or SQL/ORM query construction. Prompt: "Identify negative security tests that should exist for the following files{on branch {branch} if applicable}: {file list}. Focus on exploit paths that tests could catch before production — authorization bypass, injection, broken isolation, insecure defaults. Return test recommendations, not general threat modeling. {any additional context from user arguments}"

Extra agents named in the project config's ## Extra Agents list join this conditional-dispatch pool under the same file-list-signal selection, per ../../references/config-rule.md: dispatch one only when the file list carries a signal matching its stated specialty, and skip an entry that does not resolve to a dispatchable agent with a one-line note.

Wait for every dispatched agent to complete and collect full output for processing in Step 3.

Step 3: Merge and Prioritize

Combine findings from every dispatched agent into a unified, prioritized test plan:

  1. Classify findings —
    • han-core:test-engineer items (T1, T2, ...): map High to CRIT or HIGH depending on the code path (security, data integrity, auth = CRIT; business logic, error handling = HIGH), Medium to MED, Low to LOW.
    • han-core:edge-case-explorer items (EC1, EC2, ...): map directly — Critical to CRIT, High to HIGH, Medium to MED, Low to LOW.
    • han-core:concurrency-analyst items (C1, C2, ...) when dispatched: races on auth/billing/isolation = CRIT; realistic load contention, async error swallowing = HIGH; theoretical interleaving = MED.
    • han-core:adversarial-security-analyst items (SEC-NNN) when dispatched: every item lands at CRIT. Retain the SEC-### cross-reference in the unified item so the source is visible. CRIT ranks the finding's severity and says nothing about whose ticket the fix belongs to; the prerequisite sweep below still applies to every one of them.
  2. Apply the behavioral sweep — walk every surviving recommendation and confirm it verifies observable behavior at a public seam (caller-supplied inputs, observed outputs and side effects, interactions with collaborating objects and services). Rewrite any item that asserts on private methods, internal state, or implementation structure so it tests the same behavior through the public boundary; if no public seam exposes it, drop the item and note that the behavior should be observed at a higher level rather than pinned to internals. Collapse multiple low-level items that protect the same observable behavior into the single behavioral test that catches the same realistic failure modes.
  3. Apply the prerequisite sweep — walk every surviving recommendation and ask what has to be true before the test can be written at all. An item whose test approach needs a production-code change first, meaning an added order, a new validation, a changed return value, a schema column, or any other edit to shipped code, is not a test to write. It is a production change wearing a test's clothes, and folding it into the plan hands the implementer a code change nobody authorized. Move it out of the priority tiers into Blocked by a Production Change, recording the change it needs, the file that would carry it, and who else consumes that file, so a reader can see the blast radius and open a separate ticket. Run this sweep before IDs are assigned, and run it over security items too: an honest finding at CRIT is still out of scope if it cannot be tested without first changing shipped code.
  4. Assign unified IDs — sequential IDs: TP-001, TP-002, TP-003, etc. Include the original agent ID as a cross-reference (e.g., "TP-001 (from T3)", "TP-002 (from C1, concurrency)", "TP-003 (from SEC-001, security)").
  5. Order by priority — interleave items from every agent by priority: all CRIT items first, then HIGH, then MED, then LOW. Within each priority level, order by the agent's own ranking.
  6. Cap at 40 items — keep a maximum of 40 items total, prioritized by severity. Security items (SEC-derived) are exempt from the cap. If more than 40 non-security items exist, note how many were omitted and recommend running the skill again after addressing high-priority items.
  7. Apply the YAGNI sweep — walk every test recommendation that survived classification and apply ../../references/yagni-rule.md. Demote any test whose justification reduces to "completeness", "best practice", "for future flexibility", symmetry with another test, or hypothetical scaling/adversaries the change doesn't touch — these go to the Deferred Tests section with a Reason: YAGNI — {gate failure} and the trigger that would justify writing the test (a third real customer hits the edge case, the feature actually ships the path, a measured production failure occurs, etc.). When several recommended low-level tests can be replaced by one durable behavioral test that catches the same realistic failure modes, replace them with the single test and record the dropped low-level tests under Deferred with Reason: YAGNI — single behavioral test catches the same realistic failure modes.
Show full SKILL.md (1,364 more words)Show less

Step 4: Generate Output

Before generating, invoke han-communication:readability-guidance to source the shared readability standard into your context, then apply it as you write the plan's plain-language spine (Summary, What Needs Testing and Why, What Each Test Covers), holding the named audience: the engineer who will implement the tests. The frame governs how a fact is said, never whether a required fact appears — keep the file:line references, test levels, and TP-IDs the plan depends on.

In focused mode, do not use the template. Answer in prose: the verdict on what was asked, the reasoning that settles it, and any test worth writing, with the file:line references and test levels the engineer needs. Write no section that the question did not ask for BECAUSE a reader who asked one question should not have to search a nine-section document for its answer. Skip the rest of this step's template rules and go to Step 5.

In full mode, use the template at template.md for the output structure. The test plan leads with plain language and defers the implementation detail.

Writing rules:

Lead with behavior. These rules make the plan a human-readable overview first and an implementation outline second:

  1. Lead with the why. Write the Summary, What Needs Testing and Why, and What Each Test Covers in plain language before the ## Technical Reference region. Describe what needs testing and what would break without it, in functional terms a reader who has not seen the code can follow. Name files, types, and test levels only where they aid understanding.
  2. Summary is prose plus bullets. Open with a 2-4 sentence plain-language paragraph: what was analyzed, the overall state of coverage, where the biggest risk sits, and what to test first. No file paths, no TP-IDs, no framework jargon in the paragraph. Then the scannable orienting bullets.
  3. What Needs Testing and Why groups the work into themes. Cover the 2-4 themes a reader can hold in their head (e.g. "Authorization", "Payment edge cases", "Concurrent writes"), not every item. For each, explain in everyday terms what behavior needs coverage and why it matters — what could break, who is affected. End each theme by naming the test IDs that fall under it so the reader can jump to detail.
  4. What Each Test Covers narrates the tests in plain language. Walk the tests in priority order. Lead each line with its test ID so it cross-links to the Technical Reference, then state what behavior it protects and what would break untested — not how to write it. Summarize a long low-priority tail in one line rather than listing each.
  5. Technical Reference is supporting detail. Place the per-item Test Plan (with test level, code paths, test approach, priority justification), Deferred Tests, Dropped Edge Cases, Coverage Summary counts, and Scope under the ## Technical Reference region, below the plain-language spine. A reader who only needs to know what to test and why can stop above.

Fill in all sections:

  1. Summary — Plain-language paragraph plus orienting bullets (scope, coverage health, most significant gap, where to start, and how many items the prerequisite sweep blocked). This is the qualitative coverage assessment, promoted to the top and written for a non-author. The blocked count belongs in the Summary because a caveat buried under the Technical Reference loses every attention contest to a CRIT label above it.
  2. What Needs Testing and Why — The themes, in plain language, each ending with the test IDs it covers.
  3. What Each Test Covers — Every meaningful test as a plain-language line led by its TP-ID.
  4. Technical Reference → Test Plan — All items from Step 3, grouped by priority tier. Every test plan item should include file:line references from the agent output, and preserve agent detail — carry through the test approach, code paths, and risk assessments into the merged output.
  5. Technical Reference → Deferred Tests — Items the han-core:test-engineer excluded due to brittleness risk.
  6. Technical Reference → Dropped Edge Cases — Items the han-core:edge-case-explorer intentionally excluded.
  7. Technical Reference → Blocked by a Production Change — Items the prerequisite sweep moved out of the priority tiers, each with the production change it would need, the file that would carry it, and that file's other consumers. Write each one as separate work to be ticketed, never as a test this plan asks anyone to write.
  8. Technical Reference → Coverage Summary — Counts by priority tier.
  9. Technical Reference → Scope — Scope type, file count, branch, language, test framework, and the file list.

Step 5: Review the Output

In focused mode, skip the two reviewers below and go straight to the readability editor and the self-check that follow them. The two reviewers audit a document's structure and its plain-language layer, and a focused answer has neither BECAUSE it is a few paragraphs rather than a layered document.

In full mode, dispatch two reviewers against the generated test plan, in parallel, using the Agent tool. The plan is produced in-channel, so embed the full plan text in each agent's prompt (in place of {plan text} below). If the plan was written to a file, pass that path instead and let the agent read it.

  1. Launch han-core:information-architect agent — prompt: "Audit the following test plan for findability, orientation, and comprehension. The intended audience is an engineer or technically-literate stakeholder who needs to understand what needs testing and why before reading the implementation detail. Check: (1) Do the plain-language Summary, What Needs Testing and Why, and What Each Test Covers sections appear before the ## Technical Reference region? If the plain-language content is missing or sits below the reference detail, that is a finding. (2) Does the Summary paragraph read for someone who has not seen the code — no file paths, TP-IDs, or framework jargon? (3) Does the heading list let a scanning reader see where the plain-language overview ends and the deep Technical Reference begins? (4) Do the What Each Test Covers lines cross-link to their TP-IDs so a reader can move from plain language to detail? Return a list of structural edits; do not return an empty list unless the plan leads with plain language and defers the implementation detail.\n\n{plan text}"

  2. Launch han-core:junior-developer agent — prompt: "Review the following test plan as a generalist teammate who has not seen this code. Reading only the plain-language sections (Summary, What Needs Testing and Why, What Each Test Covers), can you understand what needs testing and why it matters without dropping into the Technical Reference? Flag any plain-language line that is unclear, leans on undefined jargon, assumes context a reader would not have, or states a test without conveying what would break if it were missing. Return a list of clarity findings with the section and line they apply to; return an empty list only if the plain-language layer stands on its own.\n\n{plan text}"

Apply every actionable edit the agents return. For findings that require author judgment (scope or audience ambiguity), surface them to the user with a recommended resolution; do not silently resolve.

Once every actionable edit is applied and the plan is final, dispatch han-communication:readability-editor (one Agent call) to audit and rewrite the plan's prose against the readability standard. If the plan was written to a file, pass the editor that file path; if it is in-channel, pass the plan text and apply the returned rewrite. Also pass the named audience: the engineer who will implement the tests; the editor reads han-communication's own canonical rule, so pass no rule path. It must preserve every fact and operate on prose regions only — never inside code fences, tables, or the TP-NNN test identifiers, which must survive unchanged so they still resolve. Apply its rewrite to the plan.

Then run the standardized readability self-check (the shared standard is in your context from han-communication:readability-guidance) over the plan's prose regions only — never inside code fences, tables, or the TP-NNN identifiers. Confirm each criterion and fix any failure before presenting:

Run the readability rule's standardized self-check, which is already in your context from the readability-guidance invocation above. Correct every failure before presenting. Its fidelity criterion is not optional: the standard governs how the content is said, and drops a required fact only when the reader asked for less and losing it would not change what they do next.

© testdouble, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in han-coding/skills/automated-test-planning of testdouble/han.

  • SKILL.md
  • references/template.md
  • scripts/detect-test-context.bats
  • scripts/detect-test-context.sh

Open the folder on GitHubat commit abba73a

Compare with similar skills

Automated Test Planning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Automated Test Planning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Automated Test Planning this skilltestdouble/han279—~6.7kAutomated safety check: PassMIT
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Test Experteinverne/dotfiles121—~2.3kAutomated safety check: PassGPL-3.0
Prd V07 Test Planningmattgierhart/PRD-driven-context-engineering180—~3.5kAutomated safety check: NotesMIT
Test Reviewsd0xdev/sd0x-harness192—~2.9kAutomated safety check: PassMIT
Risk Based Testingpetrkindlmann/qa-skills165—~5.3kAutomated safety check: PassMIT

Similar skills

  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Test Expert

    einverne/dotfiles

    Testing methodologies, test-driven development (TDD), unit and integration testing, and testing best practices across multiple frameworks.

    121 GitHub stars~2.3k tokensUpdated 29 days ago
    Testing & QAAuto-check passed
  • Prd V07 Test Planning

    mattgierhart/PRD-driven-context-engineering

    Define test cases BEFORE implementation, ensuring every API, business rule, and user journey has verifiable acceptance criteria during PRD v0.7 Build Execution.

    180 GitHub stars~3.5k tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes
  • Test Review

    sd0xdev/sd0x-harness

    Test coverage review via Codex exec. An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Risk Based Testing

    petrkindlmann/qa-skills

    Produce a risk matrix or heatmap that quantifies what could break by business impact × probability, runs failure mode analysis on the top items, and maps test coverage to risk zones.

    165 GitHub stars~5.3k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Test Planning

    petrkindlmann/qa-skills

    Build a single sprint or release test plan. An agent skill from petrkindlmann/qa-skills.

    165 GitHub stars~4.2k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed

More from testdouble/han

All 54 skills in this repo
  • HTML Summary

    testdouble/han

    Convert a stakeholder summary markdown file into a single self-contained HTML executive report — bottom line and decision asks up front, supporting detail later — styled with a Test Double-derived…

    279 GitHub stars~2.9k tokensUpdated 7 days ago
    Auto-check passed
  • Update Han plugin documentation so every skill, agent, guidance doc, index, and cross-reference is current and accurate.

    279 GitHub stars~3.4k tokensUpdated 7 days ago
    Auto-check passed
  • Guidance

    testdouble/han

    Authoritative guidance for building Claude Code skills, agents, and plugins, plus init and update steps that install and refresh the plugin-building skills in the current repository.

    279 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Han Release

    testdouble/han

    Cut a Han release: update CHANGELOG.md with the changes since the last release, bump and tag every plugin that changed as {plugin-name}--v{version} so a version-constrained dependency can resolve…

    279 GitHub stars~8.6k tokensUpdated 7 days ago
    Auto-check passed
  • Update PR Description

    testdouble/han

    Generate a PR description from the current branch's changes against a GitHub PR, using the gh CLI.

    279 GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check passed
  • Plan Implementation

    testdouble/han

    Builds a feature implementation plan from an existing feature specification (or equivalent context) through a facilitated team conversation.

    279 GitHub stars~9.5k tokensUpdated 7 days ago
    Auto-check passed

Categories

Questions about Automated Test Planning

What does Automated Test Planning do?

Produce a standalone test plan by analyzing code for test coverage gaps and edge cases. Automated Test Planning is an agent skill from testdouble/han. Produce a standalone test plan by analyzing code for test coverage gaps and edge cases.

When should I use Automated Test Planning?

Automated Test Planning fits situations like: you need to create; draft a test plan for a branch; need to analyze test coverage; need to identify what tests to write for specific files.

How do I install Automated Test Planning in Claude Code?

Run `npx skills add testdouble/han --skill automated-test-planning -a claude-code`. Or copy the skill folder (han-coding/skills/automated-test-planning in testdouble/han) into .claude/skills/automated-test-planning in your project. Claude Code loads it when a task matches its description.

How do I install Automated Test Planning in Codex?

Run `npx skills add testdouble/han --skill automated-test-planning -a codex`. Or copy the skill folder (han-coding/skills/automated-test-planning in testdouble/han) into .agents/skills/automated-test-planning in your project. Codex loads it when a task matches its description.

Can I use Automated Test Planning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add testdouble/han --skill automated-test-planning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/automated-test-planning, .gemini/skills/automated-test-planning, .github/skills/automated-test-planning and .opencode/skills/automated-test-planning in your project.

What does Automated Test Planning need to run?

Going by SKILL.md and its folder, Automated Test Planning needs a shell for the scripts in its folder and the command-line tools its instructions call (git and bash). Our summary lists: A Bash shell. Its frontmatter pre-approves these tools: Bash(git *), Bash(find *), Read, Grep, Glob, Agent, Bash(bash "${CLAUDE_PLUGIN_ROOT}/scripts/han-config-dir.sh").

Does Automated Test Planning access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Automated Test Planning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Automated Test Planning use?

Automated Test Planning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Automated Test Planning use?

About 6.7k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Automated Test Planning?

Skills that share tags, products or a category with Automated Test Planning: Designing Tests (CloudAI-X/opencode-workflow, 275 stars), Test Expert (einverne/dotfiles, 121 stars), Prd V07 Test Planning (mattgierhart/PRD-driven-context-engineering, 180 stars) and Test Review (sd0xdev/sd0x-harness, 192 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Automated Test Planning?

testdouble (a GitHub organization) maintains it in testdouble/han, which has 279 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 1, 2026.

Source: testdouble/han on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.