Agent skill

AI Test Generation

by petrkindlmann in petrkindlmann/qa-skills

Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

MITAuto-check passedTesting & QA

Install AI Test Generation

skills CLI
$ npx skills add petrkindlmann/qa-skills --skill ai-test-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install petrkindlmann/qa-skills ai-test-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-test-generation .claude/skills/ai-test-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-test-generation
GitHub stars
170
Token cost
~4.8k tokens
SKILL.md length
2,231 words
Files
2 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

  • Works in 7 steps: Extract Requirements and Entities → Risk Analysis and Invariants → Coverage Matrix → …
  • : generate tests from spec
  • SKILL.md covers Quick Route, Discovery Questions, Core Principles and The Pipeline, plus 7 more sections
  • Calls npx, git and python

What it does

AI Test Generation is an agent skill from petrkindlmann/qa-skills. Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs. Staged pipeline: requirements extraction → risk analysis → coverage matrix → scenario generation → oracle design → test code → human review, with guardrails against hallucinated APIs and weak assertions. Use when: "generate tests from spec," "tests from PRD," "tests from user story," "auto-generate test cases," "AI write tests for me." Not for: testing AI/LLM features in your product — use ai-system-testing…

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/prompt-patterns.md`).

It sits in Testing & QA, covering Test generation, User stories and PRD writing. It works with Playwright. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.

When your agent uses it

  • : generate tests from spec
  • Tests from user story
  • Auto-generate test cases
  • AI write tests for me. Not for: testing AI/LLM features in your product — use ai-system-testing

Example prompts

  • “generate tests from spec,”
  • “tests from PRD,”
  • “tests from user story,”
  • “/ai-test-generation”

Requirements

  • Node.js

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Extract Requirements and Entities
  2. Risk Analysis and Invariants
  3. Coverage Matrix
  4. Generate Candidate Scenarios
  5. Design Assertions and Oracles
  6. Generate Test Code
  7. Human Review

What it can do on your machine

Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • git
    • python
    • ruff
    • tsc

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Test Generation loads about 4.8k tokens when it runs, and up to ~8.8k if it reads all its reference files. Until then it costs about 190 tokens; SKILL.md has 2,231 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~190
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,231 words, ~4,821 tokens.

Download SKILL.mdSave it as .claude/skills/ai-test-generation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ai-test-generation
description
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs. Staged pipeline: requirements extraction → risk analysis → coverage matrix → scenario generation → oracle design → test code → human review, with guardrails against hallucinated APIs and weak assertions. Use when: "generate tests from spec," "tests from PRD," "tests from user story," "auto-generate test cases," "AI write tests for me." Not for: testing AI/LLM features in your product — use ai-system-testing. Not for: auditing a pre-existing test suite you did not just generate — use ai-qa-review (Step 7 here only reviews tests THIS pipeline produced). Related: playwright-automation, unit-testing, api-testing, qa-project-context.
license
MIT
metadata.author
kindlmann
metadata.version
2.0
metadata.category
ai-qa
<objective>
LLMs will happily emit fifty plausible-looking tests that assert nothing, target endpoints that do not exist, and duplicate each other. This skill is a staged pipeline that forces structured intermediates — assumptions, coverage matrix, oracle definitions — out of the model BEFORE any test code, so what you get is traceable, reviewable, and grounded in the real codebase instead of ad-hoc generated noise.

Before starting: Check for .agents/qa-project-context.md in the project root. It carries tech stack, test frameworks, naming conventions, selector strategy, and known risk areas that dramatically improve generated test quality. </objective>

Quick Route

The pipeline is the same for every input; only the Step 1 extraction emphasis changes. Jump to the matching row, then run Steps 2-7 unchanged.

Input typeStep 1 extractsWatch for
PRD / feature specEntities, business rules, acceptance criteria, NFRs, stated assumptionsImplicit requirements inferred from "seamless"/"fast" language
User story + ACEach AC → ≥1 happy + ≥1 negative scenarioACs that hide multiple behaviors in one line
Code diff (git diff main...HEAD)New/changed code paths, modified conditionals, removed behaviorRegression scope: test the changed paths, not the whole module
Bug reportRepro steps, expected vs actual, environmentWrite a test asserting expected — fails now, passes after fix
OpenAPI / GraphQL SDLEndpoints, schemas, required fields, enums, authValidation, auth-failure, and edge cases per endpoint, not just 200s

Playwright projects also pick an agent integration mode — see Discovery Q2.

Discovery Questions

Check .agents/qa-project-context.md first — if it exists, use it and skip anything already answered there. Then clarify:

  1. What is the input source? PRD / spec, user story + AC, code diff, bug report, or API schema. Determines Step 1 extraction emphasis (see Quick Route). For an LLM/AI feature spec, stop — generate eval datasets in ai-system-testing, not Playwright specs.

  2. What is the target test framework, and (for Playwright) which agent integration mode?

    • E2E: Playwright (preferred), Cypress. Unit: Jest, Vitest, pytest. API: Playwright APIRequestContext, Supertest, requests.
    • Playwright CLI + agents (recommended for Claude Code / Codex / Cursor): npx playwright init-agents --loop=claude scaffolds planner/generator/healer agents into .claude/agents/ as markdown. They are interactive dev tools that produce standard Playwright tests which run unchanged in CI. Token-efficient; runs inside the agent's loop.
    • Playwright MCP (npx @playwright/mcp@latest): higher overhead, right when the agent must drive a live browser interactively over a long session.
    • Neither — hand-write tests using AI as a scratch-pad helper.
  3. What project context is available? Existing test patterns, Page Objects / helpers, data factories / fixtures, CI constraints (timeout, parallelism). More context = less cleanup.

  4. What is the review workflow? Full pipeline → human review → merge (default); scenarios only → human writes code; or code → human refines iteratively.

  5. What domain knowledge is needed? Regulated industry (healthcare, finance) compliance, domain invariants (money never negative, appointments cannot overlap), known risk areas from past incidents.

Core Principles

  1. Pipeline before code. Never generate test code before establishing what to test, why, and how to verify it. The seven-step pipeline exists to prevent premature code generation that targets the wrong things.

  2. Structured intermediates are the product. The assumptions document, coverage matrix, and oracle definitions are more valuable than the test code itself. They are reviewable, traceable, and reusable.

  3. Separate what from how. Scenario generation (what to test) and oracle design (how to verify) are distinct cognitive tasks. Mixing them produces scenarios biased toward what is easy to assert, with assertions tacked on as afterthoughts.

  4. AI generates the first draft; a human reviews and refines. Never ship AI-generated tests without human review. The AI accelerates — it does not replace judgment.

  5. Context is everything. Feed the LLM your conventions, existing patterns, selector strategy, and data setup. The more context, the less cleanup.

  6. Quality over quantity. Each test has a maintenance cost. Focus on critical paths, complex logic, and known risk areas — not test count.

The Pipeline

Mandatory workflow — agents MUST follow this order:

Step 1: Extract   → Requirements, entities, business rules from input
Step 2: Analyze   → Risks, invariants, edge cases, ambiguities
Step 3: Map       → Coverage matrix (requirement → scenario → priority)
Step 4: Generate  → Candidate scenarios (happy + boundary + negative + security + a11y)
Step 5: Design    → Assertions and oracles SEPARATELY from scenarios
Step 6: Code      → Test code (only after all above exist)
Step 7: Review    → Human review with traceability back to source

Full prompt templates for every step (extraction, risk analysis, scenario, oracle, code) live in references/prompt-patterns.md. Below is the shape of each step's output.

Step 1: Extract Requirements and Entities

Parse the input into structured elements: Entities (with roles/states/attributes), Business Rules (numbered), Explicit Requirements ([REQ-N], stated in source), and Implicit Requirements ([IMP-N], inferred — flag every one for human confirmation). Separating explicit from inferred is the rule that prevents testing assumptions as if they were specifications.

Step 2: Risk Analysis and Invariants

Derive what can go wrong, what must always be true, and where the source is silent.

  • Risks — table of Risk | Likelihood | Impact | Source Requirement (e.g. race condition on stock decrement, email delay > 30s).
  • Invariants (must ALWAYS hold) — stock >= 0, order total = sum(items) + tax + shipping, user sees only their own orders.
  • Ambiguities (need human answers) — "Does free shipping apply before or after discount codes?" Capture these explicitly; do not silently pick one.
  • Edge cases derived from risks — two users buy the last item, payment succeeds but email service is down.
Step 3: Coverage Matrix

The single most important artifact — it prevents both gaps and duplicates. Map every requirement to scenarios with category, priority, and oracle type:

RequirementScenarioCategoryPriorityOracle Type
REQ-1Add single item to empty cartHappy pathP0State: cart count = 1
REQ-1Add out-of-stock itemNegativeP0UI: error message, cart unchanged
REQ-2Complete checkout with valid cardHappy pathP0State: order created, stock decremented
REQ-2Two users checkout last itemRace conditionP1One succeeds, one gets stock error
INV-1Stock never goes negativeInvariantP0Data: stock >= 0 after any operation

After building it, verify: every requirement has ≥1 happy and ≥1 negative scenario; every invariant has a direct test; every Step-2 risk has a scenario; no two rows test the same thing.

Step 4: Generate Candidate Scenarios

For each matrix row, write the full scenario in Given/When/Then with explicit test-data requirements (Given: user with 99 items in cart (max 100); When: adds one more; Then: count = 100). Cover these categories systematically:

CategoryDescription
Happy pathThe user does exactly what the feature is designed for
BoundaryEdge of valid input ranges — use the BOUNDARIES framework (references/prompt-patterns.md)
NegativeInvalid inputs, unauthorized actions
SecurityAuth bypass, injection, privilege escalation
AccessibilityScreen reader, keyboard-only, contrast
State transitionValid and invalid moves between states
ConcurrencyTwo users acting simultaneously
Step 5: Design Assertions and Oracles

Deliberately separate from Step 4. Scenarios describe behavior; oracles describe how to verify it. For each scenario, define oracles across categories — a single assertion is rarely enough to prove a behavior:

Oracle categoryAssertsExample
UI stateVisible text / element statecart badge toHaveText('1')
DataPersisted state via API/DBGET /api/cart returns 1 item, correct total
NegativeWhat should NOT happenno error toast; no navigation away
Side effectAsync/external outcomesanalytics add_to_cart fired; email in inbox < 30s

Oracle quality rules: assert business outcomes not implementation details; use the most specific assertion available (toHaveText('$29.99'), not toBeTruthy()); include negative assertions; verify data integrity, not just UI; assert accessibility (focus management, live-region announcements).

Step 6: Generate Test Code

Only after Steps 1-5 produce reviewed artifacts. Code is a mechanical translation of scenarios + oracles into framework syntax, with traceability comments linking back to the requirement and scenario:

typescript
/**
 * Scenario: SC-001 — Add single item to empty cart
 * Requirement: REQ-1 (User can add items to cart)
 * Priority: P0
 */
test('add single item to empty cart', async ({ page, testProduct }) => {
  await page.goto(`/products/${testProduct.id}`);                       // Given
  await page.getByRole('button', { name: 'Add to cart' }).click();      // When
  await expect(page.getByTestId('cart-badge')).toHaveText('1');         // Then
  await expect(page.getByTestId('error-toast')).not.toBeVisible();      // Negative oracle
});

Code generation rules: match project conventions (from qa-project-context.md); reuse existing Page Objects, fixtures, and data factories; include traceability comments (Scenario: SC-XXX, Requirement: REQ-XX); follow the project's selector strategy; put setup/teardown in fixtures, not inline.

Step 7: Human Review

Not optional — a mandatory pipeline step. This reviews the tests this pipeline just generated, before they merge. (To audit a pre-existing suite you did not just generate, use ai-qa-review instead.) Run every generated test against this checklist:

  • Traces to requirement: test → scenario → coverage row → requirement is followable.
  • Tests behavior, not implementation: survives a harmless refactor.
  • Correct abstraction level: right test type (unit vs integration vs E2E).
  • Test naming and readability: the test name states the behavior; a reader sees intent without decoding the body.
  • Test isolation / no shared state: the test creates and cleans up its own data, holds no order dependency on sibling tests, and passes when run alone or in any order.
  • Realistic test data: plausible, diverse, using example.com.
  • Meaningful assertions: matches the oracle definition; specific, not toBeTruthy().
  • Matches project conventions: naming, structure, selector strategy.
  • No flakiness risks: no hardcoded timeouts, race conditions, or order dependence.
  • Edge cases included: goes beyond the happy path.
  • Assumptions validated: Step-2 ambiguities were resolved before coding.

Review outcome per test: KEEP (merge as-is) · MODIFY (fix listed issues, then merge) · REJECT (wrong requirement, wrong abstraction, hallucinated API) · DEFER (blocked on ambiguity).

Show full SKILL.md (849 more words)Show less

Guardrails

Hard rules. Agents MUST follow them.

  • Code before coverage is forbidden. Never emit test code before Steps 1-3 (requirements, risk analysis with documented assumptions, coverage matrix) exist. If an agent skips to code: STOP, go back.
  • Assert outcomes, not implementation. expect(screen.getByRole('progressbar')).toBeVisible(), not expect(component.state.isLoading).toBe(true). expect(page.getByTestId('cart-badge')).toHaveText('1'), not expect(store.dispatch).toHaveBeenCalledWith(...).
  • Scenarios (Step 4) before oracles (Step 5), always. Scenario = WHAT happens; oracle = HOW to verify. Mixing them biases scenarios toward easy assertions.
  • Always produce the intermediates — assumptions document, uncovered ambiguities, oracle candidates, and the traceability chain — even in abbreviated form.

Flag these when detected:

  • Hallucinated APIs — endpoints, selectors, methods, or imports that do not exist in the codebase. Verify mechanically (see Verification) before human review.
  • Duplicate scenarios — same behavior, trivially different data. Consolidate or parametrize.
  • Low-value assertions — expect(response).toBeTruthy(), expect(page).toHaveURL(/.*/).
  • Missing negative cases — if every scenario is a happy path, the coverage matrix is incomplete.
  • Unrealistic test data — test@test.com, John Doe, password123. Use diverse, plausible data on example.com.

Model selection per step

Route by difficulty, not habit. Use a cheap model for mechanical extraction (Step 1) and the coverage-matrix bookkeeping (Step 3) — Haiku 4.5 or Sonnet 4.6 are plenty. Escalate to Opus 4.8 for oracle design (Step 5) and hallucination-sensitive code generation (Step 6), where a wrong inference is expensive; reach for Fable 5 only on genuinely hard reasoning (subtle invariants, regulated-domain logic). Running the strongest model on every step is wasteful; running the cheapest on Step 6 produces fabricated APIs.

Verification

Convert the "hallucinated APIs" warning into a mechanical gate. After Step 6, before human review:

  1. Resolve imports / types. TypeScript: npx tsc --noEmit — fabricated imports and wrong signatures fail here. Python: python -m pyflakes <files> or ruff check.
  2. Grep generated selectors/endpoints against the codebase. Confirm every getByTestId('...') id and every API path the test calls actually exists in source:
    bash
    grep -roE "getByTestId\('([^']+)'\)" generated/ | sed -E "s/.*'([^']+)'.*/\1/" | sort -u \
      | while read id; do grep -rq "$id" src/ || echo "MISSING testid: $id"; done
  3. Run the suite once. Tests that reference nonexistent routes/selectors fail fast; quarantine those before review rather than reviewing dead code.

Any MISSING line or tsc error is a hallucination to fix before a human spends review time.

Anti-Patterns

  1. Skipping to code. The most common failure: an agent gets a PRD and immediately writes tests. Without the coverage matrix it misses scenarios and duplicates others. The pipeline exists to prevent this.
  2. Asserting implementation instead of behavior. expect(component.state.isLoading).toBe(true) breaks on any refactor. Assert expect(screen.getByRole('progressbar')).toBeVisible() — what the user observes.
  3. Mixing scenarios and assertions. Writing "test this thing and check this value" as one step. Separate what to test from how to verify it.
  4. No project context in the prompt. Without conventions and existing patterns, you get generic tests. qa-project-context.md exists for exactly this.
  5. Over-generating. AI will write 50 tests for a simple function. Each carries maintenance cost. Use the coverage matrix to bound generation to meaningful scenarios.
  6. Copy-paste without understanding. If you cannot explain what a generated test does and why, do not merge it. Tests you do not understand become tests you cannot debug.
  7. Shipping without review. Step 7 is not optional. AI tests routinely contain hallucinated APIs, wrong selectors, incorrect business logic, and flakiness only human review catches.
  8. Ignoring the feedback loop. When AI tests catch real bugs, note the prompt patterns that worked; when they false-positive, note what went wrong. Build a project-specific library of what works.

Done When

  • All seven artifacts exist: requirements document, risk & invariants, coverage matrix, scenario set, oracle definitions, test code, and review notes with a KEEP/MODIFY/REJECT/DEFER decision per test.
  • The coverage matrix was produced and reviewed before any test code file was written.
  • Verification passed: tsc --noEmit (or language equivalent) exits 0 and the selector/endpoint grep reports zero MISSING lines.
  • Each generated test has a recorded human review decision; no test is marked KEEP without one.
  • The suite's CI job exits 0 (green).
  • Reproducibility metadata recorded: the exact model ID (e.g. claude-opus-4-8, claude-sonnet-4-6, claude-haiku-4-5-20251001), input source hash, and the version of any skill / CLI / MCP server invoked.
  • qa-project-context — Set up the context file that makes AI test generation dramatically better. Configure this first.
  • playwright-automation — Deep Playwright patterns (POM, fixtures, CI), plus the Test Agents (init-agents --loop=claude, scaffolded into .claude/agents/) and @playwright/mcp modes chosen in Discovery Q2. Generated tests live inside this framework.
  • unit-testing — Jest, Vitest, pytest patterns for unit-level generated tests.
  • api-testing — Endpoint test patterns for tests generated from OpenAPI specs.
  • test-strategy — Decide what to test and at which level before generating.
  • test-reliability — Make generated tests reliable: flake classification, healing, video receipts.
  • ai-system-testing — When the input is an LLM feature spec, generate eval datasets here (Promptfoo, DeepEval, Ragas, Braintrust) instead of Playwright specs.
  • ai-qa-review — Audit a pre-existing suite you did not just generate (test smells, testability). Step 7 here only reviews this pipeline's own output.
  • ai-bug-triage — When generated tests find bugs, classify and report them through the triage pipeline.

Reference Files (in references/)

  • prompt-patterns.md — Full prompt library aligned to the seven steps: extraction, risk analysis, scenario generation, oracle design, and code generation prompts, plus the BOUNDARIES edge-case framework used in Step 4.

© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/ai-test-generation of petrkindlmann/qa-skills.

  • SKILL.md
  • references/prompt-patterns.md

Open the folder on GitHubat commit b3bb61b

Compare with similar skills

AI Test Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Test Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Test Generation this skillpetrkindlmann/qa-skills170—~4.8kAutomated safety check: PassMIT
Prd V07 Test Planningmattgierhart/PRD-driven-context-engineering180—~3.5kAutomated safety check: NotesMIT
Review Rfcnurettincoban/ai-prd-workflow298—~1.4kAutomated safety check: PassMIT
Build Doddanshapiro/kilroy222—~2.5kAutomated safety check: PassMIT
Argent QA Flowsbbplayer-app/BBPlayer1.2k—~3.2kAutomated safety check: PassMIT
Test Scenariosphuryn/pm-skills27k—~866Automated safety check: PassMIT

Similar skills

  • Prd V07 Test Planning

    mattgierhart/PRD-driven-context-engineering

    Define test cases BEFORE implementation, ensuring every API, business rule, and user journey has verifiable acceptance criteria during PRD v0.7 Build Execution.

    180 GitHub stars~3.5k tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes
  • Review Rfc

    nurettincoban/ai-prd-workflow

    Review an implemented RFC in a fresh context against its acceptance criteria, RULES.md and the test plan, and save the review to reviews/.

    298 GitHub stars~1.4k tokensUpdated 2 days ago
    Product & Project ManagementAuto-check passed
  • Build Dod

    danshapiro/kilroy

    A skill your agent uses when converting a spec, requirements document, or goal statement into a Definition of Done with acceptance criteria and integration test scenarios

    222 GitHub stars~2.5k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Argent QA Flows

    bbplayer-app/BBPlayer

    Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria.

    1.2k GitHub stars~3.2k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Test Scenarios

    phuryn/pm-skills

    Create comprehensive test scenarios from user stories with test objectives, starting conditions, user roles, step-by-step actions, and expected outcomes.

    27k GitHub stars~866 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Test Author

    jpicklyk/task-orchestrator

    Test authoring framework for items carrying the needs-test-author trait.

    207 GitHub stars~8.2k tokensUpdated today
    Testing & QAAuto-check passed

More from petrkindlmann/qa-skills

All 45 skills in this repo
  • Accessibility Testing

    petrkindlmann/qa-skills

    Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).

    170 GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    170 GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • API Testing

    petrkindlmann/qa-skills

    Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.

    170 GitHub stars~2.7k tokensUpdated 4 mo ago
    Auto-check passed
  • CI CD Integration

    petrkindlmann/qa-skills

    Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.

    170 GitHub stars~4.8k tokensUpdated 4 mo ago
    Auto-check passed
  • Compliance Testing

    petrkindlmann/qa-skills

    Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…

    170 GitHub stars~4.6k tokensUpdated 4 mo ago
    Auto-check passed
  • Contract Testing

    petrkindlmann/qa-skills

    Implement consumer-driven contract testing with Pact-JS (v16).

    170 GitHub stars~4.1k tokensUpdated 4 mo ago
    Auto-check passed

Works with

Questions about AI Test Generation

What does AI Test Generation do?

Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs. AI Test Generation is an agent skill from petrkindlmann/qa-skills. Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

When should I use AI Test Generation?

AI Test Generation fits situations like: : generate tests from spec; tests from user story; auto-generate test cases; AI write tests for me. Not for: testing AI/LLM features in your product — use ai-system-testing.

How do I install AI Test Generation in Claude Code?

Run `npx skills add petrkindlmann/qa-skills --skill ai-test-generation -a claude-code`. Or copy the skill folder (skills/ai-test-generation in petrkindlmann/qa-skills) into .claude/skills/ai-test-generation in your project. Claude Code loads it when a task matches its description.

How do I install AI Test Generation in Codex?

Run `npx skills add petrkindlmann/qa-skills --skill ai-test-generation -a codex`. Or copy the skill folder (skills/ai-test-generation in petrkindlmann/qa-skills) into .agents/skills/ai-test-generation in your project. Codex loads it when a task matches its description.

Can I use AI Test Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill ai-test-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-test-generation, .gemini/skills/ai-test-generation, .github/skills/ai-test-generation and .opencode/skills/ai-test-generation in your project.

What does AI Test Generation need to run?

Going by SKILL.md and its folder, AI Test Generation needs the command-line tools its instructions call (npx, git, python, ruff and tsc). Our summary lists: Node.js.

Does AI Test Generation access the network?

SKILL.md contains no URLs. Its commands use npx and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is AI Test Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Test Generation use?

AI Test Generation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Test Generation use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to AI Test Generation?

Skills that share tags, products or a category with AI Test Generation: Prd V07 Test Planning (mattgierhart/PRD-driven-context-engineering, 180 stars), Review Rfc (nurettincoban/ai-prd-workflow, 298 stars), Build Dod (danshapiro/kilroy, 222 stars) and Argent QA Flows (bbplayer-app/BBPlayer, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Test Generation?

petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 170 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.

Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.