Agent skill

Approval Testing Toolkit

by lexler in lexler/skill-factory

Writes snapshot-style approval tests in Python, JavaScript, TypeScript or Java, comparing output against an approved file instead of writing individual assertions.

Apache-2.0Auto-check passedTesting & QA

Install Approval Testing Toolkit

skills CLI
$ npx skills add lexler/skill-factory --skill approval-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lexler/skill-factory approval-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lexler/skill-factory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/output_skills/testing/approval-tests .claude/skills/approval-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
approval-tests
GitHub stars
239
Token cost
~1.2k tokens
SKILL.md length
505 words
Files
26 (incl. references)
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Writes snapshot-style approval tests in Python, JavaScript, TypeScript or Java, comparing output against an approved file instead of writing individual assertions.

  • Verifying a function's complex output without writing many separate assertions
  • SKILL.md covers Philosophy, Core Workflow, Core API Pattern and Key Techniques, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Adding characterization tests to legacy code that has no existing tests

What it does

The skill's premise is that a saved output file can replace a long list of assertions: you capture the real output once, review and approve it as a golden master, and later runs are compared against that approved snapshot and fail only when the output actually changes. This suits complex or large output, characterizing legacy code that has no tests yet, and covering many input combinations without writing a separate assertion for each.

Per-language reference files, java.md, nodejs.md and python.md, point into deeper references folders covering setup, the API, inline usage, logging, custom reporters and scrubbers for each of Java, Node.js and Python, plus a links.md file. The excerpt itself stops at the philosophy section, before the setup instructions.

When your agent uses it

  • Verifying a function's complex output without writing many separate assertions
  • Adding characterization tests to legacy code that has no existing tests
  • Testing many input combinations by comparing each against an approved snapshot

Example prompts

  • “Write approval tests for this report generator in Python using .approved files.”
  • “Set up approval testing for this legacy Java module so I can refactor it safely.”
  • “Add an approval test in TypeScript for the pricing calculator's output.”

What it can do on your machine

Read from SKILL.md and the folder at commit 8017333. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Approval Testing Toolkit loads about 1.2k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 505 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lexler/skill-factory at commit 8017333, republished under its Apache-2.0 licence (© lexler). 505 words, ~1,203 tokens.

Download SKILL.mdSave it as .claude/skills/approval-tests/SKILL.md (or your agent's skills folder). This skill also uses 25 other files; get the full folder from GitHub.
name
approval-tests
description
Writes approval tests (snapshot/golden master testing) for Python, JavaScript/TypeScript, or Java. Use when verifying complex output, characterization testing legacy code, testing combinations, or working with .approved/.received files.

STARTER_CHARACTER = 📸

Approval Tests

Philosophy

"A picture's worth 1000 assertions."

Approval tests verify complex output by comparing against a saved "golden master" file instead of writing individual assertions. You capture the output once, review it, approve it, and future runs compare against that approved snapshot.

You don't need to know the expected output upfront. Run the code, see what it produces, decide if it's correct. Approval is a judgment - you're confirming this is what the code should produce. Whoever writes the code reviews and approves.

Use approval tests when:

  • Output is complex - instead of 20 assertions, one approval captures everything
  • Characterizing legacy code - snapshot behavior, then refactor safely
  • Combinatorial testing - test all input combinations in one approval
  • Assertions would be tedious or brittle

Use assertions when:

  • Simple values or specific edge cases
  • Non-deterministic output that can't be scrubbed

Core Workflow

1. Write test with verify(result)
2. Run test → FAILS (no .approved file yet)
3. Creates: TestName.approved.txt (empty) + TestName.received.txt (actual output)
4. Review .received file - is this correct?
5. YES → rename/copy .received to .approved
6. Run test again → PASSES
7. Commit .approved file to version control

File naming convention:

{TestClass}.{test_method}.approved.txt   ← commit this
{TestClass}.{test_method}.received.txt   ← gitignore this

Critical rules:

  • .approved files ARE your test expectations - commit them
  • .received files are temporary - add *.received.* to .gitignore
  • Never edit .approved files by hand - always generate via test

When a test fails, a diff tool opens showing approved vs received. This is how you review changes. Reporters configure which diff tool to use.

Core API Pattern

All languages follow the same pattern:

verify(result)                    # Basic string/object verification
verify_as_json(object)            # Objects as formatted JSON
verify_all(header, items)         # Collections with labels
verify_all_combinations(fn, inputs)  # All input combinations

Non-deterministic data (timestamps, GUIDs) must be scrubbed before verification.

Key Techniques

  • Scrubbers - replace values that change between runs (timestamps, UUIDs, random numbers, ports, paths) with stable placeholders like [Date1] or guid_1. Without scrubbing, tests pass locally but fail in CI.
  • Inline approvals - expectations in source code instead of separate files. Avoids file proliferation for short output. Python uses docstrings, Java uses text blocks.
  • Storyboard - show an object at multiple points in time, like frames in a comic. Each step appears in the diff, making it easy to see how state changes. For workflows, state machines, animations. Python/Java have classes; Node.js uses string building.
  • Combinations - test all permutations of input parameters in one approval. Exhaustive coverage without writing separate tests for each case. For large sets, pairwise testing reduces millions of combinations to ~100.
  • Multiple approvals per test - calling verify() twice overwrites the same file, so only the last one is tested. Parameter-based naming creates separate files for each scenario.

See language references for implementation details.

Show full SKILL.md (136 more words)Show less

Language References

Detect language from project files, then read the appropriate reference for installation, quick start, core patterns, and links to deeper reference files:

  • python.md - Python (pyproject.toml, setup.py, requirements.txt)
  • nodejs.md - JavaScript/TypeScript (package.json)
  • java.md - Java (pom.xml, build.gradle)

Anti-Patterns

  • Don't write assertions for complex objects - use verify_as_json() instead
  • Don't commit .received files - they're temporary
  • Don't forget scrubbers for timestamps, GUIDs, random values
  • Don't over-verify - one approval per logical behavior. Large approvals hide signal in noise; unrelated changes break tests.
  • Don't hand-edit .approved files - always generate via test. Hand-edited files may not match actual code output.
  • Don't use verify_all for structured data - use verify_as_json({"items": items})
  • Don't mix approvals with assertions - the approval captures everything
  • Don't call verify() multiple times without NamerFactory - each overwrites the same file

Flaky tests across environments usually means unscrubbed dynamic data (timestamps, UUIDs, ports, paths).

© lexler, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 25 other files (references) in output_skills/testing/approval-tests of lexler/skill-factory.

  • SKILL.md
  • java.md
  • links.md
  • nodejs.md
  • python.md
  • references/java/advanced.md
  • references/java/api.md
  • references/java/inline.md
  • references/java/logging.md
  • references/java/reporters.md
  • references/java/scrubbers.md
  • references/java/setup.md
  • references/nodejs/api.md
  • references/nodejs/inline.md
  • references/nodejs/reporters.md
  • references/nodejs/scrubbers.md
  • references/nodejs/setup.md
  • references/python
  • … and 8 more

Open the folder on GitHubat commit 8017333

Compare with similar skills

Approval Testing Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Approval Testing Toolkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Approval Testing Toolkit this skilllexler/skill-factory239—~1.2kAutomated safety check: PassApache-2.0
Polyglot Test Agentboshi-xixixi/TraeSkill275—~1.7kAutomated safety check: PassMIT
Supercovsupercorp-ai/supercov1501 repos~415Automated safety check: PassMIT
Crap Analyzerswingerman/engineer154—~1.2kAutomated safety check: PassMIT
Property Based Testingforyourhealth111-pixel/Vibe-Skills3.6k—~4.8kAutomated safety check: NotesApache-2.0
Test Case ReducerArabelaTso/Skills-4-SE253—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Polyglot Test Agent

    boshi-xixixi/TraeSkill

    Generates comprehensive, workable unit tests for any programming language using a multi-agent pipeline.

    275 GitHub stars~1.7k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Supercov

    supercorp-ai/supercov

    Measures test coverage and code quality in a repository with the supercov CLI, and turns what it finds into small, focused tests or fixes.

    150 GitHub starsUsed in 1 repo~415 tokens
    Testing & QAAuto-check passed
  • Crap Analyzer

    swingerman/engineer

    A skill your agent uses to produce a risk-based refactor + test plan for recently-changed code on a diff/branch/PR by computing CRAP (complexity × untested) on changed methods.

    154 GitHub stars~1.2k tokensUpdated 15 days ago
    Testing & QAAuto-check passed
  • Property Based Testing

    foryourhealth111-pixel/Vibe-Skills

    Property-based testing with fast-check (TypeScript/JavaScript) and Hypothesis (Python).

    3.6k GitHub stars~4.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes
  • Test Case Reducer

    ArabelaTso/Skills-4-SE

    Automatically reduces bug-triggering test cases to minimal form while preserving the failure.

    253 GitHub stars~2.5k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Langgraph Testing Evaluation

    soba-labs/langchain-agent-skills

    A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…

    107 GitHub stars~2.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from lexler/skill-factory

All 25 skills in this repo
  • C4 Architecture Diagrams

    lexler/skill-factory

    Creates C4 model diagrams at every zoom level, from system landscape to code, in ASCII, Mermaid or Structurizr, for designing or documenting software architecture.

    239 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Launching Agent Teams

    lexler/skill-factory

    Plans and launches Claude Code agent teams with distinct roles, right-sized tasks and detailed spawn prompts, and says when subagents or worktrees fit better.

    239 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Claude Code Statusline Writer

    lexler/skill-factory

    Guides writing and debugging Claude Code status line scripts that read session JSON from stdin and print one line of text.

    239 GitHub stars~872 tokensUpdated 1 mo ago
    Auto-check passed
  • Catalog of obstacles, anti-patterns and patterns for working with AI coding agents, covering context management and reliability, from a published patterns collection.

    239 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Hotspots

    lexler/skill-factory

    Find where a codebase actually costs time by mining its git history (Tornhill hotspot analysis).

    239 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Nullables Testing Pattern

    lexler/skill-factory

    Teaches the Nullables pattern for testing without mocking libraries: production classes with an off switch, stubbed only at the third-party edge.

    239 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Approval Testing Toolkit

What does Approval Testing Toolkit do?

Writes snapshot-style approval tests in Python, JavaScript, TypeScript or Java, comparing output against an approved file instead of writing individual assertions. The skill's premise is that a saved output file can replace a long list of assertions: you capture the real output once, review and approve it as a golden master, and later runs are compared against that approved snapshot and fail only when the output actually changes. This suits complex or large output, characterizing legacy code that has no tests yet, and covering many input combinations without writing a separate assertion for each.

When should I use Approval Testing Toolkit?

Approval Testing Toolkit fits situations like: verifying a function's complex output without writing many separate assertions; adding characterization tests to legacy code that has no existing tests; testing many input combinations by comparing each against an approved snapshot.

How do I install Approval Testing Toolkit in Claude Code?

Run `npx skills add lexler/skill-factory --skill approval-tests -a claude-code`. Or copy the skill folder (output_skills/testing/approval-tests in lexler/skill-factory) into .claude/skills/approval-tests in your project. Claude Code loads it when a task matches its description.

How do I install Approval Testing Toolkit in Codex?

Run `npx skills add lexler/skill-factory --skill approval-tests -a codex`. Or copy the skill folder (output_skills/testing/approval-tests in lexler/skill-factory) into .agents/skills/approval-tests in your project. Codex loads it when a task matches its description.

Can I use Approval Testing Toolkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lexler/skill-factory --skill approval-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/approval-tests, .gemini/skills/approval-tests, .github/skills/approval-tests and .opencode/skills/approval-tests in your project.

What does Approval Testing Toolkit need to run?

SKILL.md names no scripts, command-line tools or credentials: Approval Testing Toolkit is instructions for the agent only.

Does Approval Testing Toolkit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Approval Testing Toolkit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Approval Testing Toolkit use?

Approval Testing Toolkit is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Approval Testing Toolkit use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.

What are the alternatives to Approval Testing Toolkit?

Skills that share tags, products or a category with Approval Testing Toolkit: Polyglot Test Agent (boshi-xixixi/TraeSkill, 275 stars), Supercov (supercorp-ai/supercov, 150 stars), Crap Analyzer (swingerman/engineer, 154 stars) and Property Based Testing (foryourhealth111-pixel/Vibe-Skills, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Approval Testing Toolkit?

lexler (a GitHub user) maintains it in lexler/skill-factory, which has 239 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on August 26, 2026.

Source: lexler/skill-factory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.