Agent skill

Test Harness

by Mathews-Tom in Mathews-Tom/armory

Generates pytest test suites with happy path, edge cases, error conditions, fixture scaffolding, mocks, async patterns.

MITAuto-check passedTesting & QA

Install Test Harness

skills CLI
$ npx skills add Mathews-Tom/armory --skill test-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory test-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-harness .claude/skills/test-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-harness
GitHub stars
328
Token cost
~3.5k tokens
SKILL.md length
1,314 words
Files
8 (incl. references)
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

Generates pytest test suites with happy path, edge cases, error conditions, fixture scaffolding, mocks, async patterns.

  • Works in 5 steps: Reconnaissance → Test Case Enumeration → Fixture Design → …
  • : generate tests
  • SKILL.md covers Reference Files, Prerequisites, Workflow and Output Format, plus 7 more sections
  • Runs Python scripts from its folder; calls git and npm

What it does

Test Harness is an agent skill from Mathews-Tom/armory. Generates pytest test suites with happy path, edge cases, error conditions, fixture scaffolding, mocks, async patterns. Triggers on: "generate tests", "write tests for", "test this function", "create test suite", "pytest for", "unit tests for", "mock strategy for".

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `evals/cases.yaml`, `evals/fixtures/sample_module.py` and `references/async-testing.md`).

It sits in Testing & QA, covering Test generation and Unit testing. It works with pytest. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • : generate tests
  • Write tests for
  • Test this function
  • Create test suite

Example prompts

  • “generate tests”
  • “write tests for”
  • “test this function”
  • “/test-harness”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Reconnaissance
  2. Test Case Enumeration
  3. Fixture Design
  4. Mock Strategy
  5. Output

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Harness loads about 3.5k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 1,314 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 1,314 words, ~3,472 tokens.

Download SKILL.mdSave it as .claude/skills/test-harness/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
test-harness
description
Generates pytest test suites with happy path, edge cases, error conditions, fixture scaffolding, mocks, async patterns. Triggers on: "generate tests", "write tests for", "test this function", "create test suite", "pytest for", "unit tests for", "mock strategy for".
metadata.version
1.1.1
metadata.category
development
metadata.tags
testing, pytest, test-generation, python
metadata.difficulty
intermediate
metadata.phase
build

Test Harness

Systematic test suite generation that transforms source code into comprehensive, runnable pytest files. Analyzes function signatures, dependency graphs, and complexity hotspots to produce tests covering happy paths, boundary conditions, error states, and async flows — with properly scoped fixtures and focused mocks.

Reference Files

FileContentsLoad When
references/pytest-patterns.mdFixture scopes, parametrize, marks, conftest layout, built-in fixturesAlways
references/mock-strategies.mdMock decision tree, patch boundaries, assertions, anti-patternsTarget has external dependencies
references/async-testing.mdpytest-asyncio modes, event loop fixtures, async mockingTarget contains async code
references/fixture-design.mdFactory fixtures, yield teardown, scope selection, compositionTest requires non-trivial setup
references/coverage-targets.mdThreshold table, branch vs line, pytest-cov config, exclusion patternsCoverage assessment requested

Prerequisites

  • pytest >= 7.0
  • Python >= 3.10
  • pytest-asyncio — required only when generating async tests
  • pytest-mock — optional, provides mocker fixture as alternative to unittest.mock

Workflow

Phase 1: Reconnaissance

Before writing a single test, build a model of the target code:

  1. Identify scope — What functions, classes, or modules need tests? If unspecified, check for recent modifications: git diff --name-only HEAD~5
  2. Read function signatures — Parameters, types, return types, defaults. Every parameter is a test dimension.
  3. Map dependencies — Which calls go to external systems (DB, API, filesystem, clock)? These are mock candidates.
  4. Detect complexity hotspots — Functions with high branch counts, deep nesting, or multiple return paths need more test cases.
  5. Check existing tests — If tests already exist, understand what they cover. Do not duplicate; extend.
  6. Read project conventions — Check CLAUDE.md, conftest.py, pytest.ini/pyproject.toml for fixtures, markers, and test organization patterns already in use.
Phase 2: Test Case Enumeration

For each function under test, enumerate cases across four categories:

CategoryWhat to TestExample
Happy pathExpected inputs produce expected outputsadd(2, 3) returns 5
BoundaryEdge values at limits of valid inputEmpty string, zero, max int, single element
ErrorInvalid inputs trigger proper exceptionsNone where str expected, negative index
StateState transitions produce correct side effectsObject moves from pending to active

For each case, note:

  • Input values (concrete, not abstract)
  • Expected output or exception
  • Required setup (fixtures)
  • Required mocks (external calls to suppress)

Parametrize cases that share the same test logic but differ only in input/output values.

Phase 3: Fixture Design
  1. Identify shared setup — If 3+ tests need the same object, extract a fixture.

  2. Select scope — Use the narrowest scope that avoids redundant setup:

    ScopeUse WhenExample
    functionDefault. Each test gets fresh stateMost unit tests
    classTests within a class share expensive setupDB connection per test class
    moduleAll tests in a file share setupLoaded config file
    sessionEntire test run shares setupDocker container startup
  3. Design teardown — Use yield fixtures when cleanup is needed. Never leave side effects (temp files, DB rows, monkey-patches) after a test.

  4. Identify conftest candidates — Fixtures used across multiple test files belong in conftest.py. Fixtures used in one file stay in that file.

Phase 4: Mock Strategy
  1. Decide what to mock — Mock external dependencies only:

    • Network calls (API, database, message queues)
    • Filesystem operations (when testing logic, not I/O)
    • Time-dependent behavior (datetime.now, time.sleep)
    • Random/non-deterministic behavior
  2. Decide what NOT to mock — Never mock:

    • The function under test
    • Pure functions called by the target (test them through the target)
    • Data structures and value objects
  3. Choose mock level — Patch at the import boundary of the module under test, not at the definition site. @patch('mymodule.requests.get'), not @patch('requests.get').

  4. Add mock assertions — Every mock should assert it was called with expected arguments and the expected number of times. Mocks without assertions are coverage holes.

Phase 5: Output

Generate the test file following this structure:

  1. Imports (pytest, mocks, target module)
  2. Constants and test data
  3. Fixtures (ordered by scope: session > module > class > function)
  4. Test classes or functions grouped by target function
  5. Parametrized tests where applicable

Output Format

text
# tests/test_{module}.py

import pytest
from unittest.mock import Mock, patch, MagicMock

from {module} import {target_function, TargetClass}


# ============================================================
# Fixtures
# ============================================================

@pytest.fixture
def valid_input():
    """Standard valid input for happy path tests."""
    return {concrete values}


@pytest.fixture
def mock_database():
    """Mock database connection."""
    with patch("{module}.db_connection") as mock_db:
        mock_db.query.return_value = [{expected data}]
        yield mock_db


# ============================================================
# {target_function} Tests
# ============================================================

class TestTargetFunction:
    """Tests for {target_function}."""

    def test_happy_path(self, valid_input):
        """Returns expected result for valid input."""
        result = target_function(valid_input)
        assert result == {expected}

    @pytest.mark.parametrize(
        "input_val, expected",
        [
            ({boundary_1}, {expected_1}),
            ({boundary_2}, {expected_2}),
            ({boundary_3}, {expected_3}),
        ],
        ids=["empty", "single", "maximum"],
    )
    def test_boundary_conditions(self, input_val, expected):
        """Handles boundary inputs correctly."""
        assert target_function(input_val) == expected

    def test_invalid_input_raises(self):
        """Raises TypeError for invalid input."""
        with pytest.raises(TypeError, match="expected str"):
            target_function(None)

    def test_external_call(self, mock_database):
        """Calls database with correct query."""
        target_function("lookup_key")
        mock_database.query.assert_called_once_with("SELECT * FROM t WHERE key = %s", ("lookup_key",))

Configuring Scope

ModeScopeDepthWhen to Use
quickSingle functionHappy path + 1 error caseRapid iteration, TDD red-green cycle
standardFile or classHappy + boundary + error + mocksDefault for most requests
comprehensiveModule or packageAll categories + async + parametrized matrixPre-release, critical path code

Calibration Rules

  1. Test isolation is non-negotiable. Every test must pass when run alone and in any order. No test may depend on the side effects of another test.
  2. Mock discipline. Mock external dependencies, not internal logic. Over-mocking produces tests that pass when the code is broken. Under-mocking produces tests that fail when the network is down.
  3. Concrete over abstract. Test data must be concrete values, not placeholders. "alice@example.com" not "test_email". 42 not "some_number". Concrete values catch type mismatches that abstract placeholders mask.
  4. One assertion focus per test. A test should verify one behavior. Multiple assertions are acceptable when they verify different aspects of the same behavior (e.g., return value AND side effect), but not when they verify unrelated behaviors.
  5. Parametrize, don't duplicate. If two tests differ only in input/output values, combine them with @pytest.mark.parametrize. Use ids for readable test names.
  6. Match project conventions. If the project uses conftest.py fixtures, class-based tests, or specific markers, follow those patterns. Do not introduce a conflicting test style.
Show full SKILL.md (467 more words)Show less

Error Handling

ProblemResolution
Target function has no type hintsInfer types from usage patterns, default values, and docstrings. Note uncertainty in test docstring.
Target has deeply nested dependenciesMock at the nearest boundary to the function under test. Do not mock transitive dependencies individually.
No existing test infrastructure (no conftest, no pytest config)Generate a minimal conftest.py alongside the test file. Note the addition in output.
Target code is untestable (global state, hidden dependencies)Flag the design issue in the output. Generate tests for what is testable. Suggest refactoring to improve testability.
Async code detected but pytest-asyncio not installedNote the dependency requirement. Generate async test stubs with @pytest.mark.asyncio and instruct user to install.
Target module cannot be importedReport the import error. Do not generate tests for unimportable code.

When NOT to Generate Tests

Push back if:

  • The code is auto-generated (protobuf, OpenAPI client, ORM models) — test the generator or the schema, not the output
  • The request is for UI/E2E tests — this skill generates unit and integration tests only
  • The code has no clear behavior to test (pure configuration, constant definitions)
  • The user wants tests for third-party library code — test your usage of the library, not the library itself

Rationalizations

RationalizationReality
"Manual testing is sufficient"Manual testing doesn't run in CI, doesn't catch regressions, and doesn't scale with the codebase
"This code is too simple to test"Simple code becomes complex code — tests document expected behavior and catch regressions from future changes
"I'll add tests later"Tests are specifications; without them, code behavior is undefined and later never comes
"Mocking everything makes the test fast"Over-mocked tests pass when the real system fails — mock at boundaries, not deep in the call chain
"100% coverage means the code is correct"Coverage measures execution, not correctness — a test that runs code without meaningful assertions adds no value
"The happy path test is enough"Edge cases and error paths cause most production incidents — happy-path-only testing is false confidence

Red Flags

  • Tests that only cover the happy path with no edge cases or error paths
  • Test names that describe implementation ("test_calls_function") instead of behavior ("test_returns_404_when_not_found")
  • More than two mocks per test — indicates the unit under test is too coupled
  • Tests that depend on execution order or shared mutable state
  • Assertions on implementation details (mock call counts) instead of observable behavior
  • Skipping integration tests because "unit tests cover it"

Verification

  • Tests follow Arrange-Act-Assert structure with clear phase separation
  • Test names describe behavior: test_<unit>_<scenario>_<expected_outcome>
  • Edge cases covered: empty input, boundary values, error paths, null/None
  • Coverage meets thresholds: 80% overall, 90% new code, 95% critical paths
  • All tests pass: pytest / npm test exits 0 with output captured
  • No test depends on execution order — can run in any sequence
  • Mocks used only at boundaries (external APIs, system clock, filesystem in unit tests)

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in skills/test-harness of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml
  • evals/fixtures/sample_module.py
  • references/async-testing.md
  • references/coverage-targets.md
  • references/fixture-design.md
  • references/mock-strategies.md
  • references/pytest-patterns.md

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Test Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Harness this skillMathews-Tom/armory328—~3.5kAutomated safety check: PassMIT
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence
Test Coverage Reviewareed1192/finance-news-aggregator149—~2.6kAutomated safety check: PassMIT
Write Test389ds/389-ds-base294—~1.8kAutomated safety check: PassCustom licence
MoAI TDD Workflowmodu-ai/moai-adk1.2k—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Coverage Review

    areed1192/finance-news-aggregator

    Audit, plan, write, and verify unit tests for Python projects using pytest.

    149 GitHub stars~2.6k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Write Test

    389ds/389-ds-base

    Add or extend a pytest integration test for 389 Directory Server under dirsrvtests/.

    294 GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • MoAI TDD Workflow

    modu-ai/moai-adk

    Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.

    1.2k GitHub stars~3.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Testing

    meleantonio/ChernyCode

    Testing conventions using pytest. An agent skill from meleantonio/ChernyCode.

    516 GitHub stars~408 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    328 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    328 GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    328 GitHub stars~2.8k tokensUpdated 3 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    328 GitHub stars~3.1k tokensUpdated 3 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    328 GitHub stars~4.9k tokensUpdated 3 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    328 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed

Works with

Categories

Questions about Test Harness

What does Test Harness do?

Generates pytest test suites with happy path, edge cases, error conditions, fixture scaffolding, mocks, async patterns. Test Harness is an agent skill from Mathews-Tom/armory. Generates pytest test suites with happy path, edge cases, error conditions, fixture scaffolding, mocks, async patterns.

When should I use Test Harness?

Test Harness fits situations like: : generate tests; write tests for; test this function; create test suite.

How do I install Test Harness in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill test-harness -a claude-code`. Or copy the skill folder (skills/test-harness in Mathews-Tom/armory) into .claude/skills/test-harness in your project. Claude Code loads it when a task matches its description.

How do I install Test Harness in Codex?

Run `npx skills add Mathews-Tom/armory --skill test-harness -a codex`. Or copy the skill folder (skills/test-harness in Mathews-Tom/armory) into .agents/skills/test-harness in your project. Codex loads it when a task matches its description.

Can I use Test Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill test-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-harness, .gemini/skills/test-harness, .github/skills/test-harness and .opencode/skills/test-harness in your project.

What does Test Harness need to run?

Going by SKILL.md and its folder, Test Harness needs Python for the scripts in its folder and the command-line tools its instructions call (git and npm). Our summary lists: Python 3; Docker.

Does Test Harness access the network?

SKILL.md contains no URLs. Its commands use git and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Harness use?

Test Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Harness use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9k tokens, read only when the agent opens those files.

What are the alternatives to Test Harness?

Skills that share tags, products or a category with Test Harness: Adk Verify Snippets (google/adk-python, 22k stars), Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars), Test Coverage Review (areed1192/finance-news-aggregator, 149 stars) and Write Test (389ds/389-ds-base, 294 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Harness?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 328 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.