Agent skill

Testing Py

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when…

MITAuto-check passedTesting & QA

Install Testing Py

skills CLI
$ npx skills add ericrisco/rsc-harness --skill testing-py -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness testing-py --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/testing-py .claude/skills/testing-py && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-py
GitHub stars
167
Token cost
~3.2k tokens
SKILL.md length
1,447 words
Files
6 (incl. scripts, references)
Skills in repo
227
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when…

  • Works in 4 steps: Line-only coverage reports 95% while… → Mocks patched at the wrong place never… → autospec-less mocks pass even when you… → …
  • Fixing Python tests with pytest — especially when the suite is green but bugs still ship
  • SKILL.md covers Which test do I need?, Layout and config, Fixtures and Mocking, plus 5 more sections
  • Runs Shell scripts from its folder; calls pytest

What it does

Testing Py is an agent skill from ericrisco/rsc-harness. Use when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when coverage is high but catches nothing, when fixture scope leaks state, or when adding property-based tests for parsers and invariants. NOT browser end-to-end flows (that is e2e-testing), NOT JS/TS unit tests (that is testing-web), NOT Go tests (that is testing-go).

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/mocking.md`).

It sits in Testing & QA, covering Unit testing and End-to-end testing. It works with Python and pytest. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Fixing Python tests with pytest — especially when the suite is green but bugs still ship
  • You must decide where to patch a mocked dependency
  • Coverage is high but catches nothing
  • Fixture scope leaks state

Example prompts

  • “/testing-py”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Line-only coverage reports 95% while every else and except arm is unvisited.
  2. Mocks patched at the wrong place never apply, so the test runs the real code or asserts nothing.
  3. autospec-less mocks pass even when you typo a method name or change a signature — a false green.
  4. autouse fixtures mutate hidden global state, so tests pass alone and fail in CI order.

What it can do on your machine

Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing Py loads about 3.2k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 1,447 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,447 words, ~3,207 tokens.

Download SKILL.mdSave it as .claude/skills/testing-py/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
testing-py
description
Use when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when coverage is high but catches nothing, when fixture scope leaks state, or when adding property-based tests for parsers and invariants. NOT browser end-to-end flows (that is `e2e-testing`), NOT JS/TS unit tests (that is `testing-web`), NOT Go tests (that is `testing-go`).
tags
pytest, python, testing, coverage, hypothesis, mocking, fixtures, property-based
recommends
python, code-review, github-actions, e2e-testing
origin
risco

Testing Python

A green suite is not a tested suite. Most Python suites pass while shipping bugs because they lie in four predictable ways, and your job is to refuse each lie:

  1. Line-only coverage reports 95% while every else and except arm is unvisited.
  2. Mocks patched at the wrong place never apply, so the test runs the real code or asserts nothing.
  3. autospec-less mocks pass even when you typo a method name or change a signature — a false green.
  4. autouse fixtures mutate hidden global state, so tests pass alone and fail in CI order.

The posture, non-negotiable: pytest 9 strict markers and strict config on, branch coverage on, autospec= on every mock, and Hypothesis property tests for any pure logic with an invariant. Assert on behavior, never on a mock's internals.

Versions you are on (2026-06): pytest 9.0.2, coverage.py 7.14.1, pytest-cov 7.x (needs coverage ≥ 7.10.6), Hypothesis 6.155.x.

Which test do I need?

Pick the cheapest test that still catches the bug.

SituationTest typeReach for
Pure function, many input shapes, an invariant holdsproperty testHypothesis @given(...)
Same logic, a small fixed matrix of inputs/outputsparametrized test@pytest.mark.parametrize
Code calls an external boundary (HTTP, DB, SDK, clock)unit test + mock at the boundaryautospec=True / mocker
Several real components must work togetherintegration testreal objects, fixtures, tmp_path
One test needs an env var / attr swapped and revertedinline patchmonkeypatch

Default to property + parametrized for logic; reserve mocks for boundaries you do not own.

Layout and config

tests/ mirrors the package tree. Put each fixture in the nearest conftest.py, not always the root one — proximity is documentation. Drive everything from pyproject.toml:

toml
[tool.pytest.ini_options]
testpaths = ["tests"]
addopts = "--strict-markers --strict-config --cov=mypkg --cov-branch --cov-report=term-missing --cov-fail-under=85"
markers = [
  "slow: deselect with -m 'not slow'",
  "integration: touches a real boundary",
]

[tool.coverage.run]
branch = true            # measure both arms of every conditional, not just lines
source = ["mypkg"]

[tool.coverage.report]
show_missing = true
fail_under = 85
exclude_lines = [
  "if TYPE_CHECKING:",
  "raise NotImplementedError",
  "def __repr__",
  "pragma: no cover",
]

--strict-markers turns a typo'd @pytest.mark.slwo into an error instead of a silently-skipped test. --strict-config rejects unknown config keys. pytest 9 also exposes strict_markers, strict_config, strict_xfail, and strict_parametrization_ids as config aliases — treat strict as the default, not an upgrade.

Fixtures

Scope is a contract about reuse and isolation. Default to function and widen only when setup is genuinely expensive and the object is genuinely immutable.

ScopeRecreatedUse when
function (default)every testanything mutable — the safe default
classper test classa group sharing read-only setup
moduleonce per filecostly read-only object (a parsed config)
packageonce per dirshared fixture for a sub-tree
sessiononce per runa process-wide resource you never mutate

A yield fixture's pre-yield code is setup, post-yield is teardown — and teardown runs even when the test fails, so it is where cleanup belongs.

python
# Bad: everything in the root conftest.py, plus a hidden autouse reset.
# Nobody can trace why a test depends on the DB, and the reset masks state leaks.
@pytest.fixture(autouse=True)
def reset_db():
    db.truncate_all()

# Good: scoped fixture in tests/api/conftest.py, explicit, documented.
# The docstring shows up in `pytest --fixtures`, so it is discoverable.
@pytest.fixture
def api_client(tmp_path):
    """An isolated API client backed by a throwaway sqlite file."""
    client = ApiClient(db_path=tmp_path / "test.db")
    yield client
    client.close()  # runs even if the test raises

Parametrize a fixture when several tests should each run against multiple backings:

python
@pytest.fixture(params=["sqlite", "memory"])
def store(request):
    return make_store(request.param)  # every test using `store` runs twice

Use autouse=True only for something every test in scope truly needs (a frozen clock, a captured log). A hidden global dependency makes a failure impossible to read.

Mocking

Patch the name where it is looked up, not where it is defined. This is the single most common mocking bug. If mymod does from httpx import get, then get now lives at mymod.get; patching httpx.get changes a name mymod no longer reads, so the mock never applies and the test silently hits the network.

python
# Bad: patches the definition site; mymod already bound its own `get` at import.
with patch("httpx.get") as m:        # no effect on mymod.get
    result = mymod.fetch(url)

# Good: patch the name in the module under test, and autospec it.
with patch("mymod.get", autospec=True) as m:
    m.return_value.json.return_value = {"ok": True}
    result = mymod.fetch(url)
    m.assert_called_once_with(url, timeout=5)

autospec=True is mandatory. Without it a typo'd method or a wrong arg count still passes — the mock invents any attribute you touch. With autospec the mock enforces the real object's signature, so a refactor that renames a method or drops a parameter fails the test like it should. Use create_autospec(obj) when you need a standalone fake of a class or callable.

Choosing the swap tool:

  • monkeypatch (built into pytest, no install): env vars, attributes, chdir. monkeypatch.setenv / delenv / setattr / chdir auto-revert after the test. Use it for "swap one thing for this test."
  • mocker (pytest-mock): unittest.mock with automatic undo and call tracking. Less boilerplate for return_value / side_effect and assert_called_*. Use it for "fake a callable and inspect how it was called."

Both are correct; pick by what you are swapping. Faking time, HTTP, DB, the filesystem, side_effect, asserting calls, and AsyncMock for async code live in references/mocking.md.

Coverage that means something

Run it with branch coverage and a floor:

bash
pytest --cov=mypkg --cov-branch --cov-report=term-missing --cov-fail-under=85

Line coverage lies because executing a line is not the same as testing its outcomes. A function with if x: a() else: b() reaches 100% line coverage from a single test that only takes the if arm — the else ships untested. --cov-branch (or branch = true) forces both arms to be exercised, which is exactly where happy-path-only suites leak bugs. term-missing prints the unhit line and branch numbers so you know what to write next. Exclude only code that is correctly never run by tests (if TYPE_CHECKING:, raise NotImplementedError, __repr__) — never exclude a branch just to make the number go up.

Show full SKILL.md (679 more words)Show less

Mutation: does the suite actually notice?

Branch coverage tells you which code ran. It cannot tell you whether any test would have noticed if that code were wrong — and a test with no assertion, or an assertion that cannot fail, raises coverage without detecting anything. Mutation testing closes that gap by planting bugs on purpose: if the suite still passes, the mutant survived and you have found a test that asserts nothing.

bash
# pyproject.toml → [tool.mutmut]  source_paths = ["src/"]
mutmut run                    # whole source tree
mutmut run "mypkg.core*"      # scope it — this is the normal case
mutmut results                # what survived

Reach for the real tool, not a hand-written mutant list: mutmut generates mutants from the syntax tree, so it cannot apply one to code that moved and cannot report one it never ran.

  • Scale it to risk. This is not a default toll — it is minutes-to-hours on a real tree. Run it where a bug is expensive (money, auth, data loss, concurrency, a public API) or where a green suite keeps shipping bugs anyway. Always scope to the code you changed.
  • A survivor is not automatically a failure. Some mutants are semantically equivalent to the original and cannot be killed. Classify those with the reason ("equivalent: both branches return the same value for every reachable input"). Never add a test that asserts non-behaviour just to kill one — that is the same gaming that coverage-chasing is.
  • A survivor that is a real bug gets a test, not an excuse. Write the assertion that kills it, then rerun.
  • If you hand-roll a runner, prove it executed every mutant. Clear __pycache__ and set PYTHONDONTWRITEBYTECODE per mutant: two same-size mutants written in the same second can share a bytecode cache, and the runner then reports a kill for code it never ran. That defect can only ever inflate the score, so it will never show up as a red run — which is why the proof has to be built in rather than assumed.

Property testing

Reach for Hypothesis when a fixed example matrix can't cover the input space: parsers, serializers, encoders, math, anything with an invariant. You assert a property that must hold for all inputs and let Hypothesis hunt counterexamples.

python
from hypothesis import given, assume, strategies as st

@given(st.text())
def test_encode_decode_roundtrip(s):
    assume("\x00" not in s)            # skip inputs the encoder legitimately rejects
    assert decode(encode(s)) == s      # roundtrip: the core property

The property shapes worth knowing: roundtrip (decode(encode(x)) == x), idempotence (f(f(x)) == f(x)), and invariant (a sorted list stays the same length and is ordered). On failure Hypothesis shrinks to the simplest failing input and stores it in the .hypothesis/ example DB, so the next run replays that exact case until you fix it — commit the DB or persist it in CI to keep regressions pinned. Use assume() to discard inputs that aren't meant to be valid, and bound slow strategies so the deadline doesn't flake. Strategy catalog, composite strategies, settings/profiles, @example, and stateful testing live in references/property-testing.md.

Anti-patterns

Anti-patternWhy it is wrongDo instead
patch(...) without autospec=/spec=false green on a renamed or mis-called methodalways autospec=True (or create_autospec)
Patching the definition site (httpx.get)the mock never applies to the module under testpatch the use site (mymod.get)
time.sleep(...) to "fix" flakinessslow and still flakyfake the clock / poll on a condition
Asserting on mock._mock_calls or private attrstest breaks on refactor, not on bugsassert on the return value / observable behavior
Session-scoped mutable fixture with no resetstate leaks across tests, order-dependent failuresfunction scope, or explicit teardown
--cov without --cov-branch100% lines while branches are untestedalways --cov-branch / branch = true
One giant test asserting 12 thingsfirst failure hides the rest; can't tell what brokeone behavior per test, parametrize the matrix
@given with no assume/bounds on slow codedeadline timeouts, flaky property testsbound the strategy, raise/disable deadline deliberately
assert True or bare pytest.skip()a test that can never fail or never runsassert a real outcome; skip(reason=...)

Verify

Before declaring a suite done, run the static gate over the test tree. It flags the skill's own banned patterns — mocks without autospec/spec, --cov without --cov-branch, time.sleep in tests, and no-op tests:

bash
scripts/verify.sh tests/

It is read-only, prints file:line for each hit, exits non-zero on any finding, and exits 0 on a clean or empty tree. A clean run is necessary, not sufficient — it cannot tell you the branches are actually tested, only that the obvious lies are absent.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/testing-py of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/mocking.md
  • references/property-testing.md
  • scripts/verify.sh

Open the folder on GitHubat commit e3d5b33

Compare with similar skills

Testing Py next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing Py compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing Py this skillericrisco/rsc-harness167—~3.2kAutomated safety check: PassMIT
Testingnteract/nteract179—~2.2kAutomated safety check: PassBSD-3-Clause
Python Testing Patternsjh941213/my-cc-harness12617 repos~5.4kAutomated safety check: PassNone
Specx Testsmaksimzayats/specx202—~1.9kAutomated safety check: PassMIT
Python Testingmacalbert/envilder138—~3.1kAutomated safety check: PassMIT
Testing Patternssoftspark/ai-toolkit179—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Testing

    nteract/nteract

    Run tests, verify changes, and collect diagnostics. An agent skill from nteract/nteract.

    179 GitHub stars~2.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Python Testing Patterns

    jh941213/my-cc-harness

    Implement comprehensive testing strategies with pytest, fixtures, mocking, and test-driven development.

    126 GitHub starsUsed in 17 repos~5.4k tokens
    Testing & QAAuto-check passed
  • Specx Tests

    maksimzayats/specx

    Add or refine tests for specx Python services. An agent skill from maksimzayats/specx.

    202 GitHub stars~1.9k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Python Testing

    macalbert/envilder

    Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks.

    138 GitHub stars~3.1k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Testing Patterns

    softspark/ai-toolkit

    Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.

    179 GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Error Explanation Generator

    ArabelaTso/Skills-4-SE

    Explains test failures and provides actionable debugging guidance.

    253 GitHub stars~3.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from ericrisco/rsc-harness

All 227 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    167 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    167 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    167 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    167 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    167 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Testing Py

What does Testing Py do?

A skill your agent uses when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when…. Testing Py is an agent skill from ericrisco/rsc-harness. Use when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when coverage is high but catches nothing, when fixture scope leaks state, or when adding property-based tests for parsers and invariants.

When should I use Testing Py?

Testing Py fits situations like: fixing Python tests with pytest — especially when the suite is green but bugs still ship; you must decide where to patch a mocked dependency; coverage is high but catches nothing; fixture scope leaks state.

How do I install Testing Py in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill testing-py -a claude-code`. Or copy the skill folder (skills/testing-py in ericrisco/rsc-harness) into .claude/skills/testing-py in your project. Claude Code loads it when a task matches its description.

How do I install Testing Py in Codex?

Run `npx skills add ericrisco/rsc-harness --skill testing-py -a codex`. Or copy the skill folder (skills/testing-py in ericrisco/rsc-harness) into .agents/skills/testing-py in your project. Codex loads it when a task matches its description.

Can I use Testing Py in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill testing-py -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-py, .gemini/skills/testing-py, .github/skills/testing-py and .opencode/skills/testing-py in your project.

What does Testing Py need to run?

Going by SKILL.md and its folder, Testing Py needs a shell for the scripts in its folder and the command-line tools its instructions call (pytest). Our summary lists: Python 3; A Bash shell.

Does Testing Py access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Testing Py safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Testing Py use?

Testing Py is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing Py use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Testing Py?

Skills that share tags, products or a category with Testing Py: Testing (nteract/nteract, 179 stars), Python Testing Patterns (jh941213/my-cc-harness, 126 stars), Specx Tests (maksimzayats/specx, 202 stars) and Python Testing (macalbert/envilder, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing Py?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.