Testing
nteract/nteract
Run tests, verify changes, and collect diagnostics. An agent skill from nteract/nteract.
A skill your agent uses when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when…
$ npx skills add ericrisco/rsc-harness --skill testing-py -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness testing-py --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/testing-py .claude/skills/testing-py && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "testing-py" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/testing-py into .claude/skills/testing-py/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-py", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/testing-pyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill testing-py -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness testing-py --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/testing-py .agents/skills/testing-py && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "testing-py" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/testing-py into .agents/skills/testing-py/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-py", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill testing-py -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness testing-py --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/testing-py .cursor/skills/testing-py && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "testing-py" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/testing-py into .cursor/skills/testing-py/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-py", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/testing-py--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill testing-py -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness testing-py --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/testing-py .gemini/skills/testing-py && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "testing-py" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/testing-py into .gemini/skills/testing-py/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-py", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness testing-pyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill testing-py -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/testing-py .github/skills/testing-py && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "testing-py" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/testing-py into .github/skills/testing-py/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-py", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill testing-py -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness testing-py --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/testing-py .opencode/skills/testing-py && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "testing-py" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/testing-py into .opencode/skills/testing-py/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-py", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
testing-pyA skill your agent uses when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when…
Testing Py is an agent skill from ericrisco/rsc-harness. Use when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when coverage is high but catches nothing, when fixture scope leaks state, or when adding property-based tests for parsers and invariants. NOT browser end-to-end flows (that is e2e-testing), NOT JS/TS unit tests (that is testing-web), NOT Go tests (that is testing-go).
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/mocking.md`).
It sits in Testing & QA, covering Unit testing and End-to-end testing. It works with Python and pytest. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
pytestFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Testing Py loads about 3.2k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 1,447 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,447 words, ~3,207 tokens.
.claude/skills/testing-py/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.A green suite is not a tested suite. Most Python suites pass while shipping bugs because they lie in four predictable ways, and your job is to refuse each lie:
else and except arm is unvisited.The posture, non-negotiable: pytest 9 strict markers and strict config on, branch coverage on, autospec= on every mock, and Hypothesis property tests for any pure logic with an invariant. Assert on behavior, never on a mock's internals.
Versions you are on (2026-06): pytest 9.0.2, coverage.py 7.14.1, pytest-cov 7.x (needs coverage ≥ 7.10.6), Hypothesis 6.155.x.
Pick the cheapest test that still catches the bug.
| Situation | Test type | Reach for |
|---|---|---|
| Pure function, many input shapes, an invariant holds | property test | Hypothesis @given(...) |
| Same logic, a small fixed matrix of inputs/outputs | parametrized test | @pytest.mark.parametrize |
| Code calls an external boundary (HTTP, DB, SDK, clock) | unit test + mock at the boundary | autospec=True / mocker |
| Several real components must work together | integration test | real objects, fixtures, tmp_path |
| One test needs an env var / attr swapped and reverted | inline patch | monkeypatch |
Default to property + parametrized for logic; reserve mocks for boundaries you do not own.
tests/ mirrors the package tree. Put each fixture in the nearest conftest.py, not always the root one — proximity is documentation. Drive everything from pyproject.toml:
[tool.pytest.ini_options]
testpaths = ["tests"]
addopts = "--strict-markers --strict-config --cov=mypkg --cov-branch --cov-report=term-missing --cov-fail-under=85"
markers = [
"slow: deselect with -m 'not slow'",
"integration: touches a real boundary",
]
[tool.coverage.run]
branch = true # measure both arms of every conditional, not just lines
source = ["mypkg"]
[tool.coverage.report]
show_missing = true
fail_under = 85
exclude_lines = [
"if TYPE_CHECKING:",
"raise NotImplementedError",
"def __repr__",
"pragma: no cover",
]--strict-markers turns a typo'd @pytest.mark.slwo into an error instead of a silently-skipped test. --strict-config rejects unknown config keys. pytest 9 also exposes strict_markers, strict_config, strict_xfail, and strict_parametrization_ids as config aliases — treat strict as the default, not an upgrade.
Scope is a contract about reuse and isolation. Default to function and widen only when setup is genuinely expensive and the object is genuinely immutable.
| Scope | Recreated | Use when |
|---|---|---|
function (default) | every test | anything mutable — the safe default |
class | per test class | a group sharing read-only setup |
module | once per file | costly read-only object (a parsed config) |
package | once per dir | shared fixture for a sub-tree |
session | once per run | a process-wide resource you never mutate |
A yield fixture's pre-yield code is setup, post-yield is teardown — and teardown runs even when the test fails, so it is where cleanup belongs.
# Bad: everything in the root conftest.py, plus a hidden autouse reset.
# Nobody can trace why a test depends on the DB, and the reset masks state leaks.
@pytest.fixture(autouse=True)
def reset_db():
db.truncate_all()
# Good: scoped fixture in tests/api/conftest.py, explicit, documented.
# The docstring shows up in `pytest --fixtures`, so it is discoverable.
@pytest.fixture
def api_client(tmp_path):
"""An isolated API client backed by a throwaway sqlite file."""
client = ApiClient(db_path=tmp_path / "test.db")
yield client
client.close() # runs even if the test raisesParametrize a fixture when several tests should each run against multiple backings:
@pytest.fixture(params=["sqlite", "memory"])
def store(request):
return make_store(request.param) # every test using `store` runs twiceUse autouse=True only for something every test in scope truly needs (a frozen clock, a captured log). A hidden global dependency makes a failure impossible to read.
Patch the name where it is looked up, not where it is defined. This is the single most common mocking bug. If mymod does from httpx import get, then get now lives at mymod.get; patching httpx.get changes a name mymod no longer reads, so the mock never applies and the test silently hits the network.
# Bad: patches the definition site; mymod already bound its own `get` at import.
with patch("httpx.get") as m: # no effect on mymod.get
result = mymod.fetch(url)
# Good: patch the name in the module under test, and autospec it.
with patch("mymod.get", autospec=True) as m:
m.return_value.json.return_value = {"ok": True}
result = mymod.fetch(url)
m.assert_called_once_with(url, timeout=5)autospec=True is mandatory. Without it a typo'd method or a wrong arg count still passes — the mock invents any attribute you touch. With autospec the mock enforces the real object's signature, so a refactor that renames a method or drops a parameter fails the test like it should. Use create_autospec(obj) when you need a standalone fake of a class or callable.
Choosing the swap tool:
monkeypatch (built into pytest, no install): env vars, attributes, chdir. monkeypatch.setenv / delenv / setattr / chdir auto-revert after the test. Use it for "swap one thing for this test."mocker (pytest-mock): unittest.mock with automatic undo and call tracking. Less boilerplate for return_value / side_effect and assert_called_*. Use it for "fake a callable and inspect how it was called."Both are correct; pick by what you are swapping. Faking time, HTTP, DB, the filesystem, side_effect, asserting calls, and AsyncMock for async code live in references/mocking.md.
Run it with branch coverage and a floor:
pytest --cov=mypkg --cov-branch --cov-report=term-missing --cov-fail-under=85Line coverage lies because executing a line is not the same as testing its outcomes. A function with if x: a() else: b() reaches 100% line coverage from a single test that only takes the if arm — the else ships untested. --cov-branch (or branch = true) forces both arms to be exercised, which is exactly where happy-path-only suites leak bugs. term-missing prints the unhit line and branch numbers so you know what to write next. Exclude only code that is correctly never run by tests (if TYPE_CHECKING:, raise NotImplementedError, __repr__) — never exclude a branch just to make the number go up.
Branch coverage tells you which code ran. It cannot tell you whether any test would have noticed if that code were wrong — and a test with no assertion, or an assertion that cannot fail, raises coverage without detecting anything. Mutation testing closes that gap by planting bugs on purpose: if the suite still passes, the mutant survived and you have found a test that asserts nothing.
# pyproject.toml → [tool.mutmut] source_paths = ["src/"]
mutmut run # whole source tree
mutmut run "mypkg.core*" # scope it — this is the normal case
mutmut results # what survivedReach for the real tool, not a hand-written mutant list: mutmut generates mutants from the syntax tree, so it cannot apply one to code that moved and cannot report one it never ran.
__pycache__ and set PYTHONDONTWRITEBYTECODE per mutant: two same-size mutants written in the same second can share a bytecode cache, and the runner then reports a kill for code it never ran. That defect can only ever inflate the score, so it will never show up as a red run — which is why the proof has to be built in rather than assumed.Reach for Hypothesis when a fixed example matrix can't cover the input space: parsers, serializers, encoders, math, anything with an invariant. You assert a property that must hold for all inputs and let Hypothesis hunt counterexamples.
from hypothesis import given, assume, strategies as st
@given(st.text())
def test_encode_decode_roundtrip(s):
assume("\x00" not in s) # skip inputs the encoder legitimately rejects
assert decode(encode(s)) == s # roundtrip: the core propertyThe property shapes worth knowing: roundtrip (decode(encode(x)) == x), idempotence (f(f(x)) == f(x)), and invariant (a sorted list stays the same length and is ordered). On failure Hypothesis shrinks to the simplest failing input and stores it in the .hypothesis/ example DB, so the next run replays that exact case until you fix it — commit the DB or persist it in CI to keep regressions pinned. Use assume() to discard inputs that aren't meant to be valid, and bound slow strategies so the deadline doesn't flake. Strategy catalog, composite strategies, settings/profiles, @example, and stateful testing live in references/property-testing.md.
| Anti-pattern | Why it is wrong | Do instead |
|---|---|---|
patch(...) without autospec=/spec= | false green on a renamed or mis-called method | always autospec=True (or create_autospec) |
Patching the definition site (httpx.get) | the mock never applies to the module under test | patch the use site (mymod.get) |
time.sleep(...) to "fix" flakiness | slow and still flaky | fake the clock / poll on a condition |
Asserting on mock._mock_calls or private attrs | test breaks on refactor, not on bugs | assert on the return value / observable behavior |
| Session-scoped mutable fixture with no reset | state leaks across tests, order-dependent failures | function scope, or explicit teardown |
--cov without --cov-branch | 100% lines while branches are untested | always --cov-branch / branch = true |
| One giant test asserting 12 things | first failure hides the rest; can't tell what broke | one behavior per test, parametrize the matrix |
@given with no assume/bounds on slow code | deadline timeouts, flaky property tests | bound the strategy, raise/disable deadline deliberately |
assert True or bare pytest.skip() | a test that can never fail or never runs | assert a real outcome; skip(reason=...) |
Before declaring a suite done, run the static gate over the test tree. It flags the skill's own banned patterns — mocks without autospec/spec, --cov without --cov-branch, time.sleep in tests, and no-op tests:
scripts/verify.sh tests/It is read-only, prints file:line for each hit, exits non-zero on any finding, and exits 0 on a clean or empty tree. A clean run is necessary, not sufficient — it cannot tell you the branches are actually tested, only that the obvious lies are absent.
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/testing-py of ericrisco/rsc-harness.
Open the folder on GitHubat commit e3d5b33
Testing Py next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Testing Py this skillericrisco/rsc-harness | 167 | — | ~3.2k | Automated safety check: Pass | MIT | |
| Testingnteract/nteract | 179 | — | ~2.2k | Automated safety check: Pass | BSD-3-Clause | |
| Python Testing Patternsjh941213/my-cc-harness | 126 | 17 repos | ~5.4k | Automated safety check: Pass | None | |
| Specx Testsmaksimzayats/specx | 202 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Python Testingmacalbert/envilder | 138 | — | ~3.1k | Automated safety check: Pass | MIT | |
| Testing Patternssoftspark/ai-toolkit | 179 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 |
nteract/nteract
Run tests, verify changes, and collect diagnostics. An agent skill from nteract/nteract.
jh941213/my-cc-harness
Implement comprehensive testing strategies with pytest, fixtures, mocking, and test-driven development.
maksimzayats/specx
Add or refine tests for specx Python services. An agent skill from maksimzayats/specx.
macalbert/envilder
Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks.
softspark/ai-toolkit
Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.
ArabelaTso/Skills-4-SE
Explains test failures and provides actionable debugging guidance.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Categories
A skill your agent uses when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when…. Testing Py is an agent skill from ericrisco/rsc-harness. Use when writing or fixing Python tests with pytest — especially when the suite is green but bugs still ship, when you must decide where to patch a mocked dependency, when coverage is high but catches nothing, when fixture scope leaks state, or when adding property-based tests for parsers and invariants.
Testing Py fits situations like: fixing Python tests with pytest — especially when the suite is green but bugs still ship; you must decide where to patch a mocked dependency; coverage is high but catches nothing; fixture scope leaks state.
Run `npx skills add ericrisco/rsc-harness --skill testing-py -a claude-code`. Or copy the skill folder (skills/testing-py in ericrisco/rsc-harness) into .claude/skills/testing-py in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill testing-py -a codex`. Or copy the skill folder (skills/testing-py in ericrisco/rsc-harness) into .agents/skills/testing-py in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill testing-py -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-py, .gemini/skills/testing-py, .github/skills/testing-py and .opencode/skills/testing-py in your project.
Going by SKILL.md and its folder, Testing Py needs a shell for the scripts in its folder and the command-line tools its instructions call (pytest). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Testing Py is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Testing Py: Testing (nteract/nteract, 179 stars), Python Testing Patterns (jh941213/my-cc-harness, 126 stars), Specx Tests (maksimzayats/specx, 202 stars) and Python Testing (macalbert/envilder, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.