Agent skill

Testing Python

by benchflow-ai in benchflow-ai/skillsbench

Write and evaluate effective Python tests using pytest. An agent skill from benchflow-ai/skillsbench.

Apache-2.0Auto-check passedTesting & QA

Install Testing Python

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill testing-python -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench testing-python --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/fix-build-agentops/environment/skills/testing-python .claude/skills/testing-python && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-python
GitHub stars
1.8k
Token cost
~1.3k tokens
SKILL.md length
248 words
Files
1
Skills in repo
178
Repo updated
First seen
Licence
Apache-2.0

At a glance

Write and evaluate effective Python tests using pytest. An agent skill from benchflow-ai/skillsbench.

  • Reviewing test code
  • SKILL.md covers Core Principles, Test Structure, Project-Specific Rules and Fixtures, plus 5 more sections
  • Calls uv and pytest
  • Debugging test failures

What it does

Testing Python is an agent skill from benchflow-ai/skillsbench. Write and evaluate effective Python tests using pytest. Use when writing tests, reviewing test code, debugging test failures, or improving test coverage. Covers test design, fixtures, parameterization, mocking, and async testing.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Unit testing, Test coverage and Failing and flaky tests. It works with Python and pytest. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Reviewing test code
  • Debugging test failures
  • Improving test coverage

Example prompts

  • “/testing-python”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing Python loads about 1.3k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 248 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 248 words, ~1,341 tokens.

Download SKILL.mdSave it as .claude/skills/testing-python/SKILL.md (or your agent's skills folder).
name
testing-python
description
Write and evaluate effective Python tests using pytest. Use when writing tests, reviewing test code, debugging test failures, or improving test coverage. Covers test design, fixtures, parameterization, mocking, and async testing.

Writing Effective Python Tests

Core Principles

Every test should be atomic, self-contained, and test single functionality. A test that tests multiple things is harder to debug and maintain.

Test Structure

Atomic unit tests

Each test should verify a single behavior. The test name should tell you what's broken when it fails. Multiple assertions are fine when they all verify the same behavior.

python
# Good: Name tells you what's broken
def test_user_creation_sets_defaults():
    user = User(name="Alice")
    assert user.role == "member"
    assert user.id is not None
    assert user.created_at is not None

# Bad: If this fails, what behavior is broken?
def test_user():
    user = User(name="Alice")
    assert user.role == "member"
    user.promote()
    assert user.role == "admin"
    assert user.can_delete_others()
Use parameterization for variations of the same concept
python
import pytest

@pytest.mark.parametrize("input,expected", [
    ("hello", "HELLO"),
    ("World", "WORLD"),
    ("", ""),
    ("123", "123"),
])
def test_uppercase_conversion(input, expected):
    assert input.upper() == expected
Use separate tests for different functionality

Don't parameterize unrelated behaviors. If the test logic differs, write separate tests.

Project-Specific Rules

No async markers needed

This project uses asyncio_mode = "auto" globally. Write async tests without decorators:

python
# Correct
async def test_async_operation():
    result = await some_async_function()
    assert result == expected

# Wrong - don't add this
@pytest.mark.asyncio
async def test_async_operation():
    ...
Imports at module level

Put ALL imports at the top of the file:

python
# Correct
import pytest
from fastmcp import FastMCP
from fastmcp.client import Client

async def test_something():
    mcp = FastMCP("test")
    ...

# Wrong - no local imports
async def test_something():
    from fastmcp import FastMCP  # Don't do this
    ...
Use in-memory transport for testing

Pass FastMCP servers directly to clients:

python
from fastmcp import FastMCP
from fastmcp.client import Client

mcp = FastMCP("TestServer")

@mcp.tool
def greet(name: str) -> str:
    return f"Hello, {name}!"

async def test_greet_tool():
    async with Client(mcp) as client:
        result = await client.call_tool("greet", {"name": "World"})
        assert result[0].text == "Hello, World!"

Only use HTTP transport when explicitly testing network features.

Inline snapshots for complex data

Use inline-snapshot for testing JSON schemas and complex structures:

python
from inline_snapshot import snapshot

def test_schema_generation():
    schema = generate_schema(MyModel)
    assert schema == snapshot()  # Will auto-populate on first run

Commands:

  • pytest --inline-snapshot=create - populate empty snapshots
  • pytest --inline-snapshot=fix - update after intentional changes

Fixtures

Prefer function-scoped fixtures
python
@pytest.fixture
def client():
    return Client()

async def test_with_client(client):
    result = await client.ping()
    assert result is not None
Use tmp_path for file operations
python
def test_file_writing(tmp_path):
    file = tmp_path / "test.txt"
    file.write_text("content")
    assert file.read_text() == "content"

Mocking

Mock at the boundary
python
from unittest.mock import patch, AsyncMock

async def test_external_api_call():
    with patch("mymodule.external_client.fetch", new_callable=AsyncMock) as mock:
        mock.return_value = {"data": "test"}
        result = await my_function()
        assert result == {"data": "test"}
Don't mock what you own

Test your code with real implementations when possible. Mock external services, not internal classes.

Test Naming

Use descriptive names that explain the scenario:

python
# Good
def test_login_fails_with_invalid_password():
def test_user_can_update_own_profile():
def test_admin_can_delete_any_user():

# Bad
def test_login():
def test_update():
def test_delete():

Error Testing

python
import pytest

def test_raises_on_invalid_input():
    with pytest.raises(ValueError, match="must be positive"):
        calculate(-1)

async def test_async_raises():
    with pytest.raises(ConnectionError):
        await connect_to_invalid_host()

Running Tests

bash
uv run pytest -n auto              # Run all tests in parallel
uv run pytest -n auto -x           # Stop on first failure
uv run pytest path/to/test.py      # Run specific file
uv run pytest -k "test_name"       # Run tests matching pattern
uv run pytest -m "not integration" # Exclude integration tests

Checklist

Before submitting tests:

  • Each test tests one thing
  • No @pytest.mark.asyncio decorators
  • Imports at module level
  • Descriptive test names
  • Using in-memory transport (not HTTP) unless testing networking
  • Parameterization for variations of same behavior
  • Separate tests for different behaviors

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks/fix-build-agentops/environment/skills/testing-python of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Testing Python next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing Python compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing Python this skillbenchflow-ai/skillsbench1.8k—~1.3kAutomated safety check: PassApache-2.0
Squid Testing Pythoniusztinpaul/squid203—~1.3kAutomated safety check: PassApache-2.0
ONNX Runtime Test Runnermicrosoft/onnxruntime22k—~1.8kAutomated safety check: PassMIT
Test Coverage Reviewareed1192/finance-news-aggregator149—~2.6kAutomated safety check: PassMIT
ONNX Runtime GPU Transformers Testsmicrosoft/onnxruntime22k—~2.9kAutomated safety check: PassMIT
Temporal Python Testingwshobson/agents40k12 repos~1.2kAutomated safety check: PassMIT

Similar skills

  • Squid Testing Python

    iusztinpaul/squid

    Write and evaluate effective Python tests using pytest. An agent skill from iusztinpaul/squid.

    203 GitHub stars~1.3k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • ONNX Runtime Test Runner

    microsoft/onnxruntime

    Official

    Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Coverage Review

    areed1192/finance-news-aggregator

    Audit, plan, write, and verify unit tests for Python projects using pytest.

    149 GitHub stars~2.6k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Official

    Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

    22k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Temporal workflows with pytest, time-skipping, and mocking strategies.

    40k GitHub starsUsed in 12 repos~1.2k tokens
    Testing & QAAuto-check passed
  • Mutation Test Strength Audit

    buildfastwithai/gen-ai-experiments

    Measures how well a Python pytest suite catches behavior changes through diff-scoped mutation testing, then proposes and verifies tests for surviving mutants.

    785 GitHub stars~641 tokensUpdated 15 days ago
    Testing & QAAuto-check passed

More from benchflow-ai/skillsbench

All 178 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Categories

Questions about Testing Python

What does Testing Python do?

Write and evaluate effective Python tests using pytest. An agent skill from benchflow-ai/skillsbench. Testing Python is an agent skill from benchflow-ai/skillsbench. Write and evaluate effective Python tests using pytest.

When should I use Testing Python?

Testing Python fits situations like: reviewing test code; debugging test failures; improving test coverage.

How do I install Testing Python in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill testing-python -a claude-code`. Or copy the skill folder (tasks/fix-build-agentops/environment/skills/testing-python in benchflow-ai/skillsbench) into .claude/skills/testing-python in your project. Claude Code loads it when a task matches its description.

How do I install Testing Python in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill testing-python -a codex`. Or copy the skill folder (tasks/fix-build-agentops/environment/skills/testing-python in benchflow-ai/skillsbench) into .agents/skills/testing-python in your project. Codex loads it when a task matches its description.

Can I use Testing Python in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill testing-python -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-python, .gemini/skills/testing-python, .github/skills/testing-python and .opencode/skills/testing-python in your project.

What does Testing Python need to run?

Going by SKILL.md and its folder, Testing Python needs the command-line tools its instructions call (uv and pytest). Our summary lists: Python 3.

Does Testing Python access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Testing Python safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing Python use?

Testing Python is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing Python use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing Python?

Skills that share tags, products or a category with Testing Python: Squid Testing Python (iusztinpaul/squid, 203 stars), ONNX Runtime Test Runner (microsoft/onnxruntime, 22k stars), Test Coverage Review (areed1192/finance-news-aggregator, 149 stars) and ONNX Runtime GPU Transformers Tests (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing Python?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.