Agent skill

Python Testing

by macalbert in macalbert/envilder

Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks.

MITAuto-check passedTesting & QA

Install Python Testing

skills CLI
$ npx skills add macalbert/envilder --skill python-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install macalbert/envilder python-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/macalbert/envilder.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/python-testing .claude/skills/python-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
python-testing
GitHub stars
138
Token cost
~3.1k tokens
SKILL.md length
614 words
Files
3
Skills in repo
30
Repo updated
First seen
Licence
MIT

At a glance

Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks.

  • Works in 3 steps: AAA Pattern (Arrange – Act – Assert) → Test Naming Convention → Use in tests
  • E2E tests with pytest
  • SKILL.md covers Documentation Rules, Core Principles, Variable Naming (MANDATORY) and Async Testing (pytest-asyncio), plus 6 more sections
  • Calls pytest

What it does

Python Testing is an agent skill from macalbert/envilder. Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks. Use for unit, integration, or E2E tests with pytest, unittest, pytest-asyncio, or Playwright.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `examples.md` and `reference.md`).

It sits in Testing & QA, covering Unit testing, End-to-end testing and Technical documentation. It works with pytest, Python and Playwright. The repository describes itself as: One secret mapping for local dev, CI/CD, and runtime. Envilder resolves cloud secrets from your own vaults without SaaS middlemen, duplicated config, or .env drift. The licence is MIT.

When your agent uses it

  • E2E tests with pytest
  • Tasks that involve Unit testing
  • Tasks that involve End-to-end testing

Example prompts

  • “/python-testing”

Requirements

  • Python 3
  • Docker

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. AAA Pattern (Arrange – Act – Assert)
  2. Test Naming Convention
  3. Use in tests

What it can do on your machine

Read from SKILL.md and the folder at commit b6a0327. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Python Testing loads about 3.1k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 614 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from macalbert/envilder at commit b6a0327, republished under its MIT licence (© macalbert). 614 words, ~3,054 tokens.

Download SKILL.mdSave it as .claude/skills/python-testing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
python-testing
description
Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks. Use for unit, integration, or E2E tests with pytest, unittest, pytest-asyncio, or Playwright.

Testing Conventions (Python)

This skill defines the MANDATORY testing conventions for Python projects. These are rules, not guidelines.


Documentation Rules

NO Docstrings, NO Comments
  • Do NOT write docstrings - method/class names must be self-explanatory
  • Do NOT write comments except for AAA markers (# Arrange, # Act, # Assert)
  • The test name Should_X_When_Y already documents the intent

Exception: SDK public API: Code under src/sdks/*/ that is consumed by external developers (e.g. facade classes, public entry points) SHOULD have docstrings with usage examples. External users rely on IDE tooltips and help(). This exception does not apply to tests or internal helpers.

❌ FORBIDDEN
python
class TestUserService:
    
    def Should_CreateUser_When_Valid(self) -> None:
        # This creates a user  # NO explanatory comments
        user = UserFactory.build()
✅ CORRECT
python
class TestUserService:
    def Should_CreateUser_When_Valid(self) -> None:
        # Arrange
        user = UserFactory.build()
        
        # Act
        actual = self._sut.create(user)
        
        # Assert
        assert actual is not None

Core Principles

1. AAA Pattern (Arrange – Act – Assert)

ALL tests MUST follow the AAA pattern, separated by inline comments.

Rules
  • Each phase MUST be separated with comments
  • Never mix phases
  • Each comment (# Arrange, # Act, # Assert) appears AT MOST ONCE per test: if you need two actions or two asserts, write two tests
  • No if, switch, or conditional logic inside Arrange, Act, or Assert blocks
  • No try/catch/finally inside tests: use pytest fixtures with yield for teardown/cleanup
  • No # Act & Assert combined blocks: Act and Assert are ALWAYS separate
  • For exception testing, extract the action into a lambda before asserting
  • If no Arrange is needed, omit it
  • If there is no Assert, the test is invalid
pytest Example
python
def Should_CreateGroup_When_RequestIsValid(
    group_repository: Mock,
    sut: GroupService
) -> None:
    # Arrange
    request = CreateGroupRequest(
        name="Test Group",
        type=GroupType.RECURRING
    )
    group_repository.get_group_by_id.return_value = None

    # Act
    actual = sut.create_group(request)

    # Assert
    assert actual is not None
    assert actual.id is not None
    assert actual.name == "Test Group"
    group_repository.add_group.assert_called_once()
    group_repository.save.assert_called_once()

2. Test Naming Convention

Test names MUST follow exactly:

text
Should_{ExpectedBehavior}_When_{Condition}
Test Rules
  • PascalCase
  • NO natural language
  • NO vague names
  • NO missing When clause
  • test_ prefix is FORBIDDEN
pytest Discovery Configuration (MANDATORY)
toml
[tool.pytest.ini_options]
python_files = ["test_*.py"]
python_classes = ["Test*"]
python_functions = ["Should_*"]

If this config is missing → tests are wrong.


Variable Naming (MANDATORY)

PurposeName
Subject under testsut
Expected valueexpected
Actual resultactual

No creativity allowed here.


Async Testing (pytest-asyncio)

python
@pytest.mark.asyncio
async def Should_ReturnUser_When_UserExists(
    user_repository: AsyncMock,
    sut: GetUserHandler
) -> None:
    # Arrange
    expected = UserMother.create()
    user_repository.get_by_id.return_value = expected

    # Act
    actual = await sut.handle(expected.id)

    # Assert
    assert actual == expected
    user_repository.get_by_id.assert_awaited_once()

Exception Testing

Act and Assert MUST be separate. Extract the action into a lambda in the Act phase.

✅ CORRECT: separate Act and Assert
python
def Should_RaiseValueError_When_NameIsEmpty(sut: GroupService) -> None:
    # Arrange
    request = CreateGroupRequest(name="", type=GroupType.RECURRING)

    # Act
    action = lambda: sut.create_group(request)

    # Assert
    with pytest.raises(ValueError, match="Group name is required"):
        action()
❌ FORBIDDEN: combined Act & Assert
python
def Should_RaiseValueError_When_NameIsEmpty(sut: GroupService) -> None:
    # Arrange
    request = CreateGroupRequest(name="", type=GroupType.RECURRING)

    # Act & Assert   ← NEVER DO THIS
    with pytest.raises(ValueError, match="Group name is required"):
        sut.create_group(request)

Teardown & Cleanup (MANDATORY pattern)

Never use try/finally in tests. Use pytest fixtures with yield for cleanup.

✅ CORRECT: fixture with yield
python
@pytest.fixture()
def env_cleanup() -> Generator[list[str], None, None]:
    keys: list[str] = []
    yield keys
    for key in keys:
        os.environ.pop(key, None)


class TestEnvilderClient:
    def Should_SetEnvVars_When_InjectCalled(
        self, env_cleanup: list[str]
    ) -> None:
        # Arrange
        secrets = {"MY_TOKEN": "token-123"}
        env_cleanup.extend(secrets.keys())

        # Act
        EnvilderClient.inject_into_environment(secrets)

        # Assert
        assert os.environ["MY_TOKEN"] == "token-123"
❌ FORBIDDEN: try/finally in test
python
def Should_SetEnvVars_When_InjectCalled(self) -> None:
    secrets = {"MY_TOKEN": "token-123"}
    try:
        EnvilderClient.inject_into_environment(secrets)
        assert os.environ["MY_TOKEN"] == "token-123"
    finally:
        os.environ.pop("MY_TOKEN", None)

Mocking & Verification (OBLIGATORY)

If you mock something, you MUST verify it.

python
group_repository.add_group.assert_called_once()
group_repository.save.assert_called_once()
group_repository.delete_group.assert_not_called()

Async:

python
repository.save.assert_awaited_once()

No verification → test rejected.


Use Mother Pattern or Builder Pattern for creating test data. Both approaches are valid and recommended over inline object creation.

Mother Pattern
python
from dataclasses import dataclass
from uuid import UUID, uuid4
from typing import Optional


@dataclass
class Group:
    id: UUID
    name: str
    type: GroupType


class GroupMother:
    @staticmethod
    def create(
        id: Optional[UUID] = None,
        name: Optional[str] = None,
        type: Optional[GroupType] = None,
    ) -> Group:
        return Group(
            id=id or uuid4(),
            name=name or "Test Group",
            type=type or GroupType.RECURRING,
        )

Usage:

python
# Arrange
expected = GroupMother.create(name="Custom Name")
Show full SKILL.md (248 more words)Show less
Builder Pattern (polyfactory + shared Builder[T])

The shared test package provides a generic Builder[T] that wraps polyfactory to create type-safe builders for any Pydantic model. The with_* methods are generated dynamically via __getattr__.

Step 1: Define Factory + Builder
python
from polyfactory.factories.pydantic_factory import ModelFactory
from shared.factories import Builder


class GroupFactory(ModelFactory[Group]):
    __model__ = Group


class GroupBuilder(Builder[Group]):
    _factory = GroupFactory
Step 2: Use in tests
python
# Arrange - default random data
expected = GroupBuilder().build()

# Arrange - override specific fields
expected = GroupBuilder().with_name("Custom Name").with_type(GroupType.RECURRING).build()

# Arrange - build a batch
groups = GroupBuilder().with_type(GroupType.RECURRING).build_batch(5)

Anti-Patterns (PROHIBITED)

❌ Missing AAA
python
def Should_CreateGroup():
    sut = GroupService(Mock())
    sut.create_group(CreateGroupRequest(name="Test"))
❌ No mock verification
python
def Should_SaveGroup_When_Valid(sut: GroupService):
    sut.create_group(CreateGroupRequest(name="Test"))
    assert True
❌ Natural language / snake_case
python
def should_create_group_successfully():
    ...
❌ Combined Act & Assert
python
# Act & Assert   ← FORBIDDEN, always separate
with pytest.raises(ValueError):
    sut.do_something()
❌ try/catch/finally in tests
python
try:
    sut.inject(secrets)
    assert os.environ["KEY"] == "value"
finally:
    os.environ.pop("KEY", None)
❌ Conditional logic (if/switch) in Arrange, Act, or Assert
python
# Assert
if result is not None:   # ← FORBIDDEN, split into separate tests
    assert result.name == "Test"

Test Organization

Mirror Structure (MANDATORY)

Tests MUST mirror the production code structure using descriptive file naming.

Production code structure:

txt
src/apps/myapp/
├── lambda_handler.py
├── infrastructure/
│   ├── config.py
│   ├── container.py
│   └── logging/
│       ├── json_formatter.py
│       └── logger_factory.py
├── application/
│   └── handlers/
│       └── create_user.py
└── domain/
    └── entities/
        └── user.py

Test structure (hierarchical mirror):

txt
test/apps/myapp/
├── test_lambda_handler.py                    # mirrors lambda_handler.py
├── infrastructure/
│   ├── test_config.py                        # mirrors infrastructure/config.py
│   ├── test_container.py                     # mirrors infrastructure/container.py
│   └── logging/
│       ├── test_json_formatter.py            # mirrors infrastructure/logging/json_formatter.py
│       └── test_logger_factory.py            # mirrors infrastructure/logging/logger_factory.py
├── application/
│   └── handlers/
│       └── test_create_user.py               # mirrors application/handlers/create_user.py
└── domain/
    └── entities/
        └── test_user.py                      # mirrors domain/entities/user.py

Naming convention:

  • Format: {path}/test_{module}.py
  • Same folder structure as production code
  • Test files prefixed with test_
  • Exact mirror of production code hierarchy
Why hierarchical structure
  • Test files live in the same logical location as production code
  • Easy to find corresponding test file
  • Natural organization that mirrors the codebase structure
  • Clear one-to-one mapping
pytest Configuration

REQUIRED configuration in pyproject.toml:

toml
[tool.pytest.ini_options]
pythonpath = ["../../../src/apps/myapp"]  # Adjust path to your src directory
testpaths = ["."]
python_files = ["test_*.py"]
python_classes = ["Test*"]
python_functions = ["Should_*"]
asyncio_mode = "auto"
markers = [
    "acceptance: marks tests as acceptance tests (require Docker)",
    "unit: marks tests as unit tests (fast, no dependencies)",
    "integration: marks tests as integration tests",
]

⚠️ Important: The pythonpath must point to your production code directory to ensure imports work correctly from test files.

Test Classes (Optional)

Tests can be grouped in classes prefixed with Test*:

python
class TestProcessInvoiceHandler:
    async def Should_SaveRequest_When_ValidInput(self) -> None:
        # Arrange / Act / Assert
        ...
Test Markers

Use markers to categorize tests and run them selectively:

python
@pytest.mark.acceptance
class TestLambdaAcceptance:
    def Should_SaveToS3_When_LambdaInvoked(self) -> None:
        ...

@pytest.mark.unit
def Should_ValidateInput_When_EmptyName() -> None:
    ...

Run specific markers: pytest -m "not acceptance" or pytest -m unit


Final Summary

This skill enforces:

  • ✅ AAA pattern with explicit comments
  • ✅ Strict naming: Should_{ExpectedBehavior}_When_{Condition}
  • ✅ sut / actual / expected variables
  • ✅ Mock verification for all mocks
  • ✅ Mother or Builder pattern for test data (recommended)
  • ✅ Async support (pytest-asyncio)
  • ✅ Mirror structure with hierarchical organization
  • ✅ Enforceable pytest configuration

If a test doesn't follow this → it fails review.

© macalbert, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .github/skills/python-testing of macalbert/envilder.

  • SKILL.md
  • examples.md
  • reference.md

Open the folder on GitHubat commit b6a0327

Compare with similar skills

Python Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Python Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Python Testing this skillmacalbert/envilder138—~3.1kAutomated safety check: PassMIT
JS-in-HTML Testingliaohch3/claude-tap3.3k—~924Automated safety check: PassMIT
E2E TestingOpentrons/opentrons521—~3kAutomated safety check: NotesApache-2.0
Claude Code QAPramodDutta/qaskills232—~2.3kAutomated safety check: PassMIT
E2E Agent Browserjh941213/my-cc-harness126—~3.1kAutomated safety check: NotesNone
Testing Patternssoftspark/ai-toolkit179—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • JS-in-HTML Testing

    liaohch3/claude-tap

    Tests JavaScript embedded in an HTML file in two layers: pytest checks of the logic ported to Python, and Playwright runs in a real browser for the DOM.

    3.3k GitHub stars~924 tokensUpdated 15 days ago
    Testing & QAAuto-check passed
  • E2E Testing

    Opentrons/opentrons

    E2E testing conventions for Protocol Designer and Labware Library using Playwright + pytest in e2e-testing/.

    521 GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check: notes
  • Claude Code QA

    PramodDutta/qaskills

    The complete QA skill for Claude Code — turn Claude into an expert QA engineer that picks the right test type, writes reliable Playwright, Cypress, and pytest tests, eliminates flaky tests, enforces…

    232 GitHub stars~2.3k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • E2E Agent Browser

    jh941213/my-cc-harness

    E2E test automation using agent-browser CLI. An agent skill from jh941213/my-cc-harness.

    126 GitHub stars~3.1k tokensUpdated 2 mo ago
    Testing & QAAuto-check: notes
  • Testing Patterns

    softspark/ai-toolkit

    Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.

    179 GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Error Explanation Generator

    ArabelaTso/Skills-4-SE

    Explains test failures and provides actionable debugging guidance.

    253 GitHub stars~3.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from macalbert/envilder

All 30 skills in this repo
  • Code Review Perspectives

    macalbert/envilder

    Five independent analysis perspectives for code review: correctness, architecture, security, conventions, and complexity.

    138 GitHub stars~1k tokensUpdated 2 days ago
    Auto-check passed
  • Index of Architecture Decision Records (ADRs) for cross-cutting technical decisions.

    138 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Common Git

    macalbert/envilder

    Git commit messages, PR workflow, and branching strategy using Conventional Commits and Semantic Versioning.

    138 GitHub stars~991 tokensUpdated 2 days ago
    Auto-check passed
  • Common Testing Conventions

    macalbert/envilder

    Mandatory testing conventions including the narrow diagnostic exception for testing test-only code, AAA pattern, test naming, and assertions across all stacks (.NET, TypeScript, Python).

    138 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Doc Maintenance

    macalbert/envilder

    Workflow for maintaining changelogs, READMEs, and documentation files.

    138 GitHub stars~904 tokensUpdated 2 days ago
    Auto-check passed
  • Doc Sync

    macalbert/envilder

    Audit and synchronize documentation across website, READMEs, and docs/.

    138 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Python Testing

What does Python Testing do?

Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks. Python Testing is an agent skill from macalbert/envilder. Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks.

When should I use Python Testing?

Python Testing fits situations like: E2E tests with pytest; tasks that involve Unit testing; tasks that involve End-to-end testing.

How do I install Python Testing in Claude Code?

Run `npx skills add macalbert/envilder --skill python-testing -a claude-code`. Or copy the skill folder (.github/skills/python-testing in macalbert/envilder) into .claude/skills/python-testing in your project. Claude Code loads it when a task matches its description.

How do I install Python Testing in Codex?

Run `npx skills add macalbert/envilder --skill python-testing -a codex`. Or copy the skill folder (.github/skills/python-testing in macalbert/envilder) into .agents/skills/python-testing in your project. Codex loads it when a task matches its description.

Can I use Python Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add macalbert/envilder --skill python-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/python-testing, .gemini/skills/python-testing, .github/skills/python-testing and .opencode/skills/python-testing in your project.

What does Python Testing need to run?

Going by SKILL.md and its folder, Python Testing needs the command-line tools its instructions call (pytest). Our summary lists: Python 3; Docker.

Does Python Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Python Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Python Testing use?

Python Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Python Testing use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Python Testing?

Skills that share tags, products or a category with Python Testing: JS-in-HTML Testing (liaohch3/claude-tap, 3.3k stars), E2E Testing (Opentrons/opentrons, 521 stars), Claude Code QA (PramodDutta/qaskills, 232 stars) and E2E Agent Browser (jh941213/my-cc-harness, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Python Testing?

macalbert (a GitHub user) maintains it in macalbert/envilder, which has 138 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 5, 2026.

Source: macalbert/envilder on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.