Agent skill

Testing

by extra-org in extra-org/extra

How to write and run fast, deterministic, behavior-focused pytest tests that never touch real external systems.

MITAuto-check passedTesting & QA

Install Testing

skills CLI
$ npx skills add extra-org/extra --skill testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install extra-org/extra testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/extra-org/extra.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/testing .claude/skills/testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing
GitHub stars
113
Token cost
~1.2k tokens
SKILL.md length
530 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

How to write and run fast, deterministic, behavior-focused pytest tests that never touch real external systems.

  • Works in 6 steps: Pick the category (see below) for what… → Arrange with fixtures for repeated setup… → Mock external boundaries… → …
  • Changing behavior
  • SKILL.md covers Purpose, When to Use This Skill, Files to Read First and Core Principles, plus 6 more sections
  • Calls make

What it does

Testing is an agent skill from extra-org/extra. How to write and run fast, deterministic, behavior-focused pytest tests that never touch real external systems. Use whenever adding or changing behavior, or fixing a bug.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Unit testing. It works with pytest. The repository describes itself as: Turn your product into an AI-powered assistant. The licence is MIT.

When your agent uses it

  • Changing behavior
  • Tasks that involve Unit testing

Example prompts

  • “/testing”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Pick the category (see below) for what you are testing.
  2. Arrange with fixtures for repeated setup (specs, compiled graphs, fake
  3. Mock external boundaries (LLM/MCP/DB/HTTP) at the adapter seam, not deep
  4. Write behavior assertions: given input, assert observable output/effects.
  5. Cover negatives and security: invalid YAML/validation errors, missing
  6. Run make test (and make check for the full gate). Fix failures.

What it can do on your machine

Read from SKILL.md and the folder at commit 023b0a0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing loads about 1.2k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 530 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from extra-org/extra at commit 023b0a0, republished under its MIT licence (© extra-org). 530 words, ~1,249 tokens.

Download SKILL.mdSave it as .claude/skills/testing/SKILL.md (or your agent's skills folder).
name
testing
description
How to write and run fast, deterministic, behavior-focused pytest tests that never touch real external systems. Use whenever adding or changing behavior, or fixing a bug.
generated
true
source
.ai/skills/testing.md
<!--
This file is generated by tools/skills.
Do not edit this file directly.
Edit .ai/skills/testing.md and run `make generate-ai`.
-->

Skill: Testing

Purpose

Define how to write and run tests for this Python project so that every meaningful behavior is covered by fast, deterministic, behavior-focused tests that never touch real external systems.

When to Use This Skill

  • Adding or changing any behavior (tests accompany behavior — almost always).
  • Fixing a bug (add a regression test first).
  • Working on the test suite or quality gate (task 0012).
  • Reviewing whether a change is adequately tested.

Files to Read First

  • AGENTS.md
  • docs/DEVELOPMENT_WORKFLOW.md
  • pyproject.toml ([tool.pytest.ini_options])
  • The skill for the area under test (runtime, prompts, plugins, tools, etc.).

Core Principles

  • Use pytest. Tests live under tests/, mirroring src/agentplatform/.
  • Test behavior through public interfaces, not private implementation details (test internals only when there is no public seam).
  • No real external systems in unit tests: never call real LLMs, real MCP servers, real databases, or third-party APIs. Mock/fake them.
  • Integration tests only when explicitly requested, and still without real secrets or live third-party calls — use fakes/local stand-ins.
  • Deterministic: control time, randomness, and ordering.
  • Regression tests for bugs: every fixed bug gets a test that fails before the fix.

Process

  1. Pick the category (see below) for what you are testing.
  2. Arrange with fixtures for repeated setup (specs, compiled graphs, fake resolver/access plugins, fake tools). Use @pytest.mark.parametrize for input variations.
  3. Mock external boundaries (LLM/MCP/DB/HTTP) at the adapter seam, not deep inside.
  4. Write behavior assertions: given input, assert observable output/effects.
  5. Cover negatives and security: invalid YAML/validation errors, missing required prompt variables, denied protected-node access, plugin errors.
  6. Run make test (and make check for the full gate). Fix failures.

Testing categories (use the right one)

  • Unit tests — a single function/class via its public interface.
  • Integration tests — multiple layers together (e.g. validate → compile → run) with fakes; only when requested.
  • Contract tests — verify plugin contracts and the YAML schema match docs/; catch drift early.
  • Golden example tests — a known YAML config produces a known compiled graph / rendered prompt; update goldens deliberately.
  • Negative tests — invalid specs, missing variables, denied access must fail clearly with the right error.
  • Security tests — protected access fail-closed, secrets redacted, no request-state leakage.
Show full SKILL.md (182 more words)Show less

Example pytest structure (illustrative, not product tests)

python
import pytest

@pytest.fixture
def valid_spec_dict() -> dict:
    return {"version": "1.0", "app": {"name": "demo"}, "...": "..."}

@pytest.mark.parametrize("missing_key", ["version", "app", "runtime"])
def test_validation_reports_missing_top_level_key(valid_spec_dict, missing_key):
    data = dict(valid_spec_dict)
    del data[missing_key]
    result = validate_spec(data)            # public interface
    assert not result.ok
    assert any(missing_key in e.message for e in result.errors)

Do not implement real product tests in this task. Only trivial repository smoke tests (e.g. import + version) belong here until task 0001+ create code.

Checklist Before Finishing

  • New/changed behavior has tests; bugs have regression tests.
  • Tests target public interfaces and assert behavior, not internals.
  • External systems are mocked/faked; no real LLM/MCP/DB/API calls.
  • Negative and security cases covered where relevant.
  • Tests are deterministic (time/randomness/order controlled).
  • Fixtures/parametrization used to keep tests readable and DRY.
  • make test (and make check) pass, or you state why they can't run.

Common Mistakes to Avoid

  • Calling real LLMs/MCP/DBs/APIs in unit tests.
  • Asserting on private attributes/log strings instead of behavior.
  • Flaky tests depending on real time, randomness, or ordering.
  • Over-mocking until the test asserts nothing meaningful.
  • Skipping negative/security tests because the happy path passes.
  • Committing secrets or production endpoints in fixtures.

Expected Final Report

State: which test files/categories were added or changed; what behaviors and edge/negative/security cases are now covered; mocks/fakes used for external boundaries; the result of make test / make check; and any coverage gaps intentionally left for a later task.

© extra-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/testing of extra-org/extra.

Open the folder on GitHubat commit 023b0a0

Compare with similar skills

Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing this skillextra-org/extra113—~1.2kAutomated safety check: PassMIT
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence
Test GuardamElnagdy/guard-skills1.3k2 repos~2.1kAutomated safety check: PassMIT
Fla Optimization Loopfla-org/flash-linear-attention5.8k—~2.6kAutomated safety check: PassMIT
Pytest Runnersaleor/saleor23k—~251Automated safety check: PassBSD-3-Clause

Similar skills

  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Guard

    amElnagdy/guard-skills

    Reviews newly written or edited tests against nine rules that cut test bloat, such as mock-heavy checks and near-duplicate cases, before they are committed.

    1.3k GitHub starsUsed in 2 repos~2.1k tokens
    Testing & QAAuto-check passed
  • Fla Optimization Loop

    fla-org/flash-linear-attention

    Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness.

    5.8k GitHub stars~2.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Pytest Runner

    saleor/saleor

    Run pytest tests with automatic virtual environment activation. Use this skill whenever running tests, executing pytest, or when asked to "run tests", "test…

    23k GitHub stars~251 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Port Node Red Node

    oldrev/edgelinkd

    Port a Node-RED node into EdgeLinkd the way this repo does it: implement the node in Rust under crates/core/src/runtime/nodes, mirror Node-RED's mocha spec as pytest tests under tests/, register the…

    121 GitHub stars~3k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from extra-org/extra

All 14 skills in this repo
  • Architecture Review

    extra-org/extra

    Guard the project's architecture invariants. An agent skill from extra-org/extra.

    113 GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Documentation

    extra-org/extra

    Keep the repository's documentation accurate, honest, and synchronized with the code.

    113 GitHub stars~848 tokensUpdated 1 mo ago
    Auto-check passed
  • The standard for writing Python here — small, typed, explicit, testable modules with side effects pushed to the edges.

    113 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Skill Authoring

    extra-org/extra

    How to create or restructure a skill so the .ai/skills/ system stays consistent, operational, and trustworthy.

    113 GitHub stars~1k tokensUpdated 1 mo ago
    Auto-check passed
  • Code Review

    extra-org/extra

    Senior-level code review for the agent platform — architecture and boundaries first, then correctness, security, testability, and simplicity.

    113 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Git Workflow

    extra-org/extra

    Work safely with Git — keep changes scoped and reviewable, never overwrite the user's work, and produce clear history.

    113 GitHub stars~781 tokensUpdated 1 mo ago
    Auto-check: notes

Works with

Categories

Questions about Testing

What does Testing do?

How to write and run fast, deterministic, behavior-focused pytest tests that never touch real external systems. Testing is an agent skill from extra-org/extra. How to write and run fast, deterministic, behavior-focused pytest tests that never touch real external systems.

When should I use Testing?

Testing fits situations like: changing behavior; tasks that involve Unit testing.

How do I install Testing in Claude Code?

Run `npx skills add extra-org/extra --skill testing -a claude-code`. Or copy the skill folder (.claude/skills/testing in extra-org/extra) into .claude/skills/testing in your project. Claude Code loads it when a task matches its description.

How do I install Testing in Codex?

Run `npx skills add extra-org/extra --skill testing -a codex`. Or copy the skill folder (.claude/skills/testing in extra-org/extra) into .agents/skills/testing in your project. Codex loads it when a task matches its description.

Can I use Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add extra-org/extra --skill testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing, .gemini/skills/testing, .github/skills/testing and .opencode/skills/testing in your project.

What does Testing need to run?

Going by SKILL.md and its folder, Testing needs the command-line tools its instructions call (make). Our summary lists: Python 3.

Does Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing use?

Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing use?

About 1.2k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing?

Skills that share tags, products or a category with Testing: Adk Verify Snippets (google/adk-python, 22k stars), Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars), Test Guard (amElnagdy/guard-skills, 1.3k stars) and Fla Optimization Loop (fla-org/flash-linear-attention, 5.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing?

extra-org (a GitHub organization) maintains it in extra-org/extra, which has 113 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on September 7, 2026.

Source: extra-org/extra on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.