---
name: testing
description: "How to write and run fast, deterministic, behavior-focused pytest tests that never touch real external systems. Use whenever adding or changing behavior, or fixing a bug."
generated: true
source: .ai/skills/testing.md
---

<!--
This file is generated by tools/skills.
Do not edit this file directly.
Edit .ai/skills/testing.md and run `make generate-ai`.
-->

# Skill: Testing

## Purpose

Define how to write and run tests for this Python project so that every
meaningful behavior is covered by fast, deterministic, behavior-focused tests
that never touch real external systems.

## When to Use This Skill

- Adding or changing any behavior (tests accompany behavior — almost always).
- Fixing a bug (add a regression test first).
- Working on the test suite or quality gate (task `0012`).
- Reviewing whether a change is adequately tested.

## Files to Read First

- `AGENTS.md`
- `docs/DEVELOPMENT_WORKFLOW.md`
- `pyproject.toml` (`[tool.pytest.ini_options]`)
- The skill for the area under test (runtime, prompts, plugins, tools, etc.).

## Core Principles

- **Use `pytest`.** Tests live under `tests/`, mirroring `src/agentplatform/`.
- **Test behavior through public interfaces**, not private implementation
  details (test internals only when there is no public seam).
- **No real external systems in unit tests:** never call real LLMs, real MCP
  servers, real databases, or third-party APIs. Mock/fake them.
- **Integration tests only when explicitly requested**, and still without real
  secrets or live third-party calls — use fakes/local stand-ins.
- **Deterministic:** control time, randomness, and ordering.
- **Regression tests for bugs:** every fixed bug gets a test that fails before
  the fix.

## Process

1. **Pick the category** (see below) for what you are testing.
2. **Arrange with fixtures** for repeated setup (specs, compiled graphs, fake
   resolver/access plugins, fake tools). Use `@pytest.mark.parametrize` for input variations.
3. **Mock external boundaries** (LLM/MCP/DB/HTTP) at the adapter seam, not deep
   inside.
4. **Write behavior assertions:** given input, assert observable output/effects.
5. **Cover negatives and security:** invalid YAML/validation errors, missing
   required prompt variables, denied protected-node access, plugin errors.
6. **Run** `make test` (and `make check` for the full gate). Fix failures.

## Testing categories (use the right one)

- **Unit tests** — a single function/class via its public interface.
- **Integration tests** — multiple layers together (e.g. validate → compile →
  run) with fakes; only when requested.
- **Contract tests** — verify plugin contracts and the YAML schema match
  `docs/`; catch drift early.
- **Golden example tests** — a known YAML config produces a known compiled
  graph / rendered prompt; update goldens deliberately.
- **Negative tests** — invalid specs, missing variables, denied access must
  fail clearly with the right error.
- **Security tests** — protected access fail-closed, secrets redacted, no
  request-state leakage.

## Example pytest structure (illustrative, not product tests)

```python
import pytest

@pytest.fixture
def valid_spec_dict() -> dict:
    return {"version": "1.0", "app": {"name": "demo"}, "...": "..."}

@pytest.mark.parametrize("missing_key", ["version", "app", "runtime"])
def test_validation_reports_missing_top_level_key(valid_spec_dict, missing_key):
    data = dict(valid_spec_dict)
    del data[missing_key]
    result = validate_spec(data)            # public interface
    assert not result.ok
    assert any(missing_key in e.message for e in result.errors)
```

> Do not implement real product tests in this task. Only trivial repository
> smoke tests (e.g. import + version) belong here until task `0001`+ create code.

## Checklist Before Finishing

- [ ] New/changed behavior has tests; bugs have regression tests.
- [ ] Tests target public interfaces and assert behavior, not internals.
- [ ] External systems are mocked/faked; no real LLM/MCP/DB/API calls.
- [ ] Negative and security cases covered where relevant.
- [ ] Tests are deterministic (time/randomness/order controlled).
- [ ] Fixtures/parametrization used to keep tests readable and DRY.
- [ ] `make test` (and `make check`) pass, or you state why they can't run.

## Common Mistakes to Avoid

- Calling real LLMs/MCP/DBs/APIs in unit tests.
- Asserting on private attributes/log strings instead of behavior.
- Flaky tests depending on real time, randomness, or ordering.
- Over-mocking until the test asserts nothing meaningful.
- Skipping negative/security tests because the happy path passes.
- Committing secrets or production endpoints in fixtures.

## Expected Final Report

State: which test files/categories were added or changed; what behaviors and
edge/negative/security cases are now covered; mocks/fakes used for external
boundaries; the result of `make test` / `make check`; and any coverage gaps
intentionally left for a later task.
