---
name: test-design
description: Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and agreement with NumPy — with expected values generated by numpy (python/gen_*.py) instead of written by hand. Use this skill whenever tests are written or extended in Matft: implementing a new function or fixing a bug test-first (TDD), adding coverage for an existing function, reviewing whether tests are sufficient, or turning a reported numpy mismatch into a regression test — e.g. "write tests for X", "add test cases", "improve coverage", "is this tested enough?", "check it matches numpy", or in Japanese「テスト書いて」「テストケース追加して」「網羅性を上げて」「境界値のテスト」「型ごとのテスト」「Numpy と一致するか確認して」「TDD で実装して」— even if the word "skill" is never mentioned. For benchmark runs and the docs performance table use the benchmark skill; for image visual checks use image-visual-check.
---

# Designing Matft tests

Matft promises "behaves like NumPy". Most of the bugs found so far were not in the main path but at the edges:
integer wraparound, zero-length dimensions corrupting the heap, slice views reading past their data,
complex imaginary parts copied from the real part, vDSP behaving differently on x86_64, NaN dropped by a kernel.
A test that only checks `f([1, 2, 3])` in `.Float` catches none of these.
This skill turns "write tests" into a systematic pass over the viewpoints where Matft actually breaks,
with every expected value coming from numpy so the test cannot encode a misunderstanding.

TDD is required in this repository (CLAUDE.md): write the tests, watch them fail, then implement.

## Workflow

### 1. Pin down the numpy specification

Before writing anything, run the numpy counterpart in Python and read its docs for the target function:
signature and defaults, the output dtype rule, the output shape for each `axis`/`keepdims`,
what it does with NaN, inf, empty input, negative axes, out-of-range parameters.
Also note where Matft deliberately differs (see "Matft conventions" below) so you don't report them as bugs,
and check `Sources/` for the existing signature if the function exists.

Use the venv from `python/README.md` (`.venv/bin/python`). If it does not exist, create it as the README says.

### 2. Make a test plan over the viewpoints

Read `references/viewpoints.md` and go through every viewpoint for the target. For each one, decide
**cover** (list the concrete cases) or **N/A** (with a one-line reason, e.g. "no axis parameter").
Write the plan as a short table in your reply before writing code. The table is what makes the coverage
reviewable: a missing row is visible, an unexplained gap is not.

| Viewpoint | Cases | Notes |
|---|---|---|
| Values / boundaries | 0, ±1, negatives, ties, max/min of each int type | |
| dtype | Float, Double, Int, UInt8, Bool, Complex | output dtype == np.result_type |
| ... | ... | |

Size the plan to the change: a new public function gets the full pass; a one-line bug fix gets the failing
case plus the neighbouring viewpoints that the same code path touches (e.g. a fix in a reduction kernel → axes, layouts, dtypes).

### 3. Generate the expected values with numpy

Expected values are never computed by hand or copied from Matft's own output — that would make the test
agree with the bug. Choose one of:

- **Generated test file (default for anything with more than a handful of cases).** Add cases to an existing
  generator (`python/gen_numpy_gaps_coverage.py` for numpy-level functions, `gen_fft_audio_coverage.py`,
  `gen_image_coverage.py`), or create `python/gen_<area>_coverage.py` from `assets/gen_template.py`.
  Each case writes the Swift expression and the numpy expression side by side, so a reviewer can check them
  against each other. The output file starts with `// Generated by ... Do not edit by hand.`
  Register a new generator in the table of `python/README.md`. Details: `references/generator.md`.
- **Hand-written XCTest** for things a generator expresses badly: in-place mutation and aliasing, thrown errors,
  view identity, or a single regression case. Still compute the numbers in Python and paste them with the
  numpy expression as a comment (`// numpy: np.diff(a, n=4, axis=1).shape -> (3, 0)`).

Compare arrays with `XCTAssertClose` (= `np.testing.assert_allclose`, NaN matches NaN, inf must match exactly)
and pass `checkType: true` — a wrong output dtype is a numpy mismatch too. For exact Int / Bool results use
`rtol: 0, atol: 0`. Avoid `XCTAssertEqual(MfArray, MfArray)`: `==` ignores the mftype, treats floats within 1e-5
as equal, wraps integers into the type (UInt8 index 299 "equals" 43) and never matches NaN, so it hides exactly
the bugs this skill is looking for. `XCTAssertEqual` is fine for shapes and Swift scalars. Tolerances: Float `rtol/atol ≈ 1e-5..1e-6`, Double `1e-10..1e-12`;
loosen only with a comment explaining why (e.g. accumulated FFT error).

### 4. Red

Run only the new tests and confirm they fail for the expected reason (not a compile error in the test itself):

```sh
swift test --filter MatftTests.<ClassName>
```

If a case unexpectedly passes before the implementation, it is not testing the change — tighten it.
If an existing function fails a new case, you found a bug: keep the test, fix the bug in the same PR (TDD),
and mention it in the PR description.

### 5. Green, then the full suite

Implement the minimum to pass, then run `swift test` (all). Then check the platform viewpoints that apply
(`references/viewpoints.md` §Platform): x86_64 for new vDSP usage, WASI for code with a fallback path.

### 6. Performance (hot paths only)

Add a performance case only when the function is a hot path: an elementwise kernel, reduction, sort/search,
indexing/setter, conversion, linear algebra, FFT, or anything that scales with a large array. Skip it for
small helpers, creation of small arrays, and error handling — and say so in the plan.
Procedure (`references/viewpoints.md` §Performance):
add a `measureWithWarmup` test in `Tests/PerformanceTests/<Area>PefTests.swift` using `PerfFixtures`,
and register the same expression in `CASES` of `scripts/benchmark.py` with its numpy counterpart.
Include a non-contiguous (transposed) input variant if the kernel has a separate strided path.
Measuring and reporting the numbers is the benchmark skill's job.

### 7. Report

End with the plan table updated to what was actually covered, the test files touched, the Red → Green
results (counts of failures before and passes after), and any bugs found or cases left N/A.

## Matft conventions (not bugs)

These differ from numpy on purpose. Write the expected value in Matft's convention and say so in a comment.

- A reduction over all axes returns shape `[1]`, not a 0-d scalar (the generator helpers convert 0-d to `[1]`).
- Every type except `.Double` is stored as Float32: integers above 2^24 lose precision unless `.Double`;
  Bool is stored as 0/1 Float. Integer results out of range wrap like numpy's fixed-width ints.
- `MfType.result_type` follows numpy for array–array integer promotion; otherwise the higher `priority` wins
  (integers do not widen when combined with Float).
- Invalid arguments usually hit `precondition` (a crash), which XCTest cannot catch. Only test errors of APIs
  that `throws`; for precondition paths, note the intended behavior in the plan instead of testing it.

## Helpers you should reuse (Tests/MatftTests/TestHelpers.swift)

- `XCTAssertClose(actual, expected, rtol:, atol:, checkType:)` — assert_allclose with the worst index in the message.
- `layoutVariants(a)` — the same logical array as row/column major, offset view, prefix view, strided view,
  reversed view and transposed view. Loop over it: `for (name, x) in layoutVariants(A) { ... name }`.
- `rowValues(x)` — values in row-major order as `[Double]`.
