Agent skill

Test Design

by jjjkkkjjj in jjjkkkjjj/Matft

Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and…

BSD-3-ClauseAuto-check passedTesting & QA

Install Test Design

skills CLI
$ npx skills add jjjkkkjjj/Matft --skill test-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jjjkkkjjj/Matft test-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jjjkkkjjj/Matft.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-design .claude/skills/test-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-design
GitHub stars
147
Token cost
~2k tokens
SKILL.md length
1,011 words
Files
5 (incl. references, assets)
Skills in repo
5
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and…

  • Works in 7 steps: Pin down the numpy specification → Make a test plan over the viewpoints → Generate the expected values with numpy → …
  • Tests are written
  • SKILL.md covers Workflow, Matft conventions (not bugs) and Helpers you should reuse…
  • Runs Python scripts from its folder; calls swift

What it does

Test Design is an agent skill from jjjkkkjjj/Matft. Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and agreement with NumPy — with expected values generated by numpy (python/gen.py) instead of written by hand. Use this skill whenever tests are written or extended in Matft: implementing a new function or fixing a bug test-first (TDD), adding coverage for an existing function, reviewing whether tests are sufficient, or…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files and assets (for example `assets/gen_template.py`, `evals/evals.json` and `references/generator.md`).

It sits in Testing & QA, covering Test generation and Test-driven development. It works with NumPy and Python. The repository describes itself as: Numpy-like library in swift. (Multi-dimensional Array, ndarray, matrix and vector library). The licence is BSD-3-Clause.

When your agent uses it

  • Tests are written
  • Extended in Matft: implementing a new function
  • Fixing a bug test-first (TDD)
  • Adding coverage for an existing function

Example prompts

  • “write tests for X”
  • “add test cases”
  • “improve coverage”
  • “/test-design”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Pin down the numpy specification
  2. Make a test plan over the viewpoints
  3. Generate the expected values with numpy
  4. Red
  5. Green, then the full suite
  6. Performance (hot paths only)
  7. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 618dcfc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • swift

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Design loads about 2k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 240 tokens; SKILL.md has 1,011 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~240
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jjjkkkjjj/Matft at commit 618dcfc, republished under its BSD-3-Clause licence (© jjjkkkjjj). 1,011 words, ~1,977 tokens.

Download SKILL.mdSave it as .claude/skills/test-design/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
test-design
description
Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and agreement with NumPy — with expected values generated by numpy (python/gen_*.py) instead of written by hand. Use this skill whenever tests are written or extended in Matft: implementing a new function or fixing a bug test-first (TDD), adding coverage for an existing function, reviewing whether tests are sufficient, or turning a reported numpy mismatch into a regression test — e.g. "write tests for X", "add test cases", "improve coverage", "is this tested enough?", "check it matches numpy", or in Japanese「テスト書いて」「テストケース追加して」「網羅性を上げて」「境界値のテスト」「型ごとのテスト」「Numpy と一致するか確認して」「TDD で実装して」— even if the word "skill" is never mentioned. For benchmark runs and the docs performance table use the benchmark skill; for image visual checks use image-visual-check.

Designing Matft tests

Matft promises "behaves like NumPy". Most of the bugs found so far were not in the main path but at the edges: integer wraparound, zero-length dimensions corrupting the heap, slice views reading past their data, complex imaginary parts copied from the real part, vDSP behaving differently on x86_64, NaN dropped by a kernel. A test that only checks f([1, 2, 3]) in .Float catches none of these. This skill turns "write tests" into a systematic pass over the viewpoints where Matft actually breaks, with every expected value coming from numpy so the test cannot encode a misunderstanding.

TDD is required in this repository (CLAUDE.md): write the tests, watch them fail, then implement.

Workflow

1. Pin down the numpy specification

Before writing anything, run the numpy counterpart in Python and read its docs for the target function: signature and defaults, the output dtype rule, the output shape for each axis/keepdims, what it does with NaN, inf, empty input, negative axes, out-of-range parameters. Also note where Matft deliberately differs (see "Matft conventions" below) so you don't report them as bugs, and check Sources/ for the existing signature if the function exists.

Use the venv from python/README.md (.venv/bin/python). If it does not exist, create it as the README says.

2. Make a test plan over the viewpoints

Read references/viewpoints.md and go through every viewpoint for the target. For each one, decide cover (list the concrete cases) or N/A (with a one-line reason, e.g. "no axis parameter"). Write the plan as a short table in your reply before writing code. The table is what makes the coverage reviewable: a missing row is visible, an unexplained gap is not.

ViewpointCasesNotes
Values / boundaries0, ±1, negatives, ties, max/min of each int type
dtypeFloat, Double, Int, UInt8, Bool, Complexoutput dtype == np.result_type
......

Size the plan to the change: a new public function gets the full pass; a one-line bug fix gets the failing case plus the neighbouring viewpoints that the same code path touches (e.g. a fix in a reduction kernel → axes, layouts, dtypes).

3. Generate the expected values with numpy

Expected values are never computed by hand or copied from Matft's own output — that would make the test agree with the bug. Choose one of:

  • Generated test file (default for anything with more than a handful of cases). Add cases to an existing generator (python/gen_numpy_gaps_coverage.py for numpy-level functions, gen_fft_audio_coverage.py, gen_image_coverage.py), or create python/gen_<area>_coverage.py from assets/gen_template.py. Each case writes the Swift expression and the numpy expression side by side, so a reviewer can check them against each other. The output file starts with // Generated by ... Do not edit by hand. Register a new generator in the table of python/README.md. Details: references/generator.md.
  • Hand-written XCTest for things a generator expresses badly: in-place mutation and aliasing, thrown errors, view identity, or a single regression case. Still compute the numbers in Python and paste them with the numpy expression as a comment (// numpy: np.diff(a, n=4, axis=1).shape -> (3, 0)).

Compare arrays with XCTAssertClose (= np.testing.assert_allclose, NaN matches NaN, inf must match exactly) and pass checkType: true — a wrong output dtype is a numpy mismatch too. For exact Int / Bool results use rtol: 0, atol: 0. Avoid XCTAssertEqual(MfArray, MfArray): == ignores the mftype, treats floats within 1e-5 as equal, wraps integers into the type (UInt8 index 299 "equals" 43) and never matches NaN, so it hides exactly the bugs this skill is looking for. XCTAssertEqual is fine for shapes and Swift scalars. Tolerances: Float rtol/atol ≈ 1e-5..1e-6, Double 1e-10..1e-12; loosen only with a comment explaining why (e.g. accumulated FFT error).

Show full SKILL.md (421 more words)Show less
4. Red

Run only the new tests and confirm they fail for the expected reason (not a compile error in the test itself):

sh
swift test --filter MatftTests.<ClassName>

If a case unexpectedly passes before the implementation, it is not testing the change — tighten it. If an existing function fails a new case, you found a bug: keep the test, fix the bug in the same PR (TDD), and mention it in the PR description.

5. Green, then the full suite

Implement the minimum to pass, then run swift test (all). Then check the platform viewpoints that apply (references/viewpoints.md §Platform): x86_64 for new vDSP usage, WASI for code with a fallback path.

6. Performance (hot paths only)

Add a performance case only when the function is a hot path: an elementwise kernel, reduction, sort/search, indexing/setter, conversion, linear algebra, FFT, or anything that scales with a large array. Skip it for small helpers, creation of small arrays, and error handling — and say so in the plan. Procedure (references/viewpoints.md §Performance): add a measureWithWarmup test in Tests/PerformanceTests/<Area>PefTests.swift using PerfFixtures, and register the same expression in CASES of scripts/benchmark.py with its numpy counterpart. Include a non-contiguous (transposed) input variant if the kernel has a separate strided path. Measuring and reporting the numbers is the benchmark skill's job.

7. Report

End with the plan table updated to what was actually covered, the test files touched, the Red → Green results (counts of failures before and passes after), and any bugs found or cases left N/A.

Matft conventions (not bugs)

These differ from numpy on purpose. Write the expected value in Matft's convention and say so in a comment.

  • A reduction over all axes returns shape [1], not a 0-d scalar (the generator helpers convert 0-d to [1]).
  • Every type except .Double is stored as Float32: integers above 2^24 lose precision unless .Double; Bool is stored as 0/1 Float. Integer results out of range wrap like numpy's fixed-width ints.
  • MfType.result_type follows numpy for array–array integer promotion; otherwise the higher priority wins (integers do not widen when combined with Float).
  • Invalid arguments usually hit precondition (a crash), which XCTest cannot catch. Only test errors of APIs that throws; for precondition paths, note the intended behavior in the plan instead of testing it.

Helpers you should reuse (Tests/MatftTests/TestHelpers.swift)

  • XCTAssertClose(actual, expected, rtol:, atol:, checkType:) — assert_allclose with the worst index in the message.
  • layoutVariants(a) — the same logical array as row/column major, offset view, prefix view, strided view, reversed view and transposed view. Loop over it: for (name, x) in layoutVariants(A) { ... name }.
  • rowValues(x) — values in row-major order as [Double].

© jjjkkkjjj, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references, assets) in .claude/skills/test-design of jjjkkkjjj/Matft.

  • SKILL.md
  • assets/gen_template.py
  • evals/evals.json
  • references/generator.md
  • references/viewpoints.md

Open the folder on GitHubat commit 618dcfc

Compare with similar skills

Test Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Design this skilljjjkkkjjj/Matft147—~2kAutomated safety check: PassBSD-3-Clause
TDD Guidealirezarezvani/claude-skills28k—~3.4kAutomated safety check: PassMIT
TDD GuideLeoYeAI/openclaw-master-skills2.2k—~1.4kAutomated safety check: PassMIT
TDD GuideaAAaqwq/AGI-Super-Team1052 repos~1.1kAutomated safety check: PassMIT
Python Testingathola/claude-night-market341—~800Automated safety check: PassMIT
Light Experiment CodingLight0305/Light-skills640—~2.3kAutomated safety check: PassMIT

Similar skills

  • TDD Guide

    alirezarezvani/claude-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • TDD Guide

    LeoYeAI/openclaw-master-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    2.2k GitHub stars~1.4k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • TDD Guide

    aAAaqwq/AGI-Super-Team

    Test-driven development workflow with test generation, coverage analysis, and multi-framework support

    105 GitHub starsUsed in 2 repos~1.1k tokens
    Testing & QAAuto-check passed
  • Python Testing

    athola/claude-night-market

    Python testing patterns with pytest, fixtures, TDD, mocking, async and integration tests.

    341 GitHub stars~800 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Light Experiment Coding

    Light0305/Light-skills

    Builds the code for a frozen research experiment test-first, with leakage controls, seed handling and saved evidence so results can be rerun and audited.

    640 GitHub stars~2.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed

More from jjjkkkjjj/Matft

  • Docs

    jjjkkkjjj/Matft

    Procedure for writing and updating Matft's documentation (the Docusaurus site in website/ and the doc comments on the public API that become the Swift-DocC API reference).

    147 GitHub stars~2.7k tokensUpdated 14 days ago
    Auto-check passed
  • Benchmark

    jjjkkkjjj/Matft

    Procedure for benchmarking Matft's PerformanceTests against Numpy and reporting the results (and, when asked, updating the speed comparison table on the docs site, website/docs/performance.md).

    147 GitHub stars~1.7k tokensUpdated 14 days ago
    Auto-check passed
  • Image Visual Check

    jjjkkkjjj/Matft

    Procedure for adding tests for Matft's image processing (Matft.image., indexing or channel swapping on images, etc.), generating comparison images that put the result next to an OpenCV reference…

    147 GitHub stars~2.3k tokensUpdated 14 days ago
    Auto-check passed
  • Release

    jjjkkkjjj/Matft

    Procedure for releasing a new version of Matft (decide the version → check tests → write release notes → create and push an annotated tag → publish a GitHub Release).

    147 GitHub stars~1.8k tokensUpdated 14 days ago
    Auto-check passed

Works with

Categories

Questions about Test Design

What does Test Design do?

Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and…. Test Design is an agent skill from jjjkkkjjj/Matft.py) instead of written by hand.

When should I use Test Design?

Test Design fits situations like: tests are written; extended in Matft: implementing a new function; fixing a bug test-first (TDD); adding coverage for an existing function.

How do I install Test Design in Claude Code?

Run `npx skills add jjjkkkjjj/Matft --skill test-design -a claude-code`. Or copy the skill folder (.claude/skills/test-design in jjjkkkjjj/Matft) into .claude/skills/test-design in your project. Claude Code loads it when a task matches its description.

How do I install Test Design in Codex?

Run `npx skills add jjjkkkjjj/Matft --skill test-design -a codex`. Or copy the skill folder (.claude/skills/test-design in jjjkkkjjj/Matft) into .agents/skills/test-design in your project. Codex loads it when a task matches its description.

Can I use Test Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jjjkkkjjj/Matft --skill test-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-design, .gemini/skills/test-design, .github/skills/test-design and .opencode/skills/test-design in your project.

What does Test Design need to run?

Going by SKILL.md and its folder, Test Design needs Python for the scripts in its folder and the command-line tools its instructions call (swift). Our summary lists: Python 3.

Does Test Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Design use?

Test Design is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Design use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.

What are the alternatives to Test Design?

Skills that share tags, products or a category with Test Design: TDD Guide (alirezarezvani/claude-skills, 28k stars), TDD Guide (LeoYeAI/openclaw-master-skills, 2.2k stars), TDD Guide (aAAaqwq/AGI-Super-Team, 105 stars) and Python Testing (athola/claude-night-market, 341 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Design?

jjjkkkjjj (a GitHub user) maintains it in jjjkkkjjj/Matft, which has 147 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 27, 2026.

Source: jjjkkkjjj/Matft on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.