Agent skill

Benchmark

by jjjkkkjjj in jjjkkkjjj/Matft

Procedure for benchmarking Matft's PerformanceTests against Numpy and reporting the results (and, when asked, updating the speed comparison table on the docs site, website/docs/performance.md).

BSD-3-ClauseAuto-check passedFrontend & Design

Install Benchmark

skills CLI
$ npx skills add jjjkkkjjj/Matft --skill benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jjjkkkjjj/Matft benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jjjkkkjjj/Matft.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmark .claude/skills/benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark
GitHub stars
147
Token cost
~1.7k tokens
SKILL.md length
757 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Procedure for benchmarking Matft's PerformanceTests against Numpy and reporting the results (and, when asked, updating the speed comparison table on the docs site, website/docs/performance.md).

  • Works in 5 steps: Pre-checks → Measure → Report → …
  • The conversation is about measuring Matfts speed
  • SKILL.md covers 1. Pre-checks, 2. Measure, 3. Report and 4. Update the docs (only when…, plus 1 more section
  • Calls python3, git and swift

What it does

Benchmark is an agent skill from jjjkkkjjj/Matft. Procedure for benchmarking Matft's PerformanceTests against Numpy and reporting the results (and, when asked, updating the speed comparison table on the docs site, website/docs/performance.md). Use this skill whenever the conversation is about measuring Matft's speed, comparing it with Numpy, or updating the Performance section of the docs (formerly the README) — e.g. "run the benchmarks", "compare speed with Numpy", "update the perf table", "measure the effect of this optimization", or in Japanese「ベンチ回して」「Numpy…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Frontend & Design, covering Static sites and blogs, Technical documentation and iOS development. It works with NumPy. The repository describes itself as: Numpy-like library in swift. (Multi-dimensional Array, ndarray, matrix and vector library). The licence is BSD-3-Clause.

When your agent uses it

  • The conversation is about measuring Matfts speed
  • Comparing it with Numpy
  • Updating the Performance section of the docs (formerly the README) — e.g

Example prompts

  • “run the benchmarks”
  • “compare speed with Numpy”
  • “update the perf table”
  • “/benchmark”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Pre-checks
  2. Measure
  3. Report
  4. Update the docs (only when the user agrees)
  5. Adding or changing cases

What it can do on your machine

Read from SKILL.md and the folder at commit 618dcfc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • git
    • swift
    • pip3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and pip3, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark loads about 1.7k tokens when it runs. Until then it costs about 157 tokens; SKILL.md has 757 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~157
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jjjkkkjjj/Matft at commit 618dcfc, republished under its BSD-3-Clause licence (© jjjkkkjjj). 757 words, ~1,721 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark/SKILL.md (or your agent's skills folder).
name
benchmark
description
Procedure for benchmarking Matft's PerformanceTests against Numpy and reporting the results (and, when asked, updating the speed comparison table on the docs site, website/docs/performance.md). Use this skill whenever the conversation is about measuring Matft's speed, comparing it with Numpy, or updating the Performance section of the docs (formerly the README) — e.g. "run the benchmarks", "compare speed with Numpy", "update the perf table", "measure the effect of this optimization", or in Japanese「ベンチ回して」「Numpy と速度比較して」「README / ドキュメントの速度表を更新して」「perf 計測」「最適化の効果を測って」— even if the word "skill" is never mentioned.

Matft vs Numpy benchmark

All measurement is done by a single script, scripts/benchmark.py.

  • Matft side: runs Tests/PerformanceTests with swift test -c release and parses the output of XCTest measure {}.
    • Each test is measured with measureWithWarmup {} (Tests/PerformanceTests/PerfFixtures.swift), which works in two stages:
      1. Spin the block for --warmup seconds (default 0.5 s).
      2. Choose the number of calls per sample, N, so that one sample takes about --sample-time seconds (default 20 ms), then run measure {} (the same idea as number in timeit).
    • N is printed as a line MatftBench: -[<Class> <method>] number=N. The script divides each sample by N to get the time per call.
    • With one call per sample, XCTest's own per-iteration overhead dominates fast cases. For example, an operation that really takes 0.14 ms is reported as about 0.5 ms.
  • Numpy side: measures the same expression with timeit.
  • Result: compares the medians of both and writes benchmarks/results/latest.{json,md}.

On the docs site's Performance page (website/docs/performance.md), everything between <!-- BENCHMARK:START --> and <!-- BENCHMARK:END --> is generated. Do not edit it by hand.

Measure on a local Mac only. CI runners are too noisy, so never use their numbers in the docs.

1. Pre-checks

sh
git status --porcelain
python3 -c "import numpy; print(numpy.__version__)"
  • You can measure with uncommitted changes, but the report's commit field becomes -dirty. If the numbers are going into the docs, tell the user to run it after committing.
  • If numpy is missing, suggest pip3 install numpy. Do not install it yourself.
  • Ask the user to plug in the power and stop heavy work (builds, video in the browser, VMs, etc.). You can check the load with the load average from uptime and ps -Ao pcpu,comm -r | head.

2. Measure

sh
cp benchmarks/results/latest.json /tmp/matft-bench-baseline.json   # to compare with the previous run (if it exists)
python3 scripts/benchmark.py [--baseline /tmp/matft-bench-baseline.json] [--filter Bool]
  • It includes a release build, so it takes several minutes. Use a long timeout (10 minutes).
  • --filter <regex>: narrow down by case ID, e.g. Bool, Sin.
  • --skip-swift / --skip-numpy: reuse the previous JSON, e.g. to re-measure only Numpy.
  • --warmup / --sample-time: warm-up seconds and target seconds per sample on the Matft side.
  • --repeat / --number: number of samples and calls per sample for Numpy's timeit.
  • --configuration debug: measure Matft in a debug build. SwiftPM builds dependencies with the same configuration as the app, so this is the speed seen by an app built without optimization. Generic loops such as initialize(repeating:) become orders of magnitude slower at -Onone, so when you change allocation or per-element loops, measure both release and debug. Cannot be combined with --update-readme.
Measuring the effect of an optimization (A/B comparison)

Using a previously saved JSON as the baseline compares against values taken at a different time under different load, which is very noisy — background load has shifted results by ±100% or more. Measure the effect of an optimization by alternating runs of the pre-change commit with the same measurement harness.

sh
git worktree add /tmp/matft-before <commit before the change>
cp -R Tests/PerformanceTests/. /tmp/matft-before/Tests/PerformanceTests/   # align the harness and cases
cp scripts/benchmark.py /tmp/matft-before/scripts/
# Alternate before → after → before → after (--skip-numpy if only Matft matters)
(cd /tmp/matft-before && python3 scripts/benchmark.py --skip-numpy && cp benchmarks/results/latest.json /tmp/before-1.json)
python3 scripts/benchmark.py --skip-numpy && cp benchmarks/results/latest.json /tmp/after-1.json
# ...repeat for the second round. When done: git worktree remove /tmp/matft-before
  • If the cases you did not change stay within a few percent, that round is trustworthy. Use them as a noise control.
  • Finally, re-measure Numpy alone with --skip-swift so the Numpy ratios are taken under the same conditions.
Show full SKILL.md (272 more words)Show less

3. Report

Read benchmarks/results/latest.md and the summary in latest.json, and briefly summarize:

  • Cases where Matft is slower than Numpy (the ratios in bold), sorted by ratio, largest first.
  • With --baseline, improvements and regressions from the previous run (only those of ±10% or more).
  • Cases with a large summary.swift.*.rsd_steady (roughly 10% or more). Note that those numbers are noisy and less reliable.
    • Do not use rsd. Even after the warm-up, the first 2–3 samples of XCTest measure {} are 1.3–2× slower, so rsd is always inflated to 20–30%.
    • rsd_steady is the spread excluding the first 3 samples. The median itself is barely affected.

No need to paste the whole table. Show only the key points and the numbers that stand out.

4. Update the docs (only when the user agrees)

sh
python3 scripts/benchmark.py --skip-swift --skip-numpy --update-docs   # write the last results into website/docs/performance.md as-is
git diff website/docs/performance.md
  • --update-docs (formerly --update-readme) cannot be combined with --filter (it needs all cases).
  • Confirm and show that the diff stays between the markers.
  • Commit only when the user tells you to.

5. Adding or changing cases

A case must be kept in sync in three places:

  1. Add a test method to Tests/PerformanceTests/*PefTests.swift, using inputs from PerfFixtures.swift. Always measure with self.measureWithWarmup {}, not self.measure {}. self.measure {} has no warm-up and does not print the call-count line (it is treated as N=1).
  2. Add Case("<Class>.<method>", "<Category>", "<Swift expr>", "<numpy expr>") to CASES in scripts/benchmark.py. If it needs a new input, add it to both SETUP and PerfFixtures.
  3. If you changed inputs, also update the Swift / Python setup examples at the top of website/docs/performance.md.

When you change the parser or the rendering in scripts/benchmark.py, add a test to scripts/test_benchmark.py first (TDD) and check with python3 -m unittest scripts/test_benchmark.py.

© jjjkkkjjj, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/benchmark of jjjkkkjjj/Matft.

Open the folder on GitHubat commit 618dcfc

Compare with similar skills

Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark this skilljjjkkkjjj/Matft147—~1.7kAutomated safety check: PassBSD-3-Clause
Vitepress Config Creatorjeremylongshore/tons-of-skills-marketplace2.8k—~589Automated safety check: PassMIT
Orchardcore Docs WriterOrchardCMS/OrchardCore8.2k—~1.1kAutomated safety check: PassBSD-3-Clause
Tabler Menus and Redirectstabler/tabler42k—~1.2kAutomated safety check: PassMIT
Uui Documentationepam/UUI248—~1.3kAutomated safety check: PassMIT
Eisland Dev Update DocsJNTMTMTM/eIsland320—~1.4kAutomated safety check: PassGPL-3.0

Similar skills

  • Vitepress Config Creator

    jeremylongshore/tons-of-skills-marketplace

    Create vitepress config creator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~589 tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Orchardcore Docs Writer

    OrchardCMS/OrchardCore

    Authors OrchardCore documentation — MkDocs Material pages, module README docs, nav entries, admonitions, tabbed content, and redirects.

    8.2k GitHub stars~1.1k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Keeps Tabler pages reachable by updating the preview menu, docs menu and redirect table whenever a page is added, renamed, moved or removed.

    42k GitHub stars~1.2k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Helps update UUI documentation, add doc examples, configure Property Explorer, and manage component API documentation.

    248 GitHub stars~1.3k tokensUpdated 11 days ago
    DevelopmentAuto-check passed
  • Eisland Dev Update Docs

    JNTMTMTM/eIsland

    Update or create documentation in the eIsland VuePress docs site (web/eisland-web-docs).

    320 GitHub stars~1.4k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Eisland Dev Update Docs

    JNTMTMTM/eIsland

    Update or create documentation in the eIsland VuePress docs site (web/eisland-web-docs).

    320 GitHub stars~1.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from jjjkkkjjj/Matft

  • Docs

    jjjkkkjjj/Matft

    Procedure for writing and updating Matft's documentation (the Docusaurus site in website/ and the doc comments on the public API that become the Swift-DocC API reference).

    147 GitHub stars~2.7k tokensUpdated 13 days ago
    Auto-check passed
  • Test Design

    jjjkkkjjj/Matft

    Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and…

    147 GitHub stars~2k tokensUpdated 13 days ago
    Auto-check passed
  • Image Visual Check

    jjjkkkjjj/Matft

    Procedure for adding tests for Matft's image processing (Matft.image., indexing or channel swapping on images, etc.), generating comparison images that put the result next to an OpenCV reference…

    147 GitHub stars~2.3k tokensUpdated 13 days ago
    Auto-check passed
  • Release

    jjjkkkjjj/Matft

    Procedure for releasing a new version of Matft (decide the version → check tests → write release notes → create and push an annotated tag → publish a GitHub Release).

    147 GitHub stars~1.8k tokensUpdated 13 days ago
    Auto-check passed

Works with

Questions about Benchmark

What does Benchmark do?

Procedure for benchmarking Matft's PerformanceTests against Numpy and reporting the results (and, when asked, updating the speed comparison table on the docs site, website/docs/performance.md). Benchmark is an agent skill from jjjkkkjjj/Matft.md).

When should I use Benchmark?

Benchmark fits situations like: the conversation is about measuring Matfts speed; comparing it with Numpy; updating the Performance section of the docs (formerly the README) — e.g.

How do I install Benchmark in Claude Code?

Run `npx skills add jjjkkkjjj/Matft --skill benchmark -a claude-code`. Or copy the skill folder (.claude/skills/benchmark in jjjkkkjjj/Matft) into .claude/skills/benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark in Codex?

Run `npx skills add jjjkkkjjj/Matft --skill benchmark -a codex`. Or copy the skill folder (.claude/skills/benchmark in jjjkkkjjj/Matft) into .agents/skills/benchmark in your project. Codex loads it when a task matches its description.

Can I use Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jjjkkkjjj/Matft --skill benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark, .gemini/skills/benchmark, .github/skills/benchmark and .opencode/skills/benchmark in your project.

What does Benchmark need to run?

Going by SKILL.md and its folder, Benchmark needs the command-line tools its instructions call (python3, git, swift and pip3). Our summary lists: Python 3.

Does Benchmark access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark use?

Benchmark is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark?

Skills that share tags, products or a category with Benchmark: Vitepress Config Creator (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Orchardcore Docs Writer (OrchardCMS/OrchardCore, 8.2k stars), Tabler Menus and Redirects (tabler/tabler, 42k stars) and Uui Documentation (epam/UUI, 248 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark?

jjjkkkjjj (a GitHub user) maintains it in jjjkkkjjj/Matft, which has 147 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 27, 2026.

Source: jjjkkkjjj/Matft on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.