Agent skill

Benchmark Runner

by Mathews-Tom in Mathews-Tom/armory

Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps.

MITAuto-check passedTesting & QA

Install Benchmark Runner

skills CLI
$ npx skills add Mathews-Tom/armory --skill benchmark-runner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory benchmark-runner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmark-runner .claude/skills/benchmark-runner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-runner
GitHub stars
328
Token cost
~2.2k tokens
SKILL.md length
418 words
Files
6 (incl. references)
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps.

  • Works in 5 steps: Define Scope → Select Metrics → Design Test Cases → …
  • Compare performance
  • SKILL.md covers Reference Files, Prerequisites, Workflow and Output Format
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Benchmark Runner is an agent skill from Mathews-Tom/armory. Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps. Triggers on: "benchmark", "compare performance", "which is faster", "latency comparison", "run benchmark", "throughput test", "speed test".

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `evals/cases.yaml`, `references/environment-capture.md` and `references/metric-selection.md`).

It sits in Testing & QA, covering Test generation and Load testing. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • Compare performance
  • Which is faster
  • Latency comparison
  • Throughput test

Example prompts

  • “benchmark”
  • “compare performance”
  • “which is faster”
  • “/benchmark-runner”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Define Scope
  2. Select Metrics
  3. Design Test Cases
  4. Specify Environment
  5. Structure Results

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Runner loads about 2.2k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 76 tokens; SKILL.md has 418 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 418 words, ~2,244 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-runner/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
benchmark-runner
description
Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps. Triggers on: "benchmark", "compare performance", "which is faster", "latency comparison", "run benchmark", "throughput test", "speed test".
metadata.version
1.1.1
metadata.category
data
metadata.tags
benchmarking, performance, comparison, metrics
metadata.difficulty
intermediate
metadata.phase
verify

Benchmark Runner

Standardizes performance comparison methodology: metric selection, test case design, environment capture, result formatting, and tradeoff analysis. Produces reproducible benchmark reports that support informed decisions — not just "A is faster than B" but "A is faster for small inputs while B scales better."

Reference Files

FileContentsLoad When
references/metric-selection.mdMetric catalog (latency percentiles, throughput, memory, accuracy), selection criteria per task typeAlways
references/test-case-design.mdRepresentative input selection, scale variation, edge case coverage, warmup strategiesAlways
references/environment-capture.mdHardware/software context recording, reproducibility requirements, variance controlAlways
references/statistical-rigor.mdSample sizing, variance measurement, significance testing, outlier handlingResults need statistical validation

Prerequisites

  • Clear candidates to compare (at least 2)
  • Access to run or observe the candidates (code, API, or existing results)
  • Representative workload definition

Workflow

Phase 1: Define Scope
  1. What are the candidates? — Name each candidate precisely, including version. "Python dict vs Redis" is too vague. "Python 3.12 dict (in-process) vs Redis 7.2 (localhost, TCP)" is testable.
  2. What claims need validation? — "A is faster" → faster at what? For what input size? Under what load? Benchmark design flows from the specific claim.
  3. What is the decision context? — Why does this comparison matter? This determines which metrics are most important.
Phase 2: Select Metrics

Choose metrics that match the decision context:

Metric CategorySpecific MetricsWhen Important
LatencyP50, P95, P99, mean, std devUser-facing operations, API calls
Throughputops/sec, tokens/sec, MB/secBatch processing, streaming
MemoryPeak RSS, avg RSS, allocation rateResource-constrained environments
AccuracyF1, BLEU, exact match, precision/recallML models, algorithms with quality tradeoffs
Cost$/1K operations, $/hour, $/GBCloud services, API comparisons
StartupTime to first operation, cold startServerless, CLI tools

Select 2-4 metrics. More than 4 makes comparison tables unreadable.

Show full SKILL.md (135 more words)Show less
Phase 3: Design Test Cases

Create a matrix of inputs that reveal performance characteristics:

  1. Scale variation — Small, medium, large inputs. Performance often changes non-linearly with scale.
  2. Representative data — Use realistic inputs, not synthetic best-case data.
  3. Edge cases — Empty input, maximum size, adversarial input.
  4. Warmup — Exclude JIT compilation, cache warming, and connection establishment from measurements. Run N warmup iterations before recording.
Phase 4: Specify Environment

Record everything needed to reproduce the results:

  1. Hardware — CPU model, core count, RAM size, GPU model (if applicable)
  2. Software — OS version, language runtime version, dependency versions
  3. Configuration — Thread count, batch size, connection pool size, cache settings
  4. Isolation — What else was running? Background processes affect results.
Phase 5: Structure Results

Produce comparison tables with clear winners per metric, followed by tradeoff analysis.

Output Format

text
# Benchmark: {Descriptive Title}

**Date:** {YYYY-MM-DD}
**Hardware:** {CPU}, {RAM}, {GPU if applicable}
**Software:** {runtime versions}
**Configuration:** {key settings that affect results}

## Candidates

| # | Candidate | Version | Configuration |
|---|-----------|---------|---------------|
| A | {name} | {version} | {relevant config} |
| B | {name} | {version} | {relevant config} |

## Test Cases

| # | Name | Input Size | Description | Warmup | Iterations |
|---|------|------------|-------------|--------|------------|
| 1 | Small | {size} | {what it represents} | {N} | {N} |
| 2 | Medium | {size} | {what it represents} | {N} | {N} |
| 3 | Large | {size} | {what it represents} | {N} | {N} |

## Results

### Latency (ms, lower is better)

| Test Case | A (P50 / P95 / P99) | B (P50 / P95 / P99) | Winner |
|-----------|---------------------|---------------------|--------|
| Small | {values} | {values} | {A or B} |
| Medium | {values} | {values} | {A or B} |
| Large | {values} | {values} | {A or B} |

### Memory (MB, lower is better)

| Test Case | A (Peak) | B (Peak) | Winner |
|-----------|----------|----------|--------|
| Small | {value} | {value} | {A or B} |
| Medium | {value} | {value} | {A or B} |
| Large | {value} | {value} | {A or B} |

## Analysis

### Overall Winner
**{Candidate}** wins on {N} of {M} metrics across all test cases.

### Tradeoff Summary
- **Choose A when:** {conditions where A is the better choice}
- **Choose B when:** {conditions where B is the better choice}

### Caveats
- {Limitation of this benchmark}
- {Condition under which results may differ}

## Reproduction

```bash
# Environment setup
{commands to recreate the environment}

# Run benchmark
{commands to execute the benchmark}
text

## Configuring Scope

| Mode | Candidates | Depth | When to Use |
|------|-----------|-------|-------------|
| `quick` | 2 candidates, 1-2 metrics | Single test case, no statistics | Rough comparison, sanity check |
| `standard` | 2-3 candidates, 2-4 metrics | 3 test cases, mean + std dev | Default for most comparisons |
| `rigorous` | Any count, full metric suite | Multiple test cases, percentiles, significance tests | Publication, critical decisions |

## Calibration Rules

1. **Measure, don't guess.** Intuition about performance is unreliable. "Obviously
   faster" is not a benchmark result.
2. **Apples to apples.** Candidates must be compared under identical conditions.
   Different hardware, configuration, or input data invalidates the comparison.
3. **Report variance, not just means.** A mean of 50ms with std dev of 100ms is not
   the same as a mean of 50ms with std dev of 2ms. Always report spread.
4. **Warm up before measuring.** First-run performance includes JIT, cache warming,
   and connection setup. Exclude warmup iterations from results.
5. **Representative inputs only.** Benchmarking with synthetic best-case input is
   misleading. Use data that resembles production workloads.
6. **State the winner per metric, not overall.** "A is better" is lazy. "A has lower
   latency; B uses less memory" is useful.

## Error Handling

| Problem | Resolution |
|---------|------------|
| Cannot run candidates locally | Design the benchmark specification. Document what to measure and how. The user executes separately. |
| Results are noisy (high variance) | Increase iteration count. Check for background processes. Use dedicated hardware or containers for isolation. |
| Candidates serve different purposes | Acknowledge that the comparison is partial. Benchmark only the overlapping functionality. |
| No baseline exists | Establish one candidate as the baseline. Report relative performance (e.g., "B is 1.3x faster than A"). |
| Hardware context unavailable | Document what is known. Note that results may not be reproducible without full context. |

## When NOT to Benchmark

Push back if:
- The comparison is not performance-related (feature comparison → use a decision matrix or ADR instead)
- The candidates are fundamentally different tools (comparing a database to a message queue)
- The user wants to benchmark trivial operations (comparing two string concatenation methods in Python)
- Results from others already exist and conditions match — link to existing benchmarks instead

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/benchmark-runner of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml
  • references/environment-capture.md
  • references/metric-selection.md
  • references/statistical-rigor.md
  • references/test-case-design.md

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Benchmark Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Runner this skillMathews-Tom/armory328—~2.2kAutomated safety check: PassMIT
Manual Test Planningtestdouble/han279—~2.9kAutomated safety check: PassMIT
Jmeter Test Plan Creatorjeremylongshore/tons-of-skills-marketplace2.8k—~575Automated safety check: PassMIT
Test CommanderEliasOulkadi/shokunin114—~3kAutomated safety check: NotesMIT
Iterative Plan Reviewtestdouble/han279—~10kAutomated safety check: PassMIT
Azure App TestingMicrosoftDocs/Agent-Skills777—~3.6kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Manual Test Planning

    testdouble/han

    Produce a plain-language manual test plan from the context supplied to it — an executive summary, a high-level list of named tests, and a detail section per test with the steps a person follows by…

    279 GitHub stars~2.9k tokensUpdated 7 days ago
    Testing & QAAuto-check passed
  • Jmeter Test Plan Creator

    jeremylongshore/tons-of-skills-marketplace

    Create jmeter test plan creator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~575 tokensUpdated today
    Testing & QAAuto-check passed
  • Test Commander

    EliasOulkadi/shokunin

    Generate unit, integration, E2E, and visual regression tests following the Testing Trophy methodology (80% integration).

    114 GitHub stars~3k tokensUpdated 3 days ago
    Testing & QAAuto-check: notes
  • Iterative Plan Review

    testdouble/han

    Sharpens and stress-tests an existing plan file through multiple codebase-grounded review passes, editing it in place and recording every finding and iteration in cross-referenced companion files.

    279 GitHub stars~10k tokensUpdated 7 days ago
    Testing & QAAuto-check passed
  • Azure App Testing

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure App Testing development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations…

    777 GitHub stars~3.6k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Testing Strategy Builder

    aiskillstore/marketplace

    A skill your agent uses when creating comprehensive testing strategies for applications.

    430 GitHub stars~3.6k tokensUpdated today
    Testing & QAAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    328 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    328 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    328 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    328 GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    328 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    328 GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Benchmark Runner

What does Benchmark Runner do?

Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps. Benchmark Runner is an agent skill from Mathews-Tom/armory. Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps.

When should I use Benchmark Runner?

Benchmark Runner fits situations like: compare performance; which is faster; latency comparison; throughput test.

How do I install Benchmark Runner in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill benchmark-runner -a claude-code`. Or copy the skill folder (skills/benchmark-runner in Mathews-Tom/armory) into .claude/skills/benchmark-runner in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Runner in Codex?

Run `npx skills add Mathews-Tom/armory --skill benchmark-runner -a codex`. Or copy the skill folder (skills/benchmark-runner in Mathews-Tom/armory) into .agents/skills/benchmark-runner in your project. Codex loads it when a task matches its description.

Can I use Benchmark Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill benchmark-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-runner, .gemini/skills/benchmark-runner, .github/skills/benchmark-runner and .opencode/skills/benchmark-runner in your project.

What does Benchmark Runner need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark Runner is instructions for the agent only. Our summary lists: Python 3.

Does Benchmark Runner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Runner use?

Benchmark Runner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Runner use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.8k tokens, read only when the agent opens those files.

What are the alternatives to Benchmark Runner?

Skills that share tags, products or a category with Benchmark Runner: Manual Test Planning (testdouble/han, 279 stars), Jmeter Test Plan Creator (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Test Commander (EliasOulkadi/shokunin, 114 stars) and Iterative Plan Review (testdouble/han, 279 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Runner?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 328 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.