Manual Test Planning
testdouble/han
Produce a plain-language manual test plan from the context supplied to it — an executive summary, a high-level list of named tests, and a detail section per test with the steps a person follows by…
Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps.
$ npx skills add Mathews-Tom/armory --skill benchmark-runner -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Mathews-Tom/armory benchmark-runner --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmark-runner .claude/skills/benchmark-runner && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-runner" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/benchmark-runner into .claude/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Mathews-Tom/armory/tree/main/skills/benchmark-runnerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Mathews-Tom/armory --skill benchmark-runner -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Mathews-Tom/armory benchmark-runner --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/benchmark-runner .agents/skills/benchmark-runner && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-runner" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/benchmark-runner into .agents/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mathews-Tom/armory --skill benchmark-runner -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Mathews-Tom/armory benchmark-runner --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/benchmark-runner .cursor/skills/benchmark-runner && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-runner" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/benchmark-runner into .cursor/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Mathews-Tom/armory.git --path skills/benchmark-runner--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Mathews-Tom/armory --skill benchmark-runner -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Mathews-Tom/armory benchmark-runner --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/benchmark-runner .gemini/skills/benchmark-runner && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-runner" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/benchmark-runner into .gemini/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Mathews-Tom/armory benchmark-runnerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Mathews-Tom/armory --skill benchmark-runner -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/benchmark-runner .github/skills/benchmark-runner && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-runner" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/benchmark-runner into .github/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mathews-Tom/armory --skill benchmark-runner -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Mathews-Tom/armory benchmark-runner --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/benchmark-runner .opencode/skills/benchmark-runner && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-runner" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/benchmark-runner into .opencode/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-runnerDesigns structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps.
Benchmark Runner is an agent skill from Mathews-Tom/armory. Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps. Triggers on: "benchmark", "compare performance", "which is faster", "latency comparison", "run benchmark", "throughput test", "speed test".
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `evals/cases.yaml`, `references/environment-capture.md` and `references/metric-selection.md`).
It sits in Testing & QA, covering Test generation and Load testing. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Runner loads about 2.2k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 76 tokens; SKILL.md has 418 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 418 words, ~2,244 tokens.
.claude/skills/benchmark-runner/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Standardizes performance comparison methodology: metric selection, test case design, environment capture, result formatting, and tradeoff analysis. Produces reproducible benchmark reports that support informed decisions — not just "A is faster than B" but "A is faster for small inputs while B scales better."
| File | Contents | Load When |
|---|---|---|
references/metric-selection.md | Metric catalog (latency percentiles, throughput, memory, accuracy), selection criteria per task type | Always |
references/test-case-design.md | Representative input selection, scale variation, edge case coverage, warmup strategies | Always |
references/environment-capture.md | Hardware/software context recording, reproducibility requirements, variance control | Always |
references/statistical-rigor.md | Sample sizing, variance measurement, significance testing, outlier handling | Results need statistical validation |
Choose metrics that match the decision context:
| Metric Category | Specific Metrics | When Important |
|---|---|---|
| Latency | P50, P95, P99, mean, std dev | User-facing operations, API calls |
| Throughput | ops/sec, tokens/sec, MB/sec | Batch processing, streaming |
| Memory | Peak RSS, avg RSS, allocation rate | Resource-constrained environments |
| Accuracy | F1, BLEU, exact match, precision/recall | ML models, algorithms with quality tradeoffs |
| Cost | $/1K operations, $/hour, $/GB | Cloud services, API comparisons |
| Startup | Time to first operation, cold start | Serverless, CLI tools |
Select 2-4 metrics. More than 4 makes comparison tables unreadable.
Create a matrix of inputs that reveal performance characteristics:
Record everything needed to reproduce the results:
Produce comparison tables with clear winners per metric, followed by tradeoff analysis.
# Benchmark: {Descriptive Title}
**Date:** {YYYY-MM-DD}
**Hardware:** {CPU}, {RAM}, {GPU if applicable}
**Software:** {runtime versions}
**Configuration:** {key settings that affect results}
## Candidates
| # | Candidate | Version | Configuration |
|---|-----------|---------|---------------|
| A | {name} | {version} | {relevant config} |
| B | {name} | {version} | {relevant config} |
## Test Cases
| # | Name | Input Size | Description | Warmup | Iterations |
|---|------|------------|-------------|--------|------------|
| 1 | Small | {size} | {what it represents} | {N} | {N} |
| 2 | Medium | {size} | {what it represents} | {N} | {N} |
| 3 | Large | {size} | {what it represents} | {N} | {N} |
## Results
### Latency (ms, lower is better)
| Test Case | A (P50 / P95 / P99) | B (P50 / P95 / P99) | Winner |
|-----------|---------------------|---------------------|--------|
| Small | {values} | {values} | {A or B} |
| Medium | {values} | {values} | {A or B} |
| Large | {values} | {values} | {A or B} |
### Memory (MB, lower is better)
| Test Case | A (Peak) | B (Peak) | Winner |
|-----------|----------|----------|--------|
| Small | {value} | {value} | {A or B} |
| Medium | {value} | {value} | {A or B} |
| Large | {value} | {value} | {A or B} |
## Analysis
### Overall Winner
**{Candidate}** wins on {N} of {M} metrics across all test cases.
### Tradeoff Summary
- **Choose A when:** {conditions where A is the better choice}
- **Choose B when:** {conditions where B is the better choice}
### Caveats
- {Limitation of this benchmark}
- {Condition under which results may differ}
## Reproduction
```bash
# Environment setup
{commands to recreate the environment}
# Run benchmark
{commands to execute the benchmark}
## Configuring Scope
| Mode | Candidates | Depth | When to Use |
|------|-----------|-------|-------------|
| `quick` | 2 candidates, 1-2 metrics | Single test case, no statistics | Rough comparison, sanity check |
| `standard` | 2-3 candidates, 2-4 metrics | 3 test cases, mean + std dev | Default for most comparisons |
| `rigorous` | Any count, full metric suite | Multiple test cases, percentiles, significance tests | Publication, critical decisions |
## Calibration Rules
1. **Measure, don't guess.** Intuition about performance is unreliable. "Obviously
faster" is not a benchmark result.
2. **Apples to apples.** Candidates must be compared under identical conditions.
Different hardware, configuration, or input data invalidates the comparison.
3. **Report variance, not just means.** A mean of 50ms with std dev of 100ms is not
the same as a mean of 50ms with std dev of 2ms. Always report spread.
4. **Warm up before measuring.** First-run performance includes JIT, cache warming,
and connection setup. Exclude warmup iterations from results.
5. **Representative inputs only.** Benchmarking with synthetic best-case input is
misleading. Use data that resembles production workloads.
6. **State the winner per metric, not overall.** "A is better" is lazy. "A has lower
latency; B uses less memory" is useful.
## Error Handling
| Problem | Resolution |
|---------|------------|
| Cannot run candidates locally | Design the benchmark specification. Document what to measure and how. The user executes separately. |
| Results are noisy (high variance) | Increase iteration count. Check for background processes. Use dedicated hardware or containers for isolation. |
| Candidates serve different purposes | Acknowledge that the comparison is partial. Benchmark only the overlapping functionality. |
| No baseline exists | Establish one candidate as the baseline. Report relative performance (e.g., "B is 1.3x faster than A"). |
| Hardware context unavailable | Document what is known. Note that results may not be reproducible without full context. |
## When NOT to Benchmark
Push back if:
- The comparison is not performance-related (feature comparison → use a decision matrix or ADR instead)
- The candidates are fundamentally different tools (comparing a database to a message queue)
- The user wants to benchmark trivial operations (comparing two string concatenation methods in Python)
- Results from others already exist and conditions match — link to existing benchmarks instead© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/benchmark-runner of Mathews-Tom/armory.
Open the folder on GitHubat commit 4594fb7
Benchmark Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Runner this skillMathews-Tom/armory | 328 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Manual Test Planningtestdouble/han | 279 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Jmeter Test Plan Creatorjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~575 | Automated safety check: Pass | MIT | |
| Test CommanderEliasOulkadi/shokunin | 114 | — | ~3k | Automated safety check: Notes | MIT | |
| Iterative Plan Reviewtestdouble/han | 279 | — | ~10k | Automated safety check: Pass | MIT | |
| Azure App TestingMicrosoftDocs/Agent-Skills | 777 | — | ~3.6k | Automated safety check: Pass | CC-BY-4.0 |
testdouble/han
Produce a plain-language manual test plan from the context supplied to it — an executive summary, a high-level list of named tests, and a detail section per test with the steps a person follows by…
jeremylongshore/tons-of-skills-marketplace
Create jmeter test plan creator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.
EliasOulkadi/shokunin
Generate unit, integration, E2E, and visual regression tests following the Testing Trophy methodology (80% integration).
testdouble/han
Sharpens and stress-tests an existing plan file through multiple codebase-grounded review passes, editing it in place and recording every finding and iteration in cross-referenced companion files.
MicrosoftDocs/Agent-Skills
Expert knowledge for Azure App Testing development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations…
aiskillstore/marketplace
A skill your agent uses when creating comprehensive testing strategies for applications.
Mathews-Tom/armory
Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.
Mathews-Tom/armory
Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.
Mathews-Tom/armory
A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…
Mathews-Tom/armory
Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.
Mathews-Tom/armory
Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.
Mathews-Tom/armory
Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…
Categories
Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps. Benchmark Runner is an agent skill from Mathews-Tom/armory. Designs structured benchmarks comparing algorithms, models, or implementations with metrics, test cases, hardware context, and reproduction steps.
Benchmark Runner fits situations like: compare performance; which is faster; latency comparison; throughput test.
Run `npx skills add Mathews-Tom/armory --skill benchmark-runner -a claude-code`. Or copy the skill folder (skills/benchmark-runner in Mathews-Tom/armory) into .claude/skills/benchmark-runner in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Mathews-Tom/armory --skill benchmark-runner -a codex`. Or copy the skill folder (skills/benchmark-runner in Mathews-Tom/armory) into .agents/skills/benchmark-runner in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill benchmark-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-runner, .gemini/skills/benchmark-runner, .github/skills/benchmark-runner and .opencode/skills/benchmark-runner in your project.
SKILL.md names no scripts, command-line tools or credentials: Benchmark Runner is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Benchmark Runner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Benchmark Runner: Manual Test Planning (testdouble/han, 279 stars), Jmeter Test Plan Creator (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Test Commander (EliasOulkadi/shokunin, 114 stars) and Iterative Plan Review (testdouble/han, 279 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 328 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.
Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.