Agent skill

Prowl Performance Benchmark Runner

by onevcat in onevcat/Prowl

Runs Prowl's benchmark layers for a branch or release against recorded baselines, only when explicitly asked and the machine is quiet.

Custom licenceAuto-check passedDevelopment

Install Prowl Performance Benchmark Runner

skills CLI
$ npx skills add onevcat/Prowl --skill run-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install onevcat/Prowl run-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/onevcat/Prowl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/run-benchmark .claude/skills/run-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-benchmark
GitHub stars
638
Token cost
~2.4k tokens
SKILL.md length
1,184 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
Custom licence

At a glance

Runs Prowl's benchmark layers for a branch or release against recorded baselines, only when explicitly asked and the machine is quiet.

  • Works in 5 steps: Preflight → Classify intent, then always confirm → Run the confirmed mode → …
  • Checking whether a branch regressed an existing performance suite
  • SKILL.md covers Invocation contract, Step 0 — Preflight, Step 1 — Classify intent, then… and Step 2 — Run the confirmed mode, plus 3 more sections
  • Calls make, git and jq

What it does

It never runs on its own after implementation work; a run needs an explicit request, and every timing run assumes one dedicated, quiet local machine building a comparable series over time, since numbers from CI runners or a loaded host can't enter that baseline. Before running it checks that the git tree is clean and HEAD is committed, since an A/B run switches the checkout to the baseline commit, and it checks system load against the core count, warning or refusing the timed run above a chosen threshold of competing load.

It classifies the request into one of three modes by reading the user's own wording plus which files changed relative to the merge base with main: a hot-path check comparing the branch against its baseline when it touches files mapped to existing suites, a new-optimization run that adds a benchmark first when no suite covers the change yet, and a release-level sweep running the full test and benchmark suite against the trailing series plus a live layer. Whichever mode is detected, it confirms with the user before running rather than acting on the classification alone.

When your agent uses it

  • Checking whether a branch regressed an existing performance suite
  • Adding a benchmark for a new optimization before comparing it
  • Running a full pre-release performance sweep

Example prompts

  • “Run the hot-path benchmark to check if my change regressed anything.”
  • “Benchmark this new caching optimization against the baseline.”
  • “Run a full release benchmark sweep before we ship.”

Requirements

  • A clean, committed git tree
  • A dedicated, quiet local machine for timing runs

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Preflight
  2. Classify intent, then always confirm
  3. Run the confirmed mode
  4. Compare and judge
  5. Persistent report

What it can do on your machine

Read from SKILL.md and the folder at commit 386b8d0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make
    • git
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prowl Performance Benchmark Runner loads about 2.4k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,184 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 1,184 words (~2,417 tokens).

“Run this skill only after an explicit user request. Implementing, fixing, or reviewing performance-adjacent code does not by itself authorize a benchmark run.”

— opening of SKILL.md by onevcat, Custom licence
name
run-benchmark

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .claude/skills/run-benchmark of onevcat/Prowl.

Open the folder on GitHubat commit 386b8d0

Compare with similar skills

Prowl Performance Benchmark Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prowl Performance Benchmark Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prowl Performance Benchmark Runner this skillonevcat/Prowl638—~2.4kAutomated safety check: PassCustom licence
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
ExecuTorch Binary Size Reductionpytorch/executorch5.1k—~793Automated safety check: PassCustom licence
Thorough Code Reviewpretend1111/claude-desktop-app4941 repos~502Automated safety check: PassCustom licence
Climber Step Minimizationben-manes/caffeine18k—~3kAutomated safety check: NotesApache-2.0
Self-Improvement Tournament Loopzereight/gitlab-mcp2k1 repos~1.8kAutomated safety check: WarnMIT

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Measures and shrinks the ExecuTorch runtime binary by building a size test, analyzing it with bloaty and landing each reduction as its own pull request.

    5.1k GitHub stars~793 tokensUpdated today
    DevelopmentAuto-check passed
  • Thorough Code Review

    pretend1111/claude-desktop-app

    Reviews code or recent changes for bugs, security issues, performance problems and maintainability, reporting findings by severity with the reason and a fix.

    494 GitHub starsUsed in 1 repo~502 tokens
    DevelopmentAuto-check passed
  • Climber Step Minimization

    ben-manes/caffeine

    Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire.

    18k GitHub stars~3k tokensUpdated 2 days ago
    DevelopmentAuto-check: notes
  • Runs an autonomous evolutionary loop that improves a codebase against a measurable benchmark, using agent roles, tournament selection and recorded history until a stop condition.

    2k GitHub starsUsed in 1 repo~1.8k tokens
    DevelopmentAuto-check: warnings
  • Pi En Progressive

    share-skills/pi

    PI Cognitive AI. An agent skill from share-skills/pi.

    108 GitHub stars~9.3k tokensUpdated 3 mo ago
    DevelopmentAuto-check: warnings

More from onevcat/Prowl

All 10 skills in this repo
  • Drives and verifies the native macOS Prowl Debug UI through accessibility automation, proving behavior with before-and-after evidence instead of reasoning from code.

    638 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Prowl Self-Verify

    onevcat/Prowl

    Runs an opt-in end-to-end verification of Prowl changes in a separate debug instance, writing a verification contract first and reporting pass, fail, skipped or inconclusive for every scenario.

    638 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Authors, validates, runs and participates in Prowl Agent Workflows, bundles that orchestrate several live coding agents inside the Prowl macOS app.

    638 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Catches an app's string catalog up with the code once per release: translating new text, pruning unused entries, and flagging unlocalized strings.

    638 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Write AI Doc

    onevcat/Prowl

    Create and maintain curated docs-ai/ records for substantial features and non-trivial, decision-shaping fixes (numbered entries with 000-plan.md before implementation and 001-action.md after).

    638 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Keeps an agent-facing documentation set accurate against the implementation through small, diff-driven edits since a committed baseline.

    638 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Prowl Performance Benchmark Runner

What does Prowl Performance Benchmark Runner do?

Runs Prowl's benchmark layers for a branch or release against recorded baselines, only when explicitly asked and the machine is quiet. It never runs on its own after implementation work; a run needs an explicit request, and every timing run assumes one dedicated, quiet local machine building a comparable series over time, since numbers from CI runners or a loaded host can't enter that baseline. Before running it checks that the git tree is clean and HEAD is committed, since an A/B run switches the checkout to the baseline commit, and it checks system load against the core count, warning or refusing the timed run above a chosen threshold of competing load.

When should I use Prowl Performance Benchmark Runner?

Prowl Performance Benchmark Runner fits situations like: checking whether a branch regressed an existing performance suite; adding a benchmark for a new optimization before comparing it; running a full pre-release performance sweep.

How do I install Prowl Performance Benchmark Runner in Claude Code?

Run `npx skills add onevcat/Prowl --skill run-benchmark -a claude-code`. Or copy the skill folder (.claude/skills/run-benchmark in onevcat/Prowl) into .claude/skills/run-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Prowl Performance Benchmark Runner in Codex?

Run `npx skills add onevcat/Prowl --skill run-benchmark -a codex`. Or copy the skill folder (.claude/skills/run-benchmark in onevcat/Prowl) into .agents/skills/run-benchmark in your project. Codex loads it when a task matches its description.

Can I use Prowl Performance Benchmark Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add onevcat/Prowl --skill run-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-benchmark, .gemini/skills/run-benchmark, .github/skills/run-benchmark and .opencode/skills/run-benchmark in your project.

What does Prowl Performance Benchmark Runner need to run?

Going by SKILL.md and its folder, Prowl Performance Benchmark Runner needs the command-line tools its instructions call (make, git and jq). Our summary lists: A clean, committed git tree; A dedicated, quiet local machine for timing runs.

Does Prowl Performance Benchmark Runner access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Prowl Performance Benchmark Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prowl Performance Benchmark Runner use?

Prowl Performance Benchmark Runner has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Prowl Performance Benchmark Runner use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prowl Performance Benchmark Runner?

Skills that share tags, products or a category with Prowl Performance Benchmark Runner: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), ExecuTorch Binary Size Reduction (pytorch/executorch, 5.1k stars), Thorough Code Review (pretend1111/claude-desktop-app, 494 stars) and Climber Step Minimization (ben-manes/caffeine, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prowl Performance Benchmark Runner?

onevcat (a GitHub user) maintains it in onevcat/Prowl, which has 638 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 6, 2026.

Source: onevcat/Prowl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.