Agent skill

Benchmark Optimization Loop

by affaan-m in affaan-m/ECC

Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…

MITAuto-check passed

Install Benchmark Optimization Loop

skills CLI
$ npx skills add affaan-m/ECC --skill benchmark-optimization-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC benchmark-optimization-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmark-optimization-loop .claude/skills/benchmark-optimization-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-optimization-loop
GitHub stars
274k
Used in
1 other repo
Token cost
~664 tokens
SKILL.md length
274 words
Files
1
Skills in repo
657
Repo updated
First seen
Licence
MIT

At a glance

Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…

  • Works in 8 steps: Measure the baseline. → Identify bottlenecks from evidence. → Generate variants that test one… → …
  • Asked to speed something up
  • SKILL.md covers Required Baseline, Loop, Variant Table and Recursive Search, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Benchmark Optimization Loop is an agent skill from affaan-m/ECC. Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest safe variant with reproducible commands. Use when asked to speed something up, try many variants, run recursive optimization, benchmark latency/throughput/cost, or pick the best implementation by repeated measured tests.

Its SKILL.md is about 660 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Asked to speed something up
  • Try many variants
  • Run recursive optimization
  • Benchmark latency/throughput/cost

Example prompts

  • “make it faster”
  • “/benchmark-optimization-loop”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Measure the baseline.
  2. Identify bottlenecks from evidence.
  3. Generate variants that test one hypothesis each.
  4. Run variants with the same input shape.
  5. Reject variants that fail correctness, safety, or reproducibility.
  6. Promote the fastest safe variant.
  7. Codify the winning path in a script, command, test, config, or doc.
  8. Rerun the baseline and winner to confirm the delta.

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Optimization Loop loads about 664 tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 274 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~664

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 274 words, ~664 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-optimization-loop/SKILL.md (or your agent's skills folder).
name
benchmark-optimization-loop
description
Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest safe variant with reproducible commands. Use when asked to speed something up, try many variants, run recursive optimization, benchmark latency/throughput/cost, or pick the best implementation by repeated measured tests.
license
MIT
metadata.origin
ECC
tools
Read, Write, Edit, Bash, Grep, Glob

Benchmark Optimization Loop

Use this skill to convert "make it 20x faster" or "try 50 recursive optimizations" into a bounded measured loop that can actually improve a system.

Required Baseline

Do not optimize until these exist:

  • the operation being optimized;
  • the correctness gate that must stay green;
  • the metric: wall time, p95 latency, rows/sec, cost/run, memory, error rate;
  • the current baseline;
  • the search budget: max variants, max time, max spend, max data impact.

If the user asks for an unrealistic target, keep the ambition but make the loop bounded and measurable.

Loop

  1. Measure the baseline.
  2. Identify bottlenecks from evidence.
  3. Generate variants that test one hypothesis each.
  4. Run variants with the same input shape.
  5. Reject variants that fail correctness, safety, or reproducibility.
  6. Promote the fastest safe variant.
  7. Codify the winning path in a script, command, test, config, or doc.
  8. Rerun the baseline and winner to confirm the delta.

Variant Table

Track variants like this:

text
Variant | Hypothesis | Command | Time | Correct? | Notes
baseline | current path | npm run job | 120s | yes | stable
batch-500 | fewer round trips | npm run job -- --batch 500 | 42s | yes | winner
parallel-8 | more workers | npm run job -- --workers 8 | 31s | no | rate limited

For recursive or hyperparameter work:

  • persist every run to a ledger;
  • compare against the prior accepted winner, not only the previous run;
  • keep a holdout or replay check;
  • stop when improvement is within noise, correctness fails, cost exceeds the budget, or the search starts changing more variables than it can explain.

Use phrases like "best measured safe variant" instead of "global optimum" unless the search space was actually exhaustive.

Promotion Gate

A variant cannot become the new default until:

  • correctness tests pass;
  • the performance delta is repeated or explained;
  • rollback is obvious;
  • the change is encoded in source control or a durable runbook;
  • the final summary includes exact commands and measurements.

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/benchmark-optimization-loop of affaan-m/ECC.

Open the folder on GitHubat commit ef648e0

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Benchmark Optimization Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Optimization Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Optimization Loop this skillaffaan-m/ECC274k1 repos~664Automated safety check: PassMIT
Convertremotion-dev/remotion62k—~247Automated safety check: PassCustom licence
Cost Benchmarkruvnet/ruflo74k—~745Automated safety check: NotesMIT
Benchmarkandroidx/androidx6.1k—~1.1kAutomated safety check: PassApache-2.0
Benchmarksamchon/typia5.9k—~1.2kAutomated safety check: PassMIT
Benchmarkingnubjs/nub4.4k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Convert

    remotion-dev/remotion

    Official

    Start the local @remotion/convert app and open it in the Codex browser.

    62k GitHub stars~247 tokensUpdated today
    Media & CreativeAuto-check passed
  • Cost Benchmark

    ruvnet/ruflo

    Run the corpus benchmark — booster locally, optional Gemini/Sonnet/Opus baselines — and persist a verifiable measured-vs-claimed table

    74k GitHub stars~745 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed
  • Benchmark

    samchon/typia

    Defines typia benchmark fixture integrity, result reporting, and publication safeguards.

    5.9k GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Benchmarking

    nubjs/nub

    Comparative install-benchmarking methodology for nub vs npm/pnpm/bun — cold/warm protocol, genuine-cold cache isolation, load-robust measurement, and the anti-juicing honesty bar.

    4.4k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Establishes page load, Core Web Vitals and resource-size baselines, then compares before and after on every pull request to track performance trends over time.

    136k GitHub stars~7.2k tokensUpdated today
    Frontend & DesignAuto-check: notes

More from affaan-m/ECC

All 657 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    274k GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Videodb

    affaan-m/ECC

    Ingest, index, search, edit, and monitor video and audio with the VideoDB Python SDK — upload from files, URLs, or RTSP feeds, build spoken and scene indexes with timestamped search and playable…

    274k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    274k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    274k GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    274k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    274k GitHub stars~3.5k tokensUpdated 2 days ago
    Auto-check passed

Questions about Benchmark Optimization Loop

What does Benchmark Optimization Loop do?

Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…. Benchmark Optimization Loop is an agent skill from affaan-m/ECC. Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest safe variant with reproducible commands.

When should I use Benchmark Optimization Loop?

Benchmark Optimization Loop fits situations like: asked to speed something up; try many variants; run recursive optimization; benchmark latency/throughput/cost.

How do I install Benchmark Optimization Loop in Claude Code?

Run `npx skills add affaan-m/ECC --skill benchmark-optimization-loop -a claude-code`. Or copy the skill folder (skills/benchmark-optimization-loop in affaan-m/ECC) into .claude/skills/benchmark-optimization-loop in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Optimization Loop in Codex?

Run `npx skills add affaan-m/ECC --skill benchmark-optimization-loop -a codex`. Or copy the skill folder (skills/benchmark-optimization-loop in affaan-m/ECC) into .agents/skills/benchmark-optimization-loop in your project. Codex loads it when a task matches its description.

Can I use Benchmark Optimization Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill benchmark-optimization-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-optimization-loop, .gemini/skills/benchmark-optimization-loop, .github/skills/benchmark-optimization-loop and .opencode/skills/benchmark-optimization-loop in your project.

What does Benchmark Optimization Loop need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark Optimization Loop is instructions for the agent only.

Does Benchmark Optimization Loop access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark Optimization Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Optimization Loop use?

Benchmark Optimization Loop is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Optimization Loop use?

About 664 tokens (SKILL.md is roughly 2.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Optimization Loop?

Skills that share tags, products or a category with Benchmark Optimization Loop: Convert (remotion-dev/remotion, 62k stars), Cost Benchmark (ruvnet/ruflo, 74k stars), Benchmark (androidx/androidx, 6.1k stars) and Benchmark (samchon/typia, 5.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Optimization Loop?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 274,360 GitHub stars. The repository holds 657 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.