Benchmark Optimization Loop
affaan-m/ECC
Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…
A skill your agent uses when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past…
$ npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install closedloop-ai/claude-plugins measurement-discipline --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/closedloop-core/skills/measurement-discipline .claude/skills/measurement-discipline && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "measurement-discipline" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/closedloop-core/skills/measurement-discipline into .claude/skills/measurement-discipline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "measurement-discipline", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/closedloop-core/skills/measurement-disciplineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install closedloop-ai/claude-plugins measurement-discipline --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/closedloop-core/skills/measurement-discipline .agents/skills/measurement-discipline && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "measurement-discipline" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/closedloop-core/skills/measurement-discipline into .agents/skills/measurement-discipline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "measurement-discipline", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install closedloop-ai/claude-plugins measurement-discipline --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/closedloop-core/skills/measurement-discipline .cursor/skills/measurement-discipline && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "measurement-discipline" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/closedloop-core/skills/measurement-discipline into .cursor/skills/measurement-discipline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "measurement-discipline", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/closedloop-ai/claude-plugins.git --path plugins/closedloop-core/skills/measurement-discipline--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install closedloop-ai/claude-plugins measurement-discipline --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/closedloop-core/skills/measurement-discipline .gemini/skills/measurement-discipline && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "measurement-discipline" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/closedloop-core/skills/measurement-discipline into .gemini/skills/measurement-discipline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "measurement-discipline", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install closedloop-ai/claude-plugins measurement-disciplineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/closedloop-core/skills/measurement-discipline .github/skills/measurement-discipline && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "measurement-discipline" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/closedloop-core/skills/measurement-discipline into .github/skills/measurement-discipline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "measurement-discipline", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install closedloop-ai/claude-plugins measurement-discipline --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/closedloop-core/skills/measurement-discipline .opencode/skills/measurement-discipline && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "measurement-discipline" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/closedloop-core/skills/measurement-discipline into .opencode/skills/measurement-discipline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "measurement-discipline", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
measurement-disciplineA skill your agent uses when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past…
Measurement Discipline is an agent skill from closedloop-ai/claude-plugins. Use when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past optimization attempts. Enforces one-variable experiments, noise floors measured from repeated runs, a required mechanism behind any claimed win, and writing refutations down so failed attempts are not re-litigated.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Open-source Claude Code plugins for multi-agent software delivery. Plan-first SDLC workflow, code review, LLM quality judges, and self-learning — grounded in your codebase… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 4600742. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Measurement Discipline loads about 2.5k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 1,150 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from closedloop-ai/claude-plugins at commit 4600742, republished under its Apache-2.0 licence (© closedloop-ai). 1,150 words, ~2,455 tokens.
.claude/skills/measurement-discipline/SKILL.md (or your agent's skills folder).Apply the rules below to the measurement or optimization task at hand: baseline before touching anything, change one variable, keep a control where the harness could drift, require a mechanism for any claimed win, and record a loss as REFUTED rather than letting it go unwritten. If the project keeps a measurement log, search it by symptom before proposing or diagnosing anything, then append the resulting entry, success or refutation, as part of finishing the task.
Use this as team practice, or paste it into an AI agent's instructions; it is written as instructions either way. It applies to anything measurable: query latency, cache hit rates, build times, bundle size, model parameters, cost per request. Its purpose is to make every optimization question cost one search instead of one rediscovery, and to make sure the answer you find was honestly obtained.
1. Baseline before you touch anything, and get the noise floor from repeated runs. One before-number and one after-number cannot tell a real effect from a slow afternoon. Discard the first run as warmup, take at least five or six more, record every one, and state the median and the spread. The spread is your noise floor; nothing smaller is a result.
2. Change one variable. When two things change together and the number moves, you have learned nothing about either, and you will be back here next quarter to learn nothing again. If you must bundle changes to ship, still measure them one at a time first.
3. Keep a control when the harness could drift. Machines get busy, caches warm, background jobs start; a control on a path your change cannot possibly touch tells you how much of the delta was the world rather than you. If the control moves as much as the candidate, your noise floor is wider than you thought, and you say so.
4. A win must beat the noise floor AND come with a mechanism. A delta without a mechanism is a coincidence you will re-litigate, because nobody, including you in six months, can tell it apart from luck. Name what changed: call count went from 1 + N to 2, the plan switched to an index scan, the allocation left the hot path. No mechanism means the experiment is not finished.
5. Product performance goals need product-level evidence. When the stated goal is user-facing speed, such as page load time or a route's production latency, a helper microbenchmark is support evidence, not acceptance by itself. Tie each experiment back to the product surface and stage it is supposed to improve. If the local helper delta is negligible while the page or route remains slow, keep investigating the next bottleneck or record why that ticket's scope cannot move the product metric. When production telemetry exists, such as Datadog RUM/APM page or route spans, record the exact query/window and the baseline or comparison values available there. If access or export is unavailable, record that as an evidence gap instead of substituting local-only numbers for production behavior.
6. A loss is recorded as REFUTED, with its conditions and a re-check trigger. Refuted work that is not written down gets paid for a second and third time by the next person with the same good idea. Record what you tried, the numbers, why it failed, and the conditions under which the answer would change. Results are tied to conditions: a 10x change in data volume, traffic, or hardware reopens them; silently trusting an old result does not.
7. The log is append-only and dated, and every number carries the command that produced it. An unreproducible number is an opinion with decimal places. Never edit a past entry to agree with a newer finding; append the newer finding and say what it supersedes. Pin anything that moves under you: input set, revision, seed, data snapshot, sample size.
8. Label projections as projections. An extrapolation that reads like a measurement will be quoted back at you as one, usually in a decision meeting. Write "projection at this measured rate" and show the rate it came from.
9. Search the log by SYMPTOM before you propose anything, and before you explain anything. Rediscovery is the expensive failure mode, and prior findings are almost always filed under the symptom that prompted them, in prose, not under the name you would give your fix. The trigger is not "am I about to change something", it is "am I about to reach a conclusion about an area that may already have one", so diagnosis counts too: following evidence is the activity most likely to re-derive a recorded finding, and one search for the symptom ends what would otherwise be twenty steps reconstructing something already written down.
10. Decisions cite measurements, and reversing one is fine when you name what changed. Finding a prior decision does not end the argument; it ends proposing in ignorance. Either follow it or reverse it deliberately and record the reversal with the delta that justifies it (the population grew from 5 to 33, the job went from nightly to every fifteen minutes). Flipping a recorded decision silently leaves the next reader to re-derive the original reasoning and flip it back.
11. Do not stop after the first measured win; stop at the floor. A performance task is not complete merely because one change helped. After an adopted improvement, re-measure the target path, identify the next dominant bottleneck or absence of one, and run the next in-scope one-variable experiment when a credible mechanism remains. Stop when consecutive in-scope experiments land within noise, the remaining opportunity is negligible against the measured latency budget, or the next plausible mechanism would cross an explicit scope, safety, product, or requirements boundary. Record that stopping condition.
Date and a one-line title with the verdict in it. The single variable. Baseline with the exact command and every run. The change. The after numbers. The control. The mechanism. The verdict against the noise floor. The decision, plus proof that temporary scaffolding was removed. A closing block separating measured from not measured, with the re-check trigger.
## 2026-03-14: batch the per-row lookup in the order export, ADOPTED
One variable. Only the lookup changed; same input file, same machine, same warm cache.
1. Baseline. `bench export orders`, 1 discarded warmup then 6 warm runs:
812 / 798 / 826 / 805 / 819 / 803 ms. Median 807 ms, spread 28 ms (3.5%). That is the noise floor.
2. Change. Per-row lookup replaced by one keyed batch fetch. Diff touches one function.
3. After. Same command, 6 warm runs: 214 / 209 / 221 / 213 / 210 / 216 ms. Median 213 ms.
Control (`bench export invoices`, a path that does not use the lookup): 402 ms then 397 ms, flat.
4. Mechanism. Query count went from 1 + N (N = 4,180) to 2. Measured round trip is 0.14 ms,
so 4,178 removed round trips account for 585 ms of the 594 ms saved. Not a coincidence.
5. Verdict: 3.8x, twenty times the 28 ms noise floor. ADOPT.
Labelling. Measured: all timings, the query counts, the control. Not measured: cold cache, and
exports above 50k rows. Re-check if row counts grow 10x, where the single batch may become the bound.## 2026-03-16: index on the merge-date column, REFUTED
One variable. Index created on a scratch copy of the database only; no schema file, no code, no migration.
1. Baseline. `bench report throughput`, 1 discarded warmup then 6 warm runs:
19.6 / 18.6 / 19.5 / 19.6 / 20.3 / 18.6 ms. Median 19.55 ms, spread 1.7 ms.
2. After index. 12 warm runs (second batch interleaved with the control to separate the index
from time-based drift): median 19.3 ms, range 18.3 to 20.1. Delta 0.25 ms (1.3%), inside the spread.
3. Mechanism. The query plan shows the lookup already resolved by the primary key on the join
column, narrowing to under 5 rows before the aggregate. The date index is not on that access
path and is never consulted, so no mechanism existed by which it could help.
4. Control. An untouched report drifted 46.7 to 43.0 to 41.9 ms across the same session and
stayed there after the index was dropped. That 10% wander is the real noise floor today,
several times the effect under test. Reported, not smoothed away.
5. Verdict: WITHIN NOISE. REFUTED. Index dropped; index list and schema dump match the shipped
state, integrity check ok, no tracked file changed.
Labelling. Measured: everything above, on tables of 2,110 and 26 rows. Not measured: production
scale. Re-check trigger: if either table grows 10x, re-run rather than citing this entry.Before proposing or diagnosing, search the log for the symptom and for the mechanism you are about to name, and say what you searched for. Claim nothing without a number, and never present an estimate in the shape of a measurement. Run the baseline before the first edit. Change one thing. Report every run, not the best one. If the result is within noise, say "within noise" and write the refutation; a refutation is a completed piece of work, not a failure to report. Remove your scaffolding and prove it is gone. Append the entry the same session, because the entry is the deliverable and the code change is only sometimes.
© closedloop-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/closedloop-core/skills/measurement-discipline of closedloop-ai/claude-plugins.
Open the folder on GitHubat commit 4600742
Measurement Discipline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Measurement Discipline this skillclosedloop-ai/claude-plugins | 122 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Benchmark Optimization Loopaffaan-m/ECC | 276k | 1 repos | ~664 | Automated safety check: Pass | MIT | |
| Sglang Diffusion Benchmark Profilesgl-project/sglang | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Linkedin Profile Optimizersickn33/agentic-awesome-skills | 47k | 1 repos | ~2.6k | Automated safety check: Pass | MIT | |
| Linkedin Profile Optimizerdavila7/claude-code-templates | 32k | 2 repos | ~2.6k | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 276k | 3 repos | ~654 | Automated safety check: Pass | MIT |
affaan-m/ECC
Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…
sgl-project/sglang
A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.
sickn33/agentic-awesome-skills
High-intent expert for LinkedIn profile checks and SEO optimization.
davila7/claude-code-templates
Optimize a LinkedIn profile for searchability, recruiter visibility, and engagement.
affaan-m/ECC
Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…
EveryInc/compound-engineering-plugin
Optimizes a named target with a measured loop, attributing a workload's cost or scoring variants and keeping winners, on a dedicated branch with a disk log.
closedloop-ai/claude-plugins
Run Codex to review a plan file and return structured feedback with a verdict.
closedloop-ai/claude-plugins
Check if critic reviews are still valid before re-running Phase 2.5 critics.
closedloop-ai/claude-plugins
Check if cross-repo coordinator results can be reused, avoiding redundant Sonnet agent launches.
closedloop-ai/claude-plugins
Check for a cached plan-evaluation.json result before launching the plan-evaluator agent.
closedloop-ai/claude-plugins
This skill should be used when needing to locate files within the Claude Code plugins cache directory (~/.claude/plugins/cache).
closedloop-ai/claude-plugins
Finish a vibe session in symphony-alpha and hand it to whoever picks it up next, usually design and then engineering.
A skill your agent uses when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past…. Measurement Discipline is an agent skill from closedloop-ai/claude-plugins. Use when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past optimization attempts.
Measurement Discipline fits situations like: running a performance experiment; evaluating whether a change actually helped; checking past optimization attempts.
Run `npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a claude-code`. Or copy the skill folder (plugins/closedloop-core/skills/measurement-discipline in closedloop-ai/claude-plugins) into .claude/skills/measurement-discipline in your project. Claude Code loads it when a task matches its description.
Run `npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a codex`. Or copy the skill folder (plugins/closedloop-core/skills/measurement-discipline in closedloop-ai/claude-plugins) into .agents/skills/measurement-discipline in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/measurement-discipline, .gemini/skills/measurement-discipline, .github/skills/measurement-discipline and .opencode/skills/measurement-discipline in your project.
SKILL.md names no scripts, command-line tools or credentials: Measurement Discipline is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Measurement Discipline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Measurement Discipline: Benchmark Optimization Loop (affaan-m/ECC, 276k stars), Sglang Diffusion Benchmark Profile (sgl-project/sglang, 37k stars), Linkedin Profile Optimizer (sickn33/agentic-awesome-skills, 47k stars) and Linkedin Profile Optimizer (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
closedloop-ai (a GitHub organization) maintains it in closedloop-ai/claude-plugins, which has 122 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on October 8, 2026.
Source: closedloop-ai/claude-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.