Agent skill

Measurement Discipline

by closedloop-ai in closedloop-ai/claude-plugins

A skill your agent uses when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past…

Apache-2.0Auto-check passed

Install Measurement Discipline

skills CLI
$ npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install closedloop-ai/claude-plugins measurement-discipline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/closedloop-core/skills/measurement-discipline .claude/skills/measurement-discipline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
measurement-discipline
GitHub stars
122
Token cost
~2.5k tokens
SKILL.md length
1,150 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past…

  • Running a performance experiment
  • SKILL.md covers When invoked, The rules, What an entry contains and Worked example: a success, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Evaluating whether a change actually helped

What it does

Measurement Discipline is an agent skill from closedloop-ai/claude-plugins. Use when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past optimization attempts. Enforces one-variable experiments, noise floors measured from repeated runs, a required mechanism behind any claimed win, and writing refutations down so failed attempts are not re-litigated.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Open-source Claude Code plugins for multi-agent software delivery. Plan-first SDLC workflow, code review, LLM quality judges, and self-learning — grounded in your codebase… The licence is Apache-2.0.

When your agent uses it

  • Running a performance experiment
  • Evaluating whether a change actually helped
  • Checking past optimization attempts

Example prompts

  • “/measurement-discipline”

What it can do on your machine

Read from SKILL.md and the folder at commit 4600742. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Measurement Discipline loads about 2.5k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 1,150 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from closedloop-ai/claude-plugins at commit 4600742, republished under its Apache-2.0 licence (© closedloop-ai). 1,150 words, ~2,455 tokens.

Download SKILL.mdSave it as .claude/skills/measurement-discipline/SKILL.md (or your agent's skills folder).
name
measurement-discipline
description
Use when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past optimization attempts. Enforces one-variable experiments, noise floors measured from repeated runs, a required mechanism behind any claimed win, and writing refutations down so failed attempts are not re-litigated.

Measure One Thing, Then Write It Down

When invoked

Apply the rules below to the measurement or optimization task at hand: baseline before touching anything, change one variable, keep a control where the harness could drift, require a mechanism for any claimed win, and record a loss as REFUTED rather than letting it go unwritten. If the project keeps a measurement log, search it by symptom before proposing or diagnosing anything, then append the resulting entry, success or refutation, as part of finishing the task.

Use this as team practice, or paste it into an AI agent's instructions; it is written as instructions either way. It applies to anything measurable: query latency, cache hit rates, build times, bundle size, model parameters, cost per request. Its purpose is to make every optimization question cost one search instead of one rediscovery, and to make sure the answer you find was honestly obtained.

The rules

1. Baseline before you touch anything, and get the noise floor from repeated runs. One before-number and one after-number cannot tell a real effect from a slow afternoon. Discard the first run as warmup, take at least five or six more, record every one, and state the median and the spread. The spread is your noise floor; nothing smaller is a result.

2. Change one variable. When two things change together and the number moves, you have learned nothing about either, and you will be back here next quarter to learn nothing again. If you must bundle changes to ship, still measure them one at a time first.

3. Keep a control when the harness could drift. Machines get busy, caches warm, background jobs start; a control on a path your change cannot possibly touch tells you how much of the delta was the world rather than you. If the control moves as much as the candidate, your noise floor is wider than you thought, and you say so.

4. A win must beat the noise floor AND come with a mechanism. A delta without a mechanism is a coincidence you will re-litigate, because nobody, including you in six months, can tell it apart from luck. Name what changed: call count went from 1 + N to 2, the plan switched to an index scan, the allocation left the hot path. No mechanism means the experiment is not finished.

5. Product performance goals need product-level evidence. When the stated goal is user-facing speed, such as page load time or a route's production latency, a helper microbenchmark is support evidence, not acceptance by itself. Tie each experiment back to the product surface and stage it is supposed to improve. If the local helper delta is negligible while the page or route remains slow, keep investigating the next bottleneck or record why that ticket's scope cannot move the product metric. When production telemetry exists, such as Datadog RUM/APM page or route spans, record the exact query/window and the baseline or comparison values available there. If access or export is unavailable, record that as an evidence gap instead of substituting local-only numbers for production behavior.

6. A loss is recorded as REFUTED, with its conditions and a re-check trigger. Refuted work that is not written down gets paid for a second and third time by the next person with the same good idea. Record what you tried, the numbers, why it failed, and the conditions under which the answer would change. Results are tied to conditions: a 10x change in data volume, traffic, or hardware reopens them; silently trusting an old result does not.

7. The log is append-only and dated, and every number carries the command that produced it. An unreproducible number is an opinion with decimal places. Never edit a past entry to agree with a newer finding; append the newer finding and say what it supersedes. Pin anything that moves under you: input set, revision, seed, data snapshot, sample size.

8. Label projections as projections. An extrapolation that reads like a measurement will be quoted back at you as one, usually in a decision meeting. Write "projection at this measured rate" and show the rate it came from.

9. Search the log by SYMPTOM before you propose anything, and before you explain anything. Rediscovery is the expensive failure mode, and prior findings are almost always filed under the symptom that prompted them, in prose, not under the name you would give your fix. The trigger is not "am I about to change something", it is "am I about to reach a conclusion about an area that may already have one", so diagnosis counts too: following evidence is the activity most likely to re-derive a recorded finding, and one search for the symptom ends what would otherwise be twenty steps reconstructing something already written down.

Show full SKILL.md (355 more words)Show less

10. Decisions cite measurements, and reversing one is fine when you name what changed. Finding a prior decision does not end the argument; it ends proposing in ignorance. Either follow it or reverse it deliberately and record the reversal with the delta that justifies it (the population grew from 5 to 33, the job went from nightly to every fifteen minutes). Flipping a recorded decision silently leaves the next reader to re-derive the original reasoning and flip it back.

11. Do not stop after the first measured win; stop at the floor. A performance task is not complete merely because one change helped. After an adopted improvement, re-measure the target path, identify the next dominant bottleneck or absence of one, and run the next in-scope one-variable experiment when a credible mechanism remains. Stop when consecutive in-scope experiments land within noise, the remaining opportunity is negligible against the measured latency budget, or the next plausible mechanism would cross an explicit scope, safety, product, or requirements boundary. Record that stopping condition.

What an entry contains

Date and a one-line title with the verdict in it. The single variable. Baseline with the exact command and every run. The change. The after numbers. The control. The mechanism. The verdict against the noise floor. The decision, plus proof that temporary scaffolding was removed. A closing block separating measured from not measured, with the re-check trigger.

Worked example: a success

## 2026-03-14: batch the per-row lookup in the order export, ADOPTED

One variable. Only the lookup changed; same input file, same machine, same warm cache.
1. Baseline. `bench export orders`, 1 discarded warmup then 6 warm runs:
   812 / 798 / 826 / 805 / 819 / 803 ms. Median 807 ms, spread 28 ms (3.5%). That is the noise floor.
2. Change. Per-row lookup replaced by one keyed batch fetch. Diff touches one function.
3. After. Same command, 6 warm runs: 214 / 209 / 221 / 213 / 210 / 216 ms. Median 213 ms.
   Control (`bench export invoices`, a path that does not use the lookup): 402 ms then 397 ms, flat.
4. Mechanism. Query count went from 1 + N (N = 4,180) to 2. Measured round trip is 0.14 ms,
   so 4,178 removed round trips account for 585 ms of the 594 ms saved. Not a coincidence.
5. Verdict: 3.8x, twenty times the 28 ms noise floor. ADOPT.
Labelling. Measured: all timings, the query counts, the control. Not measured: cold cache, and
exports above 50k rows. Re-check if row counts grow 10x, where the single batch may become the bound.

Worked example: a refutation

## 2026-03-16: index on the merge-date column, REFUTED

One variable. Index created on a scratch copy of the database only; no schema file, no code, no migration.
1. Baseline. `bench report throughput`, 1 discarded warmup then 6 warm runs:
   19.6 / 18.6 / 19.5 / 19.6 / 20.3 / 18.6 ms. Median 19.55 ms, spread 1.7 ms.
2. After index. 12 warm runs (second batch interleaved with the control to separate the index
   from time-based drift): median 19.3 ms, range 18.3 to 20.1. Delta 0.25 ms (1.3%), inside the spread.
3. Mechanism. The query plan shows the lookup already resolved by the primary key on the join
   column, narrowing to under 5 rows before the aggregate. The date index is not on that access
   path and is never consulted, so no mechanism existed by which it could help.
4. Control. An untouched report drifted 46.7 to 43.0 to 41.9 ms across the same session and
   stayed there after the index was dropped. That 10% wander is the real noise floor today,
   several times the effect under test. Reported, not smoothed away.
5. Verdict: WITHIN NOISE. REFUTED. Index dropped; index list and schema dump match the shipped
   state, integrity check ok, no tracked file changed.
Labelling. Measured: everything above, on tables of 2,110 and 26 rows. Not measured: production
scale. Re-check trigger: if either table grows 10x, re-run rather than citing this entry.

If you are an AI agent

Before proposing or diagnosing, search the log for the symptom and for the mechanism you are about to name, and say what you searched for. Claim nothing without a number, and never present an estimate in the shape of a measurement. Run the baseline before the first edit. Change one thing. Report every run, not the best one. If the result is within noise, say "within noise" and write the refutation; a refutation is a completed piece of work, not a failure to report. Remove your scaffolding and prove it is gone. Append the entry the same session, because the entry is the deliverable and the code change is only sometimes.

© closedloop-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/closedloop-core/skills/measurement-discipline of closedloop-ai/claude-plugins.

Open the folder on GitHubat commit 4600742

Compare with similar skills

Measurement Discipline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Measurement Discipline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Measurement Discipline this skillclosedloop-ai/claude-plugins122—~2.5kAutomated safety check: PassApache-2.0
Benchmark Optimization Loopaffaan-m/ECC276k1 repos~664Automated safety check: PassMIT
Sglang Diffusion Benchmark Profilesgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
Linkedin Profile Optimizersickn33/agentic-awesome-skills47k1 repos~2.6kAutomated safety check: PassMIT
Linkedin Profile Optimizerdavila7/claude-code-templates32k2 repos~2.6kAutomated safety check: PassMIT
Benchmarkaffaan-m/ECC276k3 repos~654Automated safety check: PassMIT

Similar skills

  • Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…

    276k GitHub starsUsed in 1 repo~664 tokens
    Auto-check passed
  • A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

    37k GitHub starsUsed in 2 repos~2.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Linkedin Profile Optimizer

    sickn33/agentic-awesome-skills

    High-intent expert for LinkedIn profile checks and SEO optimization.

    47k GitHub starsUsed in 1 repo~2.6k tokens
    Business, Finance & HRAuto-check passed
  • Linkedin Profile Optimizer

    davila7/claude-code-templates

    Optimize a LinkedIn profile for searchability, recruiter visibility, and engagement.

    32k GitHub starsUsed in 2 repos~2.6k tokens
    Business, Finance & HRAuto-check passed
  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    276k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Measured Optimization Loop

    EveryInc/compound-engineering-plugin

    Optimizes a named target with a measured loop, attributing a workload's cost or scoring variants and keeping winners, on a dedicated branch with a disk log.

    25k GitHub stars~2k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from closedloop-ai/claude-plugins

All 43 skills in this repo
  • Codex Review

    closedloop-ai/claude-plugins

    Run Codex to review a plan file and return structured feedback with a verdict.

    122 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check: notes
  • Critic Cache

    closedloop-ai/claude-plugins

    Check if critic reviews are still valid before re-running Phase 2.5 critics.

    122 GitHub stars~528 tokensUpdated yesterday
    Auto-check: notes
  • Cross Repo Cache

    closedloop-ai/claude-plugins

    Check if cross-repo coordinator results can be reused, avoiding redundant Sonnet agent launches.

    122 GitHub stars~683 tokensUpdated yesterday
    Auto-check: notes
  • Eval Cache

    closedloop-ai/claude-plugins

    Check for a cached plan-evaluation.json result before launching the plan-evaluator agent.

    122 GitHub stars~516 tokensUpdated yesterday
    Auto-check: notes
  • Find Plugin File

    closedloop-ai/claude-plugins

    This skill should be used when needing to locate files within the Claude Code plugins cache directory (~/.claude/plugins/cache).

    122 GitHub stars~812 tokensUpdated yesterday
    Auto-check passed
  • Handoff

    closedloop-ai/claude-plugins

    Finish a vibe session in symphony-alpha and hand it to whoever picks it up next, usually design and then engineering.

    122 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed

Questions about Measurement Discipline

What does Measurement Discipline do?

A skill your agent uses when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past…. Measurement Discipline is an agent skill from closedloop-ai/claude-plugins. Use when optimizing, benchmarking, profiling, or running a performance experiment; when evaluating whether a change actually helped; or when recording or checking past optimization attempts.

When should I use Measurement Discipline?

Measurement Discipline fits situations like: running a performance experiment; evaluating whether a change actually helped; checking past optimization attempts.

How do I install Measurement Discipline in Claude Code?

Run `npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a claude-code`. Or copy the skill folder (plugins/closedloop-core/skills/measurement-discipline in closedloop-ai/claude-plugins) into .claude/skills/measurement-discipline in your project. Claude Code loads it when a task matches its description.

How do I install Measurement Discipline in Codex?

Run `npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a codex`. Or copy the skill folder (plugins/closedloop-core/skills/measurement-discipline in closedloop-ai/claude-plugins) into .agents/skills/measurement-discipline in your project. Codex loads it when a task matches its description.

Can I use Measurement Discipline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add closedloop-ai/claude-plugins --skill measurement-discipline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/measurement-discipline, .gemini/skills/measurement-discipline, .github/skills/measurement-discipline and .opencode/skills/measurement-discipline in your project.

What does Measurement Discipline need to run?

SKILL.md names no scripts, command-line tools or credentials: Measurement Discipline is instructions for the agent only.

Does Measurement Discipline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Measurement Discipline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Measurement Discipline use?

Measurement Discipline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Measurement Discipline use?

About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Measurement Discipline?

Skills that share tags, products or a category with Measurement Discipline: Benchmark Optimization Loop (affaan-m/ECC, 276k stars), Sglang Diffusion Benchmark Profile (sgl-project/sglang, 37k stars), Linkedin Profile Optimizer (sickn33/agentic-awesome-skills, 47k stars) and Linkedin Profile Optimizer (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Measurement Discipline?

closedloop-ai (a GitHub organization) maintains it in closedloop-ai/claude-plugins, which has 122 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on October 8, 2026.

Source: closedloop-ai/claude-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.