Self-Improvement Tournament Loop
zereight/gitlab-mcp
Runs an autonomous evolutionary loop that improves a codebase against a measurable benchmark, using agent roles, tournament selection and recorded history until a stop condition.
Optimizes a named target with a measured loop, attributing a workload's cost or scoring variants and keeping winners, on a dedicated branch with a disk log.
$ npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install EveryInc/compound-engineering-plugin ce-optimize --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/EveryInc/compound-engineering-plugin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ce-optimize .claude/skills/ce-optimize && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ce-optimize" agent skill from https://github.com/EveryInc/compound-engineering-plugin/tree/main/skills/ce-optimize into .claude/skills/ce-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ce-optimize", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/EveryInc/compound-engineering-plugin/tree/main/skills/ce-optimizeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install EveryInc/compound-engineering-plugin ce-optimize --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EveryInc/compound-engineering-plugin.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ce-optimize .agents/skills/ce-optimize && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ce-optimize" agent skill from https://github.com/EveryInc/compound-engineering-plugin/tree/main/skills/ce-optimize into .agents/skills/ce-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ce-optimize", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install EveryInc/compound-engineering-plugin ce-optimize --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EveryInc/compound-engineering-plugin.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ce-optimize .cursor/skills/ce-optimize && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ce-optimize" agent skill from https://github.com/EveryInc/compound-engineering-plugin/tree/main/skills/ce-optimize into .cursor/skills/ce-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ce-optimize", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/EveryInc/compound-engineering-plugin.git --path skills/ce-optimize--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install EveryInc/compound-engineering-plugin ce-optimize --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EveryInc/compound-engineering-plugin.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ce-optimize .gemini/skills/ce-optimize && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ce-optimize" agent skill from https://github.com/EveryInc/compound-engineering-plugin/tree/main/skills/ce-optimize into .gemini/skills/ce-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ce-optimize", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install EveryInc/compound-engineering-plugin ce-optimizeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/EveryInc/compound-engineering-plugin.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ce-optimize .github/skills/ce-optimize && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ce-optimize" agent skill from https://github.com/EveryInc/compound-engineering-plugin/tree/main/skills/ce-optimize into .github/skills/ce-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ce-optimize", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install EveryInc/compound-engineering-plugin ce-optimize --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EveryInc/compound-engineering-plugin.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ce-optimize .opencode/skills/ce-optimize && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ce-optimize" agent skill from https://github.com/EveryInc/compound-engineering-plugin/tree/main/skills/ce-optimize into .opencode/skills/ce-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ce-optimize", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ce-optimizeOptimizes a named target with a measured loop, attributing a workload's cost or scoring variants and keeping winners, on a dedicated branch with a disk log.
Use it when a working system's metric should improve and the winning change is not yet known. The agent either attributes the cost of a named workload before searching for implementations, or scores a space of variants and keeps the best, without needing a profile. Confirmed improvements end up as keep or revert commits on an `optimize/` branch named after the spec, with a log written to disk, and you take over at wrap-up.
Invoking the skill authorizes reading the repo, building the measurement harness and, after an approval gate in Phase 1, isolated experiments. The agent asks when spend is uncapped, when a new dependency appears, when wrap-up would push or open a PR, or when you must choose among post-completion options. It reports in ordinary task language, treats unknown duration or cost as unknown, and counts a step as done only after it ran. The folder ships spec and experiment-log schemas, example specs, prompt templates, `decide.mjs` and `experiment-worktree.sh`; ce-debug and ce-work cover diagnosis and known changes.
Read from SKILL.md and the folder at commit cef001f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (JavaScript and Shell, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Measured Optimization Loop loads about 2k tokens when it runs, and up to ~38k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 1,167 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from EveryInc/compound-engineering-plugin at commit cef001f, republished under its MIT licence (© EveryInc). 1,167 words, ~1,970 tokens.
.claude/skills/ce-optimize/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.Outcome: confirmed improvements to the named target live on an optimize/<spec-name> branch, with a disk log. The user takes over at wrap-up.
Intent: the next action is the cheapest step that would change what gets implemented. Attribute the cost of a named workload before searching implementations. Search and keep a scored variant space without requiring a profile.
Done when: a stopping criterion was met, every declared required target is met or another stop was met first, the final state is written and verified on disk, and the user has been given the post-completion options. If the run instead stopped at a check it could not pass, say what blocked it.
Invoking this skill authorizes reading the repo, building the harness, and (after the Phase 1 approval gate) isolated experiments and keep/revert commits on optimize/<spec-name>. Ask when spend is uncapped, when a new dependency appears, when wrap-up would push or open a PR, or when only the user can choose among the post-completion options. Do not ask again to run the next experiment inside those limits.
Independent calls and dispatches that do not depend on each other go in one response. Serialize only real dependencies.
Report findings, user decisions, blockers, and results. During longer work, give occasional updates on what was learned and what remains. Routine preparation and phase or batch transitions need no separate announcement. Keep accounting in the log and final recap unless it affects a current decision.
Explain the target, evidence, and decision in ordinary task language. Workflow labels (such as 'harness' or 'parallel readiness') belong in artifacts unless the user asks about those mechanics. State unknown duration or cost as unknown; caps are limits, not forecasts.
A step is done only after it ran. Describing a measurement, dispatch, or checkpoint is not doing it. Do not end a turn while in-scope work remains merely described.
Use the host's blocking question tool already in the current tool list (match by capability, not by a host-specific name). Presence in the current tool list is proof the tool exists; never call a user-facing question tool to discover whether it exists. If a matching tool is listed but unloaded, use the host's tool-discovery primitive to load that capability: do not search for another host's tool name. Fall back to numbered options on the host's chat surface only when no such tool is in the list or a real question call errors. Never skip the question silently.
Resolve <root> the first time you compose a path under it. Reading learnings under <root>/solutions/ counts as composing one. Give any subagent the resolved path, not the config.
<!-- ce-docs-root:start -->
Resolve the CE artifact root <root> before composing any artifact path.
docs_root from <repo-root>/.compound-engineering/config.yaml only (<repo-root> = git rev-parse --show-toplevel). Do not read it from config.local.yaml. Unset -> <root> is docs, exactly as before..git/. Otherwise stop with an error naming docs_root and the value -- never fall back to docs.<root> as the sole artifact location: create it if absent, compose each path as <root>/<subdir> with this skill's own subdirectory, and never also read docs.<!-- ce-docs-root:end -->
The experiment log on disk is the source of truth. Write order is measure, write, verify, then show the user. Read references/persistence.md now for checkpoints CP-0 through CP-5, the file layout, and resume. The phases below mark where each checkpoint falls.
Four phases run in order. Each one names the reference it cannot start without. A fresh run skips none of them: a harder optimization spends longer in a phase, it does not run fewer phases.
A resume is not a fresh run. On a resume, re-enter Phase 0 only far enough to detect the run and to recover any result.yaml markers the log is missing. Then continue from the phase the log records: skip the work the log proves finished, and re-enter any approval check it does not. A checkpoint proves the work that produced it, never a user decision: the log holds no record of approval, so a resume that has not seen the user approve presents the Phase 1 gate again.
Phase 0: Setup. The input is a goal, or a path to a spec YAML. It comes from the user or from a calling skill. If neither supplied one, ask: "What would you like to optimize? Describe the goal, or provide a path to an optimization spec YAML file." Load or build the spec and save it (CP-0): read references/spec.md. Then search prior learnings, detect run identity, and create the branch and scratch space. Read references/measurement.md for the rest of Phase 0 and Phase 1.
Phase 1: Measurement scaffolding. Build or validate the harness, write the baseline (CP-1), probe parallelism, check the worktree budget. Two gates stop the run:
scope.mutable or scope.immutable has uncommitted changes. The reference defines the check and what to ask for.stopping.max_total_cost_usd caps the whole run, while the judge cap counts only scoring. With neither set, spend is uncapped apart from the iteration and hour limits. Offer proceed, fix issues, and adjust spec. Adjusting the spec is only available while the log holds nothing derived from it (no hypothesis backlog and no experiments) and it sends the run back through Phase 1 so the baseline matches the new spec. Once anything derived from the spec is on file, the spec is fixed for the run. Do not enter Phase 2 until the user explicitly approves. Then re-read the spec and baseline from disk.Phase 2: Hypothesis generation. Analyze the current approach, rank the hypotheses, record the backlog (CP-2). Do not dispatch an implementation experiment while a cheaper locating measurement would change keep or skip. Read references/loop.md for this phase and Phase 3. One gate: dependency pre-approval. Collect every new dependency across all hypotheses and present the full list for bulk approval. A dependency the user does not approve stays in the backlog, is skipped in batch selection, and comes back at wrap-up.
Phase 3: Optimization loop. Select a batch, dispatch experiments, persist each result as it lands (CP-3), evaluate with scripts/decide.mjs, update state and the digest (CP-4), then check whether to stop. Stop as soon as any one of seven criteria holds: every declared required target is met, max iterations, max hours, a spend cap reached, plateau, a user interrupt, or no runnable hypothesis left. references/loop.md states each one exactly. Otherwise start the next batch.
Phase 4: Wrap-up. Read references/wrap-up.md for the deferred hypotheses, the summary, what is preserved, cleanup, and the post-completion options to present. CP-5 marks the log final. Write it only after the user picks an option that does not return to Phase 3. Two options do return: Continue, and approving a deferred dependency.
© EveryInc, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 19 other files (scripts, references) in skills/ce-optimize of EveryInc/compound-engineering-plugin.
Open the folder on GitHubat commit cef001f
Measured Optimization Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Measured Optimization Loop this skillEveryInc/compound-engineering-plugin | 25k | — | ~2k | Automated safety check: Pass | MIT | |
| Self-Improvement Tournament Loopzereight/gitlab-mcp | 2k | 1 repos | ~1.8k | Automated safety check: Warn | MIT | |
| Hiverllm-org/hive | 216 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Session Profilertamdogood/builder-essential-skills | 218 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Build Dotnet Agentnewrelic/newrelic-dotnet-agent | 117 | — | ~871 | Automated safety check: Pass | Apache-2.0 | |
| Effective Harnessesliangdabiao/exa-research-mcp-skill | 109 | — | ~1.5k | Automated safety check: Pass | None |
zereight/gitlab-mcp
Runs an autonomous evolutionary loop that improves a codebase against a measurable benchmark, using agent roles, tournament selection and recorded history until a stop condition.
rllm-org/hive
Run the hive experiment loop — autonomous iteration on a shared task.
tamdogood/builder-essential-skills
Profile and debug Hermes sessions from their JSONL transcripts.
newrelic/newrelic-dotnet-agent
Build the New Relic .NET agent locally and validate that code changes compile.
liangdabiao/exa-research-mcp-skill
Long-running agent project harness for Codex, OpenClaw, Claude Code, and other coding agents.
UniClipboard/UniClipboard
Agent Loop that runs local pre-flight CI checks, auto-fixes issues, creates/pushes the PR, monitors CI, and loops until all checks pass.
EveryInc/compound-engineering-plugin
Records one solved and verified problem as a durable learning in the repository, but only when the reasoning is not already clear from the final code, tests or docs.
EveryInc/compound-engineering-plugin
Audits a repo's stored learnings against the current codebase, fixes stale, overlapping or superseded docs and reports on every document.
EveryInc/compound-engineering-plugin
Builds a throwaway prototype at just the fidelity needed to settle a specific how-it-should-work-or-feel question, before committing to an approach other work will treat as fixed.
EveryInc/compound-engineering-plugin
Checks Compound Engineering plugin health and repo-local config, or scaffolds a Compound Pack when you ask for one by id.
EveryInc/compound-engineering-plugin
Watches an open GitHub pull request over time, routing review comments and CI failures to other skills until the PR is ready to merge.
EveryInc/compound-engineering-plugin
Turns a vague or ambitious feature idea into a requirements-only plan through dialogue with you, sized to the work, before any code is written.
Categories
Optimizes a named target with a measured loop, attributing a workload's cost or scoring variants and keeping winners, on a dedicated branch with a disk log. Use it when a working system's metric should improve and the winning change is not yet known. The agent either attributes the cost of a named workload before searching for implementations, or scores a space of variants and keeps the best, without needing a profile.
Measured Optimization Loop fits situations like: reducing the runtime or cost of a specific working workload when the fix is not obvious; scoring several variants of a prompt, config or approach and keeping the winner; running experiments in isolated worktrees with keep-or-revert commits.
Run `npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a claude-code`. Or copy the skill folder (skills/ce-optimize in EveryInc/compound-engineering-plugin) into .claude/skills/ce-optimize in your project. Claude Code loads it when a task matches its description.
Run `npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a codex`. Or copy the skill folder (skills/ce-optimize in EveryInc/compound-engineering-plugin) into .agents/skills/ce-optimize in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EveryInc/compound-engineering-plugin --skill ce-optimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ce-optimize, .gemini/skills/ce-optimize, .github/skills/ce-optimize and .opencode/skills/ce-optimize in your project.
Going by SKILL.md and its folder, Measured Optimization Loop needs JavaScript and a shell for the scripts in its folder and the command-line tools its instructions call (git). Our summary lists: A git repository where the agent can create an optimize branch and worktrees; A measurable metric or judge for the target.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Measured Optimization Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 36k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Measured Optimization Loop: Self-Improvement Tournament Loop (zereight/gitlab-mcp, 2k stars), Hive (rllm-org/hive, 216 stars), Session Profiler (tamdogood/builder-essential-skills, 218 stars) and Build Dotnet Agent (newrelic/newrelic-dotnet-agent, 117 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
EveryInc (a GitHub organization) maintains it in EveryInc/compound-engineering-plugin, which has 25,412 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on October 7, 2026.
Source: EveryInc/compound-engineering-plugin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.