Retune a skill corpus for a new model, measurement-first: mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a…

MITAuto-check passed

Install Ce Retune

skills CLI
$ npx skills add EveryInc/compound-engineering-plugin --skill ce-retune -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EveryInc/compound-engineering-plugin ce-retune --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EveryInc/compound-engineering-plugin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ce-retune .claude/skills/ce-retune && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ce-retune
GitHub stars
25k
Token cost
~1.2k tokens
SKILL.md length
672 words
Files
8 (incl. references)
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Retune a skill corpus for a new model, measurement-first: mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a…

  • Works in 3 steps: A run archive or a harness that produces… → A build selector — the harness can point… → A repeatable task the corpus actually…
  • SKILL.md covers Phase 0: the measurement gate… and The phases
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ce Retune is an agent skill from EveryInc/compound-engineering-plugin. Retune a skill corpus for a new model, measurement-first: mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a pre-registered bar clears. Requires a benchmark harness that can A/B two builds of the corpus; refuses without one.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `agents/openai.yaml`, `references/baseline-mining.md` and `references/corpus-audit.md`).

The repository describes itself as: Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more. The licence is MIT.

Example prompts

  • “/ce-retune”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A run archive or a harness that produces one — per-run logs carrying the tool-call trace, a terminal marker, token counts, and the final…
  2. A build selector — the harness can point a run at a specific source checkout of the corpus (a --plugin-dir-style override, a configurable…
  3. A repeatable task the corpus actually executes end to end.

What it can do on your machine

Read from SKILL.md and the folder at commit 67035e9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ce Retune loads about 1.2k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 672 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EveryInc/compound-engineering-plugin at commit 67035e9, republished under its MIT licence (© EveryInc). 672 words, ~1,166 tokens.

Download SKILL.mdSave it as .claude/skills/ce-retune/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
ce-retune
description
Retune a skill corpus for a new model, measurement-first: mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a pre-registered bar clears. Requires a benchmark harness that can A/B two builds of the corpus; refuses without one.
disable-model-invocation
true
argument-hint
[target model or symptom] [path to the corpus, defaults to ./skills] [bar:<n> consecutive clean runs]

Retune a Corpus for a New Model

A corpus that degrades on a new model is a measurement problem before it is a writing problem: rewriting what looks wrong produces a plausible fix list and no way to know whether any item mattered.

Outcome: a corpus whose measured behavior on the target model clears a bar registered before any change, with the regression classes removed and each removal attributable.

Done: the bar is cleared, or the run reports the specific claim it could not support. A green test suite is not done: it proves nothing broke, not that behavior improved.

Non-goal: word reduction. Leanness and performance are separate programs that share a corpus; only one of them is the result here. Report completion, not word count.

Phase 0: the measurement gate — check this first

This skill cannot run without a way to observe behavior. Check for all three, and name whichever is missing:

  1. A run archive or a harness that produces one — per-run logs carrying the tool-call trace, a terminal marker, token counts, and the final message.
  2. A build selector — the harness can point a run at a specific source checkout of the corpus (a --plugin-dir-style override, a configurable skills path, an env var), so two builds are comparable under one runner.
  3. A repeatable task the corpus actually executes end to end.

If any is missing, stop and say so, naming what to build. Do not fall back to a static audit and present it as retuning: an audit can say what looks cuttable and never whether cutting helped. An audit-only pass is a legitimate thing to want; it is a different request.

State the target model and the harness you found before continuing.

The phases

They run in order, and each names the reference it cannot start without. Read references/workflow-shapes.md before dispatching any phase: the wrong orchestration shape is the common failure. Fan out by disjoint file ownership, never by item. Items cross files, and agents that share a file lose each other's edits.

Before assessing whether the registered bar is met or interpreting its results, read references/noise-floor.md.

  1. Mine the archive before spending a run — references/baseline-mining.md. Historical runs are a free baseline, usually larger than any experiment affordable now.
  2. Establish the noise floor — references/noise-floor.md. Run the harness against two identical copies of the corpus, same commit on both sides; whatever difference appears is the floor every later claim must clear. Register the bar now, in writing, before any change exists. A bar chosen after seeing results is not a bar.
  3. Audit the corpus adversarially — references/corpus-audit.md. One agent per skill proposes cuts; a second per skill does the opposite and defends the existing prose. The two passes require independent contexts. If the host exposes no way to run them as separate agents, report that as a blocker and stop the audit — do not argue both sides in one context and present the result as an audit.
  4. Cut in surgical passes, one problem per agent — references/cut-passes.md, and references/halt-taxonomy.md when the symptom is stalling, halting, or a run that ends while naming work it did not do. Two rules bound every pass, whatever class it is cutting. Never edit a test to make a suite green: a removed string a test pins is a finding to report, not a test to weaken. And not every stop is the enemy. Some workflows exist to stop and ask; that is the product. Sort every stop by who is actually on the other side before touching it. references/halt-taxonomy.md carries the screens that decide, so read them before cutting any stop.
  5. Measure, then let the failure choose the next fix — references/cut-passes.md again for what each failure site means and for auditing the phases the instrument never enters. Loop 4 and 5 until the registered bar clears. Then stop; a bar cleared is done. Also report what stayed unmeasured: a cleared bar never implies coverage it does not have.
  6. Ship — references/cut-passes.md carries what the commits and the write-up must preserve.

© EveryInc, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in skills/ce-retune of EveryInc/compound-engineering-plugin.

  • SKILL.md
  • agents/openai.yaml
  • references/baseline-mining.md
  • references/corpus-audit.md
  • references/cut-passes.md
  • references/halt-taxonomy.md
  • references/noise-floor.md
  • references/workflow-shapes.md

Open the folder on GitHubat commit 67035e9

Compare with similar skills

Ce Retune next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ce Retune compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ce Retune this skillEveryInc/compound-engineering-plugin25k—~1.2kAutomated safety check: PassMIT
Tests Query Corpusnetdata/netdata81k—~5.3kAutomated safety check: PassGPL-3.0
Measure Before You Fixgarrytan/gbrain31k—~2.5kAutomated safety check: PassMIT
Measure Startup RequestsTriliumNext/Trilium38k—~1.5kAutomated safety check: PassAGPL-3.0
Corpus Remeasurenubjs/nub4.4k—~2.1kAutomated safety check: PassMIT
Lebesgue Measureparcadei/Continuous-Claude-v33.9k1 repos~827Automated safety check: NotesMIT

Similar skills

  • Tests Query Corpus

    netdata/netdata

    Run, extend or review the Netdata query contract corpus (tests/query-corpus), its fixtures, independent oracles, byte pins and harness.

    81k GitHub stars~5.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Measure Before You Fix

    garrytan/gbrain

    Before fixing a slow/stale/timeout alert, measure the step yourself.

    31k GitHub stars~2.5k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Measure Startup Requests

    TriliumNext/Trilium

    A skill your agent uses when measuring what the Trilium client loads at startup — "what loads at boot?", "did this change reduce the startup bundle?", "is <dependency lazy?", or any before/after…

    38k GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check passed
  • The iteration cycle for the build-jail corpus — change the harness or nub, push it, re-measure ONE package@version on ONE platform, and read the answer.

    4.4k GitHub stars~2.1k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Lebesgue Measure

    parcadei/Continuous-Claude-v3

    Problem-solving strategies for lebesgue measure in measure theory

    3.9k GitHub starsUsed in 1 repo~827 tokens
    Research & ScienceAuto-check: notes
  • Building Role Mining For Rbac Optimization

    mukul975/Anthropic-Cybersecurity-Skills

    Apply bottom-up and top-down role mining techniques, including clustering algorithms and formal concept analysis, to discover optimal RBAC roles from existing user-permission assignments…

    34k GitHub stars~2.6k tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed

More from EveryInc/compound-engineering-plugin

All 37 skills in this repo
  • Compound Learning Writer

    EveryInc/compound-engineering-plugin

    Records one solved and verified problem as a durable learning in the repository, but only when the reasoning is not already clear from the final code, tests or docs.

    25k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Compound Learnings Refresh

    EveryInc/compound-engineering-plugin

    Audits a repo's stored learnings against the current codebase, fixes stale, overlapping or superseded docs and reports on every document.

    25k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Compound Engineering Prototype

    EveryInc/compound-engineering-plugin

    Builds a throwaway prototype at just the fidelity needed to settle a specific how-it-should-work-or-feel question, before committing to an approach other work will treat as fixed.

    25k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Compound Engineering Setup

    EveryInc/compound-engineering-plugin

    Checks Compound Engineering plugin health and repo-local config, or scaffolds a Compound Pack when you ask for one by id.

    25k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • PR Babysitter

    EveryInc/compound-engineering-plugin

    Watches an open GitHub pull request over time, routing review comments and CI failures to other skills until the PR is ready to merge.

    25k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • CE Brainstorm

    EveryInc/compound-engineering-plugin

    Turns a vague or ambitious feature idea into a requirements-only plan through dialogue with you, sized to the work, before any code is written.

    25k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Ce Retune

What does Ce Retune do?

Retune a skill corpus for a new model, measurement-first: mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a…. Ce Retune is an agent skill from EveryInc/compound-engineering-plugin. Retune a skill corpus for a new model, measurement-first: mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a pre-registered bar clears.

How do I install Ce Retune in Claude Code?

Run `npx skills add EveryInc/compound-engineering-plugin --skill ce-retune -a claude-code`. Or copy the skill folder (skills/ce-retune in EveryInc/compound-engineering-plugin) into .claude/skills/ce-retune in your project. Claude Code loads it when a task matches its description.

How do I install Ce Retune in Codex?

Run `npx skills add EveryInc/compound-engineering-plugin --skill ce-retune -a codex`. Or copy the skill folder (skills/ce-retune in EveryInc/compound-engineering-plugin) into .agents/skills/ce-retune in your project. Codex loads it when a task matches its description.

Can I use Ce Retune in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EveryInc/compound-engineering-plugin --skill ce-retune -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ce-retune, .gemini/skills/ce-retune, .github/skills/ce-retune and .opencode/skills/ce-retune in your project.

What does Ce Retune need to run?

SKILL.md names no scripts, command-line tools or credentials: Ce Retune is instructions for the agent only.

Does Ce Retune access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ce Retune safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ce Retune use?

Ce Retune is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ce Retune use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.

What are the alternatives to Ce Retune?

Skills that share tags, products or a category with Ce Retune: Tests Query Corpus (netdata/netdata, 81k stars), Measure Before You Fix (garrytan/gbrain, 31k stars), Measure Startup Requests (TriliumNext/Trilium, 38k stars) and Corpus Remeasure (nubjs/nub, 4.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ce Retune?

EveryInc (a GitHub organization) maintains it in EveryInc/compound-engineering-plugin, which has 25,424 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on October 8, 2026.

Source: EveryInc/compound-engineering-plugin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.