MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…
$ npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Intelligent-Internet/zenith optimization-mission-playbook --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook .claude/skills/optimization-mission-playbook && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "optimization-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook into .claude/skills/optimization-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimization-mission-playbook", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbookType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Intelligent-Internet/zenith optimization-mission-playbook --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .agents/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook .agents/skills/optimization-mission-playbook && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "optimization-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook into .agents/skills/optimization-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimization-mission-playbook", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Intelligent-Internet/zenith optimization-mission-playbook --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook .cursor/skills/optimization-mission-playbook && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "optimization-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook into .cursor/skills/optimization-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimization-mission-playbook", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Intelligent-Internet/zenith.git --path zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Intelligent-Internet/zenith optimization-mission-playbook --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook .gemini/skills/optimization-mission-playbook && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "optimization-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook into .gemini/skills/optimization-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimization-mission-playbook", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Intelligent-Internet/zenith optimization-mission-playbookInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .github/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook .github/skills/optimization-mission-playbook && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "optimization-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook into .github/skills/optimization-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimization-mission-playbook", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Intelligent-Internet/zenith optimization-mission-playbook --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook .opencode/skills/optimization-mission-playbook && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "optimization-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook into .opencode/skills/optimization-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimization-mission-playbook", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
optimization-mission-playbookDomain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…
Optimization Mission Playbook is an agent skill from Intelligent-Internet/zenith. Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and similar metric-improvement work. Defines how to think about and run an optimization mission: establishing the ground truth, trusting the measurement, profiling to the dominant cost, estimating the ceiling, generating and pruning disposable hypotheses, exploring cheaply before verifying expensively, guarding against metric…
Its SKILL.md is about 11k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows. It works with Model Context Protocol. The repository describes itself as: Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a8d9b57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Optimization Mission Playbook loads about 11k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 5,943 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Intelligent-Internet/zenith at commit a8d9b57, republished under its Apache-2.0 licence (© Intelligent-Internet). 5,943 words, ~11,198 tokens.
.claude/skills/optimization-mission-playbook/SKILL.md (or your agent's skills folder).Load this playbook when the mission's goal is to move a metric: make something faster, smaller, cheaper, higher-scoring, more accurate, higher-throughput, better-compressed, better-ranked, or closer to a solver/objective bound. The signal is a number with a direction and a workload that produces it.
Do not use it for missions whose goal is durable behavior (a feature, a port, a migration, an API) — those are engineering missions; use engineering-mission-playbook. When a mission has both — build a benchmark harness then optimize a hot path, port a library then tune it — load both playbooks and record the boundary. Durable behavior and scaffolding may use engineering assertions; metric search should normally run as targetless optimization work, with formal assertions reserved for genuine durable commitments or externally required sign-off.
This playbook has two halves. The strategy (Posture through Anti-patterns) is how to think about an optimization mission — read it first; it is the reasoning every decision below depends on. The method (Operating Order onward) is how to express that thinking as orchestrator investigation, targetless experiment work in the current checkout, ledger review, and patches in this runtime. The method references the runtime lifecycle and task schema but does not redefine them.
An optimization mission moves a metric. You are not building toward a known end state — you are running an experiment whose answer you do not yet know, against a measurement that can mislead you. Think like a skeptical experimentalist, not a builder: most of the work is deciding what to measure, where, and whether the number is real — not deciding how to change the code.
Hold two ideas at once:
A real win moves the metric and survives that skepticism: correctness preserved, guardrails intact, comparison fair, improvement reproducible, and the gain generalizes beyond the cases it was tuned on. A number that goes up by breaking any of those is not a win.
Honest "no win" is a valid outcome when the search was credible, bounded, and recorded. Do not manufacture a win, and do not let effort spent become evidence of progress.
Most optimization confusion comes from collapsing three separate concepts. Name them apart and keep them apart for the whole mission:
When someone says "compare against the baseline," ask which of these three they mean. A measurement that uses the working point as its scoring reference, or the candidate's own output as its correctness oracle, proves nothing.
Before the first experiment, map the terrain deeply. You cannot search a space you have not measured, and in optimization the two easiest ways to fool yourself are weak correctness and weak measurement. Deep investigation must produce two reusable protocols with evidence, not assumption:
Also establish:
If the correctness protocol or metric protocol is unknown, weak, ambiguous, or unverifiable, resolving it is the first work — not a step you skip to start changing code.
Where cost hides, what the ceiling is, and what noise to distrust differ by metric family. Read your metric before you profile it:
| Family | Where the cost/loss usually hides | Typical ceiling / floor | Main noise or illusion |
|---|---|---|---|
| Latency | I/O waits, serialization, lock contention, cold caches, allocation, the critical path | Irreducible critical-path work / round-trip time | Warmup, co-tenancy, frequency scaling; use tail percentiles, not the mean |
| Throughput | Synchronization, per-item overhead, batching limits, the saturated resource | Saturation of the bottleneck resource (CPU / mem-bandwidth / I/O) | Too-short measurement window, queue warmup; often trades against latency |
| Memory | Retained allocations, duplication, over-allocation, fragmentation, caches | Irreducible live working set | Peak vs steady-state vs average; GC timing — measure the high-water mark you care about |
| Cost | Redundant calls, oversized resources, retries, idle time, data movement | Minimum required work × unit price | Pricing tiers and amortization; cheaper-per-unit but slower can raise total cost |
| Compression / size | Redundancy not exploited, format overhead | Entropy of the data | Ratio on one corpus ≠ general; watch the decode-cost and correctness trade |
| Score / quality (eval) | A subset of cases or classes; the aggregate hides per-case regressions | Irreducible error / oracle (or human) agreement ceiling | Small-N sampling noise is large; contamination; overfit to the eval set — never tune on the held-out split |
| Ranking | Head vs tail positions; specific query classes | The ideal ordering (e.g. perfect NDCG = 1) | Per-query variance, small judged pools; the aggregate hides query-class regressions |
| Solver / objective | Problem structure and instance mix; time-to-solution vs solution quality | A known bound (optimal, LP relaxation, dual bound) | Instance-specific results; seed/heuristic randomness; a win on one instance ≠ general |
Run the mission as a loop, not a straight line; new evidence sends you back up. Each pass:
The sections below are how to think inside each step.
The first question of every optimization mission is not "what do I change?" — it is "do I believe the number?" Resolve that before touching the system.
Pin the metric down. Name, direction, unit, the exact command that produces it, the workload, the aggregation, warmup handling, seeds, and what it means to take the same measurement twice. Choose the aggregation to match the metric: tail percentiles (p95/p99) when the tail is what hurts, median when single runs are noisy and skewed, mean/throughput for steady batch work. Never report a bare point estimate — always carry the spread with it.
Characterize the noise before optimizing. Run the unchanged baseline several times under realistic conditions and look at the spread. From that spread, fix a minimum detectable effect (MDE): the smallest delta you are willing to claim as real — a safe rule of thumb is a few times the run-to-run spread. Treat any change inside the noise band as zero, no matter how much you want it to be a win. If the noise is larger than the gains you expect to find, you must reduce it before optimizing, or you will spend the whole mission chasing fluctuations.
Reduce noise at the source. Stable environment, warmup runs discarded, enough repetitions, one variable changed at a time, nothing else competing for the machine, fixed seeds where the work is stochastic. Interleave baseline and candidate measurements in the same session so environmental drift cancels instead of biasing one side. Early on, a cheap and stable ruler you can read many times beats a faithful one you can only afford once — you need iteration speed more than the last decimal.
Check whether the ruler can be fooled. Ask: could this number improve without the thing it represents actually improving? Could the score move by changing how it is measured rather than what is measured? The classic gaming routes — hardcoding or special-casing known inputs, weakening or short-circuiting the check, measuring a cheaper but different workload, caching across the measurement boundary, or editing the scorer/reference/fixtures — all move the number while moving nothing real. Anything that scores the work is part of the ruler and is read-only. If a candidate's gain traces to the ruler rather than the work, it is not a gain; it is a leak.
The workload and the correctness cases are not given to you; they are designed, and the strength of that design caps how much any result can mean. A weak case set produces confident-looking verdicts worth nothing — a "2× faster" measured on one small convenient input, or a "still correct" checked on three happy-path examples.
Name the bias and resist it: a minimal benchmark and a handful of happy-path checks are less work to build and more likely to pass, so there is a constant pull toward them. You will be tempted to take the easy case set precisely because designing the comprehensive one is harder. "It passed the cases I wrote" is near-zero evidence when you wrote easy cases. Design the test to surface failure, not to confirm success — a case set that cannot distinguish a real improvement from a benchmark-only trick or a hidden correctness regression is not a test.
Build the case space by enumerating axes deliberately, then covering them:
For correctness, cover behaviors and failure modes against the oracle, not just the happy path, and add adversarial cases aimed at exactly where this candidate is most likely to break: the inputs it special-cases, the regime where its approximation degrades, the boundary its fast path skips. The cases an optimization is most likely to break are precisely the ones a lazy test omits — target them first.
For speed, measure a representative mix, not one easy input, and include the large and pathological cases where an optimization can collapse: a cache that helps the common case but thrashes on the adversarial one, a fast path that silently falls back to something slower, a change that wins at small N and loses at large N. Keep per-case timing — an aggregate over an easy mix hides every one of these.
A narrow case set is also what makes gaming profitable: hardcoding and special-casing only pay off when the cases are few and known. Comprehensive, adversarial coverage is the correctness defense and the anti-gaming defense at the same time.
You can only win where the cost actually is. Before forming any hypothesis, profile the current system and find the dominant term — the part of the work that consumes most of the metric.
Amdahl is the law here. If a part accounts for fraction p of the metric, then optimizing only that part — even to zero — cannot improve the whole by more than 1 / (1 - p). So the first job is to find p. Spending effort on a 5% term caps your win at ~5% no matter how clever the change.
Hot is not the same as leverage. A path can be hot (lots of time/cost) yet already near-optimal, leaving little to gain. Leverage is cost × how much of it you can plausibly remove. Rank targets by leverage, not by raw heat.
Profile with the right instrument for the metric. A sampling/time profiler for latency; an allocation profiler and heap snapshots for memory; counters and instrumentation at the bottleneck for throughput; per-call/per-token/per-request attribution for cost; a per-case breakdown for score and ranking. Profile on a workload that matches what is actually scored or used — a profile of a toy input points you at the wrong target.
Hypotheses come from the profile, not from brainstorming in the abstract. If you have an "optimization idea" before you have a profile, you are guessing.
The bottleneck moves. Every time you remove one, re-measure — the next dominant term is rarely where you expected, and the gain you just made changed the breakdown. Profiling is not a one-time step at the start; it is the engine that drives every round. Beware optimizing yesterday's bottleneck.
Before you start climbing, estimate the top. You rarely need a precise bound; an order-of-magnitude estimate changes your decisions.
Estimate two bounds. A floor on the irreducible work — the cost of just reading the input, a physical or information-theoretic limit, a known complexity bound, a dual/relaxation bound, or the number a known-better system already achieves. And a round ceiling from Amdahl — "if I drove the current dominant term to zero, the metric would become X." The gap between where you are and these bounds is the real size of the opportunity.
Use the ceiling to make two cheap decisions:
If the target is already close to its floor, redirect to the next dominant term rather than squeezing a near-exhausted one.
Generate from the profile and from how the system works. Common families, roughly in order of how often they pay off: do less (skip, early-exit, prune, deduplicate), do it once (cache, memoize, hoist, precompute), do it cheaper (better algorithm/complexity, better data layout and locality), do it together (batch, vectorize, coalesce I/O), do it approximately (lower precision or an approximation where the guardrail allows), and add concurrency only after the serial work is lean — spreading waste across workers just spreads the waste.
Rank by expected value: leverage on the dominant term × probability it works × 1 / cost to try. Pursue the cheap, high-leverage ones first, and prefer experiments designed to disconfirm fast — if an idea is wrong, you want to know cheaply and early.
Pre-register the kill criterion. Before running, state what "this didn't work" looks like (e.g. "if it isn't at least X better on the proxy, drop it"). This stops you from moving the goalposts after you have grown attached to a candidate.
Hold every hypothesis loosely and prefer breadth before depth. Try several distinct families before committing budget to deepening one. The most expensive mistake in optimization is staying with a direction because you already invested in it — sunk cost is the enemy of search. Measure, and let the data redirect you.
You have a finite budget. Spend it like a bandit: wide, cheap exploration to find what has promise, then concentrated, expensive verification on the few survivors.
Climb a fidelity ladder. Micro-benchmark or proxy workload → representative subset → full faithful measurement with correctness checked. Each rung costs more and decides more. Match the rung to the stakes: a quick proxy is enough to prune a hypothesis or rank a batch roughly; it is never enough to promote one. Reserve the full, faithful, costly measurement — real workload, enough repetitions, correctness in full, per-case evidence — for candidates that could actually be selected.
Know what a proxy may and may not conclude. A cheap check can rule a hypothesis out (no signal → kill) and order candidates approximately. It cannot confirm a promotable win. And the proxy itself can lie: periodically confirm that it still correlates with the real metric on at least one candidate. A proxy that has drifted from the real metric is just a second ruler that needs trusting.
Avoid both failure modes: full-verifying every idea burns the budget on dead ends; trusting a cheap proxy as if it were the real metric promotes a mirage. Spend more measurement where the decision is close (tie-breaks, promotion) and less where the outcome is already obvious.
The metric is a proxy for what you actually want. The moment you optimize hard against it, the gap between "scores better" and "is better" is exactly where self-deception lives (when a measure becomes a target, it stops being a good measure). Put every apparent win through four questions, and treat failure on any one as disqualifying — a win that fails a question is not a smaller win, it is not a win:
Optimization is almost always multi-objective. Faster often costs memory; smaller often costs accuracy; a higher score often costs latency, complexity, or compatibility. You are moving along a frontier, not climbing one axis.
Carry the other axes explicitly. Make correctness and every guardrail — accuracy floor, memory ceiling, latency budget, cost limit, compatibility promise — a budget you check on every candidate, not at the end. A candidate that is worse on no axis and better on one dominates and is a clean win; anything else is a trade that requires a decision.
Watch the hidden costs the headline metric does not show: maintainability, complexity, numerical stability, portability, build/cold-start time, and operational risk. State the trade in plain terms ("+8% throughput for +15% memory and a more fragile code path") and decide whether it is acceptable, rather than letting one number hide what it cost to move it. Record the trades you accept.
Hold one reference to measure against and keep it stable. Every candidate must be measured apples-to-apples against it — same workload, environment, build, cache state, and conditions — and must be reconstructible and reversible so you can re-measure, compare, or roll back cleanly. A measurement taken under drifted conditions is not evidence; it is noise wearing a number.
When a win passes all four tests of §6, lock it in as the new working point and optimize from there. Keep the working point separate from the scoring reference (see "Three Things You Must Never Conflate"): one is where the next experiment starts, the other is the fixed yardstick the result is judged against. Promoting a win advances the former and must never quietly move the latter.
Stopping is a decision made from evidence, not from exhaustion. Stop when:
A credible, bounded search that ends in "no further win" is a successful mission with an honest result, and should be recorded as such — what was tried, what was ruled out, and why. An unbounded search that never decides is not a success; it is a failure to stop.
These kill more optimization missions than any missing trick:
The strategy above is how to think. This half is how to express it in the runtime without turning ordinary metric search into a contract pipeline.
Default optimization shape:
orchestrator investigates correctness protocol + metric protocol
-> targetless experiment work in the current checkout
-> attention with ledgers + current workspace state
-> orchestrator decides accept/reject/keep/rerun
-> next round or doneA disposable hypothesis is a targetless work task, not an assertion. It owns
no contract and needs no gate. It runs against the current checkout. Later
experiments may overwrite earlier experiment changes, so every experiment must
leave a ledger detailed enough for the orchestrator to decide whether to keep
the current result, reject it, or run another experiment.
Formal assertions, validators, gates, and promotion chains are not the default optimization method. Use them only when the mission also has a durable engineering deliverable, an externally required sign-off, or a high-risk trust boundary that genuinely needs the formal runtime shape.
When the strategy above says "promote", read it as the orchestrator accepting the current candidate result and recording the new working point, not as a default promotion task.
On a PLAN wake:
submit_plan.
This means both the correctness protocol and the metric protocol. This is
orchestrator planning work, not a default DAG node.work tasks from the current working point. Each
task must cite the optimization method, mutate the current checkout directly,
write a ledger, and request attention.benchmark-infra task only when shared measurement helper files must
be created before experiments can run. This is an exception, not the default.submit_plan.On a DECIDE wake (attention):
accept,
reject, keep_as_evidence, rerun, or inspect.Do not create a validator/gate/promotion chain just to continue optimization.
Before submit_plan, produce the ground-truth optimization method with evidence,
not assumption. This investigation must be deep enough to make later experiment
evidence meaningful: correctness must have a real oracle and failure-oriented
case set, and the metric must have a comparable, reproducible measurement
protocol. Use bounded read-only investigator lanes when the orchestrator needs
help without taking on too much context.
Minimum optimization method fields:
optimization_method_ref:
correctness_protocol_ref:
correctness_oracle:
correctness_case_set:
correctness_tolerances:
correctness_invariants:
correctness_missing_output_policy:
correctness_per_case_failure_policy:
candidate_specific_correctness_evidence:
metric_protocol_ref:
metric:
direction:
unit:
metric_command:
metric_workloads:
metric_repetitions:
metric_aggregation:
metric_warmup_or_cold_state:
metric_seeds:
metric_timeout_and_resource_limits:
metric_noise_band_or_mde:
metric_per_case_reporting:
scoring_reference:
working_point:
candidate_binding:
public_benchmark_inventory:
sample_limitations:
baseline_protocol:
hotpath_profile:
ceiling:
guardrails:
protected_paths:
stop_rule:Required content:
Public benchmarks, sample tests, and repository-provided demo workloads are a starting point, not the whole score unless the mission explicitly says so.
Profiling, baseline collection, oracle-finding, correctness case-set design, and
metric protocol calibration belong in PLAN investigation or, when measurement
must run in the workspace, in a clearly-labeled targetless work task. They are
correctness/measurement setup, not an optimization candidate.
Default experiment tasks are targetless, workspace-mutating, and attention-producing:
tasks:
- id: exp-cache
type: work
targets: []
skill: latency-experiment-worker # project-authored
body: "Try the cache hypothesis from the current working point in the current checkout. Use optimization method opt-method-001. Write a ledger with changed files, correctness evidence, metric evidence, and any deviations, then call end_node(done=True, report=..., request_attention=True)."
depends_on: []
- id: exp-batch
type: work
targets: []
skill: batching-experiment-worker # project-authored
body: "Try the batching hypothesis from the current working point in the current checkout. Use optimization method opt-method-001. Write a ledger with changed files, correctness evidence, metric evidence, and any deviations, then call end_node(done=True, report=..., request_attention=True)."
depends_on: [exp-cache]Sequence experiments with depends_on when they touch the same workspace state
or when later experiments should build on earlier changes. For broad exploration,
prefer one experiment per attention cycle unless the experiments are truly
independent and either completion order is acceptable.
Only add a benchmark infrastructure task when the orchestrator cannot define a usable metric protocol, correctness protocol, or shared measurement helper without creating files that later experiments must run.
tasks:
- id: benchmark-infra
type: work
targets: []
skill: latency-benchmark-infra # project-authored
body: "Create reusable benchmark/correctness helper files for optimization method opt-method-001. Do not optimize solution code. Report created files, commands, protected paths, and any remaining gaps."
depends_on: []
- id: exp-cache
type: work
targets: []
skill: latency-experiment-worker # project-authored
body: "Try the cache hypothesis using optimization method opt-method-001 and the benchmark-infra files. Mutate the current checkout, write a ledger, and request attention."
depends_on: [benchmark-infra]This task is not a validator, gate, or experiment candidate. If it changes anything other than measurement infrastructure, split the work or reject it.
Each experiment worker must leave a ledger in its report. The ledger is more important than optimistic prose.
Minimum report fields:
candidate_id:
task_id:
parent_ref:
candidate_ref:
workspace_before_ref:
workspace_after_ref:
changed_files:
optimization_method_ref:
correctness_protocol_ref:
metric_protocol_ref:
metric:
baseline:
candidate:
delta:
noise_band:
commands:
correctness_per_case_result:
metric_per_case_result:
guardrails:
protected_paths_touched:
recommendation: accept | reject | keep_as_evidence | rerun | inconclusive
notes:Rules for workers:
end_node(done=True, report=<ledger>, request_attention=True) when the
ledger is usable.Use done=False only when the worker could not produce a usable ledger.
Use one checker only when the orchestrator is about to accept a result and the evidence is
not yet strong enough. The checker is still a targetless work task by default,
not a validate task and not a gate.
add:
- id: check-exp-cache
type: work
targets: []
skill: latency-candidate-checker # project-authored
body: "Inspect the current checkout or cited candidate ref, rerun optimization method opt-method-001, check protected paths, correctness, and metric comparability, and report accept/reject/rerun."
depends_on: [exp-cache]Checker report minimum:
candidate_id:
candidate_ref:
optimization_method_ref:
correctness_protocol_ref:
metric_protocol_ref:
metric_rerun:
correctness_rerun:
guardrails:
protected_paths_touched:
recommendation: accept | reject | rerun | inconclusive
blocking_issues:Use a formal validate task only when the mission also has contract assertions
that genuinely need validation.
Acceptance is an orchestrator decision. It is not a default runtime task chain.
Before accepting the current result:
After accepting:
Do not author VAL-* or EXP-* assertions by default.
Use formal assertions only when:
validate/gate
semantics.Runtime obligations to respect when declaring an assertion:
work task that
owns it for the life of the mission;add_items requires the matching contract/<ID>.md file on disk before the
patch;If formal assertions are introduced, use contract-review for those assertions
only. Do not demand an assertion per hypothesis.
Patch another round when:
Patch with more targetless work tasks. Keep the round small enough that the
orchestrator can compare candidates in one attention wake.
Prefer distinct hypotheses over many variants of one idea until evidence says a family is worth deepening.
end_node(done=True, report=..., request_attention=True).benchmark-validator and scrutiny-validator are
for formal contract exceptions, not routine optimization search.investigator for
protocol/profile/oracle/binding/ceiling/case-set and root-cause analysis lanes;
contract-review only when formal assertions are added.At an experiment checkpoint: read the ledger and candidate ref; apply correctness/integrity-first triage; update the scoreboard from ledgers, not prose; continue planned breadth unless evidence or budget justifies selecting; patch a new round only for distinct move families.
At an acceptance checkpoint: inspect the current checkout or cited candidate ref, rerun or check when needed, accept only when the evidence clears the mission's risk bar, then record the new working point.
Patch the root cause, mapping the strategy's anti-patterns to a fix:
invalid, patch
protected-file rules or measurement setup.Stop on evidence (strategy §9): ceiling reached, diminishing returns over clean rounds, budget spent, value below cost, or the environment cannot measure credibly.
Store only durable, reusable facts in the project memory carrier (AGENTS.md,
MEMORY.md, or project-root notes the tasks cite):
Keep one-off runs, raw logs, and per-experiment ledgers in attempts/decisions, not in the durable carrier. When a discovered fact changes how later work should be planned or measured, update memory before dispatching more work.
work tasks.end_node(done=True, report=..., request_attention=True).EXP-* contract is planned.work task.© Intelligent-Internet, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook of Intelligent-Internet/zenith.
Open the folder on GitHubat commit a8d9b57
Optimization Mission Playbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Optimization Mission Playbook this skillIntelligent-Internet/zenith | 338 | — | ~11k | Automated safety check: Pass | Apache-2.0 | |
| MCP Server Builderanthropics/skills | 180k | 63 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| MCP Server BuildershareAI-lab/learn-claude-code | 78k | 4 repos | ~1.2k | Automated safety check: Pass | MIT | |
| MCP Integration for Pluginsanthropics/claude-plugins-official | 38k | 11 repos | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| MemPalace Memory SearchMemPalace/mempalace | 59k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Crush Configurationcharmbracelet/crush | 29k | — | ~3.7k | Automated safety check: Pass | Custom licence |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
shareAI-lab/learn-claude-code
Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.
anthropics/claude-plugins-official
Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.
MemPalace/mempalace
Mines project files and conversation exports into a local, searchable memory palace and recalls past work by semantic search through the mempalace CLI.
charmbracelet/crush
Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.
mksglu/context-mode
Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.
Intelligent-Internet/zenith
A skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…
Intelligent-Internet/zenith
Benchmark validation procedure for one assigned benchmark-related target.
Intelligent-Internet/zenith
Adversarial scrutiny procedure for engineering validation assignments.
Intelligent-Internet/zenith
Real-surface validation coordinator for engineering validation assignments.
Intelligent-Internet/zenith
Automates browser and Electron app interactions for user-flow validation.
Works with
Categories
Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…. Optimization Mission Playbook is an agent skill from Intelligent-Internet/zenith. Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and similar metric-improvement work.
Optimization Mission Playbook fits situations like: agent Workflows work in your project.
Run `npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a claude-code`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook in Intelligent-Internet/zenith) into .claude/skills/optimization-mission-playbook in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a codex`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/optimization-mission-playbook in Intelligent-Internet/zenith) into .agents/skills/optimization-mission-playbook in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Intelligent-Internet/zenith --skill optimization-mission-playbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimization-mission-playbook, .gemini/skills/optimization-mission-playbook, .github/skills/optimization-mission-playbook and .opencode/skills/optimization-mission-playbook in your project.
SKILL.md names no scripts, command-line tools or credentials: Optimization Mission Playbook is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Optimization Mission Playbook is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 11k tokens (SKILL.md is roughly 45k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Optimization Mission Playbook: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars) and MemPalace Memory Search (MemPalace/mempalace, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Intelligent-Internet (a GitHub organization) maintains it in Intelligent-Internet/zenith, which has 338 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 6, 2026.
Source: Intelligent-Internet/zenith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.