Agent skill

Spec Optimize

by leo-kuang-ai in leo-kuang-ai/spec-first

Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.

MITAuto-check passedAI & LLM Engineering

Install Spec Optimize

skills CLI
$ npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install leo-kuang-ai/spec-first spec-optimize --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spec-optimize .claude/skills/spec-optimize && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spec-optimize
GitHub stars
107
Token cost
~13k tokens
SKILL.md length
6,127 words
Files
32 (incl. scripts, references)
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.

  • Works in 5 steps: Setup → Measurement Scaffolding → Hypothesis Generation → …
  • Optimizing clustering quality
  • SKILL.md covers Workflow Contract Summary, Scenario Capability, Interaction Method and Input, plus 10 more sections
  • Runs Shell and JavaScript scripts from its folder; calls bash, git and codex

What it does

Spec Optimize is an agent skill from leo-kuang-ai/spec-first. Run metric-driven iterative optimization loops. Define a measurable goal, build measurement scaffolding, then run parallel experiments that try many approaches, measure each against hard gates and/or LLM-as-judge quality scores, keep improvements, and converge toward the best solution. Use when optimizing clustering quality, search relevance, build performance, prompt quality, or any measurable outcome that benefits from systematic experimentation. Inspired by Karpathy's autoresearch, generalized for multi-file…

Its SKILL.md is about 13k tokens, which your agent loads only when the skill is triggered. The skill folder holds 38 other files, including scripts and reference files (for example `README.md`, `evals/README.md` and `evals/cases/debug-request-routes-out.yaml`).

It sits in AI & LLM Engineering, covering Project scaffolding, LLM evaluation and Search implementation. The repository describes itself as: 仓库原生 AI Coding Harness —— 把一次性 AI 对话变成可治理、可验证、可沉淀的工程闭环 · spec-first.cn. The licence is MIT.

When your agent uses it

  • Optimizing clustering quality
  • Search relevance
  • Build performance
  • Any measurable outcome that benefits from systematic experimentation

Example prompts

  • “/spec-optimize”

Requirements

  • Node.js
  • A Bash shell

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Setup
  2. Measurement Scaffolding
  3. Hypothesis Generation
  4. Optimization Loop
  5. Wrap-Up

What it can do on your machine

Read from SKILL.md and the folder at commit 74655dc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell and JavaScript, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • git
    • codex
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Spec Optimize loads about 13k tokens when it runs, and up to ~34k if it reads all its reference files. Until then it costs about 141 tokens; SKILL.md has 6,127 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~141
When it runs · the whole SKILL.md, loaded when a task matches
~13k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~34k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from leo-kuang-ai/spec-first at commit 74655dc, republished under its MIT licence (© leo-kuang-ai). 6,127 words, ~12,834 tokens.

Download SKILL.mdSave it as .claude/skills/spec-optimize/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.
name
spec-optimize
description
Run metric-driven iterative optimization loops. Define a measurable goal, build measurement scaffolding, then run parallel experiments that try many approaches, measure each against hard gates and/or LLM-as-judge quality scores, keep improvements, and converge toward the best solution. Use when optimizing clustering quality, search relevance, build performance, prompt quality, or any measurable outcome that benefits from systematic experimentation. Inspired by Karpathy's autoresearch, generalized for multi-file code changes and non-ML domains.
argument-hint
[path to optimization spec YAML, or describe the optimization goal]

Iterative Optimization Loop

Run metric-driven iterative optimization. Define a goal, build measurement scaffolding, then run parallel experiments that converge toward the best solution.

Workflow Contract Summary

When To Use

Use when a measurable outcome can improve through iterative experiments, hard gates, and/or LLM-as-judge scoring.

When Not To Use

Do not use for ordinary implementation, vague improvement requests without a metric, debugging without a feedback loop, or unbounded spend/concurrency. Name the destination when routing out: a bug with no stable repro and no measurement loop is spec-debug's work (or spec-work for settled implementation) — diagnosing the bug inside this workflow to turn it into an optimization goal is adopting the wrong workflow, not adapting it; route out explicitly instead of drafting a spec around the diagnosis.

Inputs

An optimization spec or goal, mutable/immutable scope, measurement command or scaffold plan, budget limits, experiment settings, repository instructions, and baseline evidence.

Outputs

A measurement scaffold and experiment log, scored experiment results, kept/rejected variants, final integrated changes when appropriate, and post-run recommendations.

Artifacts

Run state under .spec-first/workflows/spec-optimize/<spec-name>/, experiment worktrees/results, strategy digests, and no hidden workflow state outside the documented log.

Failure Modes

Missing metric, missing measurement command, unsafe scope, excessive or uncapped budget, failed baseline, write verification failure, or unavailable dispatch/worktree backend.

Workflow

Validate the spec and budget, establish the baseline, run bounded experiments, measure and write results immediately, select winners, integrate only verified improvements, and summarize evidence.

Downstream Consumers

Code review、benchmark maintainer、在性能/相关性变更时参与的 release reviewer,以及检查 experiment logs 的人工审查者。

Scenario Capability

Follows docs/contracts/workflows/scenario-capability-matrix.md (default). Overrides: none

Interaction Method

Use the platform's blocking question tool: AskUserQuestion in Claude Code (call ToolSearch with select:AskUserQuestion first if its schema isn't loaded) or request_user_input in Codex. Fall back to numbered options in chat only when no blocking tool exists in the harness or the call errors (e.g., Codex edit modes) — not because a schema load is required. Never silently skip the question.

Input

<optimization_input> #<invocation arguments supplied by the current host> </optimization_input>

If the input above is empty, ask: "What would you like to optimize? Describe the goal, or provide a path to an optimization spec YAML file."

Optimization Spec Schema

Reference the spec schema for validation:

references/optimize-spec-schema.yaml

Experiment Log Schema

Reference the experiment log schema for state management:

references/experiment-log-schema.yaml

Quick Start

For a first run, optimize for signal and safety, not maximum throughput:

  • Start from references/example-hard-spec.yaml when the metric is objective and cheap to measure
  • Use references/example-judge-spec.yaml only when actual quality requires semantic judgment
  • Prefer execution.mode: serial and execution.max_concurrent: 1
  • Cap the first run with stopping.max_iterations: 4 and stopping.max_hours: 1
  • Avoid new dependencies until the baseline and measurement harness are trusted
  • For judge mode, start with sample_size: 10, batch_size: 5, and max_total_cost_usd: 5

For a friendly overview of what this skill is for, when to use hard metrics vs LLM-as-judge, and example kickoff prompts, see:

references/usage-guide.md

Admission And Budget Gate

Do not run spec-optimize as an expensive substitute for ordinary work. Before Phase 1, confirm the run has all of these:

  • A repeatable measurement target: metric.primary.type, metric.primary.name, and metric.primary.direction
  • At least one cheap degenerate gate that rejects broken variants before judging or ranking
  • A measurement command, or a concrete plan to build the harness before any experiment is dispatched
  • Explicit scope.mutable and scope.immutable boundaries
  • Explicit experiment budget: stopping.max_iterations, stopping.max_hours, and stopping.plateau_iterations
  • Execution budget: execution.mode and execution.max_concurrent
  • For judge mode, a finite metric.judge.max_total_cost_usd, unless the user explicitly approves uncapped spend

If any item is missing, stop and help the user create a safe spec or route the work to the current host's plan/work/debug entrypoint instead. Do not continue with an open-ended optimization loop.

Measurement-Only Calibration Mode

When the invocation or approved spec selects mode:measurement-only, load references/measurement-only-calibration.md. This mode compares an explicitly identified baseline and candidate against the same repeatable task/corpus. It must run an A/A noise floor before A/B, use a pre-registered acceptance threshold, classify broken runs separately from regressions, and produce only a measurement artifact plus stop/defer guidance.

Before the first measurement, materialize the approved frozen inputs as a run-local measurement-admission-input.json, then run node scripts/measurement-admission.cjs admit --input <path>. Persist the returned normalized admission and admission_sha256 beside the measurement artifact and bind every A/A and A/B attempt to that digest. A rejected admission stops before invoking the measurement command. After A/A, run the helper's allow-ab command with the normalized admission, digest, attempts, and observed noise floor; A/B is forbidden unless it returns ab_allowed: true.

Measurement-only mode does not mutate either arm, any Skill package, the measurement harness, or promotion metadata. It does not select a winner for integration, invoke spec-write-skill, or authorize commit/landing. If arm identity, corpus identity, a repeatable harness, or the pre-registered threshold is missing, stop before measurement instead of inventing it after seeing data.

First-run specs should default to execution.mode: serial, execution.max_concurrent: 1, stopping.max_iterations: 4, stopping.max_hours: 1, stopping.plateau_iterations: 3, and max_runner_up_merges_per_batch: 0. Treat higher-throughput settings as opt-in. If a provided spec asks for execution.max_concurrent > 4, stopping.max_iterations > 30, stopping.max_hours > 4, or uncapped judge spend, surface those costs in the approval gate before running the baseline.

Runtime Context Exclusion

Follow docs/contracts/context-governance.md: ordinary Optimize context excludes .spec-first/audits/**, .spec-first/governance/**, and generated mirrors (.claude/**, .codex/**, .agents/skills/**, .cursor/skills/**, .cursor/spec-first/**, .cursor/mcp.json, .kiro/skills/**, .kiro/agents/**, .kiro/spec-first/**, .kiro/settings/**, .qoder/commands/spec-*.md, .qoder/commands/spec/**, .qoder/skills/**, .qoder/agents/**, .qoder/spec-first/**, .qoder/settings.local.json) by default. Optimization run state under .spec-first/workflows/spec-optimize/** is local scratch for this workflow only; pass compact strategy/result summaries to agents and avoid broad runtime/audit/governance scans unless the metric explicitly targets runtime/setup/audit/governance behavior. Cursor-native .cursor/rules/** / .cursor/agents/**, Kiro-native .kiro/specs/**, and Qoder-native .qoder/rules/** are advisory input only when explicitly named.

Evidence Utilization Boundary

Optimization may consume prior direct-read summaries, degraded reason counts, source-confirmed session evidence, review summaries, test results, and quality-gate reports as diagnostic context. Treat these as baseline diagnostics, not optimization targets by default. The optimization target, metric, winner selection, mutable scope, and final integration remain owned by the approved optimization spec and measured results, not by an external tool. Optimize must not run external-tool refresh, hooks, watchers, or daemons as part of ordinary evidence handling.

Dispatch And Backend Boundary

Optimization dispatch is optional. Before any learnings researcher, repo analyst, experiment worker, Codex delegation, or parallel worker run, record:

yaml
worker_dispatch_authorization: authorized | missing
capability_probe: not_applicable | attempted | unavailable
worker_dispatch_capability: available | missing | unknown
worker_context_isolation: isolated | inherited | unknown
worker_model_override: supported | unsupported | unknown
worker_bounded_parallelism: supported | unsupported | unknown

workflow invocation does not authorize dispatch。Approved optimization spec、baseline approval、execution.mode: parallel、预算、权限设置或 runtime readiness 都不是派发授权。只有当前用户或可见 upstream handoff 明确请求 subagent、delegated work、persona 或 parallel work 时才可派发。缺授权时不得探测 tool schema,固定为 capability_probe: not_applicable + worker_dispatch_capability: unknown,强制采用 serial inline/local execution 并记录 dispatch_authorization_missing。只有授权后才把 current-session registry/schema 作为 provider_untrusted evidence 检查:确认缺失时记录 subagent_capability_missing;surface 不可用、schema 不完整或候选不唯一时记录 worker_capability_unproven,均同样降级。隔离、模型覆盖和有界并发只取 live facts;required isolation 未满足时保持依赖 gate 打开,model unknown 时继承,parallelism unknown 时串行。记录 worker_dispatch_outcome。Fallback 可以继续使用串行 worktree,但不得声称 parallel experiment 或 independent worker coverage。

Parallel experiments additionally require both dispatch facts above, explicit execution.mode, bounded execution.max_concurrent, clean mutable/immutable scope, and the worktree readiness probes below. Worktree-backed mutation happens in experiment worktrees; Codex delegation must fall back after repeated failures when the serial/local path can continue. The orchestrator owns final integration: selecting kept experiments, merging or cherry-picking winners, reverting non-winners, cleaning worktrees, updating experiment logs, and presenting post-completion actions. Workers never stage, commit, merge, push, or mutate the authoritative experiment log.


Persistence Discipline

CRITICAL: The experiment log on disk is the single source of truth. The conversation context is NOT durable storage. Results that exist only in the conversation WILL be lost.

The files under .spec-first/workflows/spec-optimize/<spec-name>/ are local scratch state. They are ignored by git, so they survive local resumes on the same machine but are not preserved by commits, branches, or pushes unless the user exports them separately.

This skill runs for hours. Context windows compact, sessions crash, and agents restart. Every piece of state that matters MUST live on disk, not in the agent's memory.

If you produce a results table in the conversation without writing those results to disk first, you have a bug. The conversation is for the user's benefit. The experiment log file is for durability.

Core Rules
  1. Write each experiment result to disk IMMEDIATELY after measurement — not after the batch, not after evaluation, IMMEDIATELY. Append the experiment entry to the experiment log file the moment its metrics are known, before evaluating the next experiment. This is the #1 crash-safety rule.

  2. VERIFY every critical write — after writing the experiment log, read the file back and confirm the entry is present. This catches silent write failures. Do not proceed to the next experiment until verification passes.

  3. Re-read from disk at every phase boundary and before every decision — never trust in-memory state across phase transitions, batch boundaries, or after any operation that might have taken significant time. Re-read the experiment log and strategy digest from disk.

  4. The experiment log is append-only during Phase 3 — never rewrite the full file. Append new experiment entries. Update the best section in place only when a new best is found. This prevents data loss if a write is interrupted.

  5. Per-experiment result markers for crash recovery — each experiment writes a result.yaml marker in its worktree immediately after measurement. On resume, scan for these markers to recover experiments that were measured but not yet logged.

  6. Strategy digest is written after every batch, before generating new hypotheses — the agent reads the digest (not its memory) when deciding what to try next. The strategy digest is derived, reconstructable state; the experiment log remains the canonical resume and audit source. Persist every hypothesis decision, kept/reverted result, metric, and audit field needed to reconstruct the digest in the experiment log before replacing or deleting the digest.

  7. Never present results to the user without writing them to disk first — the pattern is: measure -> write to disk -> verify -> THEN show the user. Not the reverse.

Mandatory Disk Checkpoints

These are non-negotiable write-then-verify steps. At each checkpoint, the agent MUST write the specified file and then read it back to confirm the write succeeded.

CheckpointFile WrittenPhase
CP-0: Spec savedspec.yamlPhase 0, after user approval
CP-1: Baseline recordedexperiment-log.yaml (initial with baseline)Phase 1, after baseline measurement
CP-2: Hypothesis backlog savedexperiment-log.yaml (hypothesis_backlog section)Phase 2, after hypothesis generation
CP-3: Each experiment resultexperiment-log.yaml (append experiment entry)Phase 3.3, immediately after each measurement
CP-4: Batch summaryexperiment-log.yaml (outcomes + best) + strategy-digest.mdPhase 3.5, after batch evaluation
CP-5: Final summaryexperiment-log.yaml (final state)Phase 4, at wrap-up

Format of a verification step:

  1. Write the file using the native file-write tool
  2. Read the file back using the native file-read tool
  3. Confirm the expected content is present
  4. If verification fails, retry the write. If it fails twice, alert the user.
File Locations (all under .spec-first/workflows/spec-optimize/<spec-name>/)
FilePurposeWritten When
spec.yamlOptimization spec (immutable during run)Phase 0 (CP-0)
experiment-log.yamlFull history of all experimentsInitialized at CP-1, appended at CP-3, updated at CP-4
strategy-digest.mdCompressed learnings for hypothesis generationWritten at CP-4 after each batch
<worktree>/result.yamlPer-experiment crash-recovery markerImmediately after measurement, before CP-3
On Resume

When Phase 0.4 detects an existing run:

  1. Read the experiment log from disk — this is the ground truth
  2. Scan worktree directories for result.yaml markers not yet in the log
  3. Recover any measured-but-unlogged experiments
  4. Continue from where the log left off

Phase 0: Setup

0.1 Determine Input Type

Check whether the input is:

  • A spec file path (ends in .yaml or .yml): read and validate it
  • A description of the optimization goal: help the user create a spec interactively
  • Not an optimization input at all — a bug report, an unstable repro, or a "fix it" request with no measurable outcome: do not absorb it by fixing the bug or converting the diagnosis into a spec here. Name spec-debug (unstable/unresolved failures) or spec-work (settled implementation) in the reply, route out, and stop this workflow.
0.2 Load or Create Spec

If spec file provided:

  1. Read the YAML spec file. The orchestrating agent parses YAML natively -- no shell script parsing.
  2. Validate the spec against every rule in the validation_rules section of references/optimize-spec-schema.yaml. That section is the single source of truth for what a valid spec requires; do not rely on a remembered subset. Conditional rules such as exclusive-resource serial execution, singleton-rubric requirements, uncapped judge spend approval, high-throughput approval, and stopping criteria live there.
  3. If validation fails, report errors and ask the user to fix them

If description provided:

  1. Analyze the project to understand what can be measured

  2. Detect whether the optimization target is qualitative or quantitative — this determines type: hard vs type: judge and is the single most important spec decision:

    Use type: hard when:

    • The metric is a scalar number with a clear "better" direction
    • The metric is objectively measurable (build time, test pass rate, latency, memory usage)
    • No human judgment is needed to evaluate "is this result actually good?"
    • Examples: reduce build time, increase test coverage, reduce API latency, decrease bundle size

    Use type: judge when:

    • The quality of the output requires semantic understanding to evaluate
    • A human reviewer would need to look at the results to say "this is better"
    • Proxy metrics exist but can mislead (e.g., "more clusters" does not mean "better clusters")
    • The optimization could produce degenerate solutions that look good on paper
    • Examples: clustering quality, search relevance, summarization quality, code readability, UX copy, recommendation relevance

    IMPORTANT: If the target is qualitative, strongly recommend type: judge. Explain that hard metrics alone will optimize proxy numbers without checking actual quality. Show the user the three-tier approach:

    • Degenerate gates (hard, cheap, fast): catch obviously broken solutions — e.g., "all items in 1 cluster" or "0% coverage". Run first. If gates fail, skip the expensive judge step.
    • LLM-as-judge (the actual optimization target): sample outputs, score them against a rubric, aggregate. This is what the loop optimizes.
    • Diagnostics (logged, not gated): distribution stats, counts, timing — useful for understanding WHY a judge score changed.

    If the user insists on type: hard for a qualitative target, proceed but warn that the results may optimize a misleading proxy.

  3. Design the sampling strategy (for type: judge):

    Guide the user through defining stratified sampling. The key question is: "What parts of the output space do you need to check quality on?"

    Walk through these questions:

    • What does one "item" look like? (a cluster, a search result page, a summary, etc.)
    • What are the natural size/quality strata? (e.g., large clusters vs small clusters vs singletons)
    • Where are quality failures most likely? (e.g., very large clusters may be degenerate merges; singletons may be missed groupings)
    • What total sample size balances cost vs signal? (default: 30 items, adjust based on output volume)

    Example stratified sampling for clustering:

    yaml
    stratification:
      - bucket: "top_by_size"     # largest clusters — check for degenerate mega-clusters
        count: 10
      - bucket: "mid_range"       # middle of non-solo cluster size range — representative quality
        count: 10
      - bucket: "small_clusters"  # clusters with 2-3 items — check if connections are real
        count: 10
    singleton_sample: 15          # singletons — check for false negatives (items that should cluster)

    The sampling strategy is domain-specific. For search relevance, strata might be "top-3 results", "results 4-10", "tail results". For summarization, strata might be "short documents", "long documents", "multi-topic documents".

    Singleton evaluation is critical when the goal involves coverage — sampling singletons with the singleton rubric checks whether the system is missing obvious groupings.

  4. Design the rubric (for type: judge):

    Help the user define the scoring rubric. A good rubric:

    • Has a 1-5 scale (or similar) with concrete descriptions for each level
    • Includes supplementary fields that help diagnose issues (e.g., distinct_topics, outlier_count)
    • Is specific enough that two judges would give similar scores
    • Does NOT assume bigger/more is better — "3 items per cluster average" is not inherently good or bad

    Example for clustering:

    yaml
    rubric: |
      Rate this cluster 1-5:
      - 5: All items clearly about the same issue/feature
      - 4: Strong theme, minor outliers
      - 3: Related but covers 2-3 sub-topics that could reasonably be split
      - 2: Weak connection — items share superficial similarity only
      - 1: Unrelated items grouped together
      Also report: distinct_topics (integer), outlier_count (integer)
  5. Guide the user through the remaining spec fields:

    • What degenerate cases should be rejected? (gates — e.g., "solo_pct <= 0.95" catches all-singletons, "max_cluster_size <= 500" catches mega-clusters)
    • What command runs the measurement?
    • What files can be modified? What is immutable?
    • Any constraints or dependencies?
    • If this is the first run: recommend execution.mode: serial, execution.max_concurrent: 1, stopping.max_iterations: 4, and stopping.max_hours: 1
    • If type: judge: recommend sample_size: 10, batch_size: 5, and max_total_cost_usd: 5 until the rubric and harness are trusted
  6. Write the spec to .spec-first/workflows/spec-optimize/<spec-name>/spec.yaml

  7. Present the spec to the user for approval before proceeding

0.3 Search Prior Learnings

Read references/agents/learnings-researcher.md. Dispatch a generic subagent seeded with that local prompt only when the Dispatch And Backend Boundary permits it; otherwise search inline with the same bounded scope and record the matching fallback reason. Do not dispatch a standalone agent by type/name. If relevant learnings exist, incorporate them into the approach.

0.4 Run Identity Detection

Check if optimize/<spec-name> branch already exists:

bash
git rev-parse --verify "optimize/<spec-name>" 2>/dev/null

If branch exists, check for an existing experiment log at .spec-first/workflows/spec-optimize/<spec-name>/experiment-log.yaml.

Present the user with a choice via the platform question tool:

  • Resume: read ALL state from the experiment log on disk (do not rely on any in-memory context from a prior session). Recover any measured-but-unlogged experiments by scanning worktree directories for result.yaml markers. Continue from the last iteration number in the log.
  • Fresh start: archive the old branch to optimize-archive/<spec-name>/archived-<timestamp>, clear the experiment log, start from scratch
0.5 Create Optimization Branch and Scratch Space
bash
git checkout -b "optimize/<spec-name>"  # or switch to existing if resuming

Create scratch directory:

bash
mkdir -p .spec-first/workflows/spec-optimize/<spec-name>/

Phase 1: Measurement Scaffolding

This phase is a HARD GATE. The user must approve baseline and parallel readiness before Phase 2.

Bundled scripts. Phases 1 and 3 call helper scripts that ship in this skill's scripts/ directory (measure.sh, parallel-probe.sh, experiment-worktree.sh). The Bash tool's working directory is the user's project, not the skill directory, so a bare scripts/<name> path will not resolve — invoke each by the skill's own absolute path. Every runnable block below already sets SKILL_DIR inline (shell state does not persist between Bash tool calls, so each block must carry it); replace the <absolute path ...> placeholder with the directory you loaded this spec-optimize SKILL.md from before running. The shape:

bash
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
bash "$SKILL_DIR/scripts/<name>"
1.1 Clean-Tree Gate

Verify no uncommitted changes to files within scope.mutable or scope.immutable:

bash
git status --porcelain

Filter the output against the scope paths. If any in-scope files have uncommitted changes:

  • Report which files are dirty
  • Ask the user to commit or stash before proceeding
  • Do NOT continue until the working tree is clean for in-scope files
1.2 Build or Validate Measurement Harness

Resolve scripts/measure.sh, scripts/parallel-probe.sh, and scripts/experiment-worktree.sh relative to this skill's loaded directory. The measurement working directory remains the project directory named by the optimization spec.

Before the first measurement command, freeze and display this run-local execution envelope:

yaml
measurement_execution_authorization: authorized | missing
measurement_command: <exact command string from the approved spec>
measurement_working_directory: <resolved absolute cwd>
measurement_environment_names: [<inherited or overlaid variable names; never secret values>]
measurement_expected_effects: [read-only | writes-project-files | writes-local-state | network | other]

An approved optimization spec, clean tree, executable harness, baseline approval, or shell permission does not set this fact. Require the current user or visible upstream handoff to authorize the displayed command, resolved cwd, environment names, and expected effects before any invocation of measure.sh, direct measurement command, or parallel probe that executes it. When missing, return measurement_execution_authorization_missing with zero measurement command executions. If any frozen field changes later, invalidate the authorization and present the new envelope before continuing. This is a workflow-level effect gate; measure.sh remains a bounded executor and does not infer semantic authorization from the command text.

If user provides a measurement harness (the measurement.command already exists):

  1. Run it once via the measurement script:
    bash
    SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
    bash "$SKILL_DIR/scripts/measure.sh" "<measurement.command>" <timeout_seconds> "<measurement.working_directory or .>"
  2. Validate the JSON output:
    • Contains keys for all degenerate gate metric names
    • Contains keys for all diagnostic metric names
    • Values are numeric or boolean as expected
  3. If validation fails, report what is missing and ask the user to fix the harness

If agent must build the harness:

  1. Analyze the codebase to understand the current approach and what should be measured
  2. Build an evaluation script (e.g., evaluate.py, evaluate.sh, or equivalent)
  3. Add the evaluation script path to scope.immutable -- the experiment agent must not modify it
  4. Run it once and validate the output
  5. Present the harness and its output to the user for review
1.3 Establish Baseline

Run the measurement harness on the current code.

If stability mode is repeat:

  1. Run the harness repeat_count times
  2. Aggregate results using the configured aggregation method (median, mean, min, max)
  3. Calculate variance across runs
  4. If variance exceeds noise_threshold, warn the user and suggest increasing repeat_count

Record the baseline in the experiment log:

yaml
baseline:
  timestamp: "<current ISO 8601 timestamp>"
  gates:
    <gate_name>: <value>
    ...
  diagnostics:
    <diagnostic_name>: <value>
    ...

If primary type is judge, also run the judge evaluation on baseline output to establish the starting judge score.

1.4 Parallelism Readiness Probe

Run the parallelism probe script:

bash
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
bash "$SKILL_DIR/scripts/parallel-probe.sh" "<project_directory>" "<measurement.command>" "<measurement.working_directory>" <shared_files...>

Read the JSON output. Present any blockers to the user with suggested mitigations. Treat the probe as intentionally narrow: it should inspect the measurement command, the measurement working directory, and explicitly declared shared files, not the entire repository.

1.5 Worktree Budget Check

Count existing worktrees:

bash
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
bash "$SKILL_DIR/scripts/experiment-worktree.sh" count

If count + execution.max_concurrent would exceed 12:

  • Warn the user
  • Suggest cleaning up existing worktrees or reducing max_concurrent
  • Do NOT block -- the user may proceed at their own risk
1.6 Write Baseline to Disk (CP-1)

MANDATORY CHECKPOINT. Before presenting results to the user, write the initial experiment log with baseline metrics to disk:

  1. Create the experiment log file at .spec-first/workflows/spec-optimize/<spec-name>/experiment-log.yaml
  2. Include all required top-level sections from references/experiment-log-schema.yaml: spec, run_id, started_at, baseline, experiments, and best
  3. Seed experiments as an empty array and seed best from the baseline snapshot (use iteration: 0, baseline metrics, and baseline judge scores if present) so later phases have a valid current-best state to compare against
  4. Optionally seed hypothesis_backlog: [] here as well so the log shape is stable before Phase 2 populates it
  5. Verify: read the file back and confirm the required sections are present and the baseline values match
  6. Only THEN present results to the user
1.7 User Approval Gate

Present to the user via the platform question tool:

  • Baseline metrics: all gate values, diagnostic values, and judge scores (if applicable)
  • Experiment log location: show the file path so the user knows where results are saved
  • Parallel readiness: probe results, any blockers, mitigations applied
  • Clean-tree status: confirmed clean
  • Worktree budget: current count and projected usage
  • Judge budget: estimated per-experiment judge cost and configured max_total_cost_usd cap (or an explicit note that spend is uncapped)

Options:

  1. Proceed -- approve baseline and parallel config, move to Phase 2
  2. Adjust spec -- modify spec settings before proceeding
  3. Fix issues -- user needs to resolve blockers first

Do NOT proceed to Phase 2 until the user explicitly approves.

If primary type is judge and max_total_cost_usd is null, call that out as uncapped spend and require explicit approval before proceeding.

State re-read: After gate approval, re-read the spec and baseline from disk. Do not carry stale in-memory values forward.


Phase 2: Hypothesis Generation

Show full SKILL.md (2,577 more words)Show less
2.1 Analyze Current Approach

Read the code within scope.mutable to understand:

  • The current implementation approach
  • Obvious improvement opportunities
  • Constraints and dependencies between components

Optionally read references/agents/repo-research-analyst.md for deeper codebase analysis if the scope is large or unfamiliar. Dispatch a generic subagent only when the Dispatch And Backend Boundary permits it; otherwise apply the same bounded analysis inline or serially. Do not dispatch a standalone agent by type/name. Before either path, derive a run-local stack/architecture/conventions orientation from the current target repo/worktree and record its source identity and dirty state. Pass that orientation with direct source refs to repo-research-analyst, requesting only question-specific scopes such as patterns. Never reuse it across runs, branches, or worktrees. On resume, compare the checkpoint's source identity with the current tree; if it changed, re-baseline affected measurements or stop with an explicit source-drift limitation instead of carrying old grounding forward. If the current sources cannot be read, record the concrete degraded fact and narrow optimization claims.

2.2 Generate Hypothesis List

Generate an initial set of hypotheses. Each hypothesis should have:

  • Description: what to try
  • Category: one of the standard categories (signal-extraction, structural-signals, embedding, algorithm, preprocessing, parameter-tuning, architecture, data-handling) or a domain-specific category
  • Priority: high, medium, or low based on expected impact and feasibility
  • Required dependencies: any new packages or tools needed

Include user-provided hypotheses if any were given as input.

Aim for 10-30 hypotheses in the initial backlog. More can be generated during the loop based on learnings.

2.3 Dependency Pre-Approval

Collect all unique new dependencies across all hypotheses.

If any hypotheses require new dependencies:

  1. Present the full dependency list to the user via the platform question tool
  2. Ask for bulk approval
  3. Mark each hypothesis's dep_status as approved or needs_approval

Hypotheses with unapproved dependencies remain in the backlog but are skipped during batch selection. They are re-presented at wrap-up for potential approval.

2.4 Record Hypothesis Backlog (CP-2)

MANDATORY CHECKPOINT. Write the initial backlog to the experiment log file and verify:

yaml
hypothesis_backlog:
  - description: "Remove template boilerplate before embedding"
    category: "signal-extraction"
    priority: high
    dep_status: approved
    required_deps: []
  - description: "Try HDBSCAN clustering algorithm"
    category: "algorithm"
    priority: medium
    dep_status: needs_approval
    required_deps: ["scikit-learn"]

Phase 3: Optimization Loop

This phase repeats in batches until a stopping criterion is met.

3.1 Batch Selection

Select hypotheses for this batch:

  • Build a runnable backlog by excluding hypotheses with dep_status: needs_approval
  • If execution.mode is serial, force batch_size = 1
  • Otherwise, batch_size = min(runnable_backlog_size, execution.max_concurrent)
  • Prefer diversity: select from different categories when possible
  • Within a category, select by priority (high first)

If the backlog is empty and no new hypotheses can be generated, proceed to Phase 4 (wrap-up). If the backlog is non-empty but no runnable hypotheses remain because everything needs approval or is otherwise blocked, proceed to Phase 4 so the user can approve dependencies instead of spinning forever.

3.2 Execute Experiments

For each hypothesis in the batch, use the effective run mode. If either dispatch fact is missing, override the worker mode to serial inline/local execution for this run, retain the spec's requested mode as an unmet capability note, and run exactly one experiment to completion before selecting the next hypothesis. Only when the package-local boundary permits dispatch may execution.mode: parallel dispatch a batch concurrently.

Bounded dispatch. For authorized dispatch only, do not assume the host will accept all concurrent subagents at once; the active-subagent cap varies by host and profile and is independent of execution.max_concurrent (which caps worktrees, a separate budget). Queue the selected experiments, dispatch only as many as the host accepts, and when a capacity or active-agent-limit error appears, treat it as backpressure — retry the queued experiment after a slot frees rather than marking it failed. Mark an experiment failed only when dispatch fails for a non-capacity reason or a successfully dispatched experiment errors/times out.

The Phase 3 blocks below each set SKILL_DIR inline as well (the loaded spec-optimize skill directory; see the Bundled scripts note in Phase 1) — shell state does not persist from Phase 1, so each block carries its own assignment.

Worktree backend:

  1. Create experiment worktree:
    bash
    SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
    WORKTREE_PATH=$(bash "$SKILL_DIR/scripts/experiment-worktree.sh" create "<spec_name>" <exp_index> "optimize/<spec_name>" <shared_files...>)  # creates .worktrees/optimize-<spec_name>-exp-<NNN>/
  2. Apply port parameterization if configured (set env vars for the measurement script)
  3. Fill the experiment prompt template (references/experiment-prompt-template.md) with:
    • Iteration number, spec name
    • Hypothesis description and category
    • Current best and baseline metrics
    • Mutable and immutable scope
    • Constraints and approved dependencies
    • Rolling window of last 10 experiments (concise summaries)
  4. With authorized dispatch, dispatch a subagent with the filled prompt in the experiment worktree; otherwise execute the filled prompt inline in that worktree before moving to the next hypothesis

Codex backend:

  1. Check environment guard -- do NOT delegate if already inside a Codex sandbox:
    bash
    # If these exist, we're already in Codex -- fall back to subagent
    test -n "${CODEX_SANDBOX:-}" || test -n "${CODEX_SESSION_ID:-}" || test ! -w .git
  2. Fill the experiment prompt template
  3. Write the filled prompt to a temp file
  4. Dispatch via Codex only when the package-local dispatch boundary is satisfied; otherwise use the serial local/worktree path:
    bash
    cat /tmp/optimize-exp-XXXXX.txt | codex exec --skip-git-repo-check - 2>&1
  5. Security posture: use the user's selection (ask once per session if not set in spec)
3.3 Collect and Persist Results

Process experiments as they complete — do NOT wait for the entire batch to finish before writing results.

For each completed experiment, immediately:

  1. Run measurement in the experiment's worktree:

    bash
    SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
    bash "$SKILL_DIR/scripts/measure.sh" "<measurement.command>" <timeout_seconds> "<worktree_path>/<measurement.working_directory or .>" <env_vars...>
    • If stability mode is repeat, run the measurement harness repeat_count times in that working directory and aggregate the results exactly as in Phase 1 before evaluating gates or ranking the experiment.
    • Use the aggregated metrics as the experiment's score; if variance exceeds noise_threshold, record that in learnings so the operator knows the result is noisy.
  2. Write crash-recovery marker — immediately after measurement, write result.yaml in the experiment worktree containing the raw metrics. This ensures the measurement is recoverable even if the agent crashes before updating the main log.

  3. Read raw JSON output from the measurement script

  4. Evaluate degenerate gates:

    • For each gate in metric.degenerate_gates, parse the operator and threshold
    • Compare the metric value against the threshold
    • If ANY gate fails: mark outcome as degenerate, skip judge evaluation, save money
  5. If gates pass AND primary type is judge:

    • Read the experiment's output (cluster assignments, search results, etc.)
    • Apply stratified sampling per metric.judge.stratification config (using sample_seed)
    • Group samples into batches of metric.judge.batch_size
    • Fill the judge prompt template (references/judge-prompt-template.md) for each batch
    • When the package-local dispatch boundary is satisfied, dispatch the ceil(sample_size / batch_size) judge sub-agents using the same bounded scheduler as Phase 3.2. Otherwise evaluate the same batches serially inline, record the matching fallback reason, and do not claim independent judge coverage. Judge work is a separate budget from experiment worktrees in either path.
    • Each dispatched sub-agent or inline judge batch returns structured JSON scores
    • Aggregate scores: compute the configured primary judge field from metric.judge.scoring.primary (which should match metric.primary.name) plus any scoring.secondary values
    • If singleton_sample > 0: evaluate singleton batches through the same authorized-dispatch or serial-inline path
  6. If gates pass AND primary type is hard:

    • Use the metric value directly from the measurement output
  7. IMMEDIATELY append to experiment log on disk (CP-3) — do not defer this to batch evaluation. Write the experiment entry (iteration, hypothesis, outcome, metrics, learnings) to .spec-first/workflows/spec-optimize/<spec-name>/experiment-log.yaml right now. Use the transitional outcome measured once the experiment has valid metrics but has not yet been compared to the current best. Update the outcome to kept, reverted, or another terminal state in the evaluation step, but the raw metrics are on disk and safe from context compaction.

  8. VERIFY the write (CP-3 verification) — read the experiment log back from disk and confirm the entry just written is present. If verification fails, retry the write. Do NOT proceed to the next experiment until this entry is confirmed on disk.

Why immediately + verify? The agent's context window is NOT a durable store. Context compaction, session crashes, and restarts are expected during long runs. If results only exist in the agent's memory, they are lost. Karpathy's autoresearch writes to results.tsv after every single experiment — this skill must do the same with the experiment log. The verification step catches silent write failures that would otherwise lose data.

3.4 Evaluate Batch

After all experiments in the batch have been measured:

  1. Rank experiments by primary metric improvement:

    • For hard metrics: compare to the current best using metric.primary.direction (maximize means higher is better, minimize means lower is better), and require the absolute improvement to exceed measurement.stability.noise_threshold before treating it as a real win
    • For judge metrics: compare the configured primary judge score (metric.judge.scoring.primary / metric.primary.name) to the current best, and require it to exceed minimum_improvement
  2. Identify the best experiment that passes all gates and improves the primary metric

  3. If best improves on current best: KEEP

    • Commit the experiment branch first so the winning diff exists as a real commit before any merge or cherry-pick
    • Include only mutable-scope changes in that commit; if no eligible diff remains, treat the experiment as non-improving and revert it
    • Merge the committed experiment branch into the optimization branch
    • Use the message optimize(<spec-name>): <hypothesis description> for the experiment commit
    • After the merge succeeds, clean up the winner's experiment worktree and branch; the integrated commit on the optimization branch is the durable artifact
    • This is now the new baseline for subsequent batches
  4. Check file-disjoint runners-up (up to max_runner_up_merges_per_batch):

    • For each runner-up that also improved, check file-level disjointness with the kept experiment
    • File-level disjointness: two experiments are disjoint if they modified completely different files. Same file = overlapping, even if different lines.
    • If disjoint: cherry-pick the runner-up onto the new baseline, re-run full measurement
    • If combined measurement is strictly better: keep the cherry-pick (outcome: runner_up_kept), then clean up that runner-up's experiment worktree and branch
    • Otherwise: revert the cherry-pick, log as "promising alone but neutral/harmful in combination" (outcome: runner_up_reverted), then clean up the runner-up's experiment worktree and branch
    • Stop after first failed combination
  5. Handle deferred deps: experiments that need unapproved dependencies get outcome deferred_needs_approval

  6. Revert all others: cleanup worktrees, log as reverted

3.5 Update State (CP-4)

MANDATORY CHECKPOINT. By this point, individual experiment results are already on disk (written in step 3.3). This step updates aggregate state and verifies.

  1. Re-read the experiment log from disk — do not trust in-memory state. The log is the source of truth.

  2. Finalize outcomes — update experiment entries from step 3.4 evaluation (mark kept, reverted, runner_up_kept, etc.). Write these outcome updates to disk immediately.

  3. Update the best section in the experiment log if a new best was found. Write to disk.

  4. Write strategy digest to .spec-first/workflows/spec-optimize/<spec-name>/strategy-digest.md:

    • Categories tried so far (with success/failure counts)
    • Key learnings from this batch and overall
    • Exploration frontier: what categories and approaches remain untried
    • Current best metrics and improvement from baseline
  5. Generate new hypotheses based on learnings:

    • Re-read the strategy digest from disk (not from memory)
    • Read the rolling window (last 10 experiments from the log on disk)
    • Do NOT read the full experiment log -- use the digest for broad context
    • Add new hypotheses to the backlog and write the updated backlog to disk
  6. Write updated hypothesis backlog to disk — the backlog section of the experiment log must reflect newly added hypotheses and removed (tested) ones.

CP-4 Verification: Read the experiment log back from disk. Confirm: (a) all experiment outcomes from this batch are finalized, (b) the best section reflects the current best, (c) the hypothesis backlog is updated. Read strategy-digest.md back and confirm it exists. Only THEN proceed to the next batch or stopping criteria check.

Checkpoint: at this point, all state for this batch is on disk. If the agent crashes and restarts, it can resume from the experiment log without loss.

3.6 Check Stopping Criteria

Stop the loop if ANY of these are true:

  • Target reached: stopping.target_reached is true, metric.primary.target is set, and the primary metric reaches that target according to metric.primary.direction (>= for maximize, <= for minimize)
  • Max iterations: total experiments run >= stopping.max_iterations
  • Max hours: wall-clock time since Phase 3 start >= stopping.max_hours
  • Judge budget exhausted: cumulative judge spend >= metric.judge.max_total_cost_usd (if set)
  • Plateau: no improvement for stopping.plateau_iterations consecutive experiments
  • Manual stop: user interrupts (save state and proceed to Phase 4)
  • Empty backlog: no hypotheses remain and no new ones can be generated

If no stopping criterion is met, proceed to the next batch (step 3.1).

3.7 Cross-Cutting Concerns

Codex failure cascade: Track consecutive Codex delegation failures. After 3 consecutive failures, auto-disable Codex for remaining experiments. Fall back to subagent dispatch only when authorization and callable capability still permit it; otherwise continue through serial inline/local execution. Log the switch and reason code.

Error handling: If an experiment's measurement command crashes, times out, or produces malformed output:

  • Log as outcome error or timeout with the error message
  • Revert the experiment (cleanup worktree)
  • The loop continues with remaining experiments in the batch

Progress reporting: After each batch, report:

  • Batch N of estimated M (based on backlog size)
  • Experiments run this batch and total
  • Current best metric and improvement from baseline
  • Cumulative judge cost (if applicable)

Crash recovery: See Persistence Discipline section. Per-experiment result.yaml markers are written in step 3.3. Individual experiment results are appended to the log immediately in step 3.3. Batch-level state (outcomes, best, digest) is written in step 3.5. On resume (Phase 0.4), the log on disk is the ground truth — scan for any result.yaml markers not yet reflected in the log.


Phase 4: Wrap-Up

4.1 Present Deferred Hypotheses

If any hypotheses were deferred due to unapproved dependencies:

  1. List them with their dependency requirements
  2. Ask the user whether to approve, skip, or save for a future run
  3. If approved: add to backlog and offer to re-enter Phase 3 for one more round
4.2 Summarize Results

Present a comprehensive summary:

Optimization: <spec-name>
Duration: <wall-clock time>
Total experiments: <count>
  Kept: <count> (including <runner_up_kept_count> runner-up merges)
  Reverted: <count>
  Degenerate: <count>
  Errors: <count>
  Deferred: <count>

Baseline -> Final:
  <primary_metric>: <baseline_value> -> <final_value> (<delta>)
  <gate_metrics>: ...
  <diagnostics>: ...

Judge cost: $<total_judge_cost_usd> (if applicable)

Key improvements:
  1. <kept experiment 1 hypothesis> (+<delta>)
  2. <kept experiment 2 hypothesis> (+<delta>)
  ...
4.3 Preserve and Offer Next Steps

The optimization branch (optimize/<spec-name>) is preserved with all commits from kept experiments. The experiment log remains in local .spec-first/workflows/spec-optimize/<spec-name>/ scratch space for resume and audit on this machine only; it does not travel with the branch because that run-state path is gitignored. The strategy digest is derived, reconstructable state and may be regenerated from the canonical experiment log.

Present post-completion options via the platform question tool:

  1. Run code review on the cumulative diff (baseline to final). Execute spec-code-review on the optimization branch, interactive or mode:agent. To land eligible fixes before the next option, apply the mechanical-apply bar below.

    Mechanical-apply bar: apply any finding with a concrete suggested_fix that is a clear, reversible improvement; push back and keep the diff when the reviewer is wrong, noting why. Defer anything whose right fix needs a design or product decision, including architecture direction, contract shape, behavior change needing sign-off, and any finding with no concrete fix to act on. Confirm evidence still matches at file:line before editing. After applying, run tests, at least targeted tests for what changed and a broader suite for multi-file edits. Do not commit or push from this step; leave the diff on the optimization branch for the Create PR option.

  2. Capture learning by executing spec-compound to document the winning strategy as an institutional learning.

  3. Create PR from the optimization branch to the default branch.

  4. Continue with more experiments: re-enter Phase 3 with the current state. State re-read first.

  5. Done -- leave the optimization branch for manual review.

4.4 Cleanup

Clean up scratch space:

bash
# Keep the experiment log for local resume/audit on this machine
# Remove the derived strategy digest only after its required audit fields are in the log
rm -f .spec-first/workflows/spec-optimize/<spec-name>/strategy-digest.md

Do NOT delete the experiment log if the user may resume locally or wants a local audit trail. If they need a durable shared artifact, summarize or export the results into a tracked path before cleanup. Do NOT delete experiment worktrees that are still being referenced.

© leo-kuang-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 31 other files (scripts, references) in skills/spec-optimize of leo-kuang-ai/spec-first.

  • SKILL.md
  • README.md
  • evals/README.md
  • evals/cases/debug-request-routes-out.yaml
  • evals/cases/empty-input-asks.yaml
  • evals/cases/no-metric-goal-gated.yaml
  • evals/cases/r2-fastest-no-measure.yaml
  • evals/eval.yaml
  • evals/examples.json
  • evals/fixtures/repos/mini-ledger/README.md
  • evals/fixtures/repos/mini-ledger/package.json
  • evals/fixtures/repos/mini-ledger/src/server.js
  • evals/fixtures/scripts/asks-a-question.sh
  • evals/fixtures/scripts/check-no-metric-gate.sh
  • … and 18 more

Open the folder on GitHubat commit 74655dc

Compare with similar skills

Spec Optimize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spec Optimize compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spec Optimize this skillleo-kuang-ai/spec-first107—~13kAutomated safety check: PassMIT
Clawpathy AutoresearchClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT
Agents Best PracticesDenisSergeevitch/agents-best-practices2.4k—~7.4kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Autocontext Knowledge Creatorgreyhaven-ai/autocontext1.3k—~964Automated safety check: PassApache-2.0
Axiom Eval Writeropenclaw/clawhub9.5k—~4.1kAutomated safety check: WarnMIT

Similar skills

  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agents Best Practices

    DenisSergeevitch/agents-best-practices

    A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.

    2.4k GitHub stars~7.4k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Autocontext Knowledge Creator

    greyhaven-ai/autocontext

    Runs the `autoctx` CLI to improve an approach to a task over several generations, score or refine a single output and inspect what a run produced.

    1.3k GitHub stars~964 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Axiom Eval Writer

    openclaw/clawhub

    Scaffolds evaluation suites for the Axiom AI SDK: eval files, scorers, flag schemas and axiom.config.ts, generated from plain descriptions of an AI capability.

    9.5k GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: warnings
  • Skill Forge Benchmark

    AgriciDaniel/skill-forge

    Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations.

    179 GitHub stars~1.4k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from leo-kuang-ai/spec-first

All 35 skills in this repo
  • Spec App Consistency Audit

    leo-kuang-ai/spec-first

    Audit mobile App PRD/Figma/local-source consistency across page routes, KMP/Clean Architecture, components, analytics, i18n, engineering quality, and industry lenses before runtime validation; use…

    107 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Handoff

    leo-kuang-ai/spec-first

    Create a durable cross-session handoff or resume from a user-selected continuity source.

    107 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Pov

    leo-kuang-ai/spec-first

    Give a decisive, project-grounded verdict on an external input — judged against the current project, not in the abstract.

    107 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Resolve PR Feedback

    leo-kuang-ai/spec-first

    Resolve PR review feedback by evaluating validity and fixing issues with conflict-aware resolver dispatch.

    107 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check: notes
  • Spec Riffrec Feedback Analysis

    leo-kuang-ai/spec-first

    Analyze explicit Riffrec product-feedback captures, including riffrec-.zip, the Riffrec session.json + events.json + recording.webm + voice.webm bundle, or media/notes the user identifies as a…

    107 GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Compound

    leo-kuang-ai/spec-first

    Document a recently solved problem or durable project vocabulary in docs/solutions/ or CONCEPTS.md.

    107 GitHub stars~18k tokensUpdated 2 days ago
    Auto-check passed

Questions about Spec Optimize

What does Spec Optimize do?

Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first. Spec Optimize is an agent skill from leo-kuang-ai/spec-first. Run metric-driven iterative optimization loops.

When should I use Spec Optimize?

Spec Optimize fits situations like: optimizing clustering quality; search relevance; build performance; any measurable outcome that benefits from systematic experimentation.

How do I install Spec Optimize in Claude Code?

Run `npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a claude-code`. Or copy the skill folder (skills/spec-optimize in leo-kuang-ai/spec-first) into .claude/skills/spec-optimize in your project. Claude Code loads it when a task matches its description.

How do I install Spec Optimize in Codex?

Run `npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a codex`. Or copy the skill folder (skills/spec-optimize in leo-kuang-ai/spec-first) into .agents/skills/spec-optimize in your project. Codex loads it when a task matches its description.

Can I use Spec Optimize in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-optimize, .gemini/skills/spec-optimize, .github/skills/spec-optimize and .opencode/skills/spec-optimize in your project.

What does Spec Optimize need to run?

Going by SKILL.md and its folder, Spec Optimize needs a shell and JavaScript for the scripts in its folder and the command-line tools its instructions call (bash, git, codex and node). Our summary lists: Node.js; A Bash shell.

Does Spec Optimize access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Spec Optimize safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Spec Optimize use?

Spec Optimize is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spec Optimize use?

About 13k tokens (SKILL.md is roughly 51k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.

What are the alternatives to Spec Optimize?

Skills that share tags, products or a category with Spec Optimize: Clawpathy Autoresearch (ClawBio/ClawBio, 1.2k stars), Agents Best Practices (DenisSergeevitch/agents-best-practices, 2.4k stars), Looper (ksimback/looper, 710 stars) and Autocontext Knowledge Creator (greyhaven-ai/autocontext, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spec Optimize?

leo-kuang-ai (a GitHub user) maintains it in leo-kuang-ai/spec-first, which has 107 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: leo-kuang-ai/spec-first on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.