Clawpathy Autoresearch
ClawBio/ClawBio
Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.
Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.
$ npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-optimize --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spec-optimize .claude/skills/spec-optimize && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "spec-optimize" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-optimize into .claude/skills/spec-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-optimize", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-optimizeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-optimize --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/spec-optimize .agents/skills/spec-optimize && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "spec-optimize" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-optimize into .agents/skills/spec-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-optimize", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-optimize --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/spec-optimize .cursor/skills/spec-optimize && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "spec-optimize" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-optimize into .cursor/skills/spec-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-optimize", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/leo-kuang-ai/spec-first.git --path skills/spec-optimize--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-optimize --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/spec-optimize .gemini/skills/spec-optimize && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "spec-optimize" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-optimize into .gemini/skills/spec-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-optimize", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install leo-kuang-ai/spec-first spec-optimizeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/spec-optimize .github/skills/spec-optimize && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "spec-optimize" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-optimize into .github/skills/spec-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-optimize", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-optimize --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/spec-optimize .opencode/skills/spec-optimize && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "spec-optimize" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-optimize into .opencode/skills/spec-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-optimize", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
spec-optimizeRun metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.
Spec Optimize is an agent skill from leo-kuang-ai/spec-first. Run metric-driven iterative optimization loops. Define a measurable goal, build measurement scaffolding, then run parallel experiments that try many approaches, measure each against hard gates and/or LLM-as-judge quality scores, keep improvements, and converge toward the best solution. Use when optimizing clustering quality, search relevance, build performance, prompt quality, or any measurable outcome that benefits from systematic experimentation. Inspired by Karpathy's autoresearch, generalized for multi-file…
Its SKILL.md is about 13k tokens, which your agent loads only when the skill is triggered. The skill folder holds 38 other files, including scripts and reference files (for example `README.md`, `evals/README.md` and `evals/cases/debug-request-routes-out.yaml`).
It sits in AI & LLM Engineering, covering Project scaffolding, LLM evaluation and Search implementation. The repository describes itself as: 仓库原生 AI Coding Harness —— 把一次性 AI 对话变成可治理、可验证、可沉淀的工程闭环 · spec-first.cn. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 74655dc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell and JavaScript, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
bashgitcodexnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Spec Optimize loads about 13k tokens when it runs, and up to ~34k if it reads all its reference files. Until then it costs about 141 tokens; SKILL.md has 6,127 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from leo-kuang-ai/spec-first at commit 74655dc, republished under its MIT licence (© leo-kuang-ai). 6,127 words, ~12,834 tokens.
.claude/skills/spec-optimize/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.Run metric-driven iterative optimization. Define a goal, build measurement scaffolding, then run parallel experiments that converge toward the best solution.
Use when a measurable outcome can improve through iterative experiments, hard gates, and/or LLM-as-judge scoring.
Do not use for ordinary implementation, vague improvement requests without a metric, debugging without a feedback loop, or unbounded spend/concurrency. Name the destination when routing out: a bug with no stable repro and no measurement loop is spec-debug's work (or spec-work for settled implementation) — diagnosing the bug inside this workflow to turn it into an optimization goal is adopting the wrong workflow, not adapting it; route out explicitly instead of drafting a spec around the diagnosis.
An optimization spec or goal, mutable/immutable scope, measurement command or scaffold plan, budget limits, experiment settings, repository instructions, and baseline evidence.
A measurement scaffold and experiment log, scored experiment results, kept/rejected variants, final integrated changes when appropriate, and post-run recommendations.
Run state under .spec-first/workflows/spec-optimize/<spec-name>/, experiment worktrees/results, strategy digests, and no hidden workflow state outside the documented log.
Missing metric, missing measurement command, unsafe scope, excessive or uncapped budget, failed baseline, write verification failure, or unavailable dispatch/worktree backend.
Validate the spec and budget, establish the baseline, run bounded experiments, measure and write results immediately, select winners, integrate only verified improvements, and summarize evidence.
Code review、benchmark maintainer、在性能/相关性变更时参与的 release reviewer,以及检查 experiment logs 的人工审查者。
Follows docs/contracts/workflows/scenario-capability-matrix.md (default).
Overrides: none
Use the platform's blocking question tool: AskUserQuestion in Claude Code (call ToolSearch with select:AskUserQuestion first if its schema isn't loaded) or request_user_input in Codex. Fall back to numbered options in chat only when no blocking tool exists in the harness or the call errors (e.g., Codex edit modes) — not because a schema load is required. Never silently skip the question.
<optimization_input> #<invocation arguments supplied by the current host> </optimization_input>
If the input above is empty, ask: "What would you like to optimize? Describe the goal, or provide a path to an optimization spec YAML file."
Reference the spec schema for validation:
references/optimize-spec-schema.yaml
Reference the experiment log schema for state management:
references/experiment-log-schema.yaml
For a first run, optimize for signal and safety, not maximum throughput:
references/example-hard-spec.yaml when the metric is objective and cheap to measurereferences/example-judge-spec.yaml only when actual quality requires semantic judgmentexecution.mode: serial and execution.max_concurrent: 1stopping.max_iterations: 4 and stopping.max_hours: 1sample_size: 10, batch_size: 5, and max_total_cost_usd: 5For a friendly overview of what this skill is for, when to use hard metrics vs LLM-as-judge, and example kickoff prompts, see:
references/usage-guide.md
Do not run spec-optimize as an expensive substitute for ordinary work. Before Phase 1, confirm the run has all of these:
metric.primary.type, metric.primary.name, and metric.primary.directionscope.mutable and scope.immutable boundariesstopping.max_iterations, stopping.max_hours, and stopping.plateau_iterationsexecution.mode and execution.max_concurrentmetric.judge.max_total_cost_usd, unless the user explicitly approves uncapped spendIf any item is missing, stop and help the user create a safe spec or route the work to the current host's plan/work/debug entrypoint instead. Do not continue with an open-ended optimization loop.
When the invocation or approved spec selects mode:measurement-only, load
references/measurement-only-calibration.md. This mode compares an explicitly
identified baseline and candidate against the same repeatable task/corpus. It
must run an A/A noise floor before A/B, use a pre-registered acceptance
threshold, classify broken runs separately from regressions, and produce only
a measurement artifact plus stop/defer guidance.
Before the first measurement, materialize the approved frozen inputs as a
run-local measurement-admission-input.json, then run
node scripts/measurement-admission.cjs admit --input <path>. Persist the
returned normalized admission and admission_sha256 beside the measurement
artifact and bind every A/A and A/B attempt to that digest. A rejected admission
stops before invoking the measurement command. After A/A, run the helper's
allow-ab command with the normalized admission, digest, attempts, and observed
noise floor; A/B is forbidden unless it returns ab_allowed: true.
Measurement-only mode does not mutate either arm, any Skill package, the
measurement harness, or promotion metadata. It does not select a winner for
integration, invoke spec-write-skill, or authorize commit/landing. If arm
identity, corpus identity, a repeatable harness, or the pre-registered threshold
is missing, stop before measurement instead of inventing it after seeing data.
First-run specs should default to execution.mode: serial, execution.max_concurrent: 1, stopping.max_iterations: 4, stopping.max_hours: 1, stopping.plateau_iterations: 3, and max_runner_up_merges_per_batch: 0. Treat higher-throughput settings as opt-in. If a provided spec asks for execution.max_concurrent > 4, stopping.max_iterations > 30, stopping.max_hours > 4, or uncapped judge spend, surface those costs in the approval gate before running the baseline.
Follow docs/contracts/context-governance.md: ordinary Optimize context excludes .spec-first/audits/**, .spec-first/governance/**, and generated mirrors (.claude/**, .codex/**, .agents/skills/**, .cursor/skills/**, .cursor/spec-first/**, .cursor/mcp.json, .kiro/skills/**, .kiro/agents/**, .kiro/spec-first/**, .kiro/settings/**, .qoder/commands/spec-*.md, .qoder/commands/spec/**, .qoder/skills/**, .qoder/agents/**, .qoder/spec-first/**, .qoder/settings.local.json) by default. Optimization run state under .spec-first/workflows/spec-optimize/** is local scratch for this workflow only; pass compact strategy/result summaries to agents and avoid broad runtime/audit/governance scans unless the metric explicitly targets runtime/setup/audit/governance behavior. Cursor-native .cursor/rules/** / .cursor/agents/**, Kiro-native .kiro/specs/**, and Qoder-native .qoder/rules/** are advisory input only when explicitly named.
Optimization may consume prior direct-read summaries, degraded reason counts, source-confirmed session evidence, review summaries, test results, and quality-gate reports as diagnostic context. Treat these as baseline diagnostics, not optimization targets by default. The optimization target, metric, winner selection, mutable scope, and final integration remain owned by the approved optimization spec and measured results, not by an external tool. Optimize must not run external-tool refresh, hooks, watchers, or daemons as part of ordinary evidence handling.
Optimization dispatch is optional. Before any learnings researcher, repo analyst, experiment worker, Codex delegation, or parallel worker run, record:
worker_dispatch_authorization: authorized | missing
capability_probe: not_applicable | attempted | unavailable
worker_dispatch_capability: available | missing | unknown
worker_context_isolation: isolated | inherited | unknown
worker_model_override: supported | unsupported | unknown
worker_bounded_parallelism: supported | unsupported | unknownworkflow invocation does not authorize dispatch。Approved optimization spec、baseline approval、execution.mode: parallel、预算、权限设置或 runtime readiness 都不是派发授权。只有当前用户或可见 upstream handoff 明确请求 subagent、delegated work、persona 或 parallel work 时才可派发。缺授权时不得探测 tool schema,固定为 capability_probe: not_applicable + worker_dispatch_capability: unknown,强制采用 serial inline/local execution 并记录 dispatch_authorization_missing。只有授权后才把 current-session registry/schema 作为 provider_untrusted evidence 检查:确认缺失时记录 subagent_capability_missing;surface 不可用、schema 不完整或候选不唯一时记录 worker_capability_unproven,均同样降级。隔离、模型覆盖和有界并发只取 live facts;required isolation 未满足时保持依赖 gate 打开,model unknown 时继承,parallelism unknown 时串行。记录 worker_dispatch_outcome。Fallback 可以继续使用串行 worktree,但不得声称 parallel experiment 或 independent worker coverage。
Parallel experiments additionally require both dispatch facts above, explicit execution.mode, bounded execution.max_concurrent, clean mutable/immutable scope, and the worktree readiness probes below. Worktree-backed mutation happens in experiment worktrees; Codex delegation must fall back after repeated failures when the serial/local path can continue. The orchestrator owns final integration: selecting kept experiments, merging or cherry-picking winners, reverting non-winners, cleaning worktrees, updating experiment logs, and presenting post-completion actions. Workers never stage, commit, merge, push, or mutate the authoritative experiment log.
CRITICAL: The experiment log on disk is the single source of truth. The conversation context is NOT durable storage. Results that exist only in the conversation WILL be lost.
The files under .spec-first/workflows/spec-optimize/<spec-name>/ are local scratch state. They are ignored by git, so they survive local resumes on the same machine but are not preserved by commits, branches, or pushes unless the user exports them separately.
This skill runs for hours. Context windows compact, sessions crash, and agents restart. Every piece of state that matters MUST live on disk, not in the agent's memory.
If you produce a results table in the conversation without writing those results to disk first, you have a bug. The conversation is for the user's benefit. The experiment log file is for durability.
Write each experiment result to disk IMMEDIATELY after measurement — not after the batch, not after evaluation, IMMEDIATELY. Append the experiment entry to the experiment log file the moment its metrics are known, before evaluating the next experiment. This is the #1 crash-safety rule.
VERIFY every critical write — after writing the experiment log, read the file back and confirm the entry is present. This catches silent write failures. Do not proceed to the next experiment until verification passes.
Re-read from disk at every phase boundary and before every decision — never trust in-memory state across phase transitions, batch boundaries, or after any operation that might have taken significant time. Re-read the experiment log and strategy digest from disk.
The experiment log is append-only during Phase 3 — never rewrite the full file. Append new experiment entries. Update the best section in place only when a new best is found. This prevents data loss if a write is interrupted.
Per-experiment result markers for crash recovery — each experiment writes a result.yaml marker in its worktree immediately after measurement. On resume, scan for these markers to recover experiments that were measured but not yet logged.
Strategy digest is written after every batch, before generating new hypotheses — the agent reads the digest (not its memory) when deciding what to try next. The strategy digest is derived, reconstructable state; the experiment log remains the canonical resume and audit source. Persist every hypothesis decision, kept/reverted result, metric, and audit field needed to reconstruct the digest in the experiment log before replacing or deleting the digest.
Never present results to the user without writing them to disk first — the pattern is: measure -> write to disk -> verify -> THEN show the user. Not the reverse.
These are non-negotiable write-then-verify steps. At each checkpoint, the agent MUST write the specified file and then read it back to confirm the write succeeded.
| Checkpoint | File Written | Phase |
|---|---|---|
| CP-0: Spec saved | spec.yaml | Phase 0, after user approval |
| CP-1: Baseline recorded | experiment-log.yaml (initial with baseline) | Phase 1, after baseline measurement |
| CP-2: Hypothesis backlog saved | experiment-log.yaml (hypothesis_backlog section) | Phase 2, after hypothesis generation |
| CP-3: Each experiment result | experiment-log.yaml (append experiment entry) | Phase 3.3, immediately after each measurement |
| CP-4: Batch summary | experiment-log.yaml (outcomes + best) + strategy-digest.md | Phase 3.5, after batch evaluation |
| CP-5: Final summary | experiment-log.yaml (final state) | Phase 4, at wrap-up |
Format of a verification step:
.spec-first/workflows/spec-optimize/<spec-name>/)| File | Purpose | Written When |
|---|---|---|
spec.yaml | Optimization spec (immutable during run) | Phase 0 (CP-0) |
experiment-log.yaml | Full history of all experiments | Initialized at CP-1, appended at CP-3, updated at CP-4 |
strategy-digest.md | Compressed learnings for hypothesis generation | Written at CP-4 after each batch |
<worktree>/result.yaml | Per-experiment crash-recovery marker | Immediately after measurement, before CP-3 |
When Phase 0.4 detects an existing run:
result.yaml markers not yet in the logCheck whether the input is:
.yaml or .yml): read and validate itspec-debug (unstable/unresolved failures) or spec-work (settled implementation) in the reply, route out, and stop this workflow.If spec file provided:
validation_rules section of references/optimize-spec-schema.yaml. That section is the single source of truth for what a valid spec requires; do not rely on a remembered subset. Conditional rules such as exclusive-resource serial execution, singleton-rubric requirements, uncapped judge spend approval, high-throughput approval, and stopping criteria live there.If description provided:
Analyze the project to understand what can be measured
Detect whether the optimization target is qualitative or quantitative — this determines type: hard vs type: judge and is the single most important spec decision:
Use type: hard when:
Use type: judge when:
IMPORTANT: If the target is qualitative, strongly recommend type: judge. Explain that hard metrics alone will optimize proxy numbers without checking actual quality. Show the user the three-tier approach:
If the user insists on type: hard for a qualitative target, proceed but warn that the results may optimize a misleading proxy.
Design the sampling strategy (for type: judge):
Guide the user through defining stratified sampling. The key question is: "What parts of the output space do you need to check quality on?"
Walk through these questions:
Example stratified sampling for clustering:
stratification:
- bucket: "top_by_size" # largest clusters — check for degenerate mega-clusters
count: 10
- bucket: "mid_range" # middle of non-solo cluster size range — representative quality
count: 10
- bucket: "small_clusters" # clusters with 2-3 items — check if connections are real
count: 10
singleton_sample: 15 # singletons — check for false negatives (items that should cluster)The sampling strategy is domain-specific. For search relevance, strata might be "top-3 results", "results 4-10", "tail results". For summarization, strata might be "short documents", "long documents", "multi-topic documents".
Singleton evaluation is critical when the goal involves coverage — sampling singletons with the singleton rubric checks whether the system is missing obvious groupings.
Design the rubric (for type: judge):
Help the user define the scoring rubric. A good rubric:
distinct_topics, outlier_count)Example for clustering:
rubric: |
Rate this cluster 1-5:
- 5: All items clearly about the same issue/feature
- 4: Strong theme, minor outliers
- 3: Related but covers 2-3 sub-topics that could reasonably be split
- 2: Weak connection — items share superficial similarity only
- 1: Unrelated items grouped together
Also report: distinct_topics (integer), outlier_count (integer)Guide the user through the remaining spec fields:
execution.mode: serial, execution.max_concurrent: 1, stopping.max_iterations: 4, and stopping.max_hours: 1type: judge: recommend sample_size: 10, batch_size: 5, and max_total_cost_usd: 5 until the rubric and harness are trustedWrite the spec to .spec-first/workflows/spec-optimize/<spec-name>/spec.yaml
Present the spec to the user for approval before proceeding
Read references/agents/learnings-researcher.md. Dispatch a generic subagent seeded with that local prompt only when the Dispatch And Backend Boundary permits it; otherwise search inline with the same bounded scope and record the matching fallback reason. Do not dispatch a standalone agent by type/name. If relevant learnings exist, incorporate them into the approach.
Check if optimize/<spec-name> branch already exists:
git rev-parse --verify "optimize/<spec-name>" 2>/dev/nullIf branch exists, check for an existing experiment log at .spec-first/workflows/spec-optimize/<spec-name>/experiment-log.yaml.
Present the user with a choice via the platform question tool:
result.yaml markers. Continue from the last iteration number in the log.optimize-archive/<spec-name>/archived-<timestamp>, clear the experiment log, start from scratchgit checkout -b "optimize/<spec-name>" # or switch to existing if resumingCreate scratch directory:
mkdir -p .spec-first/workflows/spec-optimize/<spec-name>/This phase is a HARD GATE. The user must approve baseline and parallel readiness before Phase 2.
Bundled scripts. Phases 1 and 3 call helper scripts that ship in this skill's scripts/ directory (measure.sh, parallel-probe.sh, experiment-worktree.sh). The Bash tool's working directory is the user's project, not the skill directory, so a bare scripts/<name> path will not resolve — invoke each by the skill's own absolute path. Every runnable block below already sets SKILL_DIR inline (shell state does not persist between Bash tool calls, so each block must carry it); replace the <absolute path ...> placeholder with the directory you loaded this spec-optimize SKILL.md from before running. The shape:
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
bash "$SKILL_DIR/scripts/<name>"Verify no uncommitted changes to files within scope.mutable or scope.immutable:
git status --porcelainFilter the output against the scope paths. If any in-scope files have uncommitted changes:
Resolve scripts/measure.sh, scripts/parallel-probe.sh, and scripts/experiment-worktree.sh relative to this skill's loaded directory. The measurement working directory remains the project directory named by the optimization spec.
Before the first measurement command, freeze and display this run-local execution envelope:
measurement_execution_authorization: authorized | missing
measurement_command: <exact command string from the approved spec>
measurement_working_directory: <resolved absolute cwd>
measurement_environment_names: [<inherited or overlaid variable names; never secret values>]
measurement_expected_effects: [read-only | writes-project-files | writes-local-state | network | other]An approved optimization spec, clean tree, executable harness, baseline approval, or shell permission does not set this fact. Require the current user or visible upstream handoff to authorize the displayed command, resolved cwd, environment names, and expected effects before any invocation of measure.sh, direct measurement command, or parallel probe that executes it. When missing, return measurement_execution_authorization_missing with zero measurement command executions. If any frozen field changes later, invalidate the authorization and present the new envelope before continuing. This is a workflow-level effect gate; measure.sh remains a bounded executor and does not infer semantic authorization from the command text.
If user provides a measurement harness (the measurement.command already exists):
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
bash "$SKILL_DIR/scripts/measure.sh" "<measurement.command>" <timeout_seconds> "<measurement.working_directory or .>"If agent must build the harness:
evaluate.py, evaluate.sh, or equivalent)scope.immutable -- the experiment agent must not modify itRun the measurement harness on the current code.
If stability mode is repeat:
repeat_count timesnoise_threshold, warn the user and suggest increasing repeat_countRecord the baseline in the experiment log:
baseline:
timestamp: "<current ISO 8601 timestamp>"
gates:
<gate_name>: <value>
...
diagnostics:
<diagnostic_name>: <value>
...If primary type is judge, also run the judge evaluation on baseline output to establish the starting judge score.
Run the parallelism probe script:
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
bash "$SKILL_DIR/scripts/parallel-probe.sh" "<project_directory>" "<measurement.command>" "<measurement.working_directory>" <shared_files...>Read the JSON output. Present any blockers to the user with suggested mitigations. Treat the probe as intentionally narrow: it should inspect the measurement command, the measurement working directory, and explicitly declared shared files, not the entire repository.
Count existing worktrees:
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
bash "$SKILL_DIR/scripts/experiment-worktree.sh" countIf count + execution.max_concurrent would exceed 12:
max_concurrentMANDATORY CHECKPOINT. Before presenting results to the user, write the initial experiment log with baseline metrics to disk:
.spec-first/workflows/spec-optimize/<spec-name>/experiment-log.yamlreferences/experiment-log-schema.yaml: spec, run_id, started_at, baseline, experiments, and bestexperiments as an empty array and seed best from the baseline snapshot (use iteration: 0, baseline metrics, and baseline judge scores if present) so later phases have a valid current-best state to compare againsthypothesis_backlog: [] here as well so the log shape is stable before Phase 2 populates itPresent to the user via the platform question tool:
max_total_cost_usd cap (or an explicit note that spend is uncapped)Options:
Do NOT proceed to Phase 2 until the user explicitly approves.
If primary type is judge and max_total_cost_usd is null, call that out as uncapped spend and require explicit approval before proceeding.
State re-read: After gate approval, re-read the spec and baseline from disk. Do not carry stale in-memory values forward.
Read the code within scope.mutable to understand:
Optionally read references/agents/repo-research-analyst.md for deeper codebase analysis if the scope is large or unfamiliar. Dispatch a generic subagent only when the Dispatch And Backend Boundary permits it; otherwise apply the same bounded analysis inline or serially. Do not dispatch a standalone agent by type/name. Before either path, derive a run-local stack/architecture/conventions orientation from the current target repo/worktree and record its source identity and dirty state. Pass that orientation with direct source refs to repo-research-analyst, requesting only question-specific scopes such as patterns. Never reuse it across runs, branches, or worktrees. On resume, compare the checkpoint's source identity with the current tree; if it changed, re-baseline affected measurements or stop with an explicit source-drift limitation instead of carrying old grounding forward. If the current sources cannot be read, record the concrete degraded fact and narrow optimization claims.
Generate an initial set of hypotheses. Each hypothesis should have:
Include user-provided hypotheses if any were given as input.
Aim for 10-30 hypotheses in the initial backlog. More can be generated during the loop based on learnings.
Collect all unique new dependencies across all hypotheses.
If any hypotheses require new dependencies:
dep_status as approved or needs_approvalHypotheses with unapproved dependencies remain in the backlog but are skipped during batch selection. They are re-presented at wrap-up for potential approval.
MANDATORY CHECKPOINT. Write the initial backlog to the experiment log file and verify:
hypothesis_backlog:
- description: "Remove template boilerplate before embedding"
category: "signal-extraction"
priority: high
dep_status: approved
required_deps: []
- description: "Try HDBSCAN clustering algorithm"
category: "algorithm"
priority: medium
dep_status: needs_approval
required_deps: ["scikit-learn"]This phase repeats in batches until a stopping criterion is met.
Select hypotheses for this batch:
dep_status: needs_approvalexecution.mode is serial, force batch_size = 1batch_size = min(runnable_backlog_size, execution.max_concurrent)If the backlog is empty and no new hypotheses can be generated, proceed to Phase 4 (wrap-up). If the backlog is non-empty but no runnable hypotheses remain because everything needs approval or is otherwise blocked, proceed to Phase 4 so the user can approve dependencies instead of spinning forever.
For each hypothesis in the batch, use the effective run mode. If either dispatch fact is missing, override the worker mode to serial inline/local execution for this run, retain the spec's requested mode as an unmet capability note, and run exactly one experiment to completion before selecting the next hypothesis. Only when the package-local boundary permits dispatch may execution.mode: parallel dispatch a batch concurrently.
Bounded dispatch. For authorized dispatch only, do not assume the host will accept all concurrent subagents at once; the active-subagent cap varies by host and profile and is independent of execution.max_concurrent (which caps worktrees, a separate budget). Queue the selected experiments, dispatch only as many as the host accepts, and when a capacity or active-agent-limit error appears, treat it as backpressure — retry the queued experiment after a slot frees rather than marking it failed. Mark an experiment failed only when dispatch fails for a non-capacity reason or a successfully dispatched experiment errors/times out.
The Phase 3 blocks below each set SKILL_DIR inline as well (the loaded spec-optimize skill directory; see the Bundled scripts note in Phase 1) — shell state does not persist from Phase 1, so each block carries its own assignment.
Worktree backend:
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
WORKTREE_PATH=$(bash "$SKILL_DIR/scripts/experiment-worktree.sh" create "<spec_name>" <exp_index> "optimize/<spec_name>" <shared_files...>) # creates .worktrees/optimize-<spec_name>-exp-<NNN>/references/experiment-prompt-template.md) with:Codex backend:
# If these exist, we're already in Codex -- fall back to subagent
test -n "${CODEX_SANDBOX:-}" || test -n "${CODEX_SESSION_ID:-}" || test ! -w .gitcat /tmp/optimize-exp-XXXXX.txt | codex exec --skip-git-repo-check - 2>&1Process experiments as they complete — do NOT wait for the entire batch to finish before writing results.
For each completed experiment, immediately:
Run measurement in the experiment's worktree:
SKILL_DIR="<absolute path of the directory containing this SKILL.md>"
bash "$SKILL_DIR/scripts/measure.sh" "<measurement.command>" <timeout_seconds> "<worktree_path>/<measurement.working_directory or .>" <env_vars...>repeat, run the measurement harness repeat_count times in that working directory and aggregate the results exactly as in Phase 1 before evaluating gates or ranking the experiment.noise_threshold, record that in learnings so the operator knows the result is noisy.Write crash-recovery marker — immediately after measurement, write result.yaml in the experiment worktree containing the raw metrics. This ensures the measurement is recoverable even if the agent crashes before updating the main log.
Read raw JSON output from the measurement script
Evaluate degenerate gates:
metric.degenerate_gates, parse the operator and thresholddegenerate, skip judge evaluation, save moneyIf gates pass AND primary type is judge:
metric.judge.stratification config (using sample_seed)metric.judge.batch_sizereferences/judge-prompt-template.md) for each batchceil(sample_size / batch_size) judge sub-agents using the same bounded scheduler as Phase 3.2. Otherwise evaluate the same batches serially inline, record the matching fallback reason, and do not claim independent judge coverage. Judge work is a separate budget from experiment worktrees in either path.metric.judge.scoring.primary (which should match metric.primary.name) plus any scoring.secondary valuessingleton_sample > 0: evaluate singleton batches through the same authorized-dispatch or serial-inline pathIf gates pass AND primary type is hard:
IMMEDIATELY append to experiment log on disk (CP-3) — do not defer this to batch evaluation. Write the experiment entry (iteration, hypothesis, outcome, metrics, learnings) to .spec-first/workflows/spec-optimize/<spec-name>/experiment-log.yaml right now. Use the transitional outcome measured once the experiment has valid metrics but has not yet been compared to the current best. Update the outcome to kept, reverted, or another terminal state in the evaluation step, but the raw metrics are on disk and safe from context compaction.
VERIFY the write (CP-3 verification) — read the experiment log back from disk and confirm the entry just written is present. If verification fails, retry the write. Do NOT proceed to the next experiment until this entry is confirmed on disk.
Why immediately + verify? The agent's context window is NOT a durable store. Context compaction, session crashes, and restarts are expected during long runs. If results only exist in the agent's memory, they are lost. Karpathy's autoresearch writes to results.tsv after every single experiment — this skill must do the same with the experiment log. The verification step catches silent write failures that would otherwise lose data.
After all experiments in the batch have been measured:
Rank experiments by primary metric improvement:
metric.primary.direction (maximize means higher is better, minimize means lower is better), and require the absolute improvement to exceed measurement.stability.noise_threshold before treating it as a real winmetric.judge.scoring.primary / metric.primary.name) to the current best, and require it to exceed minimum_improvementIdentify the best experiment that passes all gates and improves the primary metric
If best improves on current best: KEEP
optimize(<spec-name>): <hypothesis description> for the experiment commitCheck file-disjoint runners-up (up to max_runner_up_merges_per_batch):
runner_up_kept), then clean up that runner-up's experiment worktree and branchrunner_up_reverted), then clean up the runner-up's experiment worktree and branchHandle deferred deps: experiments that need unapproved dependencies get outcome deferred_needs_approval
Revert all others: cleanup worktrees, log as reverted
MANDATORY CHECKPOINT. By this point, individual experiment results are already on disk (written in step 3.3). This step updates aggregate state and verifies.
Re-read the experiment log from disk — do not trust in-memory state. The log is the source of truth.
Finalize outcomes — update experiment entries from step 3.4 evaluation (mark kept, reverted, runner_up_kept, etc.). Write these outcome updates to disk immediately.
Update the best section in the experiment log if a new best was found. Write to disk.
Write strategy digest to .spec-first/workflows/spec-optimize/<spec-name>/strategy-digest.md:
Generate new hypotheses based on learnings:
Write updated hypothesis backlog to disk — the backlog section of the experiment log must reflect newly added hypotheses and removed (tested) ones.
CP-4 Verification: Read the experiment log back from disk. Confirm: (a) all experiment outcomes from this batch are finalized, (b) the best section reflects the current best, (c) the hypothesis backlog is updated. Read strategy-digest.md back and confirm it exists. Only THEN proceed to the next batch or stopping criteria check.
Checkpoint: at this point, all state for this batch is on disk. If the agent crashes and restarts, it can resume from the experiment log without loss.
Stop the loop if ANY of these are true:
stopping.target_reached is true, metric.primary.target is set, and the primary metric reaches that target according to metric.primary.direction (>= for maximize, <= for minimize)stopping.max_iterationsstopping.max_hoursmetric.judge.max_total_cost_usd (if set)stopping.plateau_iterations consecutive experimentsIf no stopping criterion is met, proceed to the next batch (step 3.1).
Codex failure cascade: Track consecutive Codex delegation failures. After 3 consecutive failures, auto-disable Codex for remaining experiments. Fall back to subagent dispatch only when authorization and callable capability still permit it; otherwise continue through serial inline/local execution. Log the switch and reason code.
Error handling: If an experiment's measurement command crashes, times out, or produces malformed output:
error or timeout with the error messageProgress reporting: After each batch, report:
Crash recovery: See Persistence Discipline section. Per-experiment result.yaml markers are written in step 3.3. Individual experiment results are appended to the log immediately in step 3.3. Batch-level state (outcomes, best, digest) is written in step 3.5. On resume (Phase 0.4), the log on disk is the ground truth — scan for any result.yaml markers not yet reflected in the log.
If any hypotheses were deferred due to unapproved dependencies:
Present a comprehensive summary:
Optimization: <spec-name>
Duration: <wall-clock time>
Total experiments: <count>
Kept: <count> (including <runner_up_kept_count> runner-up merges)
Reverted: <count>
Degenerate: <count>
Errors: <count>
Deferred: <count>
Baseline -> Final:
<primary_metric>: <baseline_value> -> <final_value> (<delta>)
<gate_metrics>: ...
<diagnostics>: ...
Judge cost: $<total_judge_cost_usd> (if applicable)
Key improvements:
1. <kept experiment 1 hypothesis> (+<delta>)
2. <kept experiment 2 hypothesis> (+<delta>)
...The optimization branch (optimize/<spec-name>) is preserved with all commits from kept experiments.
The experiment log remains in local .spec-first/workflows/spec-optimize/<spec-name>/ scratch space for resume and audit on this machine only; it does not travel with the branch because that run-state path is gitignored. The strategy digest is derived, reconstructable state and may be regenerated from the canonical experiment log.
Present post-completion options via the platform question tool:
Run code review on the cumulative diff (baseline to final). Execute spec-code-review on the optimization branch, interactive or mode:agent. To land eligible fixes before the next option, apply the mechanical-apply bar below.
Mechanical-apply bar: apply any finding with a concrete suggested_fix that is a clear, reversible improvement; push back and keep the diff when the reviewer is wrong, noting why. Defer anything whose right fix needs a design or product decision, including architecture direction, contract shape, behavior change needing sign-off, and any finding with no concrete fix to act on. Confirm evidence still matches at file:line before editing. After applying, run tests, at least targeted tests for what changed and a broader suite for multi-file edits. Do not commit or push from this step; leave the diff on the optimization branch for the Create PR option.
Capture learning by executing spec-compound to document the winning strategy as an institutional learning.
Create PR from the optimization branch to the default branch.
Continue with more experiments: re-enter Phase 3 with the current state. State re-read first.
Done -- leave the optimization branch for manual review.
Clean up scratch space:
# Keep the experiment log for local resume/audit on this machine
# Remove the derived strategy digest only after its required audit fields are in the log
rm -f .spec-first/workflows/spec-optimize/<spec-name>/strategy-digest.mdDo NOT delete the experiment log if the user may resume locally or wants a local audit trail. If they need a durable shared artifact, summarize or export the results into a tracked path before cleanup. Do NOT delete experiment worktrees that are still being referenced.
© leo-kuang-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 31 other files (scripts, references) in skills/spec-optimize of leo-kuang-ai/spec-first.
Open the folder on GitHubat commit 74655dc
Spec Optimize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Spec Optimize this skillleo-kuang-ai/spec-first | 107 | — | ~13k | Automated safety check: Pass | MIT | |
| Clawpathy AutoresearchClawBio/ClawBio | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Agents Best PracticesDenisSergeevitch/agents-best-practices | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Autocontext Knowledge Creatorgreyhaven-ai/autocontext | 1.3k | — | ~964 | Automated safety check: Pass | Apache-2.0 | |
| Axiom Eval Writeropenclaw/clawhub | 9.5k | — | ~4.1k | Automated safety check: Warn | MIT |
ClawBio/ClawBio
Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.
DenisSergeevitch/agents-best-practices
A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
greyhaven-ai/autocontext
Runs the `autoctx` CLI to improve an approach to a task over several generations, score or refine a single output and inspect what a run produced.
openclaw/clawhub
Scaffolds evaluation suites for the Axiom AI SDK: eval files, scorers, flag schemas and axiom.config.ts, generated from plain descriptions of an AI capability.
AgriciDaniel/skill-forge
Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations.
leo-kuang-ai/spec-first
Audit mobile App PRD/Figma/local-source consistency across page routes, KMP/Clean Architecture, components, analytics, i18n, engineering quality, and industry lenses before runtime validation; use…
leo-kuang-ai/spec-first
Create a durable cross-session handoff or resume from a user-selected continuity source.
leo-kuang-ai/spec-first
Give a decisive, project-grounded verdict on an external input — judged against the current project, not in the abstract.
leo-kuang-ai/spec-first
Resolve PR review feedback by evaluating validity and fixing issues with conflict-aware resolver dispatch.
leo-kuang-ai/spec-first
Analyze explicit Riffrec product-feedback captures, including riffrec-.zip, the Riffrec session.json + events.json + recording.webm + voice.webm bundle, or media/notes the user identifies as a…
leo-kuang-ai/spec-first
Document a recently solved problem or durable project vocabulary in docs/solutions/ or CONCEPTS.md.
Categories
Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first. Spec Optimize is an agent skill from leo-kuang-ai/spec-first. Run metric-driven iterative optimization loops.
Spec Optimize fits situations like: optimizing clustering quality; search relevance; build performance; any measurable outcome that benefits from systematic experimentation.
Run `npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a claude-code`. Or copy the skill folder (skills/spec-optimize in leo-kuang-ai/spec-first) into .claude/skills/spec-optimize in your project. Claude Code loads it when a task matches its description.
Run `npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a codex`. Or copy the skill folder (skills/spec-optimize in leo-kuang-ai/spec-first) into .agents/skills/spec-optimize in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add leo-kuang-ai/spec-first --skill spec-optimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-optimize, .gemini/skills/spec-optimize, .github/skills/spec-optimize and .opencode/skills/spec-optimize in your project.
Going by SKILL.md and its folder, Spec Optimize needs a shell and JavaScript for the scripts in its folder and the command-line tools its instructions call (bash, git, codex and node). Our summary lists: Node.js; A Bash shell.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Spec Optimize is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 13k tokens (SKILL.md is roughly 51k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Spec Optimize: Clawpathy Autoresearch (ClawBio/ClawBio, 1.2k stars), Agents Best Practices (DenisSergeevitch/agents-best-practices, 2.4k stars), Looper (ksimback/looper, 710 stars) and Autocontext Knowledge Creator (greyhaven-ai/autocontext, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
leo-kuang-ai (a GitHub user) maintains it in leo-kuang-ai/spec-first, which has 107 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.
Source: leo-kuang-ai/spec-first on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.