PR Babysitter
openinterpreter/openinterpreter
Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.
Execute a settled plan, validated task pack, spec path, or concrete implementation request end-to-end.
$ npx skills add leo-kuang-ai/spec-first --skill spec-work -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-work --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spec-work .claude/skills/spec-work && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "spec-work" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-work into .claude/skills/spec-work/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-work", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-workType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add leo-kuang-ai/spec-first --skill spec-work -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-work --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/spec-work .agents/skills/spec-work && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "spec-work" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-work into .agents/skills/spec-work/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-work", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add leo-kuang-ai/spec-first --skill spec-work -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-work --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/spec-work .cursor/skills/spec-work && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "spec-work" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-work into .cursor/skills/spec-work/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-work", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/leo-kuang-ai/spec-first.git --path skills/spec-work--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add leo-kuang-ai/spec-first --skill spec-work -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-work --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/spec-work .gemini/skills/spec-work && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "spec-work" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-work into .gemini/skills/spec-work/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-work", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install leo-kuang-ai/spec-first spec-workInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add leo-kuang-ai/spec-first --skill spec-work -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/spec-work .github/skills/spec-work && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "spec-work" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-work into .github/skills/spec-work/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-work", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add leo-kuang-ai/spec-first --skill spec-work -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install leo-kuang-ai/spec-first spec-work --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/spec-work .opencode/skills/spec-work && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "spec-work" agent skill from https://github.com/leo-kuang-ai/spec-first/tree/master/skills/spec-work into .opencode/skills/spec-work/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-work", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
spec-workExecute a settled plan, validated task pack, spec path, or concrete implementation request end-to-end.
Spec Work is an agent skill from leo-kuang-ai/spec-first. Execute a settled plan, validated task pack, spec path, or concrete implementation request end-to-end. Use spec-debug for open-ended bugs; stop when target repo, scope, source ownership, or required authorization is unresolved. Not for explicitly requested GitHub PR-review-feedback handling — route those to spec-resolve-pr-feedback.
Its SKILL.md is about 9.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 66 other files, including scripts and reference files (for example `evals/eval.yaml`, `evals/examples.json` and `evals/skillup/cases/drifted-pack-rejects.yaml`).
It sits in Development, covering Pull requests. It works with GitHub. The repository describes itself as: 仓库原生 AI Coding Harness —— 把一次性 AI 对话变成可治理、可验证、可沉淀的工程闭环 · spec-first.cn. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 74655dc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
rgFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Spec Work loads about 9.4k tokens when it runs, and up to ~43k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 4,545 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from leo-kuang-ai/spec-first at commit 74655dc, republished under its MIT licence (© leo-kuang-ai). 4,545 words, ~9,354 tokens.
.claude/skills/spec-work/SKILL.md (or your agent's skills folder). This skill also uses 58 other files; get the full folder from GitHub.Execute work efficiently while maintaining quality and finishing features.
This command takes a work document (plan or specification) or a bare prompt describing the work, and executes it systematically. The focus is on shipping complete features by understanding requirements quickly, following existing patterns, and maintaining quality throughout.
spec-code-review, caller-owned LFG/goal flows, commit/PR/release workflows, spec-compound, and human reviewers.| Reference | Trigger | If unread/unavailable |
|---|---|---|
| Work intake and task pack | Shallow metadata says type: task-pack. | Do not execute the pack; return validation/regeneration handoff. |
| Non-code execution | Metadata says execution: knowledge-work. | Do not enter code/shipping lifecycle; report the missing production route. |
| Execution strategy | Before first write/test/review-fix, task tracking, worker dispatch, commit, or landing. | Lock repo/source/dirty facts inline; use inline/serial; no commit or landing claim. |
| Execution engines | A structured plan/task pack or explicit request makes goal/dynamic/worker engine selection relevant. | Use inline; do not infer a callable non-default engine. |
| Feedback and tests | Before behavior mutation, test design, or verification coverage claim. | Run the narrowest known check and do not claim system-wide coverage. |
| Implementation quality | Before durable-surface mutation or phase-boundary simplification. | Do not add a new durable surface; return to the plan owner if current source fit is unresolved. |
| Shipping workflow | All implementation tasks are accounted for and quality/closeout begins. | No completion/lifecycle/commit/landing claim. |
| Review findings followup | A completed review returned actionable caller-owned findings. | Preserve in-band findings/limitations; do not rerun or silently drop them. |
| Tracker defer | Residual gate explicitly selects external tracker deferral. | Return structured no_sink; do not lose residuals or infer external authority. |
Follows docs/contracts/workflows/scenario-capability-matrix.md.
Overrides: high-risk
foreign-residual-workspace -> blocked-action-required: stop before source writes, behavior-bearing tests that rely on suspect local artifacts, review fixes, commits, lifecycle mutation, or PR-ready claims until the named cleanup/init action runs or the user explicitly accepts degraded evidence.fallback-only: use bounded direct source, test, log, diff, and user-provided evidence; disclose the missing capability and do not claim unconfirmed impact or coverage.non-git-build-workspace coverage gaps -> partial: keep work inside explicit target_repo/covered roots and directly inspect any uncovered build module before changing or claiming behavior there.<input_document> #<invocation arguments supplied by the current host> </input_document>
First, parse a leading mode token. If <input_document> begins with mode:return-to-caller (or the legacy aliases mode:caller-owned-tail / caller:lfg), strip that token before anything else: the remainder of the string is the plan path, and this run executes in Return-to-Caller Mode (see § Return-to-Caller Mode) — implement and locally verify only, then return the structured envelope instead of running the standalone shipping tail. Classify the stripped plan path with the rules below. A mode token with no following path is an error: report it rather than treating mode:return-to-caller as a bare prompt.
Determine how to proceed based on what was provided in <input_document> (after any mode token is stripped).
The classification order is mode token -> file metadata -> task pack -> unified plan -> legacy plan / knowledge-work -> bare prompt. Do not classify by filename alone.
File document (input is a path to an existing plan, specification, or task pack): read only the metadata first — YAML frontmatter for Markdown, or visible header metadata for HTML.
If Markdown metadata carries type: task-pack, do not read the full task-pack body before this classification. Load references/work-intake-and-task-pack.md and follow its deterministic validation, source-plan replay, semantic-fit, Task Pack Contract/Waves, drift, task review, failure handoff, and lifecycle rules. A validated task pack is a first-class execution input, but its source plan remains authoritative and the task pack stays status: derived. This branch replaces the remaining file classification; after successful intake, continue Phase 1 with the resolved source plan plus validated contract.
For a declared unified artifact, validate critical metadata before classification. Duplicate critical metadata, missing artifact_readiness or execution, or a conflict between visible HTML metadata and content shape is invalid. Fail closed to a spec-plan <plan-path> repair handoff; do not normalize, merge, or guess the intended value. This validation applies only after the artifact declares artifact_contract: spec-unified-plan/v1; a truly legacy plan with no unified contract remains on the compatibility path below.
Before readiness classification, inspect lifecycle status. Only status: active is eligible for a new implementation run. completed, partially-shipped, or superseded returns source-plan-non-active and must not enter spec-work, generate execution tasks, or be treated as implementation-ready even when artifact_readiness: implementation-ready remains in historical metadata. A legacy plan with no managed status may continue only with source-plan-lifecycle-unmanaged recorded as a limitation. A discovered mismatch between plan status and source reality is a finding, not authorization. "The plan says completed but the code doesn't have it yet", "the user asked to finish it", or a recorded limitation note never converts a non-active plan into an executable one — return source-plan-non-active with the mismatch as a finding for the plan owner instead of implementing.
artifact_contract: spec-unified-plan/v1, classify artifact_readiness before reading the body.artifact_readiness: requirements-only -> stop and tell the user this Product Contract needs spec-plan enrichment before implementation. Offer the exact spec-plan <plan-path> handoff.artifact_readiness: implementation-ready plus execution: code -> continue to Phase 1 using the unified-plan reader strategy below.execution: knowledge-work to the non-code carve-out; otherwise ask the user to return to spec-plan to produce an implementation-ready code plan.active, in_progress, completed, done) are invalid readiness values. Stop and ask for plan repair rather than guessing.execution: knowledge-work, this is a non-code plan — read references/non-code-execution.md and follow that carve-out instead of the rest of this workflow.execution: code) -> continue to Phase 1 and run the normal code lifecycle. Exception: metadata that carries a progress-like artifact_readiness value (active, in_progress, completed, done) without declaring the unified contract is still an invalid readiness marker, not a legacy plan — stop and ask for plan repair exactly as in the declared-contract branch instead of silently entering the code lifecycle.Blank invocation latest-plan discovery: when <input_document> is blank, glob docs/plans/*.md and docs/plans/*.html, inspect metadata for the newest candidates, and only auto-select an active plan that is artifact_readiness: implementation-ready plus execution: code or an eligible legacy code plan. Never select completed, partially-shipped, or superseded. Stop instead of silently executing when the newest matching artifact is requirements-only, non-active, execution: knowledge-work, an approach-plan, or an unclassified universal/answer-seeking output. Ask for an explicit path or a spec-plan enrichment step. Superseded sibling: if a requirements-only candidate has a same-basename file in the other format (<basename>.md / <basename>.html) that is active and implementation-ready, a format conversion left the requirements-only copy stale — select the active implementation-ready sibling and execute it rather than stopping.
Bare prompt (input is a description of work, not a file path):
Open-ended symptom check. A symptom report with no error text, no stack, no failing test, and no concrete reproducible behavior ("it feels off", "something is wrong somewhere, fix it") is an open-ended diagnosis request, not an implementation prompt: name spec-debug as the owning workflow and route there instead of entering implementation here. Finding and fixing the bug yourself inside this workflow is doing spec-debug's job in the wrong place, even when the diagnosis is quick.
Scan the work area
Assess complexity and route
| Complexity | Signals | Action |
|---|---|---|
| Trivial | 1-2 files, no behavioral change (typo, config, rename) | Proceed to Phase 1 step 2 (execution boundary), then implement directly — no task list, no execution loop. Apply Test Discovery if the change touches behavior-bearing code |
| Small / Medium | Clear scope, under ~10 files | Build a task list from discovery. Proceed to Phase 1 step 2 |
| Large | Cross-cutting, architectural decisions, 10+ files, touches auth/payments/migrations | Inform the user this would benefit from spec-brainstorm or spec-plan to surface edge cases and scope boundaries. Honor their choice. If proceeding, build a task list and continue to Phase 1 step 2 |
Read Plan and Clarify (skip if arriving from Phase 0 with a bare prompt)
source_plan as the plan read below and use only the machine-readable Task Pack Contract/execution_waves for task creation. Follow references/work-intake-and-task-pack.md; do not re-split from the source plan or human-readable cards.Goal Capsule, Verification Contract, Definition of Done, the Implementation Units heading list, and only the active U-ID section plus referenced R/F/AE/KTD excerpts. Read appendices or unrelated U-IDs only when the active unit cites them. To build the map: in markdown scan headings (rg -n '^#{1,3} ' <plan> — top-level sections plus ### U<N>. units); in HTML scan the <h1>–<h3> heading elements and their anchor ids. Match on the stable section names / unit IDs (Goal Capsule, Verification Contract, ### U<N>., …), ignoring HTML wrapper tags — not on a format-specific pattern..md, .html) carry the same section names and IDs; HTML just wraps them in semantic elements (<section>, <article>, etc.).Implementation Units, Work Breakdown, Requirements (or legacy Requirements Trace), Files, Test Scenarios, or Verification, use those as the primary source material for executionExecution note on each implementation unit — these carry the plan's natural-language execution direction for that unit (for example, start from failing proof, characterize legacy behavior, or prefer smoke/runtime verification). Note them when creating tasks, but do not reduce them to keyword matching.Deferred to Implementation or Implementation-Time Unknowns section — these are questions the planner intentionally left for you to resolve during execution. Note them before starting so they inform your approach rather than surprising you mid-taskScope Boundaries section — these are explicit non-goals. Refer back to them if implementation starts pulling you toward adjacent workspec-write-tasks once as an optional path. Never auto-compile it and never block direct execution solely because a task pack would help.Execution notereferences/shipping-workflow.md: after the completion gates close, the tail owner may use the deterministic helper to change a Markdown source plan from active to completed. This marker is not progress or completion evidence. Leaf workers, reviewers, and subagents never mutate plan status. Legacy - [ ] / - [x] marks remain ignored; per-unit completion is determined from current source and verification evidence.Establish Execution Boundary And Strategy
Read references/execution-strategy.md before the first write, behavior-bearing test, review fix, commit, or landing action. It is the owner for branch/worktree, task tracking, worker dispatch, parallel safety, integration, commit, and landing details.
Hard anchors remain here:
target_repo or per-task repo scope. Artifact --repo is not mutation authority.spec-plan/task regeneration.commit_authorization; commit does not imply landing_authorization. Without them, keep verified changes uncommitted and do not push/open a PR.STOP — before the first behavior-bearing mutation, read references/feedback-and-tests.md. It owns smallest feedback loop, vertical slicing, proof/characterization, test discovery, system-wide checks, and not-run replacement evidence. For a trivial non-behavioral edit, use the narrow obvious check and do not load or restate the full reference.
STOP — before adding or materially changing a durable surface, read references/implementation-quality.md. Durable surfaces include dependencies, files, abstractions, helpers/wrappers/adapters, public/schema/runtime/provider/source-of-truth boundaries, workflow handoffs, generators, skills/agents, and artifact contracts. Recheck current source with reuse / extend / compose / new; if the active plan/task did not authorize the needed architecture decision, stop back instead of designing it during implementation. Ordinary bounded edits to an already-owned surface do not emit an architecture matrix or decision note.
Apply the reference and record one run-local boundary: target_repo, current HEAD/branch, pre-existing dirty paths and overlap, canonical source owner, allowed/changed paths, scope-changing discoveries, worker dispatch authorization/capability/isolation, and separate mutation/commit/landing authorization. Branch or worktree mutation requires explicit authority; never pull/switch/create/rename merely because a plan exists. Default-branch commit still requires explicit confirmation.
A necessary discovered file may join the changed set only when direct evidence shows it completes existing scope. Acceptance/public-contract/architecture/provider/repo/source-owner expansion returns to spec-plan or task regeneration. Unknown isolation follows shared-directory rules; missing dispatch authorization/capability runs inline. Workers never commit.
Create Task List (skip if Phase 0 already built one, or if Phase 0 routed as Trivial)
Task Pack Contract.tasks and execution_waves; preserve task_id, dependencies, source refs, declared files, stop_if, and review intent.Choose Execution Engine, then Strategy
Read Execution engines only when plan shape or explicit direction makes a non-default engine relevant. Inline is the portable default. A non-default engine needs explicit authorization and current-session semantic capability, preserves task-pack checkpoints/structured returns, and never changes tail ownership.
Before worker dispatch, inherit the full boundary from references/execution-strategy.md and record worker_dispatch_authorization, capability_probe, worker_dispatch_capability, worker_context_isolation, worker_model_override, and worker_bounded_parallelism, then normalize the path as worker_dispatch_outcome. Missing authorization forbids discovery and fixes capability_probe: not_applicable plus capability unknown. Only after authorization may the current-session registry/schema be consumed as provider_untrusted evidence. Use serial execution for dependencies, overlapping files/contracts/schema/config/lockfiles/generated outputs, shared environment singletons, or unknown bounded parallelism. Stop parallelizing after broad unplanned edits, repeated conflicts, or out-of-scope failures.
Give each worker a bounded unit packet rather than the whole plan: Goal Capsule/DoD, active unit, relevant R/F/AE/KTD and Verification excerpts, files/patterns/scenarios/execution note, plus the triggered feedback/implementation-quality rules. Require changed paths and evidence fields in the return. The orchestrator verifies the actual tree, detects collisions/overwrites, integrates in dependency order, reruns authoritative checks, updates task state, and records one of dispatch_authorization_missing, subagent_capability_missing, or worker_capability_unproven, plus any isolation limitations.
| 红旗念头 | 停下来做什么 |
|---|---|
| 「测试大概会过,先声明完成」 | 跑匹配当前 slice 的真实验证,读取 exit/log,再声明 passed 或记录 not-run reason。 |
| 「计划写了 new wrapper,照着建就行」 | 读 current source,按 reuse / extend / compose / new 重查 owner;无 translation/sequencing/safety/evidence 边界的 wrapper 不创建。 |
| 「相邻代码顺手一起清理」 | 回到 active plan/task 与实际 changed set;非必要 debt 进入既有 residual/defer sink。 |
| 「临时文件或 orphan 留着不影响」 | 清理本次 run 造成的 orphaned source、test、reference、log 或 runtime artifact,并复跑对应 feedback loop。 |
这是注意力提醒,不是 gate,也不替代 LLM 判断;最终是否停下、如何处理仍由你按当前证据决定。
Task Execution Loop
Execute one dependency-ready unit/task at a time, or one bounded disjoint wave when dispatch/isolation facts allow it:
stop_if when applicable; drift or stop conditions halt this task and dependents before mutation.reuse / extend / compose / new, and stop back on unapproved architecture/scope.review_gate: required with bounded spec-code-review mode:agent before dependent waves. Caller-owned fixes rerun affected verification; at most one follow-up review is allowed. Blocking/degraded round two stops; non-blocking P2/P3 remains run-local residual work.commit_authorization: authorized.Execution notes are intent, not enums. Proof-first requires observing the expected failure before production change; characterization records existing behavior without declaring it correct. Trivial rename/config/style/generated/manual-only work may use an explicit replacement check. Never add duplicate tests merely to demonstrate ceremony, over-implement beyond the active slice, or claim coverage for a check that did not run.
Commit Checkpoint
Follow references/execution-strategy.md § Commit Authorization. Local implementation and green tests do not authorize a commit. Workers never commit. Without explicit commit authorization, keep verified changes uncommitted and report coherent commit candidates; with authorization, the orchestrator stages only run-owned files and commits only a verified logical unit.
Follow Existing Patterns
Test Continuously
Simplify as You Go
At a behavior-cluster/dependency-wave boundary, read references/implementation-quality.md § Simplification At Phase Boundaries. Classify findings as remove-now, minimality-debt, protected, or architecture-mismatch; do not default to extract-helper, delete security/data-integrity/a11y/observability/required-verification code for lower LOC, or widen scope to pay unrelated debt.
If spec-simplify-code is available, invoke it at phase boundaries (especially before Phase 3 when the diff is >=30 lines) with the same classification and protected-surface constraints. Otherwise, perform the bounded pass inline. Rerun the same feedback loop for every remove-now or authorized architecture correction.
Figma Design Sync (if applicable)
For UI work with Figma designs:
references/agents/figma-design-sync.md and dispatch a generic subagent seeded with that local prompt to compare implementation against the Figma design. Do not dispatch a standalone agent by type/name.Frontend Design Guidance (if applicable)
For UI tasks without a Figma design -- where the implementation touches view, template, component, layout, or page files, creates user-visible routes, or the plan contains explicit UI/frontend/design language:
Track Progress
stop_if and return to spec-write-tasks/spec-plan, and for direct-plan input return to the plan owner when acceptance, architecture, ownership, or verification scope changesWhen all Phase 2 tasks are complete and execution transitions to quality check, you must read references/shipping-workflow.md for the full shipping workflow. Do not skip this.
Code review: one portable path. Review with spec-code-review, which self-sizes (lite roster for small low-risk code-only diffs, full roster otherwise). No harness-native review detection and no escalation tiers — the size/sensitive-surface judgment lives inside spec-code-review. Skip dedicated review only for a purely mechanical diff (formatting, dep-bumps, lint-only, generated). Full rules (autonomous Residual Gate, infra fallback) in shipping-workflow.md.
Review is two steps — review, then fix. spec-work's mode:agent invocation is report-only: it returns JSON findings and does not edit the checkout, commit, or apply fixes. This statement is scoped to the orchestrated invocation below; other explicit spec-code-review entry modes retain their own contract.
spec-code-review skill (invocation command in references/review-findings-followup.md § Fallback). Use mode:agent in orchestrated workflows; pass plan:<path> when you have a plan, base:<ref> when the merge base is known, and depth:full when a deep/thorough review was explicitly requested.references/review-findings-followup.md. Filter eligibility on JSON only and batch by file. Use authorized fix workers or inline fallback; the orchestrator integrates and tests. Commit only with commit_authorization: authorized.shipping-workflow.md (autonomous sessions auto-accept + record residuals; interactive sessions ask).mode:return-to-caller <plan-path> (legacy alias: mode:caller-owned-tail) is
reserved for orchestrators such as lfg that own simplification, code review,
PR creation, and CI watching after implementation. In this mode spec-work
performs implementation and local verification only, then returns a structured
summary instead of running the standalone shipping tail.
Return:
status: complete, blocked, or failedplan_path: direct plan path, or the validated task pack's authoritative source_plantask_pack_path and pinned task_pack_digest when task-pack intake was used; otherwise nullchanged_filesu_ids_attemptedu_ids_completedtask_ids_attempted and task_ids_completed when Task Cards drove executionverification_resultsverification_evidence: one entry per attempted behavior-bearing unit, plus any non-behavioral unit where tests were intentionally skipped. Each entry states the unit/task, behavior_changed, existing_tests_inspected, tests_added_or_changed, tests used unchanged, red failure or characterization observed when applicable, verification commands/results, and any exception reason. For units executed by subagents, this entry is assembled from each worker's returned evidence (Phase 1 Step 4), not reconstructed from the diff — the red-before-implementation observation exists only in the worker's report.verification_run_summary_ref: repo-relative verification-run-summary.v1 ref produced from the commands this work run actually executed, or null with an explicit limitation when no structured summary could be writtenverified_worktree_fingerprint: the complete spec-work-working-tree-fingerprint/v1 object produced by scripts/working-tree-fingerprint.cjs (resolved from this skill's own SKILL_DIR) after this invocation's final required verification and immediately before return. It covers HEAD, tracked/staged/unstaged diff, untracked paths, and untracked bytes. A behavior-bearing status: complete return requires it; non-behavior returns still include it whenever the helper can run, so callers can apply freshness gates uniformly, and may omit it only together with the documented deliberate non-behavior exception. If the helper cannot run (missing runtime asset, no git, no Node), record a fingerprint-helper-unavailable blocker naming the concrete cause — never fabricate the object or omit it silently.honest_closeout_verdict: verified, degraded, or unsupported, together with the validator overall_reason_coderun_artifact_path: repo-relative spec-work-run-artifact/v2 path when a durable trigger wrote one; otherwise nullrun_artifact_reason_code: the matched durable trigger, no-trigger-matched, or the producer's concrete not-written reasonclaim_limitations: structured limitations for not-run checks, unsupported claim refs, review-evidence materialization failure, provider-bounded evidence, or other claim ceilingsblockersbehavior_change: whether behavior-bearing code changedcommit_authorization: missing and landing_authorization: missing for this mode; Return-to-Caller does not consume either exitplan_status_completion_candidate: the repo-relative direct docs/plans/*.md source plan that the caller may complete after its own shipping gates, or null when lifecycle mutation is not applicableplan_status_completion_degraded_reason: null when a candidate is present; otherwise one of html-plan-lifecycle-degraded, legacy-plan-lifecycle-degraded, read-compatible-status-unmanaged, or source-plan-path-lifecycle-degraded. Duplicate, malformed, or invalid lifecycle metadata is a blocker, not a degraded result.standalone_shipping_skipped: trueReturn status: complete only when every in-scope unit/task is accounted for and completed, task-pack pins still match when applicable, every required task review is closed, blockers is empty, and every required verification result is passed or explicitly not applicable with a reason. Behavior-bearing work also requires the verification evidence and verified_worktree_fingerprint above or a deliberate non-behavior exception. Failed, degraded, not-run, vague, stale, or missing required verification/review cannot return complete.
If a previous return-to-caller run implemented code but omitted evidence, or the caller re-enters after caller-owned simplification/review fixes, the later same-plan invocation must use the idempotency path instead of reimplementing. Re-read the current plan and tree, rerun the complete applicable Verification Contract against the current working tree, create a fresh verification-run-summary ref for commands executed by this invocation, and capture a new verified_worktree_fingerprint only after those checks finish. Never reuse the earlier run summary or fingerprint as final-tree evidence.
standalone_shipping_skipped: true 只表示 caller owns simplify、full review、plan lifecycle 与 landing tail;它不跳过本次 work 已执行命令的 structured closeout。按照 references/shipping-workflow.md 的 Step 5.1 记录 run summary、校验 honest closeout,并按 durable trigger 返回 run artifact path/reason。不得把 plan 中列出的候选命令、worker 的自然语言“tests pass”或 session-temp review path 当成这些字段的 confirmed evidence。
Engine selection (references/execution-engines.md) still applies in this mode,
but only for implementation. In return-to-caller mode do not emit a copyable
goal/workflow prompt — a manual paste step strands the caller; run
inline/authorized workers or return a blocker instead. Any goal/workflow engine used here
must not commit, push, open a PR, run the owner workflow tail, or bypass the caller-owned
gates. Return-to-Caller never invokes plan-status complete; it returns only the
completion candidate, and the caller owns the eventual shipping closeout.
© leo-kuang-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 58 other files (scripts, references) in skills/spec-work of leo-kuang-ai/spec-first.
Open the folder on GitHubat commit 74655dc
Spec Work next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Spec Work this skillleo-kuang-ai/spec-first | 107 | — | ~9.4k | Automated safety check: Pass | MIT | |
| PR Babysitteropeninterpreter/openinterpreter | 69k | 3 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| MAUI PR Performance Analysisdotnet/maui | 23k | — | ~2.4k | Automated safety check: Pass | MIT | |
| Docs AuthoringTracecatHQ/tracecat | 3.8k | — | ~3.1k | Automated safety check: Notes | AGPL-3.0 | |
| Pull Requestnoh-rs/nohrs | 156 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Gitshot Image Uploadertrekawek/coffee-gb | 1.2k | — | ~855 | Automated safety check: Pass | MIT |
openinterpreter/openinterpreter
Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.
dotnet/maui
Interprets pinned managed benchmark evidence for a dotnet/maui pull request and writes a narrative for the performance review workflow, without running or publishing anything.
TracecatHQ/tracecat
A skill your agent uses when adding or updating documentation pages in an existing docs site.
noh-rs/nohrs
Take a change from working tree to a merge-ready pull request, then keep iterating until the CI checks and AI reviewers (CodeRabbit, cubic) all pass.
trekawek/coffee-gb
Uploads screenshots and other images from the terminal and returns markdown-ready URLs to embed in GitHub pull requests, issues and comments.
roryeckel/wyoming_openai
GitHub issues, pull requests, bug reports, scope questions, and support threads.
leo-kuang-ai/spec-first
Audit mobile App PRD/Figma/local-source consistency across page routes, KMP/Clean Architecture, components, analytics, i18n, engineering quality, and industry lenses before runtime validation; use…
leo-kuang-ai/spec-first
Create a durable cross-session handoff or resume from a user-selected continuity source.
leo-kuang-ai/spec-first
Give a decisive, project-grounded verdict on an external input — judged against the current project, not in the abstract.
leo-kuang-ai/spec-first
Resolve PR review feedback by evaluating validity and fixing issues with conflict-aware resolver dispatch.
leo-kuang-ai/spec-first
Analyze explicit Riffrec product-feedback captures, including riffrec-.zip, the Riffrec session.json + events.json + recording.webm + voice.webm bundle, or media/notes the user identifies as a…
leo-kuang-ai/spec-first
Document a recently solved problem or durable project vocabulary in docs/solutions/ or CONCEPTS.md.
Works with
Categories
Execute a settled plan, validated task pack, spec path, or concrete implementation request end-to-end. Spec Work is an agent skill from leo-kuang-ai/spec-first. Execute a settled plan, validated task pack, spec path, or concrete implementation request end-to-end.
Spec Work fits situations like: tasks that involve Pull requests.
Run `npx skills add leo-kuang-ai/spec-first --skill spec-work -a claude-code`. Or copy the skill folder (skills/spec-work in leo-kuang-ai/spec-first) into .claude/skills/spec-work in your project. Claude Code loads it when a task matches its description.
Run `npx skills add leo-kuang-ai/spec-first --skill spec-work -a codex`. Or copy the skill folder (skills/spec-work in leo-kuang-ai/spec-first) into .agents/skills/spec-work in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add leo-kuang-ai/spec-first --skill spec-work -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-work, .gemini/skills/spec-work, .github/skills/spec-work and .opencode/skills/spec-work in your project.
Going by SKILL.md and its folder, Spec Work needs the command-line tools its instructions call (rg).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Spec Work is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.4k tokens (SKILL.md is roughly 37k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 34k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Spec Work: PR Babysitter (openinterpreter/openinterpreter, 69k stars), MAUI PR Performance Analysis (dotnet/maui, 23k stars), Docs Authoring (TracecatHQ/tracecat, 3.8k stars) and Pull Request (noh-rs/nohrs, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
leo-kuang-ai (a GitHub user) maintains it in leo-kuang-ai/spec-first, which has 107 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.
Source: leo-kuang-ai/spec-first on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.