Agent skill

Implement

by jpicklyk in jpicklyk/task-orchestrator

End-to-end workflow for taking MCP work items from backlog to merged PR.

MITAuto-check passedAgent Workflows

Install Implement

skills CLI
$ npx skills add jpicklyk/task-orchestrator --skill implement -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jpicklyk/task-orchestrator implement --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/implement .claude/skills/implement && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
implement
GitHub stars
207
Token cost
~17k tokens
SKILL.md length
8,156 words
Files
5 (incl. references)
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

End-to-end workflow for taking MCP work items from backlog to merged PR.

  • Works in 6 steps: Assess the Work → Prepare the Branch and Worktree → Queue Phase: Planning → …
  • A user says implement this
  • SKILL.md covers Step 1 — Assess the Work, Step 2 — Prepare the Branch…, Step 3 — Queue Phase: Planning and Step 4 — Work Phase:…, plus 5 more sections
  • Runs Python scripts from its folder; calls git and gh

What it does

Implement is an agent skill from jpicklyk/task-orchestrator. End-to-end workflow for taking MCP work items from backlog to merged PR. Handles git branching, schema-driven planning, implementation, independent review, and PR creation. Composes spec-quality, review-quality, and schema-workflow skills into a single pipeline. Use when a user says "implement this", "work on this item", "fix these bugs", "pick up the next task", "create a PR for this", "go through the backlog", or references specific MCP item IDs for implementation.

Its SKILL.md is about 17k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `WORKTREE.md`, `references/dispatch-contract-template.md` and `references/patch-anchored.py`).

It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol and Git. The repository describes itself as: Server-enforced workflow discipline for AI agents. An MCP server providing persistent work items, dependency graphs, quality gates, and actor attribution. Schemas define what… The licence is MIT.

When your agent uses it

  • A user says implement this
  • Work on this item
  • Pick up the next task
  • Create a PR for this

Example prompts

  • “implement this”
  • “work on this item”
  • “fix these bugs”
  • “/implement”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Assess the Work
  2. Prepare the Branch and Worktree
  3. Queue Phase: Planning
  4. Work Phase: Implementation
  5. Review Phase
  6. Finalize and PR

What it can do on your machine

Read from SKILL.md and the folder at commit b688ea0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Implement loads about 17k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 8,156 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~17k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jpicklyk/task-orchestrator at commit b688ea0, republished under its MIT licence (© jpicklyk). 8,156 words, ~16,749 tokens.

Download SKILL.mdSave it as .claude/skills/implement/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
implement
description
End-to-end workflow for taking MCP work items from backlog to merged PR. Handles git branching, schema-driven planning, implementation, independent review, and PR creation. Composes spec-quality, review-quality, and schema-workflow skills into a single pipeline. Use when a user says "implement this", "work on this item", "fix these bugs", "pick up the next task", "create a PR for this", "go through the backlog", or references specific MCP item IDs for implementation.
user-invocable
true

Implement

End-to-end workflow for taking MCP work items from queue to PR. This skill composes the schema-driven planning (spec-quality), implementation, review (review-quality), and git/PR workflow into a single pipeline.

Usage:

  • /implement <item-id> — work on a specific item
  • /implement — with context about what to work on
  • Can process single items or multiple items in batch

Step 1 — Assess the Work

Load the item(s) and determine the execution tier and interaction mode.

For each item, call get_context(itemId=...) to understand:

  • Current role and gate status
  • Schema tag (feature-implementation, bug-fix, etc.)
  • Existing notes already filled
  • Dependencies and blocked status

When the user or an item references a PR or commit as already landed, verify its merge state before relying on it — gh pr view <n> --json state,mergeCommit plus git log origin/main — and confirm the working tree is current (Step 2, base-freshness precondition) before any pass that checks whether a change is present in the codebase.

Execution tier — classify by this table (canonical source shared with the task-orchestrator:orchestrate skill; edit the fragment, not this copy):

<!-- BEGIN GENERATED:tier-classification | source: claude-plugins/task-orchestrator/_fragments/tier-classification.md · regen: node claude-plugins/task-orchestrator/_fragments/generate.mjs -->
CriteriaTierPipeline
1-2 files, known fix, no migration/new APIDirectOrchestrator edits, tests, reviews inline
3-10 files, single logical unit, clear or explorable scopeDelegatedSingle subagent, separate review agent
11+ files, multiple independent work streams, dependency edgesParallelWorktree agents, full pipeline

Force-UP signals (bump tier regardless of file count):

  • Database migration → min Delegated
  • New public API surface → min Delegated
  • Multiple independent work streams → Parallel
  • User says "let's plan" / collaborative language → min Delegated

Force-DOWN signals:

  • User says "just fix it" / "quick" → Direct (unless complexity contradicts)
  • Schema tag is default or absent → eligible for Direct
<!-- END GENERATED:tier-classification -->

Delegated with ≥2 ready items, or Parallel → run /task-orchestrator:run-wave.

If the item has no schema tag, apply quick-fix for Direct tier or leave untagged for Delegated/Parallel (the default schema catches these).

Trait application on classification. When the tier resolves to Delegated or Parallel and the workspace defines a delegated trait (it appears in availableTraits on create responses), apply it before any dispatch — traits: "delegated" at item creation, or manage_items(operation="update", items=[{itemId: "<uuid>", traits: "delegated"}]) for an existing item. This makes the orchestrator-filled delegation-metadata note schema-visible instead of convention-only. Direct tier: do not apply it — nothing is delegated.

Model-selection traits. Seat dispatch defaults come from the delegated trait (planner opus, implementer sonnet, test-author sonnet, reviewer opus). Override them per item, at the same point:

  • complex-implementation — implementer on opus. Apply for architecture-heavy or multi-file synthesis work: new public API, cross-layer seams, subtle state machines or concurrency.
  • high-stakes — planner and reviewer on fable. Apply sparingly, where a wrong plan or a missed review finding is expensive: security predicates, auth, data-loss paths, hard-to-reverse design decisions. Fable draws on its own usage bucket, so it is never a default. needs-security-review does NOT imply it.

Read the resolved per-seat profile from get_context / query_items(operation="schema") (dispatchBySeat) and pass its model explicitly on the dispatch.

Test-author trigger rule. bug-fix.default_traits already includes needs-test-author — no action needed, it applies automatically. For feature-task items, apply needs-test-author per item (manage_items(operation="update", items=[{itemId: "<uuid>", traits: "needs-test-author"}])) when acceptance criteria involve a predicate, algorithm, parser, validator, or state transition, or when any force-ON signal is present: new public API surface, a database migration, a security predicate, or a prior vacuous-test finding in this item's area. Other schema tags (feature-implementation, plugin-change, quick-fix, and the global floor) do not carry the trait. Direct tier: apply the trait only in its temporal-only degraded mode, and only on bug-fixes — the test-plan gate and red-first rule still apply, test-manifest declares a single actor, and there is no separate test-author dispatch. Other Direct-tier items are exempt regardless of the criteria above — the tier is too small to separate authorship into a second dispatch.

Interaction mode — orthogonal to tier:

SignalMode
User says "work with me on", "let's plan", or similar collaborative languageCollaborative — user participates in planning and key decisions
Scope is clear, no user participation neededAutonomous — agent handles the pipeline
Unclear scope or ambiguous complexityAsk the user

When processing multiple items, evaluate whether related items (e.g., bugs in the same module, fixes that touch the same files) should be grouped into a single branch and PR. Group when the changes are cohesive and independent fixes would create merge conflicts. Keep items separate when they're unrelated or when isolation makes review cleaner.


Step 2 — Prepare the Branch and Worktree

Sync local main before any implementation begins.

bash
git checkout main
git pull origin main --tags

Base-freshness precondition. Before creating any branch or worktree, and before any source-verification pass (a grep or file read that decides whether a change is present), confirm the intended base is current:

bash
git fetch origin
git log --oneline -1 origin/main
git rev-parse --short HEAD

The two SHAs must match, or HEAD must be a descendant of origin/main on the intended branch. A source-verification pass must never run against a working tree that has not been confirmed current — this includes the orchestrator's own session worktree. If the tree is behind, read the file directly from the remote (git show origin/main:<path>, from PowerShell on Windows — the Bash tool's MSYS layer mangles the ref:path colon) or create a fresh worktree from origin/main (fallback block below). A stale base manufactures absent-looking evidence and absence reads as a discovery, not an error: a 6-commit-stale tree reported four already-adopted proposals as missing (2026-08-04), and a one-commit-stale dispatch re-implemented existing runtime into a conflicting 17-file commit (2026-05-01) — trend 760be80d.

The branching/worktree strategy depends on tier:

Direct tier (orchestrator implements 1–2 files inline) — create a working branch on the main directory:

bash
git checkout -b <branch-name>

Delegated tier (single subagent) — same as Direct: orchestrator creates the branch on the main directory, the subagent works against it. No worktree.

If the main checkout is unavailable (another branch checked out, uncommitted changes present) or the orchestrator itself runs from a worktree — Direct and Delegated tiers both: do not touch the occupied checkout. Create a dedicated worktree instead:

bash
git worktree add .claude/worktrees/<slug> -b <branch-name> origin/main

Work there using absolute paths and git -C (never cd); after the PR merges, remove the worktree and delete the branch. Sync local main via git fetch origin main:main while main is not checked out anywhere, or a normal git pull from the main checkout when it is.

Parallel tier (parent feature with multiple children) — create a single feature worktree that all child agents share:

bash
FEATURE_SLUG=<short-feature-description>          # e.g. issue-117-followup
FEATURE_BRANCH=feat/$FEATURE_SLUG
FEATURE_WORKTREE=.claude/worktrees/feat-$FEATURE_SLUG

# Resume detection — if the branch/worktree already exist (orchestrator restart
# mid-feature), reuse them rather than recreating:
if git show-ref --verify --quiet "refs/heads/$FEATURE_BRANCH"; then
  echo "Resuming existing feature branch $FEATURE_BRANCH"
else
  git branch "$FEATURE_BRANCH" main
fi
if [ ! -d "$FEATURE_WORKTREE" ]; then
  git worktree add "$FEATURE_WORKTREE" "$FEATURE_BRANCH"
fi

All child-task agents will be dispatched into this shared worktree (Step 4). The feature branch is pushed and PR'd once, when the parent feature reaches terminal (Step 6).

Plan file as dispatch contract (3+ children). For every Parallel-tier run, instantiate references/dispatch-contract-template.md into the wave's plan file before the first dispatch, fill every placeholder, and have every Step 3, 4, 4b and 5 dispatch prompt reference it under the template's conflict rule instead of restating design or process details per prompt. The template states the rules; this table only says what each slot is for — do not paraphrase a slot's rule into a prompt. Under /task-orchestrator:run-wave, the Header and File ownership rows are filled from the planner's explain output plus the planner-v1 stage's own returns rather than hand-typed, and the run reuses this same worktree by passing --mode shared --worktree $FEATURE_WORKTREE --branch $FEATURE_BRANCH — the contract file and its slot order are unchanged either way.

Write it to <main-checkout>\plans\<slug>.md and hand agents that absolute path, recorded once in the contract's Header Contract path line. The plans/ directory is gitignored, so it exists only in the main checkout and never in the feature worktree; a prompt that pins a worktree working directory and then names a relative plans/<slug>.md points at a path that is not there.

SlotWhy it exists
HeaderOne place agents read branch, worktree path, the write root every agent write must fall under, the contract's own absolute path, base SHA, PR boundary, scratchpad and the gradle lock helper, so no prompt repeats them (write root: proposal d1484a3e).
Conflict ruleMakes the plan file authoritative over prompt text, and names the per-item note that outranks the plan for its own dimension.
ItemsFreezes each item's scope anchor so an agent cannot re-litigate scope mid-wave.
Planning seat return templateFixes the fields each planning seat returns, so streams compose without re-reading every plan.
Commit disciplineThe shared index makes a bare commit unsafe: pathspec staging and commit, post-commit verification, and the recovery step (proposal 568e7f7e / #301).
Compile self-checkOne pinned invocation through the lock helper, how its exit code is read, the foreign-file early exit, the owned-file retry bound (proposal b82537e4 / #302), and the single statement of who runs gradle in the wave.
File ownershipEach agent's entire writable scope — including any declared edit to an existing test file — with cross-stream overlaps declared up front rather than discovered mid-wave.
Test author protocolThe blindness, oracle and scope rules that keep test authorship independent of the implementation.
Contract-change sweepAfter a contract tightens, call sites and fixtures the changed item never names break; the sweep makes finding and construction-only-repairing them a step rather than a reminder (proposals 82034e9a, 31a1abeb).
DocsRoutes doc edits — and the CHANGELOG.md bullet (proposal c068c943) — to one serialized seat so parallel agents never write the same doc.
NotesWho fills which note key, and the maxLength each must respect.
Review scopingReviews diff owned files, not SHA ranges; the commit map serves only the ownership check (#301).

Two runs (retros 202b5d42, 6d562acb) produced zero vocabulary deviations across up to five concurrent authors this way, with prompts roughly 60% smaller. The delivery surface is itself the point (proposal 728a3e57 / #307): in a same-day controlled comparison, the wave with no contract file recurred every previously-adopted failure class, while the waves that ran from one held the gains — a rule that lives only in skill prose does not reach the agents, because agents read the dispatch prompt and their item's schema guidance. The contract guarantees deviation-free execution, not plan correctness — pair it with Step 3's "plans are spec inputs" rule so the plan's own factual claims are verified by the implementer rather than merely followed.

Why one worktree per feature, not per child: the feature is the natural PR boundary. Per-child PRs created cross-PR test contamination and PR-body staleness during the #117 follow-up (see retro a7f6024f). Shared worktree means one commit history, one CI cycle, one PR — and the parent feature's review-checklist gives a coherent point at which to finalize.

Branch naming:

  • feat/<feature-slug> — feature-implementation parents (the integration branch)
  • fix/<short-description> — bug-fix items (Direct or Delegated tier)
  • fix/<grouped-description> — batch of related bug fixes (Delegated tier)
  • chore/<short-description> — tech debt, refactoring (Direct tier)

Step 3 — Queue Phase: Planning

This step is tier-conditional:

Direct tier: Skip this step entirely. No plan mode. No queue-phase notes. Call advance_item(trigger="start") immediately to move queue→work. The quick-fix schema has no queue-phase required notes, so the gate passes. Exception: if the item carries needs-test-author (temporal-only degraded mode, Step 1), test-plan is a queue-phase note and must be filled before this advance — see the seat-timing rule below.

Delegated tier: Fill queue-phase notes per schema. Use get_context(itemId=...) to see expectedNotes and guidancePointer. Pre-plan-workflow is optional — use only if scope needs exploration. Post-plan-workflow only if child items need materialization. Advance: advance_item(trigger="start").

Parallel tier: Full planning pipeline:

Collaborative mode:

  1. Tell the user the item is ready for planning and ask them to enter plan mode. The pre-plan-workflow and post-plan-workflow hooks fire automatically on plan mode entry and exit, handling context gathering and materialization.
  2. During planning, follow the guidancePointer for each required note — this will reference the spec-quality skill where applicable.
  3. After plan approval and post-plan materialization, advance the item: advance_item(trigger="start")

Autonomous mode:

  1. Read and follow the pre-plan-workflow skill — gather existing MCP state, check schema requirements, and understand the definition floor.
  2. Research the codebase — explore relevant files, understand current state.
  3. Fill all queue-phase notes following the guidancePointer for each. The spec-quality framework applies regardless of mode.
  4. Read and follow the post-plan-workflow skill — materialize child items if the plan calls for them.
  5. Advance the item: advance_item(trigger="start")

The gate will reject advancement if required notes are missing. If rejected, fill the missing notes and retry.

Plans are spec inputs. Every file path, tool contract, or call sequence a plan names must cite its authoritative location and be verified on disk or in source before the plan is approved — child items' specification notes are generated from the plan and inherit its errors silently (spec-quality, "Cite Contracts, Don't Restate Them"). This reminder lives here because spec-quality is reached via note guidance pointers that plan mode does not traverse.

Do not confuse this with resource-lease contention. A queue→work advance_item can also fail with applied: false, errorCode: "resource_unavailable", errorKind: "transient" — a resource a trait on this item declares (resources:) is currently held by another item. This is not a note gate failure: filling notes will not fix it. Work a different item and retry later (retryAfterMs is a hint), or report the contended key(s) to the user.

Seat timing for needs-test-author items. The queue-phase test-plan note is filled by the planning seat — the orchestrator (Direct tier) or the plan author (Delegated/Parallel tier), invoking the test-author skill's scenario-derivation and oracle-derivation sections — and this happens before advance_item(trigger="start") moves the item queue→work. This is by design: test-plan gates work entry, the same way any other required queue-phase note does. This is a distinct seat from Step 4b's test author seat, which writes test code and fills test-manifest at work phase — see Step 4b's intro for the two-seat distinction.

Planning seat (Parallel tier, one opus agent per stream). Under /task-orchestrator:run-wave this becomes the run plan's planner-v1 stage — it mirrors the same eight fields below, dispatched by the run rather than by hand; the field table stays the semantics either way. Dispatch one opus agent per stream after child items are materialized (post-plan-workflow) and before Step 4's implementer dispatch — this is the concrete, per-stream instantiation of the planning seat named above. Inputs: the stream's diagnosis/task-scope note and the CURRENT codebase, not the plan text alone. Task: verify every claim in that note against source, citing file:line for each correction; write the note if materialization left it missing, or revise it if verification finds it wrong; fill and freeze test-plan per the seat-timing rule above; then return the structured template below instead of free-form prose. Give the seat the contract by its absolute path (Step 2's Contract path line) — never a relative plans/..., which does not resolve from the worktree — so it reads the Items and Conflict rule slots it is about to extend; its return is what fills the File ownership and Planning seat return template slots for the rest of the wave.

Structured return (eight fields, each required — none is a valid value):

FieldSemantics
diagnosis-correctionsWhat the diagnosis/task-scope got wrong or missed, with current file:line — or none.
defect-class-siblingsEvery other site sharing the root-cause pattern (same cast/fallback idiom, same predicate, same wildcard or character class, same error-classification path, other readers or writers of the field), found by a named sweep (command + scope + hit count). Mark each one frozen as D# or deferred: <reason>, or return none (sweep: <command>) (proposal bb191508).
cross-stream-file-overlapsFiles this stream must write that another stream in the same wave also touches — drives wave sequencing (serialize the overlap, parallelize the rest) — or none.
missing-api-or-seamA surface the fix needs that does not exist yet, with the exact proposed NEW signature — or none.
test-plan-statusWhether test-plan is filled and frozen (filled (<n> chars)) or still open, and why.
main-filesThe stream's production files (src/main), comma-separated.
test-filesThe stream's NEW test files (author-owned), comma-separated.
red-proof-shapePer scenario in test-plan: EXISTING-SURFACE (the test targets code that exists before the fix — revert-the-fix red-proof applies directly) or NEW-SURFACE (the test targets a surface the fix itself introduces). Every NEW-SURFACE scenario needs the narrowest-revert recipe — typically revert only the call sites and keep the new type/parameter — or, when no revert can produce a behavioural red, an explicit substitute: no behavioural red possible, reviewer verifies <X>.

The fields are the budget. The return may exceed any line budget when the fields have content — trimming a field to fit a target length is the failure this stage exists to prevent, not a virtue.

Evidence. This seat produced a material correction on 15/15 items across three bug waves (retros 1b8bba5a, cc69671b, 8e6a7cf7) — every wave that ran it found at least one diagnosis error, cross-stream overlap, or missing seam the plan had missed.


Step 4 — Work Phase: Implementation

Verification commands. Throughout this step, "run tests" means running BOTH the test suite AND the project linter:

bash
./gradlew :current:test
./gradlew :current:ktlintCheck

CI enforces both — a green test run with failing lint will still block the PR. If ktlintCheck fails, run ./gradlew :current:ktlintFormat to auto-fix formatting violations, then verify with ktlintCheck again and re-run tests. Include both commands in every implementation-agent and review-agent prompt.

Who runs the lint cycle (proposal ee6f5d32, ~9 sessions of evidence):

  • Delegated tier (agent owns gradle): the implementation agent runs the ktlintCheck → ktlintFormat → re-verify cycle itself before committing. Validated 2026-07-13 (PRs #213/#214/#215): zero orchestrator fix-up commits.
  • Parallel tier: per-seat gradle ownership is stated once, in the dispatch contract's Compile self-check slot ("Who runs gradle"); this bullet only adds the lint-cycle detail that slot's table implies — agents get :current:ktlintFormat as part of their single pinned self-check, and the orchestrator additionally runs ktlintFormat before ktlintCheck after each commit batch (historically 3-5 fix-up commits per multi-phase run when skipped).

Capturing gradle's real exit code (use this pattern, not 2>&1 | tail -N). Piping gradle into tail discards gradle's exit code — tail always exits 0 on a successful read of the log, so BUILD FAILED at the end of the gradle output is reported by you as a successful run. Combined with gradle's daemon incremental cache, this can hide broken compilation through entire CI cycles (retro 568a8584: 9 silently-failing tests shipped on PR #151 before the followup audit caught it).

The reliable pattern, especially when running in run_in_background:

bash
./gradlew :current:test > /tmp/gradle-out.log 2>&1; EXIT=$?
echo "EXIT=$EXIT"
tail -30 /tmp/gradle-out.log

Redirect ordering matters: use > /tmp/log 2>&1, NOT 2>&1 > /tmp/log — the latter is evaluated left-to-right and leaks stderr (where gradle writes compile errors) to the terminal instead of the log (retro ac25db89: a compileTestKotlin failure was invisible in the captured log until the ordering was fixed).

EXIT=$? captures gradle's actual exit code before any pipe consumes it. Read the captured exit code AND the tail of the log; never trust the tail alone. This applies to every orchestrator-owned gradle invocation throughout this step.

Agent-side gradle is not this pattern. In a Parallel-tier wave, agents get exactly one lock-serialized invocation, pinned verbatim — argument form included — in the dispatch contract's Compile self-check slot, together with the rule for what to do when it fails. Do not hand an agent an ad-hoc gradle command line; point it at that slot.

Use --rerun-tasks after dependency upgrades or large refactors. Gradle's incremental compile cache retains class files from prior good builds. After a gradle/libs.versions.toml bump or a refactor that changes public API surfaces (removed methods, renamed types, sealed-class arms, generic-parameter shifts), incremental compilation can keep the OLD class files alongside source that no longer compiles, producing apparent BUILD SUCCESSFUL on stale bytecode. Run ./gradlew :current:test --rerun-tasks once after such changes to force a clean run; ordinary incremental builds are safe afterwards.

This step is tier-conditional:

Direct tier: Implement directly. Edit the files, run the test suite. No subagent dispatch. No /simplify pass. Exception — a Direct-tier bug-fix carrying needs-test-author: follow Step 4b's Direct-tier test-first-then-fix sequence (test-plan frozen at queue → regression test observed red → fix → green) instead of implementing first. Fill the session-tracking note (required by both quick-fix and default schemas) with a brief summary of what changed and test results. Advance to review: advance_item(trigger="start").

Delegated and Parallel tiers: Use get_context(itemId=...) to see work-phase expectedNotes and guidancePointer values. Fill each required note following its guidance. Follow /task-orchestrator:orchestrate (model table, return formats, UUID inclusion). The key decisions at this step are:

  • Single item (Delegated): delegate to one implementation subagent or implement directly. Subagent works in the main directory on the working branch.
  • Multiple child tasks, independent (Parallel): dispatch parallel subagents into the shared feature worktree created in Step 2. Each agent receives the worktree path and branch name. Do not use isolation: "worktree" on the Agent tool — that would spawn a separate worktree per dispatch, which is the deprecated per-child PR pattern.
  • Multiple child tasks, dependent: dispatch sequentially into the shared feature worktree. Wait for each agent's commit to land before dispatching the next.

Under /task-orchestrator:run-wave, both bullets above are executed by the run plan (Method A via Workflow, Method B via next-scheduling), reading each stage's dispatchBySeat entry for model/agent instead of the table below. The hand-dispatch form in this section and the template it points at remain the fallback for dispatches a run plan doesn't cover — a one-off fix agent, or an arbitration re-dispatch (Step 4b) — not the default path for a Parallel-tier wave.

Test-file ownership boundary. Every implementation dispatch (Delegated single agent or Parallel per-child agent) excludes src/test/** from scope — implementers do not create or modify test files. When a change surfaces a needed test update, the agent reports it in its return (or in implementation-notes) rather than editing the test itself; the test author (Step 4b) owns that file tree exclusively on items carrying needs-test-author.

Parallel dispatch into a shared feature worktree:

Agent(
  prompt="""
  Working directory: <feature-worktree-path>
  Branch (already checked out): feat/<feature-slug>
  Dispatch contract: <ABSOLUTE path from the contract's Header "Contract path" line, e.g.
  D:\Projects\task-orchestrator\plans\<slug>.md> — read it first; it wins on conflict with
  this prompt. It lives in the main checkout, not in the worktree above.
  Scope (modify ONLY these files): <explicit list> — the same list as your row in the
  contract's File ownership slot.
  Do NOT create or modify any file under src/test/** — test authoring is a separate,
  independent dispatch (Step 4b) on items carrying needs-test-author. If your change
  surfaces a needed test update, report it in your return; never edit the test yourself.

  Before returning: run the compile self-check exactly as the contract's "Compile self-check"
  slot pins it, and handle its outcome by that slot's rules. Then commit exactly as the
  "Commit discipline" slot states, including the post-commit verification.
  Do NOT run :current:test or :current:ktlintCheck — the orchestrator owns full build
  verification.
  """,
  model="sonnet",  // or dispatch.model when the item's resolved profile sets one — see "Model selection" above
  subagent_type="<dispatch.agent when the item's resolved profile names one, else general-purpose>"
  // NOTE: no isolation parameter — agents share the feature worktree
)

Why the prompt points at slots instead of carrying the commands: the self-check is fast (~3s) and catches type-mismatch / signature errors that gradle's incremental cache may otherwise mask in the orchestrator's later test run (retro 568a8584: H3's dbNow() shipped with a Result<T> vs Instant return type mismatch hidden for ~6 hours), and its ktlintFormat step is formatting-only and idempotent, preventing the recurring lint fix-up commits (proposal ee6f5d32). But the invocation's exact argument form, the foreign-file early exit and the retry bound only work if every agent gets them identically — six agents in one wave each burned three lock-serialized runs against another agent's mid-refactor file, and a re-typed nested invocation flattened a [string[]] argument (#302). A prompt that restates the command drifts from the contract; a prompt that points at the slot cannot. The same applies to the commit step: never tell an agent to "commit your changes with a descriptive message" in a shared worktree — the Commit discipline slot is the only safe form (#301).

File-edit overlap discipline: Parallel agents in a shared worktree must operate on non-overlapping files. The orchestrator scopes each agent's prompt to a specific file list (per MEMORY.md §"Parallel File-Edit Delegation"). When inherent overlap exists, dispatch sequentially.

Contract-change sweep discipline. When a child task tightens a contract — making a parameter required, adding validate() invariants, narrowing a sealed-class arm, or otherwise rejecting inputs that earlier passed — the orchestrator must sweep the rest of the codebase before advancing to review. Two recurrences (retros a7f6024f and 568a8584) showed that:

  • Pre-existing test fixtures constructed under the old contract will fail under the new one. Example: H2's WorkItem.validate() claim-field invariants broke 8+ test fixtures that constructed mixed-state items via separate Instant.now() calls (microsecond drift) or partial claim fields.
  • The failure typically surfaces on a different PR's merge commit, not the PR that introduced the contract change — the original PR's tests passed because they used the new contract correctly.

For each contract-tightening change in this run:

  1. Identify the affected tool / class / method.
  2. Grep all test files (and other call sites) for usages: grep -rn "<tool>\|<class>\|<method>".
  3. Verify every usage is consistent with the new contract. Update any that are not.
  4. Re-run the full :current:test suite (orchestrator-owned, not the agent) to confirm no fixture-vs-contract conflicts surfaced elsewhere.
  5. Fixture repairs are orchestrator-owned and construction-only. When a fixture fails under the tightened contract, the orchestrator may adjust how the fixture is constructed (fix the stale call site) but must never adjust what it asserts. Anything that would touch an expectation instead of a construction call is not a fixture repair — re-dispatch the test author (this preserves the independence the trait exists to protect). Declare these commits in the review handoff: list each orchestrator fixture-repair commit's SHA alongside the impl-range/author-range pair (see Step 4's tracking table and Step 4b's disjointness check) so review-quality can tell a declared fixture repair apart from a silent implementer edit to src/test/**.
  6. Doc-claims sweep after behavior-changing fixes: when an orchestrator-owned bug-fix changes shipped behavior after documentation was authored (e.g. a tokenizer or default flips mid-run), grep all changed docs for claims about the OLD behavior before finalizing (retro ac25db89: "case-insensitive" survived in 3 places after the fix made search case-sensitive).

This sweep is part of the orchestrator's verification step between waves, not the implementing agent's responsibility — agents are file-scoped and can't see the full fixture surface.

Model selection — always set model explicitly on every Agent dispatch. Under /task-orchestrator:run-wave, dispatchBySeat already resolves model per seat and this table is the fallback for dispatches the run plan doesn't cover:

Agent purposeModel
Implementation (production code only)model="sonnet"
Independent test authoring (Step 4b)model="sonnet"
Architecture, complex multi-file synthesismodel="opus"
MCP bulk ops, materializationmodel="haiku"

Omitting model causes the agent to inherit the orchestrator's model (typically opus), wasting tokens on sonnet-eligible implementation work.

When dispatching an item's phase owner (implementer on work, reviewer on review), read the profile for the phase being dispatched INTO: if the orchestrator already called advance_item, its dispatch field already reports the profile for newRole; if the agent will enter its own phase (agent-owned-phase protocol — dispatched while the item is still in queue), get_context returns the queue profile, not work's, so read query_items(operation="schema", itemId=...)'s per-phase dispatch.work map instead. When that profile names an agent, dispatch with subagent_type = dispatch.agent. Regardless of whether agent is set, still pass model explicitly — dispatch.model if set, else the table above. effort has no Agent-tool parameter; in Claude Code it applies only through the dispatched agent's own frontmatter, so a profile's effort is advisory unless agent names a definition carrying that effort — to change effort, point agent at a definition with that effort.

After implementation agents return:

For Parallel-tier features, agents return having committed to feat/<feature-slug> inside the shared feature worktree. Record each agent's commit SHA range alongside the child's MCP item ID — this map feeds the ownership two-range check below and the dispatch contract's Review scoping slot, where reviews are diffed by owned file rather than by SHA range:

| Child UUID | Agent ID | Pre-SHA | Post-SHA | Test-Pre-SHA | Test-Post-SHA | Changed Files |
|------------|----------|---------|----------|--------------|---------------|---------------|
| <uuid>     | <id>     | <sha>   | <sha>    | <sha>        | <sha>         | <file list>   |

Capture pre-commit SHA before dispatch (git -C <feature-worktree> rev-parse HEAD) and post-commit SHA after the agent returns. In a shared worktree this range can interleave other streams' commits and misses later fix-ups, so use it to confirm ownership, not to bound a review:

bash
git -C <feature-worktree> diff <pre-sha>..<post-sha> --name-only

Test-Pre-SHA/Test-Post-SHA are captured the same way around the test author's dispatch (Step 4b) and stay blank until that wave runs.

Disjointness check (required whenever needs-test-author applies). After both ranges are captured, verify:

bash
git -C <feature-worktree> diff <pre-sha>..<post-sha> --name-only          # impl range
git -C <feature-worktree> diff <test-pre-sha>..<test-post-sha> --name-only # author range
  • impl-range ∩ src/test/** = ∅ (the implementer touched no test files)
  • author-range ∩ src/main/** = ∅ and author-range is non-empty (the author touched only test files, and touched at least one)

A violation is reverted (git -C <feature-worktree> revert the offending commit, or a targeted git checkout of the crossed-boundary file back to the prior SHA) and recorded in implementation-notes — do not silently keep a cross-ownership commit.


Step 4b — Test Authoring

Applies only to items carrying the needs-test-author trait. Runs after implementation agents return (their commits exist on the branch/worktree) and before the orchestrator's build verification below (Direct tier inverts this ordering — see the Direct-tier bullet below) — the test author compiles against the real implemented surface, not a predicted one. Two distinct seats, do not conflate them: the planning seat already froze test-plan at queue phase, before this step, gating work entry (see Step 3's seat-timing rule); the test author seat here writes test CODE and fills test-manifest — it does not fill or re-open test-plan.

Per-tier sequencing:

  • Direct tier: temporal-only degraded mode only (see Step 1) — no separate agent, and on a bug-fix the sequence is test-first-then-fix, not "tests after implementation": (a) fill test-plan at queue phase, before any implementation begins; (b) write the regression test from the diagnosis note's reproduction steps and observe it red against the pre-fix code, recording the red evidence (assertion, observed value) in test-manifest; (c) implement the fix; (d) confirm the test is green post-fix and record that confirmation too. test-manifest declares the single actor, and the disjointness check above does not apply (one actor, one range) — see Step 5's verdict rule for how this tier is scored.
  • Delegated tier: one additional sequential dispatch on the same branch, after the implementation agent's commit lands. The test author owns its own gradle cycle on that branch (ktlintCheck → ktlintFormat → re-verify, same rule as Step 4's "Who runs the lint cycle") and runs :current:test itself before committing.
  • Parallel tier: one test-author agent per child, dispatched as a dedicated wave between the implementation wave and the orchestrator's build verification (below). Authors in this tier do NOT run :current:test or :current:ktlintCheck — only ktlintFormat plus a compile self-check, mirroring the implementation dispatch template, since the orchestrator owns full build verification for the wave.

Blindness rule. The test author may read: the item's test-plan note, public signatures, domain models, existing test conventions/docs, and the implementer's changed-file names (not content). The test author must NOT read: the implementation diff content, the implementer's own tests, implementation-notes, or session-tracking. This is what makes the separation real rather than nominal — tests probe the spec, not the implementation's behavior.

Oracle-provenance obligation. Every scenario's expected result must trace to a source declared in test-plan (spec clause, stated algorithm, external reference) — never "what the code returns" and never the ticket's own worked example.

Manifest duty. Before returning, the test author fills test-manifest (work, required): actor id, test file paths, commit SHA range, S-id→test coverage mapping (covered / not-covered:reason), probes executed with results, forbidden-pattern declaration (every assumeTrue/escape with justification, or none), and any implementer-modification-to-test-files rationale (should be none if ownership held).

Skill routing. Invoke the test-author skill before filling test-manifest — it carries the scenario-derivation, oracle-derivation, blindness, and forbidden-pattern framework this step depends on. (The planning seat also invokes it earlier, at queue phase, for the scenario-derivation sections that inform test-plan — see Step 3.)

Show full SKILL.md (3,403 more words)Show less
Arbitration — Red Author-Authored Tests

When a test the author wrote comes back red against the implementation, the test author does not decide who is at fault (test-author §8/§9) — it reports. The orchestrator arbitrates, with exactly three dispositions:

  1. Implementation wrong. Re-dispatch the implementer, scoped to src/main/** only.
  2. Test wrong. REQUIRES a spec citation before this disposition is available: the note key plus the quoted acceptance criterion plus the asserted-vs-observed values. "The implementation does X" is not evidence for this case — it is evidence for case 1. If no covering criterion exists to cite, this disposition is not available; fall through to case 3. Once cited, re-dispatch the test author, scoped to src/test/** only.
  3. Spec wrong or silent. Escalate to the user. Amend task-scope/diagnosis first, then dispatch whichever side the amendment implicates.

Record every arbitration — case number, the citation (for case 2), and the resolving commit SHA — as a bullet in the child's implementation-notes. Never weaken, skip, or @Disable a test to unblock a wave, regardless of disposition; that is exactly the failure mode this trait exists to prevent (test-author §7, §9).

This carve-out also modifies the build-verification rule below: a red author-authored test is not a broken build. Route it to arbitration above; if unresolved, hold the child in work and report — do not dispatch a generic "fix agent" against it.

Test-author dispatch prompt: point the agent at the Test author protocol slot of the wave's dispatch contract, named by the absolute Contract path from its Header (Step 2) and instantiated from references/dispatch-contract-template.md — it states the declarations block, the src/main tool ban, the keys= filter, stop-and-ask, fixture invariants, surface labels, commit form and manifest fields in full; on a Delegated-tier run with no contract file, paste that slot into the prompt instead of paraphrasing it. The declarations block is filled by a separate, non-author declarations-extractor seat, and the orchestrator scans and strips it of implementation prose before the author sees it — the slot states the seat's accuracy contract and the scan (proposal 234b50a0). Under /task-orchestrator:run-wave, the extractor and the test author are both run-plan stages, and the scan step is named scan-declarations in Method B or the run script's own scanDeclarations call in Method A — same accuracy contract and strip rule either way.

Capture the test author's pre/post commit SHAs the same way as implementation agents (Test-Pre-SHA / Test-Post-SHA in the tracking table above), then run the disjointness check.


Step 4c — Build Verification and Wrap-Up

Build verification (orchestrator-owned, serialized). After each parallel-batch completes (or between sequential children), run from the feature worktree:

bash
git -C <feature-worktree> status                            # confirm clean tree
./gradlew -p <feature-worktree> :current:test
./gradlew -p <feature-worktree> :current:ktlintCheck

A failure means a recently-committed child broke something. Dispatch a fix agent (same shared worktree) before continuing. Do not advance any child to review until the build is green — the trend memory has multiple sessions of flaky-test-hides-real-bug showing why retry-until-green is wrong.

Carve-out: a red author-authored test is not a build failure. On items carrying needs-test-author, a test failure traced to a test the test author wrote (Step 4b) is not "the build broke" — do not dispatch a generic fix agent against it. Route it to the Arbitration subsection in Step 4b instead. If arbitration doesn't resolve before this verification pass needs to conclude, hold the child in work phase and report the unresolved red test rather than forcing green through a fix-agent dispatch.

Why orchestrator owns gradle invocations: ./gradlew runs against a single Gradle daemon and a single build/ cache per project directory. Parallel gradlew test invocations against the shared feature worktree will queue at the daemon, corrupt the build cache, or hit Windows file locks. Serializing build verification at the orchestrator prevents this without slowing the agents (they're not running gradle).

For Delegated tier (single subagent), the agent commits to the working branch on the main directory. Capture the changed files via git diff main --name-only and proceed.

Post-implementation steps (run in the feature worktree for Parallel tier, or on the working branch for Direct/Delegated):

  1. Run the /simplify skill on the changed code to check for reuse, quality, and efficiency — this is a cleanup pass before review, not a review itself
  2. If /simplify made changes, the resulting test coverage work follows the ownership boundary above. On items carrying needs-test-author, re-dispatch the test author to write or update the covering tests (Step 4b) — the implementer/orchestrator does not touch src/test/** even for a simplify-driven update. On items without the trait, write or update tests inline as before. Either way, the simplify pass is still part of the work phase — all code changes require test coverage before advancing to review.
  3. Log findings as work items — any issues surfaced by /simplify or during implementation that are not immediately addressed (pre-existing tech debt, optimization opportunities, related bugs) must be logged via /task-orchestrator:create-item before moving on. Do not discard findings.
  4. CHANGELOG bullet (proposal c068c943) — written here, in the work phase, so the reviewer sees it. A user-visible change — server behaviour, the MCP or REST surface, config keys, plugin skills or hooks — adds ONE bullet under CHANGELOG.md's ## [Unreleased] section, in its ### Added / ### Changed / ### Fixed subsection (plugin skill and hook changes go under ### Plugin). The bullet describes behaviour, not the diff, and carries no volatile counts. A docs-, process- or chore-only change adds none; its PR body will state Changelog: none (<why>). Direct/Delegated: committed with the change. Parallel: the orchestrator or docs seat writes it once, after the implementation wave and before review — never an implementer (the dispatch contract's Docs slot). /prepare-release Step 8c folds these bullets into the release section.
  5. Fill all work-phase notes following their guidancePointer — focus on context that downstream agents need to know
  6. After implementation completes:
    • Subagent delegation: The agent returns after filling work-phase notes. The orchestrator then calls advance_item(trigger="start") to advance the item to the next phase.
    • Direct implementation: The orchestrator calls advance_item(trigger="start") itself after filling work-phase notes. In both cases, inspect newRole in the response to determine what comes next (see Step 5).

Step 5 — Review Phase

Before dispatching or performing review, check the item's current role. Inspect newRole from the advance_item response in the previous step:

  • If newRole is terminal: The item's schema has no review phase (lightweight lifecycle). Review dispatch is not needed — the item completed through its natural lifecycle. Proceed to Step 6. Note: feature-task items skip review by default (work→terminal directly) — the needs-task-review trait re-enables a review phase for a specific child when needed.
  • If newRole is review: Continue with review per tier below.

This step is tier-conditional:

Direct tier: Perform an inline review. Read the diff, verify correctness, confirm tests pass. If the review-phase note has a skillPointer (visible in the advance_item response or via get_context), invoke that skill for the evaluation framework before filling the review note. Write the review note, then advance to terminal: advance_item(trigger="start"). No separate review agent — the overhead exceeds the risk for 1-2 file changes with known fixes.

Delegated and Parallel tiers: Dispatch a separate review agent. The agent that implemented the code must not review its own work.

Reviewer scoping for Parallel waves. Per-child reviewers remain the default for code-bearing waves. When every child in the wave is content-only (config, skill, or doc edits — no src/main or src/test changes) AND the wave shares a pinned plan-file contract (Step 2), dispatch ONE consolidated opus reviewer over the full feature-branch diff (git diff main...<FEATURE_BRANCH>) instead of N per-child reviewers. Cross-file coherence defects — sibling skills stating contradictory rules, gate placement contradicting seat timing, an example contradicting its own rule — are visible only to a reviewer holding the whole diff; per-child reviewers structurally cannot see them, and a shared naming contract does not prevent them, since naming consistency and semantic coherence are orthogonal. Evidence: 5 blocking cross-file contradictions caught this way at 5-item scale (retro 6d562acb) and canonical-config consistency issues at 7-item scale (retro bd4ec109) — trend 028c7b5d. The consolidated reviewer fills the review-phase notes of whichever items carry a review phase (typically the parent feature's review-checklist, since feature-task children skip review by default).

If the implementation used the shared feature worktree (Parallel tier), the review agent operates in that worktree, scoped to that child's owned files — the diff of those paths from the wave's base SHA to current HEAD, per the contract's Review scoping slot. Not the <pre-sha>..<post-sha> range captured during dispatch: in a shared worktree that range interleaves other streams' commits and excludes any later fix-up, formatter or fixture-repair commit to the same files, so it both over- and under-reports. The pre/post map still serves the ownership two-range check (list item 6 below).

Review agent template (copy verbatim, fill placeholders):

You are reviewing one child task within a shared feature worktree.

- Feature worktree: <FEATURE_WORKTREE_PATH>
- Feature branch: <FEATURE_BRANCH>
- Dispatch contract: <ABSOLUTE path from its Header "Contract path" line, e.g.
  D:\Projects\task-orchestrator\plans\<slug>.md> — read it first; it wins on conflict
  with this prompt. It lives in the main checkout, not in the worktree above.
- This child's owned files (its row in the contract's File ownership slot):
  <FILE LIST>
- Your review scope is the diff of exactly those files:
  git -C <FEATURE_WORKTREE_PATH> diff <BASE_SHA>..HEAD -- <FILE LIST>
- Other children committed into this same branch, and later fix-up, formatter or
  fixture-repair commits may touch these files. Review the owned-file diff above as it
  stands now; do NOT review another child's files, and do NOT bound your review by a
  commit range.

Run ALL commands from within the feature worktree.
Read ALL files from that directory. Do NOT read from the main working directory.

Tests have already been verified green by the orchestrator after the most recent
commit batch. Run no gradle — the contract's Compile self-check slot states who runs
it, and the Review scoping slot records the build state you rely on. Focus on plan
alignment, test quality, and simplification per the review-quality skill.

Report every finding at every severity — do not self-filter to a high-severity-only
bar. Mark each finding blocking or observation and state your confidence; the
orchestrator's verdict handling is the downstream filter, not your own judgment.

If using Direct or Delegated tier (single working branch on the main directory), the review agent reads from the working branch and runs tests itself per the existing template — no worktree-specific scoping needed.

The review agent:

  1. Reads the review-quality skill
  2. Uses get_context(itemId=...) to load the item's notes and review-phase requirements
  3. Reads the changed files — on Parallel tier, the owned-file diff above; on Direct/Delegated tier, the working branch
  4. Runs gradle per the contract's "Who runs gradle" statement (Compile self-check slot): in a Parallel wave the reviewer runs none, because the orchestrator already verified the wave and recorded the build state in the Review scoping slot — for that tier this supersedes review-quality Area 1's "run the test suite first", whose recorded numbers the reviewer takes from that slot instead. On Direct/Delegated tier there is no contract and no orchestrator-run wave verification, so the reviewer runs the test suite AND the linter itself (Step 4's "Verification commands") — a PR with failing lint will not merge.
  5. Evaluates plan alignment, test quality, and simplification
  6. On items carrying needs-test-author, additionally:
    • Two-range ownership check: confirm the impl-range/author-range disjointness recorded during Step 4b actually holds by spot-checking git diff <impl-range> --name-only touches no src/test/** path and git diff <test-range> --name-only touches no src/main/** path. N/A-by-declaration on Direct-tier temporal-only items (Step 1, Step 4b): there is one actor and one range by design, so there is no second range to diff — record it as N/A: temporal-only, single actor rather than attempting the diff, and it cannot fail.
    • Temporal checks (Direct-tier temporal-only items only): confirm test-plan was filled and frozen at queue phase, before implementation began (Step 3's seat-timing rule), and — for a bug-fix — confirm the red-first evidence (the regression test observed failing against pre-fix code, per Step 4b/B3) is recorded in test-manifest.
    • Arbitration-record check: if any red author-test required arbitration (implementation wrong / test wrong / spec wrong), confirm the outcome is recorded in implementation-notes with the required detail (spec citation for a "test wrong" call: note key + quoted criterion + asserted-vs-observed).
    • assumeTrue/@Disabled scan on author-owned files: grep the test author's changed files for assumeTrue, @Disabled, or other weakening introduced after the author's first commit in this range. Any such introduction is an automatic blocking finding — independence does not permit softening a red test to unblock a wave. Verdict rule: test-independence-audit fails to not-independent (blocking) if any applicable check above fails, regardless of how green the test suite is — the ownership two-range check is not applicable, and cannot fail, on a Direct-tier temporal-only item. On a Direct-tier temporal-only item where every applicable check (temporal checks, arbitration-record, assumeTrue/@Disabled scan) passes, the verdict is independent-degraded (temporal-only) — never plain independent, matching test-author §11.
  7. Fills the review-phase notes per guidancePointer with a verdict

Handling the verdict:

VerdictAction
PassProceed to Step 6
Pass with observationsProceed to Step 6; log observations for follow-up
Fail — blocking issuesStop and report to the user with the full findings. Do not attempt to fix autonomously — bring the human into the loop.

Review failures surface issues that may indicate systemic problems worth learning from. Automatically retrying hides these signals.


Step 6 — Finalize and PR

The shape of Step 6 depends on tier.

Direct and Delegated tiers — finalize per item

After review passes:

  1. Post-dispatch commit audit (non-blocking): run git log --oneline -3 and git status --short and compare against what the orchestrator itself committed. Any commit a subagent made despite stop-boundary instructions is flagged for review here — inspect its scope before it rides into the squash-merge (a subagent commit lacks the co-author trailer; reconcile at merge time). On items carrying needs-test-author, additionally check each commit's file list for cross-ownership: an implementer commit touching src/test/**, or a test-author commit touching src/main/**, is flagged here even if it slipped past the Step 4b disjointness check — do not let it ride into the squash-merge unreconciled.

  2. Verify the working branch is committed (orchestrator commits if Direct tier; subagent committed if Delegated). Stage only the files related to the implementation:

    bash
    git add <specific-files>
    git commit -m "$(cat <<'EOF'
    <type>(<scope>): <description>
    
    <body — what changed and why, referencing the MCP item>
    
    Co-Authored-By: Claude <noreply@anthropic.com>
    EOF
    )"

    Commit types: feat for features, fix for bugs, refactor for tech debt, perf for performance, test for test-only changes, chore for maintenance.

  3. CHANGELOG check. Confirm the branch carries the [Unreleased] bullet written in Step 4c (post-implementation step 4), or that the PR body will state Changelog: none (<why>).

  4. Push the working branch:

    bash
    git push -u origin <branch-name>
  5. Create the PR:

    bash
    gh pr create --base main --title "<type>(<scope>): <description>" --body "$(cat <<'EOF'
    ## Summary
    <2-4 bullets>
    
    ## Test Results
    <test count, pass/fail, new tests>
    
    ## Review
    <verdict summary>
    
    ## Changelog
    <the [Unreleased] bullet text, or none (<why>)>
    
    ## MCP
    <item ID>
    EOF
    )"
  6. After PR merges:

    bash
    git checkout main
    git pull origin main
    git branch -D <branch-name>
  7. Advance the item to terminal:

    bash
    advance_item(transitions=[{ itemId: "<uuid>", trigger: "start" }])
  8. After the item reaches terminal, follow the retrospective hook's directive if one fires (see retrospective.mode).

Report the PR URL and a summary.

Parallel tier — finalize ONCE at parent-feature completion

For Parallel-tier features with a shared feature worktree:

For each child task (after its review passes, if it has one):

  1. For children whose schema/trait declares a review phase (e.g. needs-task-review), confirm review-checklist is filled. Children without one advance work→terminal directly — there is nothing to confirm.
  2. advance_item(itemId=<child-uuid>, trigger="start") to move work→review→terminal (or work→terminal directly) as the child's schema dictates.
  3. Do NOT push. Do NOT create a PR. The work is committed to feat/<feature-slug> inside the shared worktree; that's the integration point.

When all children reach terminal, the parent feature is ready to finalize:

  1. Fill the parent's implementation-notes and session-tracking notes (aggregating across children — distributed-tracking pattern works as today).
  2. Run final verification from the feature worktree:
    bash
    ./gradlew -p <feature-worktree-path> :current:test
    ./gradlew -p <feature-worktree-path> :current:ktlintCheck
  3. Advance the parent to review and fill review-checklist (orchestrator-authored, summarizing across all children's reviews).
  4. CHANGELOG check — as the Direct/Delegated tier's step 2: the bullet the orchestrator or docs seat wrote before review (Step 4c), or Changelog: none (<why>) in the PR body.
  5. Push the feature branch:
    bash
    git -C <feature-worktree-path> push -u origin feat/<feature-slug>
  6. Create one PR for the whole feature:
    bash
    gh pr create --base main --title "feat(<scope>): <feature description>" --body "$(cat <<'EOF'
    ## Summary
    <feature-level summary aggregating all children>
    
    ## Children completed
    - <child-1 title> (<MCP UUID>)
    - <child-2 title> (<MCP UUID>)
    ...
    
    ## Test Results
    <total test count, new tests added across feature>
    
    ## Review
    <feature-level review verdict, references each child's review-checklist>
    
    ## Changelog
    <the [Unreleased] bullet text, or none (<why>)>
    
    ## MCP
    Parent: <parent UUID>
    Children: <list of child UUIDs>
    EOF
    )"
  7. After PR merges:
    bash
    git checkout main
    git pull origin main
    git worktree remove <feature-worktree-path>
    git branch -D feat/<feature-slug>
  8. Advance the parent feature to terminal.
  9. Retrospective — the plugin's retrospective hook fires on the parent's terminal transition. In dispatch mode (see retrospective.mode in .taskorchestrator/config.yaml) it directs a background /session-retrospective automatically — follow its directive; in nudge mode, or if no directive arrives, suggest running it manually.
Why one PR at parent finalization, not per child
  • Coherent review context. The PR diff shows the whole feature, not N disjoint pieces.
  • One CI cycle per feature instead of N. Local verification (orchestrator-owned gradle runs between commits) gives equivalent regression signal during development.
  • No cross-PR contamination. Contract changes can't surface on a sibling's merge commit because there are no sibling PRs.
  • No PR-body staleness. The PR body is authored once, after the feature is done, describing what actually shipped.
  • Aggregate retrospective material. Distributed session-tracking notes across children plus parent-level aggregation gives clean retro input (validated by retro a7f6024f).

Local main always tracks origin/main — no divergence, no reset --hard needed.


Autonomous Batch Processing

When processing a Parallel-tier feature with multiple child tasks autonomously: run /task-orchestrator:run-wave over the parent, one run per topological layer — steps 1-2 below describe what that run does under the hood; they are the fallback for a manual walk-through when a run plan doesn't apply.

  1. Step 2 — One worktree, one branch. Orchestrator creates the feature worktree and feature branch (feat/<slug>) at planning time. All children share it.
  2. Step 4 — Dispatch into shared worktree. Agents dispatched without isolation: "worktree". Independent children dispatch in parallel waves; dependent children dispatch sequentially. Orchestrator scopes each agent's file list to prevent overlap.
  3. Step 4b — Test-author wave. For children carrying needs-test-author, dispatch one test-author agent per child as its own wave, after that child's implementation wave and before build verification. Authors run ktlintFormat + compile self-check only (never :current:test/:current:ktlintCheck — the orchestrator owns those next). Capture Test-Pre-SHA/Test-Post-SHA per child and run the disjointness check before proceeding.
  4. Build verification — orchestrator-owned, serialized. After each parallel wave (implementation or test-author), the orchestrator runs :current:test and :current:ktlintCheck from the feature worktree. Fix failures before advancing any child to review.
  5. Step 5 — Review per child, scoped to that child's owned files, when the child has a review phase. feature-task children skip review by default (work→terminal directly) unless the needs-task-review trait is set, or needs-test-author is set (its test-independence-audit note adds a review phase). When a review phase applies, the review agent reads from the shared worktree and diffs that child's owned files from the wave's base SHA to HEAD (git -C <feature-worktree> diff <base-sha>..HEAD -- <files>), per the contract's Review scoping slot — never a commit range, which in a shared worktree picks up other streams and misses later fix-ups. The pre/post and test-pre/post SHAs feed the ownership two-range check only.
  6. Step 6 — One PR at parent finalization. Children advance to terminal without pushing or PR'ing. Only when the parent feature itself reaches terminal does the orchestrator push feat/<slug> and open the single feature-level PR.
  7. Track child commits — maintain a table mapping child UUID → pre-commit SHA → post-commit SHA → Test-Pre-SHA → Test-Post-SHA → status (implementing / test-authoring / reviewing / done / failed). Worktree path is shared across all children.
  8. Report at the end — summarize children completed, review failures, and the single PR URL.

If any child hits a review failure, continue processing siblings (their commits are already in the feature branch). Report all failures together at the end. The orchestrator decides whether the failed child blocks parent finalization (e.g. fix-and-re-review), can be cancelled (descope from the feature), or warrants reverting its commits.

Bug-fix batches (multiple unrelated fixes) — these are NOT a Parallel-tier feature. Use Delegated tier per item, each with its own branch and PR (the legacy per-item flow). The shared-worktree pattern applies only when items share a parent feature item.


Worktree Strategy

For full setup, dispatch patterns, lifecycle, parallel validation, and test baseline management, see WORKTREE.md.

Quick reference:

TierWorktreeBranchPR scope
Direct / Delegated (single item)None — work on main directory<type>/<slug>One PR per item
Parallel (parent feature with N children)One shared feature worktreefeat/<feature-slug> (one branch for all children)One PR at parent finalization

Do NOT use isolation: "worktree" on the Agent tool for Parallel-tier child dispatches. That spawns a separate worktree per dispatch — the deprecated per-child PR pattern. For Parallel tier, the orchestrator pre-creates one shared worktree in Step 2 and dispatches each child agent into it.

When NOT to create any worktree:

  • Direct/Delegated tier (single item — work on the main directory)
  • Pure MCP operations with no file modifications
  • Orchestrator implementing directly

Resuming In-Progress Work

Tier classification happens at Step 1 even when resuming. Classify the tier from the item's tags, file scope, and note state, then resume using that tier's pipeline.

If an item is already past the queue phase (e.g., previously planned but not implemented), the skill picks up from the current state:

Current roleResume from
queue (notes filled)Step 3 — advance and proceed
queue (notes missing)Step 3 — fill missing notes
work (in progress)Step 4 — check implementation state
work (notes filled)Step 4 — advance to review
reviewStep 5 — run review
terminalAlready done — report status
open run/<runId>/state found (F7)/task-orchestrator:run-wave --resume <runId>

Always call get_context(itemId=...) first to determine exact state before resuming.

© jpicklyk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in .claude/skills/implement of jpicklyk/task-orchestrator.

  • SKILL.md
  • WORKTREE.md
  • references/dispatch-contract-template.md
  • references/patch-anchored.py
  • references/test_patch_anchored.py

Open the folder on GitHubat commit b688ea0

Compare with similar skills

Implement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Implement compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Implement this skilljpicklyk/task-orchestrator207—~17kAutomated safety check: PassMIT
Devcontainer Devstacklok/toolhive-studio170—~3.8kAutomated safety check: NotesApache-2.0
Zed Configwcygan/dotfiles196—~930Automated safety check: PassNone
Agent Deckasheshgoplani/agent-deck1k—~1.7kAutomated safety check: PassMIT
Foremergenaw103/foremerge536—~2.4kAutomated safety check: PassApache-2.0
Puppetmaster Agent Orchestrationprofessorpalmer/Puppetmaster467—~3.2kAutomated safety check: PassMIT

Similar skills

  • Devcontainer Dev

    stacklok/toolhive-studio

    Spin up and interact with ToolHive Studio's containerized dev environment (Xvfb + noVNC + DinD).

    170 GitHub stars~3.8k tokensUpdated today
    Agent WorkflowsAuto-check: notes
  • Zed Config

    wcygan/dotfiles

    Zed editor configuration expert. An agent skill from wcygan/dotfiles.

    196 GitHub stars~930 tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Agent Deck

    asheshgoplani/agent-deck

    agent-deck, the terminal session manager for AI coding agents.

    1k GitHub stars~1.7k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Foremerge

    naw103/foremerge

    Coordinate parallel coding agents with Foremerge's local Git-compatible CLI and MCP server.

    536 GitHub stars~2.4k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Puppetmaster Agent Orchestration

    professorpalmer/Puppetmaster

    Operates and supervises Puppetmaster, a multi-agent orchestrator, through its MCP tools or CLI, picking the right verb for edits, reviews, audits and long-running jobs.

    467 GitHub stars~3.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Reads and updates the Sandboxed.sh Library, a git-backed store of skills, agents, commands, tools, rules and MCP servers, through its library tools.

    515 GitHub stars~931 tokensUpdated today
    Agent WorkflowsAuto-check passed

More from jpicklyk/task-orchestrator

All 28 skills in this repo
  • Task Orchestrator Server Setup

    jpicklyk/task-orchestrator

    Walks through how to launch and reach the MCP Task Orchestrator server container: transport, REST API, port publishing, config mounts and config-sync.

    207 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Run Wave

    jpicklyk/task-orchestrator

    Resolves ready MCP work items into a run plan, shows it to you, then executes it through the Workflow tool or direct subagent dispatch, with post-run verification.

    207 GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Adopt Project Scope Migration

    jpicklyk/task-orchestrator

    Migrates an existing unscoped Task Orchestrator database to the project-scoping convention in place, creating one project anchor root and re-parenting work trees under it after a mandatory dry run.

    207 GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Bulk Task Completion

    jpicklyk/task-orchestrator

    Completes or cancels a whole feature subtree, a named list of items, or a batch of stale work items at once, previewing the impact and warning before force-completing anything active.

    207 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Task Orchestrator Item Creator

    jpicklyk/task-orchestrator

    Creates an MCP work item from conversation context, anchoring it under the right container, inferring type and priority and pre-filling the required notes.

    207 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Work Item Dependency Manager

    jpicklyk/task-orchestrator

    Views, creates, deletes and diagnoses BLOCKS, IS_BLOCKED_BY and RELATES_TO links between MCP work items, including why an item cannot start.

    207 GitHub stars~3.5k tokensUpdated today
    Auto-check passed

Questions about Implement

What does Implement do?

End-to-end workflow for taking MCP work items from backlog to merged PR. Implement is an agent skill from jpicklyk/task-orchestrator. End-to-end workflow for taking MCP work items from backlog to merged PR.

When should I use Implement?

Implement fits situations like: A user says implement this; work on this item; pick up the next task; create a PR for this.

How do I install Implement in Claude Code?

Run `npx skills add jpicklyk/task-orchestrator --skill implement -a claude-code`. Or copy the skill folder (.claude/skills/implement in jpicklyk/task-orchestrator) into .claude/skills/implement in your project. Claude Code loads it when a task matches its description.

How do I install Implement in Codex?

Run `npx skills add jpicklyk/task-orchestrator --skill implement -a codex`. Or copy the skill folder (.claude/skills/implement in jpicklyk/task-orchestrator) into .agents/skills/implement in your project. Codex loads it when a task matches its description.

Can I use Implement in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jpicklyk/task-orchestrator --skill implement -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/implement, .gemini/skills/implement, .github/skills/implement and .opencode/skills/implement in your project.

What does Implement need to run?

Going by SKILL.md and its folder, Implement needs Python for the scripts in its folder and the command-line tools its instructions call (git and gh). Our summary lists: Python 3.

Does Implement access the network?

SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Implement safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Implement use?

Implement is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Implement use?

About 17k tokens (SKILL.md is roughly 67k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.9k tokens, read only when the agent opens those files.

What are the alternatives to Implement?

Skills that share tags, products or a category with Implement: Devcontainer Dev (stacklok/toolhive-studio, 170 stars), Zed Config (wcygan/dotfiles, 196 stars), Agent Deck (asheshgoplani/agent-deck, 1k stars) and Foremerge (naw103/foremerge, 536 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Implement?

jpicklyk (a GitHub user) maintains it in jpicklyk/task-orchestrator, which has 207 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 6, 2026.

Source: jpicklyk/task-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.