Devcontainer Dev
stacklok/toolhive-studio
Spin up and interact with ToolHive Studio's containerized dev environment (Xvfb + noVNC + DinD).
End-to-end workflow for taking MCP work items from backlog to merged PR.
$ npx skills add jpicklyk/task-orchestrator --skill implement -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jpicklyk/task-orchestrator implement --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/implement .claude/skills/implement && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "implement" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/implement into .claude/skills/implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/implementType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jpicklyk/task-orchestrator --skill implement -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jpicklyk/task-orchestrator implement --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/implement .agents/skills/implement && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "implement" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/implement into .agents/skills/implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jpicklyk/task-orchestrator --skill implement -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jpicklyk/task-orchestrator implement --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/implement .cursor/skills/implement && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "implement" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/implement into .cursor/skills/implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jpicklyk/task-orchestrator.git --path .claude/skills/implement--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jpicklyk/task-orchestrator --skill implement -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jpicklyk/task-orchestrator implement --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/implement .gemini/skills/implement && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "implement" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/implement into .gemini/skills/implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jpicklyk/task-orchestrator implementInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jpicklyk/task-orchestrator --skill implement -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/implement .github/skills/implement && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "implement" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/implement into .github/skills/implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jpicklyk/task-orchestrator --skill implement -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jpicklyk/task-orchestrator implement --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/implement .opencode/skills/implement && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "implement" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/implement into .opencode/skills/implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
implementEnd-to-end workflow for taking MCP work items from backlog to merged PR.
Implement is an agent skill from jpicklyk/task-orchestrator. End-to-end workflow for taking MCP work items from backlog to merged PR. Handles git branching, schema-driven planning, implementation, independent review, and PR creation. Composes spec-quality, review-quality, and schema-workflow skills into a single pipeline. Use when a user says "implement this", "work on this item", "fix these bugs", "pick up the next task", "create a PR for this", "go through the backlog", or references specific MCP item IDs for implementation.
Its SKILL.md is about 17k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `WORKTREE.md`, `references/dispatch-contract-template.md` and `references/patch-anchored.py`).
It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol and Git. The repository describes itself as: Server-enforced workflow discipline for AI agents. An MCP server providing persistent work items, dependency graphs, quality gates, and actor attribution. Schemas define what… The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b688ea0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
gitghFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git and gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Implement loads about 17k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 8,156 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jpicklyk/task-orchestrator at commit b688ea0, republished under its MIT licence (© jpicklyk). 8,156 words, ~16,749 tokens.
.claude/skills/implement/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.End-to-end workflow for taking MCP work items from queue to PR. This skill composes the schema-driven planning (spec-quality), implementation, review (review-quality), and git/PR workflow into a single pipeline.
Usage:
/implement <item-id> — work on a specific item/implement — with context about what to work onLoad the item(s) and determine the execution tier and interaction mode.
For each item, call get_context(itemId=...) to understand:
When the user or an item references a PR or commit as already landed, verify its merge state
before relying on it — gh pr view <n> --json state,mergeCommit plus git log origin/main —
and confirm the working tree is current (Step 2, base-freshness precondition) before any pass
that checks whether a change is present in the codebase.
Execution tier — classify by this table (canonical source shared with the task-orchestrator:orchestrate skill; edit the fragment, not this copy):
<!-- BEGIN GENERATED:tier-classification | source: claude-plugins/task-orchestrator/_fragments/tier-classification.md · regen: node claude-plugins/task-orchestrator/_fragments/generate.mjs -->
| Criteria | Tier | Pipeline |
|---|---|---|
| 1-2 files, known fix, no migration/new API | Direct | Orchestrator edits, tests, reviews inline |
| 3-10 files, single logical unit, clear or explorable scope | Delegated | Single subagent, separate review agent |
| 11+ files, multiple independent work streams, dependency edges | Parallel | Worktree agents, full pipeline |
Force-UP signals (bump tier regardless of file count):
Force-DOWN signals:
default or absent → eligible for Direct<!-- END GENERATED:tier-classification -->
Delegated with ≥2 ready items, or Parallel → run /task-orchestrator:run-wave.
If the item has no schema tag, apply quick-fix for Direct tier or leave untagged for Delegated/Parallel (the default schema catches these).
Trait application on classification. When the tier resolves to Delegated or Parallel and the
workspace defines a delegated trait (it appears in availableTraits on create responses), apply
it before any dispatch — traits: "delegated" at item creation, or
manage_items(operation="update", items=[{itemId: "<uuid>", traits: "delegated"}]) for an existing
item. This makes the orchestrator-filled delegation-metadata note schema-visible instead of
convention-only. Direct tier: do not apply it — nothing is delegated.
Model-selection traits. Seat dispatch defaults come from the delegated trait (planner opus,
implementer sonnet, test-author sonnet, reviewer opus). Override them per item, at the same
point:
complex-implementation — implementer on opus. Apply for architecture-heavy or multi-file
synthesis work: new public API, cross-layer seams, subtle state machines or concurrency.high-stakes — planner and reviewer on fable. Apply sparingly, where a wrong plan or a missed
review finding is expensive: security predicates, auth, data-loss paths, hard-to-reverse design
decisions. Fable draws on its own usage bucket, so it is never a default. needs-security-review
does NOT imply it.Read the resolved per-seat profile from get_context / query_items(operation="schema")
(dispatchBySeat) and pass its model explicitly on the dispatch.
Test-author trigger rule. bug-fix.default_traits already includes needs-test-author — no
action needed, it applies automatically. For feature-task items, apply needs-test-author per
item (manage_items(operation="update", items=[{itemId: "<uuid>", traits: "needs-test-author"}]))
when acceptance criteria involve a predicate, algorithm, parser, validator, or state transition,
or when any force-ON signal is present: new public API surface, a database migration, a security
predicate, or a prior vacuous-test finding in this item's area. Other schema tags
(feature-implementation, plugin-change, quick-fix, and the global floor) do not carry the
trait. Direct tier: apply the trait only in its temporal-only degraded mode, and only on
bug-fixes — the test-plan gate and red-first rule still apply, test-manifest declares a single
actor, and there is no separate test-author dispatch. Other Direct-tier items are exempt
regardless of the criteria above — the tier is too small to separate authorship into a second
dispatch.
Interaction mode — orthogonal to tier:
| Signal | Mode |
|---|---|
| User says "work with me on", "let's plan", or similar collaborative language | Collaborative — user participates in planning and key decisions |
| Scope is clear, no user participation needed | Autonomous — agent handles the pipeline |
| Unclear scope or ambiguous complexity | Ask the user |
When processing multiple items, evaluate whether related items (e.g., bugs in the same module, fixes that touch the same files) should be grouped into a single branch and PR. Group when the changes are cohesive and independent fixes would create merge conflicts. Keep items separate when they're unrelated or when isolation makes review cleaner.
Sync local main before any implementation begins.
git checkout main
git pull origin main --tagsBase-freshness precondition. Before creating any branch or worktree, and before any source-verification pass (a grep or file read that decides whether a change is present), confirm the intended base is current:
git fetch origin
git log --oneline -1 origin/main
git rev-parse --short HEADThe two SHAs must match, or HEAD must be a descendant of origin/main on the intended branch.
A source-verification pass must never run against a working tree that has not been confirmed
current — this includes the orchestrator's own session worktree. If the tree is behind, read
the file directly from the remote (git show origin/main:<path>, from PowerShell on Windows —
the Bash tool's MSYS layer mangles the ref:path colon) or create a fresh worktree from
origin/main (fallback block below). A stale base manufactures absent-looking evidence and
absence reads as a discovery, not an error: a 6-commit-stale tree reported four already-adopted
proposals as missing (2026-08-04), and a one-commit-stale dispatch re-implemented existing
runtime into a conflicting 17-file commit (2026-05-01) — trend 760be80d.
The branching/worktree strategy depends on tier:
Direct tier (orchestrator implements 1–2 files inline) — create a working branch on the main directory:
git checkout -b <branch-name>Delegated tier (single subagent) — same as Direct: orchestrator creates the branch on the main directory, the subagent works against it. No worktree.
If the main checkout is unavailable (another branch checked out, uncommitted changes present) or the orchestrator itself runs from a worktree — Direct and Delegated tiers both: do not touch the occupied checkout. Create a dedicated worktree instead:
git worktree add .claude/worktrees/<slug> -b <branch-name> origin/mainWork there using absolute paths and git -C (never cd); after the PR merges, remove the worktree
and delete the branch. Sync local main via git fetch origin main:main while main is not
checked out anywhere, or a normal git pull from the main checkout when it is.
Parallel tier (parent feature with multiple children) — create a single feature worktree that all child agents share:
FEATURE_SLUG=<short-feature-description> # e.g. issue-117-followup
FEATURE_BRANCH=feat/$FEATURE_SLUG
FEATURE_WORKTREE=.claude/worktrees/feat-$FEATURE_SLUG
# Resume detection — if the branch/worktree already exist (orchestrator restart
# mid-feature), reuse them rather than recreating:
if git show-ref --verify --quiet "refs/heads/$FEATURE_BRANCH"; then
echo "Resuming existing feature branch $FEATURE_BRANCH"
else
git branch "$FEATURE_BRANCH" main
fi
if [ ! -d "$FEATURE_WORKTREE" ]; then
git worktree add "$FEATURE_WORKTREE" "$FEATURE_BRANCH"
fiAll child-task agents will be dispatched into this shared worktree (Step 4). The feature branch is pushed and PR'd once, when the parent feature reaches terminal (Step 6).
Plan file as dispatch contract (3+ children). For every Parallel-tier run, instantiate
references/dispatch-contract-template.md into the
wave's plan file before the first dispatch, fill every placeholder, and have every Step 3, 4, 4b
and 5 dispatch prompt reference it under the template's conflict rule instead of restating design
or process details per prompt. The template states the rules; this table only says what each
slot is for — do not paraphrase a slot's rule into a prompt. Under /task-orchestrator:run-wave,
the Header and File ownership rows are filled from the planner's explain output plus the
planner-v1 stage's own returns rather than hand-typed, and the run reuses this same worktree by
passing --mode shared --worktree $FEATURE_WORKTREE --branch $FEATURE_BRANCH — the contract file and its slot order are unchanged either way.
Write it to <main-checkout>\plans\<slug>.md and hand agents that absolute path, recorded
once in the contract's Header Contract path line. The plans/ directory is gitignored, so it
exists only in the main checkout and never in the feature worktree; a prompt that pins a worktree
working directory and then names a relative plans/<slug>.md points at a path that is not there.
| Slot | Why it exists |
|---|---|
| Header | One place agents read branch, worktree path, the write root every agent write must fall under, the contract's own absolute path, base SHA, PR boundary, scratchpad and the gradle lock helper, so no prompt repeats them (write root: proposal d1484a3e). |
| Conflict rule | Makes the plan file authoritative over prompt text, and names the per-item note that outranks the plan for its own dimension. |
| Items | Freezes each item's scope anchor so an agent cannot re-litigate scope mid-wave. |
| Planning seat return template | Fixes the fields each planning seat returns, so streams compose without re-reading every plan. |
| Commit discipline | The shared index makes a bare commit unsafe: pathspec staging and commit, post-commit verification, and the recovery step (proposal 568e7f7e / #301). |
| Compile self-check | One pinned invocation through the lock helper, how its exit code is read, the foreign-file early exit, the owned-file retry bound (proposal b82537e4 / #302), and the single statement of who runs gradle in the wave. |
| File ownership | Each agent's entire writable scope — including any declared edit to an existing test file — with cross-stream overlaps declared up front rather than discovered mid-wave. |
| Test author protocol | The blindness, oracle and scope rules that keep test authorship independent of the implementation. |
| Contract-change sweep | After a contract tightens, call sites and fixtures the changed item never names break; the sweep makes finding and construction-only-repairing them a step rather than a reminder (proposals 82034e9a, 31a1abeb). |
| Docs | Routes doc edits — and the CHANGELOG.md bullet (proposal c068c943) — to one serialized seat so parallel agents never write the same doc. |
| Notes | Who fills which note key, and the maxLength each must respect. |
| Review scoping | Reviews diff owned files, not SHA ranges; the commit map serves only the ownership check (#301). |
Two runs (retros 202b5d42, 6d562acb) produced zero vocabulary deviations across up to five
concurrent authors this way, with prompts roughly 60% smaller. The delivery surface is itself
the point (proposal 728a3e57 / #307): in a same-day controlled comparison, the wave with no
contract file recurred every previously-adopted failure class, while the waves that ran from one
held the gains — a rule that lives only in skill prose does not reach the agents, because agents
read the dispatch prompt and their item's schema guidance. The contract guarantees
deviation-free execution, not plan correctness — pair it with Step 3's "plans are spec inputs"
rule so the plan's own factual claims are verified by the implementer rather than merely
followed.
Why one worktree per feature, not per child: the feature is the natural PR boundary. Per-child PRs created cross-PR test contamination and PR-body staleness during the #117 follow-up (see retro a7f6024f). Shared worktree means one commit history, one CI cycle, one PR — and the parent feature's review-checklist gives a coherent point at which to finalize.
Branch naming:
feat/<feature-slug> — feature-implementation parents (the integration branch)fix/<short-description> — bug-fix items (Direct or Delegated tier)fix/<grouped-description> — batch of related bug fixes (Delegated tier)chore/<short-description> — tech debt, refactoring (Direct tier)This step is tier-conditional:
Direct tier: Skip this step entirely. No plan mode. No queue-phase notes. Call
advance_item(trigger="start") immediately to move queue→work. The quick-fix
schema has no queue-phase required notes, so the gate passes. Exception: if the item
carries needs-test-author (temporal-only degraded mode, Step 1), test-plan is a
queue-phase note and must be filled before this advance — see the seat-timing rule below.
Delegated tier: Fill queue-phase notes per schema. Use get_context(itemId=...)
to see expectedNotes and guidancePointer. Pre-plan-workflow is optional — use
only if scope needs exploration. Post-plan-workflow only if child items need
materialization. Advance: advance_item(trigger="start").
Parallel tier: Full planning pipeline:
Collaborative mode:
pre-plan-workflow and post-plan-workflow hooks fire automatically on
plan mode entry and exit, handling context gathering and materialization.guidancePointer for each required note — this
will reference the spec-quality skill where applicable.advance_item(trigger="start")Autonomous mode:
pre-plan-workflow skill — gather existing MCP state,
check schema requirements, and understand the definition floor.guidancePointer for each. The
spec-quality framework applies regardless of mode.post-plan-workflow skill — materialize child items
if the plan calls for them.advance_item(trigger="start")The gate will reject advancement if required notes are missing. If rejected, fill the missing notes and retry.
Plans are spec inputs. Every file path, tool contract, or call sequence a plan names must
cite its authoritative location and be verified on disk or in source before the plan is
approved — child items' specification notes are generated from the plan and inherit its errors
silently (spec-quality, "Cite Contracts, Don't Restate Them"). This reminder lives here because
spec-quality is reached via note guidance pointers that plan mode does not traverse.
Do not confuse this with resource-lease contention. A queue→work advance_item can also
fail with applied: false, errorCode: "resource_unavailable", errorKind: "transient" — a
resource a trait on this item declares (resources:) is currently held by another item. This
is not a note gate failure: filling notes will not fix it. Work a different item and retry
later (retryAfterMs is a hint), or report the contended key(s) to the user.
Seat timing for needs-test-author items. The queue-phase test-plan note is filled by the
planning seat — the orchestrator (Direct tier) or the plan author (Delegated/Parallel tier),
invoking the test-author skill's scenario-derivation and oracle-derivation sections — and this
happens before advance_item(trigger="start") moves the item queue→work. This is by design:
test-plan gates work entry, the same way any other required queue-phase note does. This is a
distinct seat from Step 4b's test author seat, which writes test code and fills
test-manifest at work phase — see Step 4b's intro for the two-seat distinction.
Planning seat (Parallel tier, one opus agent per stream). Under /task-orchestrator:run-wave
this becomes the run plan's planner-v1 stage — it mirrors the same eight fields below, dispatched
by the run rather than by hand; the field table stays the semantics either way. Dispatch one opus
agent per stream after child items are materialized (post-plan-workflow) and before Step 4's
implementer dispatch — this is the concrete, per-stream instantiation of the planning seat named
above.
Inputs: the stream's diagnosis/task-scope note and the CURRENT codebase, not the plan text
alone. Task: verify every claim in that note against source, citing file:line for each
correction; write the note if materialization left it missing, or revise it if verification
finds it wrong; fill and freeze test-plan per the seat-timing rule above; then return the
structured template below instead of free-form prose. Give the seat the contract by its absolute
path (Step 2's Contract path line) — never a relative plans/..., which does not resolve from the
worktree — so it reads the Items and Conflict rule slots it is about to extend; its return is what
fills the File ownership and Planning seat return template slots for the rest of the wave.
Structured return (eight fields, each required — none is a valid value):
| Field | Semantics |
|---|---|
diagnosis-corrections | What the diagnosis/task-scope got wrong or missed, with current file:line — or none. |
defect-class-siblings | Every other site sharing the root-cause pattern (same cast/fallback idiom, same predicate, same wildcard or character class, same error-classification path, other readers or writers of the field), found by a named sweep (command + scope + hit count). Mark each one frozen as D# or deferred: <reason>, or return none (sweep: <command>) (proposal bb191508). |
cross-stream-file-overlaps | Files this stream must write that another stream in the same wave also touches — drives wave sequencing (serialize the overlap, parallelize the rest) — or none. |
missing-api-or-seam | A surface the fix needs that does not exist yet, with the exact proposed NEW signature — or none. |
test-plan-status | Whether test-plan is filled and frozen (filled (<n> chars)) or still open, and why. |
main-files | The stream's production files (src/main), comma-separated. |
test-files | The stream's NEW test files (author-owned), comma-separated. |
red-proof-shape | Per scenario in test-plan: EXISTING-SURFACE (the test targets code that exists before the fix — revert-the-fix red-proof applies directly) or NEW-SURFACE (the test targets a surface the fix itself introduces). Every NEW-SURFACE scenario needs the narrowest-revert recipe — typically revert only the call sites and keep the new type/parameter — or, when no revert can produce a behavioural red, an explicit substitute: no behavioural red possible, reviewer verifies <X>. |
The fields are the budget. The return may exceed any line budget when the fields have content — trimming a field to fit a target length is the failure this stage exists to prevent, not a virtue.
Evidence. This seat produced a material correction on 15/15 items across three bug waves
(retros 1b8bba5a, cc69671b, 8e6a7cf7) — every wave that ran it found at least one diagnosis
error, cross-stream overlap, or missing seam the plan had missed.
Verification commands. Throughout this step, "run tests" means running BOTH the test suite AND the project linter:
./gradlew :current:test
./gradlew :current:ktlintCheckCI enforces both — a green test run with failing lint will still block the PR.
If ktlintCheck fails, run ./gradlew :current:ktlintFormat to auto-fix
formatting violations, then verify with ktlintCheck again and re-run tests.
Include both commands in every implementation-agent and review-agent prompt.
Who runs the lint cycle (proposal ee6f5d32, ~9 sessions of evidence):
ktlintCheck → ktlintFormat → re-verify cycle itself before committing.
Validated 2026-07-13 (PRs #213/#214/#215): zero orchestrator fix-up commits.:current:ktlintFormat as part of their single pinned
self-check, and the orchestrator additionally runs ktlintFormat before ktlintCheck after
each commit batch (historically 3-5 fix-up commits per multi-phase run when skipped).Capturing gradle's real exit code (use this pattern, not 2>&1 | tail -N).
Piping gradle into tail discards gradle's exit code — tail always exits 0
on a successful read of the log, so BUILD FAILED at the end of the gradle
output is reported by you as a successful run. Combined with gradle's daemon
incremental cache, this can hide broken compilation through entire CI cycles
(retro 568a8584: 9 silently-failing tests shipped on PR #151 before the
followup audit caught it).
The reliable pattern, especially when running in run_in_background:
./gradlew :current:test > /tmp/gradle-out.log 2>&1; EXIT=$?
echo "EXIT=$EXIT"
tail -30 /tmp/gradle-out.logRedirect ordering matters: use > /tmp/log 2>&1, NOT 2>&1 > /tmp/log — the
latter is evaluated left-to-right and leaks stderr (where gradle writes compile
errors) to the terminal instead of the log (retro ac25db89: a compileTestKotlin
failure was invisible in the captured log until the ordering was fixed).
EXIT=$? captures gradle's actual exit code before any pipe consumes it.
Read the captured exit code AND the tail of the log; never trust the tail
alone. This applies to every orchestrator-owned gradle invocation throughout
this step.
Agent-side gradle is not this pattern. In a Parallel-tier wave, agents get exactly one lock-serialized invocation, pinned verbatim — argument form included — in the dispatch contract's Compile self-check slot, together with the rule for what to do when it fails. Do not hand an agent an ad-hoc gradle command line; point it at that slot.
Use --rerun-tasks after dependency upgrades or large refactors. Gradle's
incremental compile cache retains class files from prior good builds. After a
gradle/libs.versions.toml bump or a refactor that changes public API surfaces
(removed methods, renamed types, sealed-class arms, generic-parameter shifts),
incremental compilation can keep the OLD class files alongside source that no
longer compiles, producing apparent BUILD SUCCESSFUL on stale bytecode. Run
./gradlew :current:test --rerun-tasks once after such changes to force a
clean run; ordinary incremental builds are safe afterwards.
This step is tier-conditional:
Direct tier: Implement directly. Edit the files, run the test suite. No subagent
dispatch. No /simplify pass. Exception — a Direct-tier bug-fix carrying
needs-test-author: follow Step 4b's Direct-tier test-first-then-fix sequence
(test-plan frozen at queue → regression test observed red → fix → green) instead of
implementing first. Fill the session-tracking note (required by both
quick-fix and default schemas) with a brief summary of what changed and test
results. Advance to review:
advance_item(trigger="start").
Delegated and Parallel tiers: Use get_context(itemId=...) to see work-phase
expectedNotes and guidancePointer values. Fill each required note following its
guidance. Follow /task-orchestrator:orchestrate (model table, return
formats, UUID inclusion). The key decisions at this step are:
isolation: "worktree" on the Agent tool — that would
spawn a separate worktree per dispatch, which is the deprecated per-child PR pattern.Under /task-orchestrator:run-wave, both bullets above are executed by the run plan (Method A
via Workflow, Method B via next-scheduling), reading each stage's dispatchBySeat entry for
model/agent instead of the table below. The hand-dispatch form in this section and the template it
points at remain the fallback for dispatches a run plan doesn't cover — a one-off fix agent, or an
arbitration re-dispatch (Step 4b) — not the default path for a Parallel-tier wave.
Test-file ownership boundary. Every implementation dispatch (Delegated single agent or
Parallel per-child agent) excludes src/test/** from scope — implementers do not create or
modify test files. When a change surfaces a needed test update, the agent reports it in its
return (or in implementation-notes) rather than editing the test itself; the test author
(Step 4b) owns that file tree exclusively on items carrying needs-test-author.
Parallel dispatch into a shared feature worktree:
Agent(
prompt="""
Working directory: <feature-worktree-path>
Branch (already checked out): feat/<feature-slug>
Dispatch contract: <ABSOLUTE path from the contract's Header "Contract path" line, e.g.
D:\Projects\task-orchestrator\plans\<slug>.md> — read it first; it wins on conflict with
this prompt. It lives in the main checkout, not in the worktree above.
Scope (modify ONLY these files): <explicit list> — the same list as your row in the
contract's File ownership slot.
Do NOT create or modify any file under src/test/** — test authoring is a separate,
independent dispatch (Step 4b) on items carrying needs-test-author. If your change
surfaces a needed test update, report it in your return; never edit the test yourself.
Before returning: run the compile self-check exactly as the contract's "Compile self-check"
slot pins it, and handle its outcome by that slot's rules. Then commit exactly as the
"Commit discipline" slot states, including the post-commit verification.
Do NOT run :current:test or :current:ktlintCheck — the orchestrator owns full build
verification.
""",
model="sonnet", // or dispatch.model when the item's resolved profile sets one — see "Model selection" above
subagent_type="<dispatch.agent when the item's resolved profile names one, else general-purpose>"
// NOTE: no isolation parameter — agents share the feature worktree
)Why the prompt points at slots instead of carrying the commands: the self-check is fast (~3s)
and catches type-mismatch / signature errors that gradle's incremental cache may otherwise mask
in the orchestrator's later test run (retro 568a8584: H3's dbNow() shipped with a
Result<T> vs Instant return type mismatch hidden for ~6 hours), and its ktlintFormat step
is formatting-only and idempotent, preventing the recurring lint fix-up commits (proposal
ee6f5d32). But the invocation's exact argument form, the foreign-file early exit and the
retry bound only work if every agent gets them identically — six agents in one wave each burned
three lock-serialized runs against another agent's mid-refactor file, and a re-typed nested
invocation flattened a [string[]] argument (#302). A prompt that restates the command drifts
from the contract; a prompt that points at the slot cannot. The same applies to the commit step:
never tell an agent to "commit your changes with a descriptive message" in a shared worktree —
the Commit discipline slot is the only safe form (#301).
File-edit overlap discipline: Parallel agents in a shared worktree must operate on non-overlapping files. The orchestrator scopes each agent's prompt to a specific file list (per MEMORY.md §"Parallel File-Edit Delegation"). When inherent overlap exists, dispatch sequentially.
Contract-change sweep discipline. When a child task tightens a contract — making a
parameter required, adding validate() invariants, narrowing a sealed-class arm, or
otherwise rejecting inputs that earlier passed — the orchestrator must sweep the rest
of the codebase before advancing to review. Two recurrences (retros a7f6024f and
568a8584) showed that:
WorkItem.validate() claim-field invariants broke 8+ test
fixtures that constructed mixed-state items via separate Instant.now() calls
(microsecond drift) or partial claim fields.For each contract-tightening change in this run:
grep -rn "<tool>\|<class>\|<method>".:current:test suite (orchestrator-owned, not the agent) to
confirm no fixture-vs-contract conflicts surfaced elsewhere.src/test/**.ac25db89: "case-insensitive"
survived in 3 places after the fix made search case-sensitive).This sweep is part of the orchestrator's verification step between waves, not the implementing agent's responsibility — agents are file-scoped and can't see the full fixture surface.
Model selection — always set model explicitly on every Agent dispatch. Under
/task-orchestrator:run-wave, dispatchBySeat already resolves model per seat and this table is
the fallback for dispatches the run plan doesn't cover:
| Agent purpose | Model |
|---|---|
| Implementation (production code only) | model="sonnet" |
| Independent test authoring (Step 4b) | model="sonnet" |
| Architecture, complex multi-file synthesis | model="opus" |
| MCP bulk ops, materialization | model="haiku" |
Omitting model causes the agent to inherit the orchestrator's model (typically
opus), wasting tokens on sonnet-eligible implementation work.
When dispatching an item's phase owner (implementer on work, reviewer on review), read the profile
for the phase being dispatched INTO: if the orchestrator already called advance_item, its
dispatch field already reports the profile for newRole; if the agent will enter its own phase
(agent-owned-phase protocol — dispatched while the item is still in queue), get_context returns
the queue profile, not work's, so read query_items(operation="schema", itemId=...)'s per-phase
dispatch.work map instead. When that profile names an agent, dispatch with
subagent_type = dispatch.agent. Regardless of whether agent is set, still pass model
explicitly — dispatch.model if set, else the table above. effort has no Agent-tool parameter;
in Claude Code it applies only through the dispatched agent's own frontmatter, so a profile's
effort is advisory unless agent names a definition carrying that effort — to change effort,
point agent at a definition with that effort.
After implementation agents return:
For Parallel-tier features, agents return having committed to feat/<feature-slug>
inside the shared feature worktree. Record each agent's commit SHA range alongside
the child's MCP item ID — this map feeds the ownership two-range check below and the
dispatch contract's Review scoping slot, where reviews are diffed by owned file rather
than by SHA range:
| Child UUID | Agent ID | Pre-SHA | Post-SHA | Test-Pre-SHA | Test-Post-SHA | Changed Files |
|------------|----------|---------|----------|--------------|---------------|---------------|
| <uuid> | <id> | <sha> | <sha> | <sha> | <sha> | <file list> |Capture pre-commit SHA before dispatch (git -C <feature-worktree> rev-parse HEAD)
and post-commit SHA after the agent returns. In a shared worktree this range can
interleave other streams' commits and misses later fix-ups, so use it to confirm ownership,
not to bound a review:
git -C <feature-worktree> diff <pre-sha>..<post-sha> --name-onlyTest-Pre-SHA/Test-Post-SHA are captured the same way around the test author's dispatch
(Step 4b) and stay blank until that wave runs.
Disjointness check (required whenever needs-test-author applies). After both ranges are
captured, verify:
git -C <feature-worktree> diff <pre-sha>..<post-sha> --name-only # impl range
git -C <feature-worktree> diff <test-pre-sha>..<test-post-sha> --name-only # author rangesrc/test/** = ∅ (the implementer touched no test files)src/main/** = ∅ and author-range is non-empty (the author touched only
test files, and touched at least one)A violation is reverted (git -C <feature-worktree> revert the offending commit, or a targeted
git checkout of the crossed-boundary file back to the prior SHA) and recorded in
implementation-notes — do not silently keep a cross-ownership commit.
Applies only to items carrying the needs-test-author trait. Runs after implementation agents
return (their commits exist on the branch/worktree) and before the orchestrator's build
verification below (Direct tier inverts this ordering — see the Direct-tier bullet below) —
the test author compiles against the real implemented surface, not a
predicted one. Two distinct seats, do not conflate them: the planning seat already froze
test-plan at queue phase, before this step, gating work entry (see Step 3's seat-timing rule);
the test author seat here writes test CODE and fills test-manifest — it does not fill or
re-open test-plan.
Per-tier sequencing:
test-plan at queue phase, before any implementation begins;
(b) write the regression test from the diagnosis note's reproduction steps and observe it
red against the pre-fix code, recording the red evidence (assertion, observed value) in
test-manifest;
(c) implement the fix;
(d) confirm the test is green post-fix and record that confirmation too.
test-manifest declares the single actor, and the disjointness check above does not apply (one
actor, one range) — see Step 5's verdict rule for how this tier is scored.ktlintCheck → ktlintFormat → re-verify, same rule as Step 4's "Who runs the lint cycle")
and runs :current:test itself before committing.:current:test or :current:ktlintCheck — only ktlintFormat plus a compile
self-check, mirroring the implementation dispatch template, since the orchestrator owns full
build verification for the wave.Blindness rule. The test author may read: the item's test-plan note, public signatures,
domain models, existing test conventions/docs, and the implementer's changed-file names (not
content). The test author must NOT read: the implementation diff content, the implementer's own
tests, implementation-notes, or session-tracking. This is what makes the separation real
rather than nominal — tests probe the spec, not the implementation's behavior.
Oracle-provenance obligation. Every scenario's expected result must trace to a source
declared in test-plan (spec clause, stated algorithm, external reference) — never "what the
code returns" and never the ticket's own worked example.
Manifest duty. Before returning, the test author fills test-manifest (work, required):
actor id, test file paths, commit SHA range, S-id→test coverage mapping (covered /
not-covered:reason), probes executed with results, forbidden-pattern declaration (every
assumeTrue/escape with justification, or none), and any implementer-modification-to-test-files
rationale (should be none if ownership held).
Skill routing. Invoke the test-author skill before filling test-manifest — it carries the
scenario-derivation, oracle-derivation, blindness, and forbidden-pattern framework this step
depends on. (The planning seat also invokes it earlier, at queue phase, for the
scenario-derivation sections that inform test-plan — see Step 3.)
When a test the author wrote comes back red against the implementation, the test author does not
decide who is at fault (test-author §8/§9) — it reports. The orchestrator arbitrates, with
exactly three dispositions:
src/main/** only.src/test/** only.task-scope/diagnosis first, then
dispatch whichever side the amendment implicates.Record every arbitration — case number, the citation (for case 2), and the resolving commit
SHA — as a bullet in the child's implementation-notes. Never weaken, skip, or @Disable a
test to unblock a wave, regardless of disposition; that is exactly the failure mode this trait
exists to prevent (test-author §7, §9).
This carve-out also modifies the build-verification rule below: a red author-authored test is not a broken build. Route it to arbitration above; if unresolved, hold the child in work and report — do not dispatch a generic "fix agent" against it.
Test-author dispatch prompt: point the agent at the Test author protocol slot of the
wave's dispatch contract, named by the absolute Contract path from its Header (Step 2) and
instantiated from
references/dispatch-contract-template.md — it
states the declarations block, the src/main tool ban, the keys= filter, stop-and-ask, fixture
invariants, surface labels, commit form and manifest fields in full; on a Delegated-tier run with
no contract file, paste that slot into the prompt instead of paraphrasing it. The declarations
block is filled by a separate, non-author declarations-extractor seat, and the orchestrator
scans and strips it of implementation prose before the author sees it — the slot states the
seat's accuracy contract and the scan (proposal 234b50a0). Under /task-orchestrator:run-wave,
the extractor and the test author are both run-plan stages, and the scan step is named
scan-declarations in Method B or the run script's own scanDeclarations call in Method A —
same accuracy contract and strip rule either way.
Capture the test author's pre/post commit SHAs the same way as implementation agents
(Test-Pre-SHA / Test-Post-SHA in the tracking table above), then run the disjointness check.
Build verification (orchestrator-owned, serialized). After each parallel-batch completes (or between sequential children), run from the feature worktree:
git -C <feature-worktree> status # confirm clean tree
./gradlew -p <feature-worktree> :current:test
./gradlew -p <feature-worktree> :current:ktlintCheckA failure means a recently-committed child broke something. Dispatch a fix agent
(same shared worktree) before continuing. Do not advance any child to review
until the build is green — the trend memory has multiple sessions of
flaky-test-hides-real-bug showing why retry-until-green is wrong.
Carve-out: a red author-authored test is not a build failure. On items carrying
needs-test-author, a test failure traced to a test the test author wrote (Step 4b) is not
"the build broke" — do not dispatch a generic fix agent against it. Route it to the
Arbitration subsection in Step 4b instead. If arbitration doesn't resolve before this
verification pass needs to conclude, hold the child in work phase and report the unresolved
red test rather than forcing green through a fix-agent dispatch.
Why orchestrator owns gradle invocations: ./gradlew runs against a single
Gradle daemon and a single build/ cache per project directory. Parallel
gradlew test invocations against the shared feature worktree will queue at the
daemon, corrupt the build cache, or hit Windows file locks. Serializing build
verification at the orchestrator prevents this without slowing the agents (they're
not running gradle).
For Delegated tier (single subagent), the agent commits to the working branch
on the main directory. Capture the changed files via
git diff main --name-only and proceed.
Post-implementation steps (run in the feature worktree for Parallel tier, or on the working branch for Direct/Delegated):
/simplify skill on the changed code to check for reuse, quality, and
efficiency — this is a cleanup pass before review, not a review itself/simplify made changes, the resulting test coverage work follows the
ownership boundary above. On items carrying needs-test-author, re-dispatch the
test author to write or update the covering tests (Step 4b) — the implementer/orchestrator
does not touch src/test/** even for a simplify-driven update. On items without the trait,
write or update tests inline as before. Either way, the simplify pass is still part of the
work phase — all code changes require test coverage before advancing to review./simplify or during
implementation that are not immediately addressed (pre-existing tech debt,
optimization opportunities, related bugs) must be logged via
/task-orchestrator:create-item before moving on. Do not discard findings.c068c943) — written here, in the work phase, so the
reviewer sees it. A user-visible change — server behaviour, the MCP or REST surface, config
keys, plugin skills or hooks — adds ONE bullet under CHANGELOG.md's ## [Unreleased]
section, in its ### Added / ### Changed / ### Fixed subsection (plugin skill and hook
changes go under ### Plugin). The bullet describes behaviour, not the diff, and carries no
volatile counts. A docs-, process- or chore-only change adds none; its PR body will state
Changelog: none (<why>). Direct/Delegated: committed with the change. Parallel: the
orchestrator or docs seat writes it once, after the implementation wave and before review —
never an implementer (the dispatch contract's Docs slot). /prepare-release Step 8c folds
these bullets into the release section.guidancePointer — focus on context
that downstream agents need to knowadvance_item(trigger="start") to advance
the item to the next phase.advance_item(trigger="start")
itself after filling work-phase notes.
In both cases, inspect newRole in the response to determine what comes next
(see Step 5).Before dispatching or performing review, check the item's current role. Inspect
newRole from the advance_item response in the previous step:
newRole is terminal: The item's schema has no review phase (lightweight
lifecycle). Review dispatch is not needed — the item completed through its natural
lifecycle. Proceed to Step 6. Note: feature-task items skip review by default
(work→terminal directly) — the needs-task-review trait re-enables a review phase
for a specific child when needed.newRole is review: Continue with review per tier below.This step is tier-conditional:
Direct tier: Perform an inline review. Read the diff, verify correctness, confirm
tests pass. If the review-phase note has a skillPointer (visible in the advance_item
response or via get_context), invoke that skill for the evaluation framework before
filling the review note. Write the review note, then advance to terminal:
advance_item(trigger="start"). No separate review agent —
the overhead exceeds the risk for 1-2 file changes with known fixes.
Delegated and Parallel tiers: Dispatch a separate review agent. The agent that implemented the code must not review its own work.
Reviewer scoping for Parallel waves. Per-child reviewers remain the default for code-bearing
waves. When every child in the wave is content-only (config, skill, or doc edits — no src/main
or src/test changes) AND the wave shares a pinned plan-file contract (Step 2), dispatch ONE
consolidated opus reviewer over the full feature-branch diff (git diff main...<FEATURE_BRANCH>)
instead of N per-child reviewers. Cross-file coherence defects — sibling skills stating
contradictory rules, gate placement contradicting seat timing, an example contradicting its own
rule — are visible only to a reviewer holding the whole diff; per-child reviewers structurally
cannot see them, and a shared naming contract does not prevent them, since naming consistency
and semantic coherence are orthogonal. Evidence: 5 blocking cross-file contradictions caught
this way at 5-item scale (retro 6d562acb) and canonical-config consistency issues at 7-item
scale (retro bd4ec109) — trend 028c7b5d. The consolidated reviewer fills the review-phase
notes of whichever items carry a review phase (typically the parent feature's
review-checklist, since feature-task children skip review by default).
If the implementation used the shared feature worktree (Parallel tier), the
review agent operates in that worktree, scoped to that child's owned files — the diff of
those paths from the wave's base SHA to current HEAD, per the contract's Review scoping
slot. Not the <pre-sha>..<post-sha> range captured during dispatch: in a shared worktree that
range interleaves other streams' commits and excludes any later fix-up, formatter or
fixture-repair commit to the same files, so it both over- and under-reports. The pre/post map
still serves the ownership two-range check (list item 6 below).
Review agent template (copy verbatim, fill placeholders):
You are reviewing one child task within a shared feature worktree.
- Feature worktree: <FEATURE_WORKTREE_PATH>
- Feature branch: <FEATURE_BRANCH>
- Dispatch contract: <ABSOLUTE path from its Header "Contract path" line, e.g.
D:\Projects\task-orchestrator\plans\<slug>.md> — read it first; it wins on conflict
with this prompt. It lives in the main checkout, not in the worktree above.
- This child's owned files (its row in the contract's File ownership slot):
<FILE LIST>
- Your review scope is the diff of exactly those files:
git -C <FEATURE_WORKTREE_PATH> diff <BASE_SHA>..HEAD -- <FILE LIST>
- Other children committed into this same branch, and later fix-up, formatter or
fixture-repair commits may touch these files. Review the owned-file diff above as it
stands now; do NOT review another child's files, and do NOT bound your review by a
commit range.
Run ALL commands from within the feature worktree.
Read ALL files from that directory. Do NOT read from the main working directory.
Tests have already been verified green by the orchestrator after the most recent
commit batch. Run no gradle — the contract's Compile self-check slot states who runs
it, and the Review scoping slot records the build state you rely on. Focus on plan
alignment, test quality, and simplification per the review-quality skill.
Report every finding at every severity — do not self-filter to a high-severity-only
bar. Mark each finding blocking or observation and state your confidence; the
orchestrator's verdict handling is the downstream filter, not your own judgment.If using Direct or Delegated tier (single working branch on the main directory), the review agent reads from the working branch and runs tests itself per the existing template — no worktree-specific scoping needed.
The review agent:
get_context(itemId=...) to load the item's notes and review-phase requirementsreview-quality Area 1's "run the test suite first", whose recorded numbers the reviewer
takes from that slot instead. On Direct/Delegated tier there is no
contract and no orchestrator-run wave verification, so the reviewer runs the test suite AND
the linter itself (Step 4's "Verification commands") — a PR with failing lint will not merge.needs-test-author, additionally:git diff <impl-range> --name-only touches
no src/test/** path and git diff <test-range> --name-only touches no src/main/** path.
N/A-by-declaration on Direct-tier temporal-only items (Step 1, Step 4b): there is one
actor and one range by design, so there is no second range to diff — record it as
N/A: temporal-only, single actor rather than attempting the diff, and it cannot fail.test-plan was filled
and frozen at queue phase, before implementation began (Step 3's seat-timing rule), and —
for a bug-fix — confirm the red-first evidence (the regression test observed failing
against pre-fix code, per Step 4b/B3) is recorded in test-manifest.implementation-notes
with the required detail (spec citation for a "test wrong" call: note key + quoted
criterion + asserted-vs-observed).assumeTrue/@Disabled scan on author-owned files: grep the test author's changed
files for assumeTrue, @Disabled, or other weakening introduced after the author's
first commit in this range. Any such introduction is an automatic blocking finding —
independence does not permit softening a red test to unblock a wave.
Verdict rule: test-independence-audit fails to not-independent (blocking) if any
applicable check above fails, regardless of how green the test suite is — the ownership
two-range check is not applicable, and cannot fail, on a Direct-tier temporal-only item.
On a Direct-tier temporal-only item where every applicable check (temporal checks,
arbitration-record, assumeTrue/@Disabled scan) passes, the verdict is
independent-degraded (temporal-only) — never plain independent, matching test-author §11.guidancePointer with a verdictHandling the verdict:
| Verdict | Action |
|---|---|
| Pass | Proceed to Step 6 |
| Pass with observations | Proceed to Step 6; log observations for follow-up |
| Fail — blocking issues | Stop and report to the user with the full findings. Do not attempt to fix autonomously — bring the human into the loop. |
Review failures surface issues that may indicate systemic problems worth learning from. Automatically retrying hides these signals.
The shape of Step 6 depends on tier.
After review passes:
Post-dispatch commit audit (non-blocking): run git log --oneline -3 and
git status --short and compare against what the orchestrator itself committed.
Any commit a subagent made despite stop-boundary instructions is flagged for
review here — inspect its scope before it rides into the squash-merge (a
subagent commit lacks the co-author trailer; reconcile at merge time). On items
carrying needs-test-author, additionally check each commit's file list for
cross-ownership: an implementer commit touching src/test/**, or a test-author
commit touching src/main/**, is flagged here even if it slipped past the
Step 4b disjointness check — do not let it ride into the squash-merge unreconciled.
Verify the working branch is committed (orchestrator commits if Direct tier; subagent committed if Delegated). Stage only the files related to the implementation:
git add <specific-files>
git commit -m "$(cat <<'EOF'
<type>(<scope>): <description>
<body — what changed and why, referencing the MCP item>
Co-Authored-By: Claude <noreply@anthropic.com>
EOF
)"Commit types: feat for features, fix for bugs, refactor for tech debt,
perf for performance, test for test-only changes, chore for maintenance.
CHANGELOG check. Confirm the branch carries the [Unreleased] bullet written in Step 4c
(post-implementation step 4), or that the PR body will state Changelog: none (<why>).
Push the working branch:
git push -u origin <branch-name>Create the PR:
gh pr create --base main --title "<type>(<scope>): <description>" --body "$(cat <<'EOF'
## Summary
<2-4 bullets>
## Test Results
<test count, pass/fail, new tests>
## Review
<verdict summary>
## Changelog
<the [Unreleased] bullet text, or none (<why>)>
## MCP
<item ID>
EOF
)"After PR merges:
git checkout main
git pull origin main
git branch -D <branch-name>Advance the item to terminal:
advance_item(transitions=[{ itemId: "<uuid>", trigger: "start" }])After the item reaches terminal, follow the retrospective hook's directive if one fires (see retrospective.mode).
Report the PR URL and a summary.
For Parallel-tier features with a shared feature worktree:
For each child task (after its review passes, if it has one):
needs-task-review),
confirm review-checklist is filled. Children without one advance work→terminal
directly — there is nothing to confirm.advance_item(itemId=<child-uuid>, trigger="start") to move work→review→terminal
(or work→terminal directly) as the child's schema dictates.feat/<feature-slug>
inside the shared worktree; that's the integration point.When all children reach terminal, the parent feature is ready to finalize:
implementation-notes and session-tracking notes (aggregating
across children — distributed-tracking pattern works as today)../gradlew -p <feature-worktree-path> :current:test
./gradlew -p <feature-worktree-path> :current:ktlintCheckreview-checklist (orchestrator-authored,
summarizing across all children's reviews).Changelog: none (<why>) in the PR body.git -C <feature-worktree-path> push -u origin feat/<feature-slug>gh pr create --base main --title "feat(<scope>): <feature description>" --body "$(cat <<'EOF'
## Summary
<feature-level summary aggregating all children>
## Children completed
- <child-1 title> (<MCP UUID>)
- <child-2 title> (<MCP UUID>)
...
## Test Results
<total test count, new tests added across feature>
## Review
<feature-level review verdict, references each child's review-checklist>
## Changelog
<the [Unreleased] bullet text, or none (<why>)>
## MCP
Parent: <parent UUID>
Children: <list of child UUIDs>
EOF
)"git checkout main
git pull origin main
git worktree remove <feature-worktree-path>
git branch -D feat/<feature-slug>dispatch mode (see retrospective.mode in .taskorchestrator/config.yaml) it directs
a background /session-retrospective automatically — follow its directive; in nudge mode,
or if no directive arrives, suggest running it manually.session-tracking notes across
children plus parent-level aggregation gives clean retro input (validated by retro
a7f6024f).Local main always tracks origin/main — no divergence, no reset --hard needed.
When processing a Parallel-tier feature with multiple child tasks autonomously: run
/task-orchestrator:run-wave over the parent, one run per topological layer — steps 1-2 below
describe what that run does under the hood; they are the fallback for a manual walk-through when a
run plan doesn't apply.
feat/<slug>) at planning time. All children share it.isolation: "worktree". Independent children dispatch in parallel waves; dependent
children dispatch sequentially. Orchestrator scopes each agent's file list to
prevent overlap.needs-test-author, dispatch one
test-author agent per child as its own wave, after that child's implementation wave and
before build verification. Authors run ktlintFormat + compile self-check only (never
:current:test/:current:ktlintCheck — the orchestrator owns those next). Capture
Test-Pre-SHA/Test-Post-SHA per child and run the disjointness check before proceeding.:current:test and
:current:ktlintCheck from the feature worktree. Fix failures before advancing
any child to review.feature-task children skip review by default (work→terminal
directly) unless the needs-task-review trait is set, or needs-test-author is set
(its test-independence-audit note adds a review phase). When a review phase applies,
the review agent reads from the shared worktree and diffs that child's owned files from the
wave's base SHA to HEAD (git -C <feature-worktree> diff <base-sha>..HEAD -- <files>),
per the contract's Review scoping slot — never a commit range, which in a shared worktree
picks up other streams and misses later fix-ups. The pre/post and test-pre/post SHAs feed the
ownership two-range check only.feat/<slug> and open the single feature-level PR.If any child hits a review failure, continue processing siblings (their commits are already in the feature branch). Report all failures together at the end. The orchestrator decides whether the failed child blocks parent finalization (e.g. fix-and-re-review), can be cancelled (descope from the feature), or warrants reverting its commits.
Bug-fix batches (multiple unrelated fixes) — these are NOT a Parallel-tier feature. Use Delegated tier per item, each with its own branch and PR (the legacy per-item flow). The shared-worktree pattern applies only when items share a parent feature item.
For full setup, dispatch patterns, lifecycle, parallel validation, and test baseline management, see WORKTREE.md.
Quick reference:
| Tier | Worktree | Branch | PR scope |
|---|---|---|---|
| Direct / Delegated (single item) | None — work on main directory | <type>/<slug> | One PR per item |
| Parallel (parent feature with N children) | One shared feature worktree | feat/<feature-slug> (one branch for all children) | One PR at parent finalization |
Do NOT use isolation: "worktree" on the Agent tool for Parallel-tier child dispatches.
That spawns a separate worktree per dispatch — the deprecated per-child PR pattern.
For Parallel tier, the orchestrator pre-creates one shared worktree in Step 2 and
dispatches each child agent into it.
When NOT to create any worktree:
Tier classification happens at Step 1 even when resuming. Classify the tier from the item's tags, file scope, and note state, then resume using that tier's pipeline.
If an item is already past the queue phase (e.g., previously planned but not implemented), the skill picks up from the current state:
| Current role | Resume from |
|---|---|
| queue (notes filled) | Step 3 — advance and proceed |
| queue (notes missing) | Step 3 — fill missing notes |
| work (in progress) | Step 4 — check implementation state |
| work (notes filled) | Step 4 — advance to review |
| review | Step 5 — run review |
| terminal | Already done — report status |
open run/<runId>/state found (F7) | /task-orchestrator:run-wave --resume <runId> |
Always call get_context(itemId=...) first to determine exact state before
resuming.
© jpicklyk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in .claude/skills/implement of jpicklyk/task-orchestrator.
Open the folder on GitHubat commit b688ea0
Implement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Implement this skilljpicklyk/task-orchestrator | 207 | — | ~17k | Automated safety check: Pass | MIT | |
| Devcontainer Devstacklok/toolhive-studio | 170 | — | ~3.8k | Automated safety check: Notes | Apache-2.0 | |
| Zed Configwcygan/dotfiles | 196 | — | ~930 | Automated safety check: Pass | None | |
| Agent Deckasheshgoplani/agent-deck | 1k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Foremergenaw103/foremerge | 536 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Puppetmaster Agent Orchestrationprofessorpalmer/Puppetmaster | 467 | — | ~3.2k | Automated safety check: Pass | MIT |
stacklok/toolhive-studio
Spin up and interact with ToolHive Studio's containerized dev environment (Xvfb + noVNC + DinD).
wcygan/dotfiles
Zed editor configuration expert. An agent skill from wcygan/dotfiles.
asheshgoplani/agent-deck
agent-deck, the terminal session manager for AI coding agents.
naw103/foremerge
Coordinate parallel coding agents with Foremerge's local Git-compatible CLI and MCP server.
professorpalmer/Puppetmaster
Operates and supervises Puppetmaster, a multi-agent orchestrator, through its MCP tools or CLI, picking the right verb for edits, reviews, audits and long-running jobs.
Th0rgal/sandboxed.sh
Reads and updates the Sandboxed.sh Library, a git-backed store of skills, agents, commands, tools, rules and MCP servers, through its library tools.
jpicklyk/task-orchestrator
Walks through how to launch and reach the MCP Task Orchestrator server container: transport, REST API, port publishing, config mounts and config-sync.
jpicklyk/task-orchestrator
Resolves ready MCP work items into a run plan, shows it to you, then executes it through the Workflow tool or direct subagent dispatch, with post-run verification.
jpicklyk/task-orchestrator
Migrates an existing unscoped Task Orchestrator database to the project-scoping convention in place, creating one project anchor root and re-parenting work trees under it after a mandatory dry run.
jpicklyk/task-orchestrator
Completes or cancels a whole feature subtree, a named list of items, or a batch of stale work items at once, previewing the impact and warning before force-completing anything active.
jpicklyk/task-orchestrator
Creates an MCP work item from conversation context, anchoring it under the right container, inferring type and priority and pre-filling the required notes.
jpicklyk/task-orchestrator
Views, creates, deletes and diagnoses BLOCKS, IS_BLOCKED_BY and RELATES_TO links between MCP work items, including why an item cannot start.
Works with
Categories
End-to-end workflow for taking MCP work items from backlog to merged PR. Implement is an agent skill from jpicklyk/task-orchestrator. End-to-end workflow for taking MCP work items from backlog to merged PR.
Implement fits situations like: A user says implement this; work on this item; pick up the next task; create a PR for this.
Run `npx skills add jpicklyk/task-orchestrator --skill implement -a claude-code`. Or copy the skill folder (.claude/skills/implement in jpicklyk/task-orchestrator) into .claude/skills/implement in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jpicklyk/task-orchestrator --skill implement -a codex`. Or copy the skill folder (.claude/skills/implement in jpicklyk/task-orchestrator) into .agents/skills/implement in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jpicklyk/task-orchestrator --skill implement -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/implement, .gemini/skills/implement, .github/skills/implement and .opencode/skills/implement in your project.
Going by SKILL.md and its folder, Implement needs Python for the scripts in its folder and the command-line tools its instructions call (git and gh). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Implement is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 17k tokens (SKILL.md is roughly 67k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Implement: Devcontainer Dev (stacklok/toolhive-studio, 170 stars), Zed Config (wcygan/dotfiles, 196 stars), Agent Deck (asheshgoplani/agent-deck, 1k stars) and Foremerge (naw103/foremerge, 536 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jpicklyk (a GitHub user) maintains it in jpicklyk/task-orchestrator, which has 207 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 6, 2026.
Source: jpicklyk/task-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.