user-story
A guided path from "build me X" to integrated, verified work on main. Two ideas run through all of
it:
The orchestrator holds the plan; executors hold implementation. Before dispatch, the orchestrator decides
what improvement is worth making, its concrete benefit, the approach, and boundaries. Researchers return facts
and bounded alternatives for that judgment. Executors use ordinary local coding judgment to implement the
selected plan, then return concrete evidence when it cannot work. They do not silently replace the agenda,
architecture, or scope.
The ceremony scales to the evidence. A deterministic done-oracle is enough when it covers the
contract. Add independent review only for a named untested seam or safety-sensitive risk.
Board mechanics live in the sidequest skill (claim lifecycle, dispatch fields, routing, publish
transaction) and its references/. Read ../sidequest/references/orchestration.md before a first
wave. This skill is the sequence and the sizing, not a second copy of those rules.
Size it first
Sizing is a call you make and state in a line, not a question you ask the user. Read these signals:
- Independently checkable pieces: one, a handful, or many.
- Surfaces crossed: one module, or store plus CLI plus MCP plus UI plus hooks plus docs.
- Reversibility: a schema, migration, on-disk format, wire format, or public API is expensive to
undo. Internal helpers are not.
- Is the approach contested? More than one defensible architecture, with real trade-offs between
them, is the single strongest reason to run a design panel.
- Blast radius: does existing behavior change for people already using this?
- Stakes: data loss, auth, money, the release path itself.
Three sizes, and what each dial does at each:
Escalate mid-flight when evidence says so. A spike that comes back reporting the seam is worse than
expected raises the size; write that in the story log so the jump is on the record and not a vibe.
De-escalating is equally fine: if the panel converges on one obvious answer, stop paying for the
panel.
Planning checkpoint
Before dispatching a substantial or ambiguous feature, put one visible, pinned contract on the story,
a planning ticket, or the ticket descriptions. It is the handoff from planning to execution, not a
second design process. For substantial or safety-sensitive work, first check architecture and feasibility
before expensive implementation or tests: use the preflight in
../sidequest/references/ticket-authoring.md. Consume evidence from the project's configured quality gate, if any, through its dedicated owner,
check genuine native baseline/candidate ownership and the oracle's real deadline/resource fit. Small
deterministic fixes keep one owner and a focused check; no mandatory planning panel. Pin:
- Outcome and explicit non-goals, so later work has a boundary to cut against.
- Smallest authority and intervention, the actual call flow and source of truth that decide behavior.
Question whether a change is needed; reuse local code, stdlib, native features, or installed dependencies;
make one minimal shared-root fix. Prefer measured deletion over hypothetical guards, forced extractions,
and unrelated cleanup. Preserve validation at trust boundaries, data-loss prevention, accessibility,
permissions, and immutable candidate/review authority.
- Surgical boundaries, the files each piece may change and every public surface it may expose or
deliberately leave alone.
- A bounded executable done-oracle per piece, the behavior it observes, and the named consumer or
regression input that would fail if the piece were wrong.
- Review budget, from the sizing table, including the exact reason an oracle alone is enough or the
lens a review ticket must cover.
- Design-reopen evidence, concrete findings that would invalidate the contract, such as a consumer
that cannot use the pinned seam, a required compatibility break, or an oracle that cannot observe the
claimed outcome.
Exact small work keeps its lightweight path: state the outcome and oracle and dispatch one ticket. A
pinned contract does not earn proposal theater. Write one authored contract for settled work; use two
or three bounded proposals only when the approach is genuinely contested.
The shape
- Frame the outcome, check this flow applies, and state the size.
- List the unknowns, then recon: sub-agents resolve what the code can answer.
- One question round for what they could not, or none.
- Design: write the contract, or run a panel and merge the winner.
- Story plus the complete backlog for every planned wave, filed before anything dispatches.
- Dispatch each ready wave in full, keep routine supervision quiet, allow declared live advice, and re-plan between waves.
- Review at the sized depth, integrate by oracle, publish, close out.
Track these with the task tools so the user can see where the feature is.
1. Frame it
Say back the outcome in a line or two, name the surfaces it touches, and state the size with the
signal that drove it ("multi-wave: this changes the on-disk format, so the migration has to land
before anything reads it"). That framing is what every later step is cut against, so a vague frame
produces vague tickets.
Drop out of this flow when it does not fit, and say so plainly:
- Operational asks (run the build, start the dev server, open the dashboard, answer from what is
already on screen): just do them.
- A trivial edit to one or two files the user named, with no investigation: edit inline.
- A single bug with a known cause: that is one ticket, not a user-story flow.
2. Unknowns first, then recon, bounded
Before any reading, write down what the request leaves unclear: the outcome itself, which existing
behavior it replaces, who consumes the result, what "done" looks like, and anything the user said in a
way that has two readings. Then sort each unknown by who can answer it. The codebase, the git history,
or an external source answers most of them; only the user answers the rest. That sort decides the next
two steps: unknowns the code can answer become exploration tickets, and unknowns only the user can
answer become the question round. Do not guess at either kind and call the guess a contract.
Recon answers exactly one question: what does the contract need to say? Stop the moment you can
write the shared interfaces, the file boundaries, and a verify command per piece.
Launch one read-only sub-agent per code-answerable unknown, in parallel, and let them run while you
write the contract skeleton. In the orchestrator, bounded recon may Read, Glob, or Grep named anchors and make one narrow location sweep. Route unfamiliar path tracing, deep investigation, or multi-angle research through the live taxonomy: use a read-only codebase-exploration ticket for repository behavior, source-lookup for a bounded external question, or evidence-research when sources need reconciliation. Give each ticket one distinct angle and run independent investigations in parallel. At
multi-wave size the angles that earn their keep are the closest existing feature traced end to end,
the extension point and who else depends on it, and the convention plus test pattern to match.
Give an exploration ticket a deliverable, not a topic. It should come back with anchors as
file:line, the seam to extend, the convention to copy, the existing verify command, and the
consumers that would break. Findings return as compressed comments of roughly one to two thousand
tokens.
Then build the contract from those findings. Do not go read every file they named. That habit is what
makes the expensive loop expensive, and it buys nothing the finding did not already carry. When a
finding is too thin to write a contract from, the fix is a sharper deliverable on the next exploration
ticket, not the orchestrator crawling the tree itself.
3. One question round, or none
The question round is what is left of the unknowns list after the sub-agents report: the ambiguities
that neither the code nor the sources could settle. Ask them together in a single AskUserQuestion
(up to four), each carrying what the investigation found so the user decides from evidence instead of
from the same uncertainty you started with. Ask when the answer changes user-visible behavior,
compatibility, migration, public API, dependencies, or expensive-to-reverse scope, and ask when the
request itself is unclear enough that two careful readers would build different things. One round
respects the user's attention and gets better answers, because they see the whole shape of the decision
at once. Asking one question per ambiguity trains them to stop reading. Asking before the sub-agents
report wastes the round on things the code would have answered.
Worth asking: a user-visible behavior with two defensible answers, a scope boundary that changes how
much gets built, a compatibility break, a data migration, anything else expensive to reverse.
Not worth asking: naming, file layout, test placement, error copy, ordering, and every other call you
can make and change later. State the assumption in the contract and keep going. Explicit phrases such
as "do your thing", "use your judgment", or "whatever you think" delegate the current feature: pick,
record the pick, and move on. They do not create a durable standing preference.
This is where the generic flow is deliberately narrowed. It waits for answers before designing;
Sidequest defaults autonomous, because a stalled feature costs the user more than a decision they can
correct at review. That default covers calls you can make and change later. It never covers an unclear
outcome: a contract guessed from a request nobody understood is a full wave of work in the wrong
direction, and one question is cheaper than that.
4. Design: contract, or a panel
The output of this step is always the same artifact: one contract. What changes with size is how
much you spend arriving at it.
When the approach is settled, write it. Three proposals for a foregone conclusion is three tickets
of cost for an answer you already had.
When the approach is genuinely contested and hard to reverse, run a design panel. Two or three
spike-investigation tickets, readonly: true, same problem, deliberately different mandates:
- Smallest change: maximum reuse of what already exists, least new surface.
- Cleanest seams: what you would build if this area took three more features after this one.
- Risk-first: what breaks, what is hard to reverse, what the migration and rollback actually cost.
They run in parallel, on category-appropriate routes, with a bounded deliverable: the seam it
introduces, the signatures it pins, file boundaries per piece, what it costs, and what it forecloses.
Cap each proposal at roughly two thousand tokens. Three uncapped design docs landing in the
orchestrator's context is the failure mode the panel is supposed to avoid.
Then do the orchestrator's planning work: judge the plans and merge them.
Pick a spine, graft the parts of the runners-up that are better than the winner's version, and write
one contract out of the result. A panel whose output is "we went with proposal B" wasted the other
two; the point is that the merged contract beats every individual proposal. Record in the story log
why the losers lost, because the next session will otherwise re-propose them.
Whatever the route, a contract that makes fan-out safe pins:
- Shared surface: the types, interfaces, and function signatures pieces hand each other, written
out. This is the whole reason parallel executors compose instead of producing five conflicting
interpretations of the same seam.
- File boundaries per piece, the blast radius each piece may touch, and any committed build output. Content-hashed output gets one rebuild ticket per wave.
- Dependency order, so
ready partitions the backlog into waves by itself.
- The exact scoped verify command per piece, runnable and deterministic. The integrator runs the full merged-tree gate once per wave.
- Shared runtime resources: fixed ports, servers, databases, fixture paths. Worktrees isolate
files, not runtime. Name the shared resource and current holder; serialize commands using it,
not entire tickets. Non-owners continue independent read/edit/commit work, record readiness, and
end the turn retaining their claim until the parent explicitly hands off. Heavy commands use
at most two workers, finite owned deadlines, and descendant cleanup, within the two-core heavy budget.
Present the chosen approach and its main trade-off in a few lines. Ask for approval only when the
choice is expensive to reverse: a schema or migration, a public API, a user-visible default, a new
dependency. Otherwise state what you picked and proceed.
Cannot pin a multi-item contract at all? File concurrent read-only investigation tickets, one
investigation ticket per independent item, then pin a separate fix wave from their compressed findings.
A shared runtime serializes writes and live reproduction, never read-only investigation; one combined
ticket is only for exactly one item or a provably single defect. Claiming a one-item contract is
unpinnable needs either a completed planning ticket that names the interface that resisted, or no
written contract surface in the request. "Feels coupled" is a reason to file planning first.