Test Scenarios
phuryn/pm-skills
Create comprehensive test scenarios from user stories with test objectives, starting conditions, user roles, step-by-step actions, and expected outcomes.
Test authoring framework for items carrying the needs-test-author trait.
$ npx skills add jpicklyk/task-orchestrator --skill test-author -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jpicklyk/task-orchestrator test-author --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-author .claude/skills/test-author && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "test-author" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/test-author into .claude/skills/test-author/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-author", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/test-authorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jpicklyk/task-orchestrator --skill test-author -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jpicklyk/task-orchestrator test-author --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/test-author .agents/skills/test-author && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "test-author" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/test-author into .agents/skills/test-author/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-author", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jpicklyk/task-orchestrator --skill test-author -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jpicklyk/task-orchestrator test-author --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/test-author .cursor/skills/test-author && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "test-author" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/test-author into .cursor/skills/test-author/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-author", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jpicklyk/task-orchestrator.git --path .claude/skills/test-author--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jpicklyk/task-orchestrator --skill test-author -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jpicklyk/task-orchestrator test-author --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/test-author .gemini/skills/test-author && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "test-author" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/test-author into .gemini/skills/test-author/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-author", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jpicklyk/task-orchestrator test-authorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jpicklyk/task-orchestrator --skill test-author -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/test-author .github/skills/test-author && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "test-author" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/test-author into .github/skills/test-author/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-author", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jpicklyk/task-orchestrator --skill test-author -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jpicklyk/task-orchestrator test-author --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/test-author .opencode/skills/test-author && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "test-author" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/test-author into .opencode/skills/test-author/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-author", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
test-authorTest authoring framework for items carrying the needs-test-author trait.
Test Author is an agent skill from jpicklyk/task-orchestrator. Test authoring framework for items carrying the needs-test-author trait. Defines scenario derivation from acceptance criteria, the oracle-derivation and blindness rules that keep test authorship independent of implementation, the adversarial probe catalog, forbidden patterns, and the test-plan/test-manifest note formats. Referenced by trait note guidance during queue-phase test-plan and work-phase test-manifest filling. Use when filling test-plan or test-manifest notes, or when asked to author or review tests…
Its SKILL.md is about 8.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/forbidden-patterns.md`).
It sits in Testing & QA, covering Test generation and User stories. The repository describes itself as: Server-enforced workflow discipline for AI agents. An MCP server providing persistent work items, dependency graphs, quality gates, and actor attribution. Schemas define what… The licence is MIT.
11 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3e83170. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Test Author loads about 8.2k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 141 tokens; SKILL.md has 4,776 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jpicklyk/task-orchestrator at commit 3e83170, republished under its MIT licence (© jpicklyk). 4,776 words, ~8,229 tokens.
.claude/skills/test-author/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.This skill defines how test authorship is separated from implementation. It exists because an
agent that writes both the code and its tests has no adversary in the loop — the tests confirm
what was built, not what was intended. The trend record backing this trait names the recurring
cost directly: vacuous positive assertions with || isEmpty() escapes, a test oracle computed
from the implementation's own formula, assumeTrue wrapped around three real production bugs so
they never turned the suite red, and coverage claimed in notes that didn't exist in the test
tree. review-quality already names the bias — "the agent that wrote the tests has an inherent
bias toward believing they're correct" — this skill is the structural fix upstream of review:
an independent authorship step, gated at the point where an oracle can still be frozen before
implementation exists to copy from.
Everything below applies whether the test author is a separate dispatch or, in a degraded mode, the same agent operating at a different, declared point in time. The separation is the point — follow it exactly even when it feels redundant with work you can already see.
When query_rules(get, rootId, key:"test-author") succeeds, the served text is authoritative for
§3–§10; this skill is the fallback and the Claude-specific addendum.
The trait needs-test-author puts two seats on an item, occupied at different phases:
test-plan — scenarios, oracle sources, and probes,
frozen before any implementation exists. This is the structural lever: an oracle written before
code exists cannot be derived from that code, no matter who writes it later.test-manifest — the actual test files, written
after implementation has landed so they compile against the real signatures, but without
reading the implementation's reasoning or its own tests.Same-agent exception (Direct tier, temporal-only): for a single-item Direct-tier dispatch
there is no second agent to send the work to. The same agent may occupy both seats, but only
under a temporal-only degraded mode — see §11. The gate still applies: test-plan must exist
and be frozen before implementation begins, and the manifest must declare the single-actor mode
explicitly rather than silently reusing the plan author's context. This is a narrower guarantee
than true two-agent separation, and test-independence-audit records it as such
(independent-degraded), never as independent.
Outside Direct tier — Delegated and Parallel — the two seats are two dispatches. Never collapse them to save a round-trip; the whole value of this trait is in the second agent's blindness.
Start from the acceptance criteria in the item's planning note (task-scope, feature-summary,
or diagnosis) or the queue-phase spec these criteria live in. Every criterion maps to at least
one scenario; a criterion with zero scenarios is a coverage gap, not an implicit pass.
Partition each criterion into happy / failure / edge, mirroring the spec-quality Test
Strategy discipline this note composes with — test-plan gives each scenario an oracle and a
probe list; task-scope's Test Strategy section should already have named the scenarios
themselves, so this step should feel like formalizing, not inventing from scratch. If it feels
like inventing from scratch, the queue-phase spec's test strategy was too thin — flag it rather
than filling the gap silently.
A criterion that only yields a happy-path scenario has not been fully partitioned — go back and ask what could violate it or sit at its boundary before treating it as covered.
Stable S-ids. Number scenarios S1, S2, S3… in the test-plan note and never renumber once
written test code refers to them — the test-manifest's S-id→test mapping is only auditable if
the ids are stable across the plan/manifest boundary. If a scenario is dropped, mark it
S4 — dropped: <reason> rather than closing the numbering gap.
Label every scenario EXISTING-SURFACE or NEW-SURFACE. Red-proof is obtained by reverting
the fix and running the item's tests. That works only where the scenario binds to a surface that
already existed: when the fix introduces the type, parameter, seam or enum constant a test
references, reverting it produces a compile failure rather than a behavioral red. Compile-red
proves the tests reference new code; it does not prove they detect wrong behavior. Across two bug
waves this hit 3 of 10 items and then 2 of 5 — and in the second wave the two affected items were
the security fix and the silent-exit fix, the two whose defects mattered most.
EXISTING-SURFACE — every declaration the scenario touches exists before the fix. A plain
revert yields behavioral red; no extra field is needed.NEW-SURFACE — the scenario binds to a declaration the fix introduces. It MUST carry a
narrowest-revert recipe: the smallest revert that still compiles and still exercises the
behavior. Typically keep the new type, parameter or enum constant; revert only its call sites
— the technique that recovered genuine behavioral red on 2 of 3 such items in bug wave 1.NEW-SURFACE with no revert that can yield behavioral red — say so explicitly and name the
substitute verification that replaces it (e.g. "reviewer reads each test body against the
implementation and confirms every asserted value traces to its oracle citation"). A substitute
declared in the plan is evidence; one produced at verification time is an excuse.The plan author assigns these labels, not the test author: the test author is blind to the
implementation by design (§4) and therefore cannot tell which surfaces are new. Where a dispatch
contract's planning seat returns a red-proof-shape field, it carries this same labelling, set
before the author is dispatched. The shape actually obtained is recorded in test-manifest (§10).
Every scenario's expected result must have a stated oracle source: the spec clause, the documented algorithm, an external reference (an RFC, a library's own documented contract), or an independently-run computation. The oracle is never "what the code returns" and never a value lifted uninspected from the ticket's worked example.
Never encode an illustration as the oracle. A worked example in a spec or ticket exists to build intuition, not to be trusted as ground truth — illustrative examples are frequently approximate, off-by-one, or simplified for readability. If a scenario's only available expected value is an illustration, recompute it independently from the stated rule before writing the assertion, and cite the rule, not the illustration, as the oracle source.
Never read the implementation to decide correctness. This is the rule the whole trait exists to enforce. If you find yourself opening the file under test to see what it currently returns and then asserting that value, stop — you have just converted the test into a change-detector for whatever the implementation happens to do today, including its bugs. This is the exact failure pattern in the trend record: an oracle computed from the implementation's own formula reproduces the formula, not the requirement, so it stays green even when the formula is wrong.
Concretely, an oracle citation looks like:
"per task-scope §Test Strategy S3: duplicate dependency edges are rejected with DUPLICATE_EDGE" — a spec clause."per RFC 7396 §2: a merge-patch null removes the key" — an external reference."computed independently: SHA-256 of the canonical byte sequence, verified against a second implementation (Python hashlib) run outside the codebase" — an independent computation.An oracle citation that instead reads "matches current behavior" or "see implementation" is
not an oracle — it is a confession that this rule was skipped, and should block the note from
being accepted as complete.
A NEW-SURFACE label (§2) does not relax any of this. When a scenario binds to a declaration
the fix introduces, the pull toward sourcing the expected value from the new code is strongest —
the surface exists nowhere else yet. It still does not qualify as an oracle. Derive the expected
result from the spec clause that motivated the new surface, or escalate per §8. A new surface with
no oracle available outside the implementation is a spec gap to raise, not a licence to read.
The test author's independence is only real if its inputs are actually restricted. This section is the enforceable half of that — the queue-phase oracle freeze (§3) is the other half.
Why this is not a reading-discipline rule. An earlier version of this section told the author
how to read implementation sources narrowly: look up the declaration, stop before the body. That
model cannot hold, because no reading tool has a declarations-only mode. Read returns a window,
Grep -A<n> returns context lines, sed -n <a>,<b>p returns a range — every one of them will put
a function body in front of an author who is trying, in good faith, to resolve a signature.
Restraint fails at the tool boundary, so the boundary moves.
The evidence. Bug wave 3 (2026-09), item a3ebd108, consumed THREE test authors:
query_notes(operation="list") with no keys= filter, received
implementation-notes in the response, and read the implementation files it named — an ingress
no reading-method rule had contemplated.-A2, never Read" remedy, widened
to sed -n <a>,<b>p ranges "to catch wrapped signatures" and read the sentinel function's body.src/main by any tool. Zero lookups; 12/12 scenarios plus 7 probes; reviewer verdict
independent.Authors 1 and 2 both self-disclosed and stopped before committing, so no contaminated artifact was ever tracked — the self-disclosure protocol worked, twice. What failed was the reading model, at a price of two burned dispatches. The rules below are the structural replacement.
They are capability rules, not care rules. A breach is a breach whether or not anything useful was seen, and it is disclosed the same way either way.
The dispatch prompt — or the Test author protocol slot of the wave's dispatch contract — MUST
paste inline and verbatim every public declaration the author needs: types and data-class
constructors with full parameter lists and defaults, function and method signatures, constants,
enum values, and any KDoc that carries an oracle or states an invariant (validate() included,
per §7 Fixture invariants). The author writes tests against that block and goes looking for
nothing further. That includes the block's harness: and runtime call order: lines: build the
application under test only as they state, per the dispatch contract's DECLARATIONS block and its
rule 4 (a NONE harness the scenarios need is stop-and-ask, §4.4).
A declarations block that is missing entirely is an orchestrator error. Say so and stop (§4.4); do not reconstruct it.
Do not open any file under src/main with ANY tool — Read, Grep, sed, cat, head, Glob
preview, an editor, or any shell command whose output includes file content. The ban attaches to
the file, not to the intent: "I only wanted the signature" does not make the call permitted,
because the tool decides what comes back, not you.
Banned for the same reason: git diff, git show, git log -p and any other view of the
implementer's changes on this branch.
src/test stays fully readable. Existing tests, fixtures and harnesses are how you match the
codebase's conventions — reading them is expected, not merely tolerated.
keys= is mandatory on every query_notes callEvery query_notes call carries an explicit keys= filter restricted to the item's queue-phase
keys — typically ["task-scope", "diagnosis", "test-plan"]. An unfiltered operation="list"
returns implementation-notes and session-tracking in the same response and is itself a
breach, whether or not their bodies were read. includeBody=true with a keys= filter is fine;
includeBody=true without one is the exact call that burned wave 3's first author.
The same holds for query_notes(operation="search"): scope it to the item, and never search for
terms that would rank implementation prose highly.
If a declaration you need is absent from the supplied block — or what was supplied does not compile against the test as planned, because a parameter was renamed, a return type reshaped, or a method the plan assumed exists does not — do not derive the corrected shape from context, from a compiler error that quotes surrounding source, or from the diff.
Send the question back: SendMessage to the orchestrator (Parallel/Delegated tier), or ask the
user (Direct tier), naming the exact declaration you need. Asking costs one round-trip; the lookup
costs the dispatch. This is the escalation path of §8, and a plan-vs-implementation drift found
this way is a finding worth recording, not an inconvenience to route around.
task-scope, feature-summary, diagnosis) and test-plan,
fetched with a keys= filter per §4.3.current/docs/, CLAUDE.md) and any external reference cited as an
oracle.src/test — conventions, fixtures, harnesses, existing assertions.git diff --name-only, or the File ownership rows of
the dispatch contract) — enough to know where tests belong, not what the files contain.Never: any file under src/main; any diff, commit or patch content; the implementer's own tests
where the dispatch identifies them as such; implementation-notes; session-tracking.
If you breach any rule above, deliberately or by accident: stop immediately, commit nothing, and
report exactly what was read and when — in your return message and in test-manifest's
arbitration record. Delete any draft written after the breach and say that you did.
This protocol is the part of the old model that worked, three times across two waves, and it is
unchanged. A disclosed breach costs a re-dispatch. An undisclosed one costs the trait: every later
independent verdict on the item becomes unverifiable.
Where red is achievable before the fix exists, the test must actually observe it.
Which scenarios "red is achievable" covers is decided by the §2 surface labels, not here. An
EXISTING-SURFACE scenario can reach behavioural red by a plain revert. A NEW-SURFACE one
cannot — reverting the fix removes the declaration the test binds to, so the run is compile-red,
which proves only that the test references new code. For those, red-first means executing the
plan's narrowest-revert recipe (keep the new type or parameter, revert its call sites) or, where
the plan declared no revert can work, the substitute verification it named instead. Read §2's
label definitions before deciding a scenario's red evidence; the shape actually obtained is a
test-manifest field (§10).
Bug-fix regression tests: write the test from the diagnosis note's reproduction steps and
confirm it fails against the pre-fix code — actually run it red, don't assume the reproduction
description implies a failing assertion. A regression test that was never seen red proves nothing
about whether it would have caught the bug; this is exactly how assumeTrue wrapping neutralizes
a would-be regression test without anyone noticing, because the wrapped test never fails at all,
pre-fix or post-fix.
Feature suites (test-after): since implementation already exists by the time the test author
writes code, true red-before-fix isn't available. Substitute a mandatory per-scenario line in the
manifest: "what specific wrong behavior would this assertion catch?" — name a plausible bug
this test would fail against (wrong value, wrong exception, silently-accepted invalid input). If
you cannot state one, the assertion is not adding coverage; strengthen it or mark the scenario
not-covered: <reason> rather than writing an assertion that would pass against almost anything.
"Nothing specific" for this line is a blocking review finding under review-quality.
Beyond the scenarios derived from acceptance criteria, run the applicable probes below against
the feature's actual input surface. Record every probe attempted in test-manifest — including
the ones that found nothing. A probe list with only findings looks like it was written after the
fact to justify existing tests; a probe list that also records clean results is evidence the
surface was actually exercised.
\ vs /, ..%2f).\\?\, \\host\share)
variants of path or identifier input, where the surface accepts path-like or URI-like input.Not every probe applies to every surface — a pure in-memory computation has no path-encoding
surface. Record the ones that don't apply as N/A: <reason> rather than omitting them silently,
so a reviewer can tell "not applicable" apart from "forgotten."
These patterns produce a green suite without verifying real behavior. Full before/after examples
for each, in Kotlin/JUnit5, are in references/forbidden-patterns.md. Treat this list as gate
criteria for test-manifest's forbidden-pattern declaration (§10) — every instance found in the
authored tests must be declared with a justification, and an undeclared instance found in review
is a blocking finding regardless of intent.
assumeTrue on non-platform conditions — assumeTrue exists to skip a test on an
environment it cannot run in (OS, missing external service). Using it to skip a test when a
behavioral precondition isn't met turns a would-be failure into a silent skip. Three real
production bugs shipped this way in this codebase's history.result == expected || result.isEmpty(), or any assertion with an
|| branch that accepts a "didn't do anything" outcome as equally valid to the intended one.
This passes whether the feature works or does nothing at all.assertNotNull(result) (or result != null, list.isNotEmpty())
standing in for a check of the actual value, size, or contents.try whose catch block logs or
ignores the exception instead of failing the test, so an assertion failure and a caught
exception both read as a pass.Distinct from the patterns above, and their mirror image: a fixture that violates a domain
invariant produces a red suite that looks like an implementation bug. The test fails for a
reason unrelated to the behavior under test, and the cost lands on the orchestrator as an
arbitration round-trip, after the author has already returned. The recorded instance: a claim
fixture built with claimedAt = Instant.now() alongside a claimExpiresAt in the past, which
WorkItem.validate() rejects outright.
Before writing a fixture or a fixture helper, read the domain type's validate() — from the
declarations supplied per §4.1, not by opening src/main — and satisfy it by construction:
claimedAt = claimExpiresAt.minus(ttl). The same shape applies to any pair the type constrains
— created/modified, start/end, offset/limit, parent depth vs. child depth.validate() by reflection or a test-only backdoor, and do not
quietly retarget the scenario at whatever state happens to be constructible.test-manifest (§10), so a reviewer can tell a fixture
that satisfies validate() by construction from one that passes by luck.If the invariant check is not among the supplied declarations, ask for it (§4.4) — it is exactly the kind of oracle-bearing declaration the dispatch is required to paste.
The patterns above are phrased around assertion text. This class hides in the fixture or the
harness instead, so a clean-looking assertion still passes whether or not the fix is present. The
recorded instances, all blind-authored and caught only in review: a "no item created" assertion
after POSTing a non-JSON body that could never create an item; a duplicate-Host case that never
reached its branch because Ktor testApplication merges repeated headers into one comma-joined
value; a cycle rejection that asserted created=0 but not that the reason was "circular".
test-plan or test-manifest names
the fixture condition that would make it fail without the fix. If none exists — the fixture
could never produce the asserted outcome anyway — fix the fixture, not the assertion.testApplication
merges duplicate headers and sends no Host by default.When a scenario's expected result is genuinely unclear — the spec is silent, two spec clauses conflict, or the public signature doesn't match what the plan assumed — the test author does not resolve it by reading the implementation to see what was built. That is exactly the shortcut §3 and §4 exist to close off.
Escalate instead. Record the ambiguity in test-manifest under the arbitration record: what
was ambiguous, what the plan said, what would need to be true for each candidate resolution.
Escalation goes to the orchestrator (Parallel/Delegated tier) or the user (Direct tier) — never
resolved unilaterally by the same agent that hit the ambiguity. The one carve-out: an ambiguity
resolvable from public non-src/main evidence (the tool's parameterSchema, src/test harnesses,
docs) may be self-resolved when the arbitration record states the evidence used; the reviewer
verifies it.
If an ambiguity must be resolved by consulting the implementation (rare, and only when
escalation is not available and the item cannot proceed otherwise), the resulting scenario is
oracle-degraded — mark it explicitly in the manifest's S-id mapping. test-independence-audit
must surface every oracle-degraded scenario individually; it cannot be waved through as part of
an overall "independent" verdict.
This is distinct from implementation-vs-test arbitration proper (red author-test → is the
implementation wrong, or the test wrong, or the spec wrong) — that triage is orchestrator-owned,
per the /implement skill's Step 4b "Arbitration — Red Author-Authored Tests" subsection, not
something the test author decides. The test author's job when a written test comes back red is
to report it (§9), not to guess which side is at fault.
If a test you authored fails against the implementation as it stands, that is a finding, not a problem to make go away. Report it — in the manifest and, if the test author is a live dispatch reporting to an orchestrator, in the return message — with the scenario id, the assertion, the observed value, and the oracle citation.
Never weaken the assertion, skip the test, or wrap it in assumeTrue to unblock a wave or a
merge. Every one of the forbidden patterns in §7 is, in practice, a rationalized version of
"this test is inconvenient right now." A red test authored independently, against a frozen
oracle, is the entire point of this trait — silencing it defeats the purpose as completely as
never having written it. Arbitration (§8, orchestrator-owned) decides whether the implementation,
the test, or the spec is wrong; the test author's job stops at an accurate, specific report.
test-manifest (work phase, required) is the test author's complete account of what was done.
Fill every field — an omitted field reads as "not done," not as "not applicable" (use an explicit
N/A: <reason> for genuinely inapplicable fields, matching the probe-catalog convention in §6).
test-plan: covered (name the test method)
or not-covered: <reason>.assumeTrue, a disjunctive assertion, or any
other §7 pattern present in the authored tests, each with a justification. An empty declaration
asserts none were used — it is itself a claim the reviewer checks.validate() clause and the derivation used. Record any invariant
that forced an escalation rather than a fixture.NEW-SURFACE (§2), which
shape was actually obtained: behavioral-red (narrowest revert: <what was reverted>), or
compile-red only — substitute verification: <what the reviewer does instead>. Where the revert
is orchestrator-run rather than author-run, record red evidence: orchestrator-run together with
the recipe you expect it to use. State WHERE the revert ran (the scratch-copy path) and at WHICH
commit it was made; an in-place revert (stash, checkout, restore, reset) in the shared worktree is
not red evidence and the reviewer treats it as absent. EXISTING-SURFACE scenarios need nothing here beyond the §5
red-first result.oracle-degraded
markers.Direct-tier temporal-only mode (§1) is the only sanctioned degradation. It applies only when there is no second agent available to dispatch. Under this mode:
test-plan still gates work entry — the oracle freeze happens before implementation, exactly
as in the two-agent case. This is the part of the guarantee that survives.test-manifest declares the single-actor mode explicitly (actor: <id> (temporal-only, same agent as implementer)) rather than presenting as if a second agent were involved.src/main and cannot un-read it. The ban is replaced by the red-first ORDERING of §5 —
the test is written from the frozen test-plan and observed red against pre-fix code BEFORE
the fix is written, which is the only separation available here. An "I avoided looking"
claim is not a substitute and must not be recorded as one.test-plan
and the planning-seat notes, not from the implementation just written; every query_notes
call still carries keys= restricted to the queue-phase keys. implementation-notes and
session-tracking stay out of the author pass even though the same actor wrote them.independent-degraded names: the oracle boundary held, the capability boundary
could not.test-independence-audit records the verdict as independent-degraded, never independent.
A temporal-only separation is real but weaker than two-agent separation — it defeats
implementation-derived oracles (the oracle was frozen before code existed) but not
implementation-derived test code, since the same agent that reads the plan later writes both
the implementation and the tests informed by whatever it just built.No other degraded mode is sanctioned. If a Delegated- or Parallel-tier item finds itself unable to actually separate the two dispatches (e.g. capacity pressure), that is a scheduling problem to raise with the orchestrator — not a reason to fall back to same-agent authorship without declaring it, and not a reason to skip declaring the degradation if it happens anyway.
© jpicklyk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in .claude/skills/test-author of jpicklyk/task-orchestrator.
Open the folder on GitHubat commit 3e83170
Test Author next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Test Author this skilljpicklyk/task-orchestrator | 207 | — | ~8.2k | Automated safety check: Pass | MIT | |
| Test Scenariosphuryn/pm-skills | 27k | — | ~866 | Automated safety check: Pass | MIT | |
| Test Case Writermohitagw15856/pm-claude-skills | 1.4k | — | ~970 | Automated safety check: Pass | MIT | |
| Argent QA Flowsbbplayer-app/BBPlayer | 1.1k | — | ~3.2k | Automated safety check: Pass | MIT | |
| AI Test Generationpetrkindlmann/qa-skills | 165 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Prd V07 Test Planningmattgierhart/PRD-driven-context-engineering | 180 | — | ~3.5k | Automated safety check: Notes | MIT |
phuryn/pm-skills
Create comprehensive test scenarios from user stories with test objectives, starting conditions, user roles, step-by-step actions, and expected outcomes.
mohitagw15856/pm-claude-skills
Turn a requirement or user story into clear, executable test cases.
bbplayer-app/BBPlayer
Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
mattgierhart/PRD-driven-context-engineering
Define test cases BEFORE implementation, ensuring every API, business rule, and user journey has verifiable acceptance criteria during PRD v0.7 Build Execution.
petrkindlmann/qa-skills
Author and maintain MANUAL and hybrid test cases and suites in TestRail, Xray (Jira), Zephyr Scale, and Qase.
jpicklyk/task-orchestrator
Walks through how to launch and reach the MCP Task Orchestrator server container: transport, REST API, port publishing, config mounts and config-sync.
jpicklyk/task-orchestrator
Resolves ready MCP work items into a run plan, shows it to you, then executes it through the Workflow tool or direct subagent dispatch, with post-run verification.
jpicklyk/task-orchestrator
Migrates an existing unscoped Task Orchestrator database to the project-scoping convention in place, creating one project anchor root and re-parenting work trees under it after a mandatory dry run.
jpicklyk/task-orchestrator
Completes or cancels a whole feature subtree, a named list of items, or a batch of stale work items at once, previewing the impact and warning before force-completing anything active.
jpicklyk/task-orchestrator
Creates an MCP work item from conversation context, anchoring it under the right container, inferring type and priority and pre-filling the required notes.
jpicklyk/task-orchestrator
Views, creates, deletes and diagnoses BLOCKS, IS_BLOCKED_BY and RELATES_TO links between MCP work items, including why an item cannot start.
Categories
Test authoring framework for items carrying the needs-test-author trait. Test Author is an agent skill from jpicklyk/task-orchestrator. Test authoring framework for items carrying the needs-test-author trait.
Test Author fits situations like: filling test-plan; test-manifest notes; asked to author; review tests independently of an implementation.
Run `npx skills add jpicklyk/task-orchestrator --skill test-author -a claude-code`. Or copy the skill folder (.claude/skills/test-author in jpicklyk/task-orchestrator) into .claude/skills/test-author in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jpicklyk/task-orchestrator --skill test-author -a codex`. Or copy the skill folder (.claude/skills/test-author in jpicklyk/task-orchestrator) into .agents/skills/test-author in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jpicklyk/task-orchestrator --skill test-author -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-author, .gemini/skills/test-author, .github/skills/test-author and .opencode/skills/test-author in your project.
Going by SKILL.md and its folder, Test Author needs the command-line tools its instructions call (git). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Test Author is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.2k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Test Author: Test Scenarios (phuryn/pm-skills, 27k stars), Test Case Writer (mohitagw15856/pm-claude-skills, 1.4k stars), Argent QA Flows (bbplayer-app/BBPlayer, 1.1k stars) and AI Test Generation (petrkindlmann/qa-skills, 165 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jpicklyk (a GitHub user) maintains it in jpicklyk/task-orchestrator, which has 207 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 8, 2026.
Source: jpicklyk/task-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.