Agent skill

Spec Quality

by jpicklyk in jpicklyk/task-orchestrator

Specification quality framework for planning. An agent skill from jpicklyk/task-orchestrator.

MITAuto-check passedTesting & QA

Install Spec Quality

skills CLI
$ npx skills add jpicklyk/task-orchestrator --skill spec-quality -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jpicklyk/task-orchestrator spec-quality --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/spec-quality .claude/skills/spec-quality && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spec-quality
GitHub stars
207
Token cost
~3.1k tokens
SKILL.md length
1,776 words
Files
2 (incl. references)
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Specification quality framework for planning. An agent skill from jpicklyk/task-orchestrator.

  • Filling feature-summary
  • SKILL.md covers Specification Disciplines, Completion Checklist and Using This Framework
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Other queue-phase specification notes for any MCP work item

What it does

Spec Quality is an agent skill from jpicklyk/task-orchestrator. Specification quality framework for planning. Defines the minimum bar for what a plan must address — alternatives, non-goals, blast radius, risk flags, and test strategy. Referenced by schema guidance fields during queue-phase note filling. Use when filling feature-summary, task-scope, diagnosis, or other queue-phase specification notes for any MCP work item.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/project-concerns.md`).

It sits in Testing & QA, covering Test strategy and MCP servers. It works with Model Context Protocol. The repository describes itself as: Server-enforced workflow discipline for AI agents. An MCP server providing persistent work items, dependency graphs, quality gates, and actor attribution. Schemas define what… The licence is MIT.

When your agent uses it

  • Filling feature-summary
  • Other queue-phase specification notes for any MCP work item

Example prompts

  • “/spec-quality”

What it can do on your machine

Read from SKILL.md and the folder at commit 3e83170. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Spec Quality loads about 3.1k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 94 tokens; SKILL.md has 1,776 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jpicklyk/task-orchestrator at commit 3e83170, republished under its MIT licence (© jpicklyk). 1,776 words, ~3,057 tokens.

Download SKILL.mdSave it as .claude/skills/spec-quality/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
spec-quality
description
Specification quality framework for planning. Defines the minimum bar for what a plan must address — alternatives, non-goals, blast radius, risk flags, and test strategy. Referenced by schema guidance fields during queue-phase note filling. Use when filling feature-summary, task-scope, diagnosis, or other queue-phase specification notes for any MCP work item.
user-invocable
false

Specification Quality Framework

This skill defines the minimum thinking floor for plans and specifications. The sections below represent what every plan must address. They are not a ceiling — if the problem demands additional analysis, add it. But these areas must not be skipped.

Where this applies at each level: for a feature-implementation parent, the queue-phase feature-summary note stays lean by design (goal, findings→tasks table, dependency edges, a pointer to non-goals — target under 2k chars) and does not carry the full disciplines below. The disciplines below apply in full to each child's task-scope note (or to specification/ diagnosis notes on schemas without a parent/child split) — that is where alternatives, blast radius, risk flags, and test strategy must actually be worked through.

The value of a spec is entirely in the thinking it forces before code is written. If a section doesn't change how you'd approach implementation, it isn't earning its place. Every sentence should either prevent a mistake or force a decision.


Specification Disciplines

These are the required areas of analysis. Each one exists because skipping it leads to a specific, recurring class of failure.

Alternatives Considered

Evaluate at least two real approaches. "Do nothing" always counts as one. For each alternative, state what it would look like and the specific trade-off that led to its rejection. If you can only think of one approach, you haven't explored the solution space — step back and look for a fundamentally different angle.

The point is not to document alternatives for posterity. It's to catch yourself before committing to an approach that has a better option sitting next to it.

Anti-pattern: strawman alternatives. "Alternative: rewrite everything from scratch. Rejected: too much work." This doesn't force any real thinking.

Open Decisions

If the spec lists open decisions, naming alternatives, or "options under consideration," resolve every one before materializing MCP work items or dispatching implementation. A spec with unresolved choices left in it is not ready to dispatch.

Mid-implementation pivots cost 5-10× the time of a pre-dispatch user round-trip. Once an agent is in a worktree, a naming change or interface change requires unwinding partial code, retracting commits, re-spec'ing the work, and re-dispatching. Resolving the same question before dispatch is a single short conversation.

The orchestrator should pause for a user round-trip rather than dispatch with ambiguity. A short pause now is cheaper than a partial implementation later.

If a decision genuinely cannot be resolved at planning time (e.g., depends on a measurement only available post-implementation), it isn't an "open decision" — it's a deliberate two-phase approach. Document it as such, scope the first phase explicitly, and create a separate work item for the second phase. Don't leave the choice latent in the spec.

Cite Contracts, Don't Restate Them

When a spec prescribes call sequences, API semantics, or any fact owned by another artifact (a tool description, a source file, another spec), cite the authoritative location and direct the implementer to verify against it before relying on it. A restated copy drifts from its source — and a spec's confident restatement reads as authority to the agent that implements it faithfully.

If the spec claims something was verified, name exactly what was checked. "Verified against source" that checked one path while prescribing another is how a partially verified claim ships a defect (two sessions of evidence: task-scope facts contradicting target files, 2026-07-27; a design prescribing advance_item trigger sequences that gate-block, 2026-07-31).

This discipline applies to the plan artifact as much as to spec notes. Specification notes are generated from the plan and inherit its errors silently, so plan-mode output that names a file path, tool contract, or call sequence must cite the authoritative location and be verified against it before any child item inherits it. The rule above was first installed for spec notes only (2026-07-31); the next occurrence (retro 6d562acb, 2026-08-04) originated one artifact upstream — the approved plan asserted a dispatch surface at ralph-iteration.md when the live file was skills/ralph/iteration-prompt.md, and a child's specification note and its first review verdict both inherited the wrong path. If a further occurrence still originates in a plan, treat that as evidence that prose rules cannot carry this load and add a mechanical check (every path a plan names must resolve on disk).

Semantic claims carry a file:line. Any sentence that states nullability, redaction, ownership, or an honest limit (what a guard does NOT cover) cites the file:line that establishes it. This holds for spec notes and for everything written from them — a CHANGELOG bullet, an API doc line, a hook or KDoc header comment. Worked case (2026-09-29, #384): the spec and a hook header said note attribution is shown only when the caller is ADMIN and the flag is false; the code redacts only when redactNoteAttribution && !isAdmin (AttributionRedactor.kt:53), so it is shown when ADMIN OR the flag is false. The error survived from spec to header because the sentence was written from intent and no line was opened. For a NEW-SURFACE claim with no code yet, cite the spec decision that fixes the behavior and mark the claim for the reviewer to confirm in the diff. Prose that is not a semantic claim needs no citation — this is not a blanket citation rule.

Verification Commands

Before writing a verification command into a spec, plan, or skill, run it twice: once on the target platform against known-good input (it must succeed) and once against known-bad input (it must fail). A command that has not been shown to fail on bad input has not been shown to verify anything — a check that silently always passes is worse than no check, because it is trusted. Record the command in the form actually executed, not a reconstruction.

Non-Goals

Name what someone might reasonably expect this work to include but that is deliberately excluded. If you cannot name a single non-goal, the scope is not tight enough.

Non-goals prevent scope creep during implementation. Without them, agents tend to gold-plate — adding adjacent improvements that weren't asked for and that introduce unplanned risk.

Show full SKILL.md (780 more words)Show less
Blast Radius

Identify every module, file, and interface affected by the change. Trace downstream consumers — if you change a repository method signature, what tools call it? If you change a domain model default, what tests assume the old value?

This analysis exists to catch "I didn't realize changing X breaks Y" before it happens. Read references/project-concerns.md for cross-cutting constraints specific to this codebase that frequently expand blast radius in non-obvious ways.

If the blast radius touches a surface with an automated budget, ceiling, or quota test (e.g. ToolTokenBudgetTest), measure current headroom while writing the spec and record both the measured values and the predicted post-change values. Do not defer the measurement to implementation — by then the scope decision is already made.

Negative and exhaustive claims must name their sweep. Any claim that something does not exist or that a list is complete ("zero callers", "no stubs", "nothing depends on this", "none found", "all sites") must be written as the command that produced it plus its scope and result count — for example grep -rn "createSuspend" current/src → 0 hits (src/main + src/test). A negative claim with no named sweep is not a finding; write "not checked" instead. A sweep that covered only part of the surface states the part it covered, and the reader treats the remainder as unchecked, not as clear. The defect-class-siblings field of /implement's planning seat return template is this rule's special case for a bug's root-cause pattern.

Risk Flags

Call out the one or two things most likely to go wrong. These might be areas of tight coupling, migration complexity, concurrency concerns, or simply parts of the codebase you don't fully understand yet.

The purpose is to focus review attention where it matters and to make uncertainty explicit rather than hidden.

Test Strategy

Every plan must include a concrete test strategy. This is not "add tests" — it's a specific accounting of what will be verified and how.

Required coverage areas:

  • Happy paths — the primary use cases the change enables. These confirm the feature works as intended under normal conditions.
  • Failure paths — what happens when inputs are invalid, dependencies are missing, or operations fail. These confirm the system fails gracefully rather than silently corrupting state or throwing unhandled exceptions.
  • Edge cases — boundary conditions specific to the change. Examples: empty collections, null/optional fields, maximum depth limits, circular references, concurrent access. Think about what a user or caller could do that you didn't explicitly design for.

For each area, name the specific scenarios you'll test. "Test edge cases" is not a strategy. "Test that circular parent references are detected and rejected with a clear error" is.

If the change modifies shared interfaces (domain models, repository contracts, tool parameters), note which existing tests may break and how you'll handle that — update them, or confirm they still pass with the new behavior.

Give every scenario a stable numbered id and an oracle source. Number scenarios sequentially (S1, S2, S3…) across all three coverage areas, and for each one record the oracle source the expected result comes from — a spec clause, a stated algorithm, or an external reference. Never "what the code returns," and never just the ticket's worked example restated as if it were independent confirmation. Stable ids let downstream artifacts map coverage back to this spec without duplicating it — most directly the needs-test-author trait's test-plan and test-manifest notes, which reference these same ids; cite the trait by name rather than restating its note schema here.


Completion Checklist

Validate spec completeness before advancing past queue phase:

  • At least 2 real alternatives evaluated (not strawmen)
  • No "open decisions" or "options under consideration" sections remain in the spec
  • At least 1 non-goal named (scope boundary explicit)
  • Downstream consumers of changed interfaces traced
  • Contracts cited (with location), not restated; any "verified" claim names what was checked
  • Every nullability, redaction, ownership or honest-limit claim cites the file:line that establishes it
  • Verification commands proven (succeed on good input, fail on bad input)
  • Automated budget/ceiling headroom measured, if the blast radius touches one
  • 1-2 concrete risk flags identified
  • Test scenarios named for happy paths, failure paths, and edge cases
  • Scenarios carry stable numbered ids (S1…) with an oracle source recorded for each
  • Shared interface breakage assessed (if applicable)

Using This Framework

This framework sets a floor. The disciplines above are the minimum required analysis. Depending on the complexity of the work, additional analysis may be warranted — performance implications, migration strategies, API compatibility concerns, or anything else that would change the implementation approach if examined carefully.

Add whatever the problem demands. The goal is a plan that lets someone implement the change confidently, understanding not just what to build but why this approach was chosen and what to watch out for.

© jpicklyk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .claude/skills/spec-quality of jpicklyk/task-orchestrator.

  • SKILL.md
  • references/project-concerns.md

Open the folder on GitHubat commit 3e83170

Compare with similar skills

Spec Quality next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spec Quality compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spec Quality this skilljpicklyk/task-orchestrator207—~3.1kAutomated safety check: PassMIT
MCP Testingadeze/raindrop-mcp188—~1.6kAutomated safety check: PassMIT
Frontmcp Testingagentfront/frontmcp146—~10kAutomated safety check: NotesApache-2.0
Qwen Code E2E TestingQwenLM/qwen-code28k—~2.1kAutomated safety check: PassApache-2.0
Adding LLM MCP ToolsTriliumNext/Trilium38k—~2.5kAutomated safety check: PassAGPL-3.0
Agentacct Workflowmikehasa/agentacct766—~1.6kAutomated safety check: PassMIT

Similar skills

  • MCP Testing

    adeze/raindrop-mcp

    MCP Testing Strategies with Vitest, Inspector, and Integration Tests

    188 GitHub stars~1.6k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Frontmcp Testing

    agentfront/frontmcp

    A skill your agent uses for anything about testing FrontMCP servers: writing or running unit, integration, and E2E tests and reaching the 95%+ coverage bar.

    146 GitHub stars~10k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Adding LLM MCP Tools

    TriliumNext/Trilium

    A skill your agent uses when adding, changing, or reviewing an LLM/MCP tool in Trilium (the defineTools definitions under packages/trilium-core/src/services/llm/tools/ —…

    38k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check passed
  • Agentacct Workflow

    mikehasa/agentacct

    A skill your agent uses when working in a repo with agentacct MCP configured, or when asked to track coding-agent work, smoke-test agentacct integrations, or report objective AI-agent task evidence.

    766 GitHub stars~1.6k tokensUpdated 5 days ago
    Testing & QAAuto-check passed
  • Chatgpt App Submission

    nteract/semiotic

    Inspect a ChatGPT Apps MCP server codebase and generate chatgpt-app-submission.json with app info suggestions, tool hint justifications, test cases, and negative test cases, then report review-check…

    2.7k GitHub stars~2.8k tokensUpdated 2 days ago
    Testing & QAAuto-check passed

More from jpicklyk/task-orchestrator

All 28 skills in this repo
  • Task Orchestrator Server Setup

    jpicklyk/task-orchestrator

    Walks through how to launch and reach the MCP Task Orchestrator server container: transport, REST API, port publishing, config mounts and config-sync.

    207 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Run Wave

    jpicklyk/task-orchestrator

    Resolves ready MCP work items into a run plan, shows it to you, then executes it through the Workflow tool or direct subagent dispatch, with post-run verification.

    207 GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Adopt Project Scope Migration

    jpicklyk/task-orchestrator

    Migrates an existing unscoped Task Orchestrator database to the project-scoping convention in place, creating one project anchor root and re-parenting work trees under it after a mandatory dry run.

    207 GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Bulk Task Completion

    jpicklyk/task-orchestrator

    Completes or cancels a whole feature subtree, a named list of items, or a batch of stale work items at once, previewing the impact and warning before force-completing anything active.

    207 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Task Orchestrator Item Creator

    jpicklyk/task-orchestrator

    Creates an MCP work item from conversation context, anchoring it under the right container, inferring type and priority and pre-filling the required notes.

    207 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Work Item Dependency Manager

    jpicklyk/task-orchestrator

    Views, creates, deletes and diagnoses BLOCKS, IS_BLOCKED_BY and RELATES_TO links between MCP work items, including why an item cannot start.

    207 GitHub stars~3.5k tokensUpdated today
    Auto-check passed

Categories

Questions about Spec Quality

What does Spec Quality do?

Specification quality framework for planning. An agent skill from jpicklyk/task-orchestrator. Spec Quality is an agent skill from jpicklyk/task-orchestrator. Specification quality framework for planning.

When should I use Spec Quality?

Spec Quality fits situations like: filling feature-summary; other queue-phase specification notes for any MCP work item.

How do I install Spec Quality in Claude Code?

Run `npx skills add jpicklyk/task-orchestrator --skill spec-quality -a claude-code`. Or copy the skill folder (.claude/skills/spec-quality in jpicklyk/task-orchestrator) into .claude/skills/spec-quality in your project. Claude Code loads it when a task matches its description.

How do I install Spec Quality in Codex?

Run `npx skills add jpicklyk/task-orchestrator --skill spec-quality -a codex`. Or copy the skill folder (.claude/skills/spec-quality in jpicklyk/task-orchestrator) into .agents/skills/spec-quality in your project. Codex loads it when a task matches its description.

Can I use Spec Quality in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jpicklyk/task-orchestrator --skill spec-quality -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-quality, .gemini/skills/spec-quality, .github/skills/spec-quality and .opencode/skills/spec-quality in your project.

What does Spec Quality need to run?

SKILL.md names no scripts, command-line tools or credentials: Spec Quality is instructions for the agent only.

Does Spec Quality access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Spec Quality safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Spec Quality use?

Spec Quality is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spec Quality use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 687 tokens, read only when the agent opens those files.

What are the alternatives to Spec Quality?

Skills that share tags, products or a category with Spec Quality: MCP Testing (adeze/raindrop-mcp, 188 stars), Frontmcp Testing (agentfront/frontmcp, 146 stars), Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars) and Adding LLM MCP Tools (TriliumNext/Trilium, 38k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spec Quality?

jpicklyk (a GitHub user) maintains it in jpicklyk/task-orchestrator, which has 207 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 8, 2026.

Source: jpicklyk/task-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.