MCP Testing
adeze/raindrop-mcp
MCP Testing Strategies with Vitest, Inspector, and Integration Tests
Specification quality framework for planning. An agent skill from jpicklyk/task-orchestrator.
$ npx skills add jpicklyk/task-orchestrator --skill spec-quality -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jpicklyk/task-orchestrator spec-quality --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/spec-quality .claude/skills/spec-quality && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "spec-quality" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/spec-quality into .claude/skills/spec-quality/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-quality", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/spec-qualityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jpicklyk/task-orchestrator --skill spec-quality -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jpicklyk/task-orchestrator spec-quality --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/spec-quality .agents/skills/spec-quality && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "spec-quality" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/spec-quality into .agents/skills/spec-quality/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-quality", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jpicklyk/task-orchestrator --skill spec-quality -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jpicklyk/task-orchestrator spec-quality --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/spec-quality .cursor/skills/spec-quality && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "spec-quality" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/spec-quality into .cursor/skills/spec-quality/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-quality", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jpicklyk/task-orchestrator.git --path .claude/skills/spec-quality--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jpicklyk/task-orchestrator --skill spec-quality -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jpicklyk/task-orchestrator spec-quality --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/spec-quality .gemini/skills/spec-quality && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "spec-quality" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/spec-quality into .gemini/skills/spec-quality/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-quality", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jpicklyk/task-orchestrator spec-qualityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jpicklyk/task-orchestrator --skill spec-quality -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/spec-quality .github/skills/spec-quality && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "spec-quality" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/spec-quality into .github/skills/spec-quality/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-quality", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jpicklyk/task-orchestrator --skill spec-quality -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jpicklyk/task-orchestrator spec-quality --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/spec-quality .opencode/skills/spec-quality && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "spec-quality" agent skill from https://github.com/jpicklyk/task-orchestrator/tree/main/.claude/skills/spec-quality into .opencode/skills/spec-quality/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spec-quality", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
spec-qualitySpecification quality framework for planning. An agent skill from jpicklyk/task-orchestrator.
Spec Quality is an agent skill from jpicklyk/task-orchestrator. Specification quality framework for planning. Defines the minimum bar for what a plan must address — alternatives, non-goals, blast radius, risk flags, and test strategy. Referenced by schema guidance fields during queue-phase note filling. Use when filling feature-summary, task-scope, diagnosis, or other queue-phase specification notes for any MCP work item.
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/project-concerns.md`).
It sits in Testing & QA, covering Test strategy and MCP servers. It works with Model Context Protocol. The repository describes itself as: Server-enforced workflow discipline for AI agents. An MCP server providing persistent work items, dependency graphs, quality gates, and actor attribution. Schemas define what… The licence is MIT.
Read from SKILL.md and the folder at commit 3e83170. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Spec Quality loads about 3.1k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 94 tokens; SKILL.md has 1,776 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jpicklyk/task-orchestrator at commit 3e83170, republished under its MIT licence (© jpicklyk). 1,776 words, ~3,057 tokens.
.claude/skills/spec-quality/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.This skill defines the minimum thinking floor for plans and specifications. The sections below represent what every plan must address. They are not a ceiling — if the problem demands additional analysis, add it. But these areas must not be skipped.
Where this applies at each level: for a feature-implementation parent, the queue-phase
feature-summary note stays lean by design (goal, findings→tasks table, dependency edges,
a pointer to non-goals — target under 2k chars) and does not carry the full disciplines below.
The disciplines below apply in full to each child's task-scope note (or to specification/
diagnosis notes on schemas without a parent/child split) — that is where alternatives,
blast radius, risk flags, and test strategy must actually be worked through.
The value of a spec is entirely in the thinking it forces before code is written. If a section doesn't change how you'd approach implementation, it isn't earning its place. Every sentence should either prevent a mistake or force a decision.
These are the required areas of analysis. Each one exists because skipping it leads to a specific, recurring class of failure.
Evaluate at least two real approaches. "Do nothing" always counts as one. For each alternative, state what it would look like and the specific trade-off that led to its rejection. If you can only think of one approach, you haven't explored the solution space — step back and look for a fundamentally different angle.
The point is not to document alternatives for posterity. It's to catch yourself before committing to an approach that has a better option sitting next to it.
Anti-pattern: strawman alternatives. "Alternative: rewrite everything from scratch. Rejected: too much work." This doesn't force any real thinking.
If the spec lists open decisions, naming alternatives, or "options under consideration," resolve every one before materializing MCP work items or dispatching implementation. A spec with unresolved choices left in it is not ready to dispatch.
Mid-implementation pivots cost 5-10× the time of a pre-dispatch user round-trip. Once an agent is in a worktree, a naming change or interface change requires unwinding partial code, retracting commits, re-spec'ing the work, and re-dispatching. Resolving the same question before dispatch is a single short conversation.
The orchestrator should pause for a user round-trip rather than dispatch with ambiguity. A short pause now is cheaper than a partial implementation later.
If a decision genuinely cannot be resolved at planning time (e.g., depends on a measurement only available post-implementation), it isn't an "open decision" — it's a deliberate two-phase approach. Document it as such, scope the first phase explicitly, and create a separate work item for the second phase. Don't leave the choice latent in the spec.
When a spec prescribes call sequences, API semantics, or any fact owned by another artifact (a tool description, a source file, another spec), cite the authoritative location and direct the implementer to verify against it before relying on it. A restated copy drifts from its source — and a spec's confident restatement reads as authority to the agent that implements it faithfully.
If the spec claims something was verified, name exactly what was checked. "Verified
against source" that checked one path while prescribing another is how a partially
verified claim ships a defect (two sessions of evidence: task-scope facts contradicting
target files, 2026-07-27; a design prescribing advance_item trigger sequences that
gate-block, 2026-07-31).
This discipline applies to the plan artifact as much as to spec notes. Specification notes
are generated from the plan and inherit its errors silently, so plan-mode output that names a
file path, tool contract, or call sequence must cite the authoritative location and be verified
against it before any child item inherits it. The rule above was first installed for spec notes
only (2026-07-31); the next occurrence (retro 6d562acb, 2026-08-04) originated one artifact
upstream — the approved plan asserted a dispatch surface at ralph-iteration.md when the live
file was skills/ralph/iteration-prompt.md, and a child's specification note and its first
review verdict both inherited the wrong path. If a further occurrence still originates in a
plan, treat that as evidence that prose rules cannot carry this load and add a mechanical check
(every path a plan names must resolve on disk).
Semantic claims carry a file:line. Any sentence that states nullability, redaction,
ownership, or an honest limit (what a guard does NOT cover) cites the file:line that
establishes it. This holds for spec notes and for everything written from them — a CHANGELOG
bullet, an API doc line, a hook or KDoc header comment. Worked case (2026-09-29, #384): the spec
and a hook header said note attribution is shown only when the caller is ADMIN and the flag is
false; the code redacts only when redactNoteAttribution && !isAdmin
(AttributionRedactor.kt:53), so it is shown when ADMIN OR the flag is false. The error
survived from spec to header because the sentence was written from intent and no line was
opened. For a NEW-SURFACE claim with no code yet, cite the spec decision that fixes the
behavior and mark the claim for the reviewer to confirm in the diff. Prose that is not a
semantic claim needs no citation — this is not a blanket citation rule.
Before writing a verification command into a spec, plan, or skill, run it twice: once on the target platform against known-good input (it must succeed) and once against known-bad input (it must fail). A command that has not been shown to fail on bad input has not been shown to verify anything — a check that silently always passes is worse than no check, because it is trusted. Record the command in the form actually executed, not a reconstruction.
Name what someone might reasonably expect this work to include but that is deliberately excluded. If you cannot name a single non-goal, the scope is not tight enough.
Non-goals prevent scope creep during implementation. Without them, agents tend to gold-plate — adding adjacent improvements that weren't asked for and that introduce unplanned risk.
Identify every module, file, and interface affected by the change. Trace downstream consumers — if you change a repository method signature, what tools call it? If you change a domain model default, what tests assume the old value?
This analysis exists to catch "I didn't realize changing X breaks Y" before it happens.
Read references/project-concerns.md for cross-cutting constraints specific to this
codebase that frequently expand blast radius in non-obvious ways.
If the blast radius touches a surface with an automated budget, ceiling, or quota test
(e.g. ToolTokenBudgetTest), measure current headroom while writing the spec and record
both the measured values and the predicted post-change values. Do not defer the
measurement to implementation — by then the scope decision is already made.
Negative and exhaustive claims must name their sweep. Any claim that something does
not exist or that a list is complete ("zero callers", "no stubs", "nothing depends on
this", "none found", "all sites") must be written as the command that produced it plus
its scope and result count — for example
grep -rn "createSuspend" current/src → 0 hits (src/main + src/test). A negative claim
with no named sweep is not a finding; write "not checked" instead. A sweep that covered
only part of the surface states the part it covered, and the reader treats the remainder
as unchecked, not as clear. The defect-class-siblings field of /implement's planning
seat return template is this rule's special case for a bug's root-cause pattern.
Call out the one or two things most likely to go wrong. These might be areas of tight coupling, migration complexity, concurrency concerns, or simply parts of the codebase you don't fully understand yet.
The purpose is to focus review attention where it matters and to make uncertainty explicit rather than hidden.
Every plan must include a concrete test strategy. This is not "add tests" — it's a specific accounting of what will be verified and how.
Required coverage areas:
For each area, name the specific scenarios you'll test. "Test edge cases" is not a strategy. "Test that circular parent references are detected and rejected with a clear error" is.
If the change modifies shared interfaces (domain models, repository contracts, tool parameters), note which existing tests may break and how you'll handle that — update them, or confirm they still pass with the new behavior.
Give every scenario a stable numbered id and an oracle source. Number scenarios
sequentially (S1, S2, S3…) across all three coverage areas, and for each one record the
oracle source the expected result comes from — a spec clause, a stated algorithm, or an
external reference. Never "what the code returns," and never just the ticket's worked
example restated as if it were independent confirmation. Stable ids let downstream
artifacts map coverage back to this spec without duplicating it — most directly the
needs-test-author trait's test-plan and test-manifest notes, which reference these
same ids; cite the trait by name rather than restating its note schema here.
Validate spec completeness before advancing past queue phase:
This framework sets a floor. The disciplines above are the minimum required analysis. Depending on the complexity of the work, additional analysis may be warranted — performance implications, migration strategies, API compatibility concerns, or anything else that would change the implementation approach if examined carefully.
Add whatever the problem demands. The goal is a plan that lets someone implement the change confidently, understanding not just what to build but why this approach was chosen and what to watch out for.
© jpicklyk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in .claude/skills/spec-quality of jpicklyk/task-orchestrator.
Open the folder on GitHubat commit 3e83170
Spec Quality next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Spec Quality this skilljpicklyk/task-orchestrator | 207 | — | ~3.1k | Automated safety check: Pass | MIT | |
| MCP Testingadeze/raindrop-mcp | 188 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Frontmcp Testingagentfront/frontmcp | 146 | — | ~10k | Automated safety check: Notes | Apache-2.0 | |
| Qwen Code E2E TestingQwenLM/qwen-code | 28k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Adding LLM MCP ToolsTriliumNext/Trilium | 38k | — | ~2.5k | Automated safety check: Pass | AGPL-3.0 | |
| Agentacct Workflowmikehasa/agentacct | 766 | — | ~1.6k | Automated safety check: Pass | MIT |
adeze/raindrop-mcp
MCP Testing Strategies with Vitest, Inspector, and Integration Tests
agentfront/frontmcp
A skill your agent uses for anything about testing FrontMCP servers: writing or running unit, integration, and E2E tests and reaching the 95%+ coverage bar.
QwenLM/qwen-code
Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.
TriliumNext/Trilium
A skill your agent uses when adding, changing, or reviewing an LLM/MCP tool in Trilium (the defineTools definitions under packages/trilium-core/src/services/llm/tools/ —…
mikehasa/agentacct
A skill your agent uses when working in a repo with agentacct MCP configured, or when asked to track coding-agent work, smoke-test agentacct integrations, or report objective AI-agent task evidence.
nteract/semiotic
Inspect a ChatGPT Apps MCP server codebase and generate chatgpt-app-submission.json with app info suggestions, tool hint justifications, test cases, and negative test cases, then report review-check…
jpicklyk/task-orchestrator
Walks through how to launch and reach the MCP Task Orchestrator server container: transport, REST API, port publishing, config mounts and config-sync.
jpicklyk/task-orchestrator
Resolves ready MCP work items into a run plan, shows it to you, then executes it through the Workflow tool or direct subagent dispatch, with post-run verification.
jpicklyk/task-orchestrator
Migrates an existing unscoped Task Orchestrator database to the project-scoping convention in place, creating one project anchor root and re-parenting work trees under it after a mandatory dry run.
jpicklyk/task-orchestrator
Completes or cancels a whole feature subtree, a named list of items, or a batch of stale work items at once, previewing the impact and warning before force-completing anything active.
jpicklyk/task-orchestrator
Creates an MCP work item from conversation context, anchoring it under the right container, inferring type and priority and pre-filling the required notes.
jpicklyk/task-orchestrator
Views, creates, deletes and diagnoses BLOCKS, IS_BLOCKED_BY and RELATES_TO links between MCP work items, including why an item cannot start.
Works with
Categories
Specification quality framework for planning. An agent skill from jpicklyk/task-orchestrator. Spec Quality is an agent skill from jpicklyk/task-orchestrator. Specification quality framework for planning.
Spec Quality fits situations like: filling feature-summary; other queue-phase specification notes for any MCP work item.
Run `npx skills add jpicklyk/task-orchestrator --skill spec-quality -a claude-code`. Or copy the skill folder (.claude/skills/spec-quality in jpicklyk/task-orchestrator) into .claude/skills/spec-quality in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jpicklyk/task-orchestrator --skill spec-quality -a codex`. Or copy the skill folder (.claude/skills/spec-quality in jpicklyk/task-orchestrator) into .agents/skills/spec-quality in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jpicklyk/task-orchestrator --skill spec-quality -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-quality, .gemini/skills/spec-quality, .github/skills/spec-quality and .opencode/skills/spec-quality in your project.
SKILL.md names no scripts, command-line tools or credentials: Spec Quality is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Spec Quality is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 687 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Spec Quality: MCP Testing (adeze/raindrop-mcp, 188 stars), Frontmcp Testing (agentfront/frontmcp, 146 stars), Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars) and Adding LLM MCP Tools (TriliumNext/Trilium, 38k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jpicklyk (a GitHub user) maintains it in jpicklyk/task-orchestrator, which has 207 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 8, 2026.
Source: jpicklyk/task-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.