MCP Testing
adeze/raindrop-mcp
MCP Testing Strategies with Vitest, Inspector, and Integration Tests
Testing patterns for MCP tool/resource handlers using createMockContext and Vitest.
$ npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cyanheads/pubmed-mcp-server api-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .claude/skills && cp -r skills-src/framework-skills/api-testing .claude/skills/api-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "api-testing" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/api-testing into .claude/skills/api-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/api-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cyanheads/pubmed-mcp-server api-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .agents/skills && cp -r skills-src/framework-skills/api-testing .agents/skills/api-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "api-testing" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/api-testing into .agents/skills/api-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cyanheads/pubmed-mcp-server api-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/framework-skills/api-testing .cursor/skills/api-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "api-testing" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/api-testing into .cursor/skills/api-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cyanheads/pubmed-mcp-server.git --path framework-skills/api-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cyanheads/pubmed-mcp-server api-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/framework-skills/api-testing .gemini/skills/api-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "api-testing" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/api-testing into .gemini/skills/api-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cyanheads/pubmed-mcp-server api-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .github/skills && cp -r skills-src/framework-skills/api-testing .github/skills/api-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "api-testing" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/api-testing into .github/skills/api-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cyanheads/pubmed-mcp-server api-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/framework-skills/api-testing .opencode/skills/api-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "api-testing" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/api-testing into .opencode/skills/api-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
api-testingTesting patterns for MCP tool/resource handlers using createMockContext and Vitest.
API Testing is an agent skill from cyanheads/pubmed-mcp-server. Testing patterns for MCP tool/resource handlers using createMockContext and Vitest. Covers mock context options, handler testing, McpError assertions, format testing, Vitest config setup, and test isolation conventions.
Its SKILL.md is about 8.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Unit testing, API testing and MCP servers. It works with Vitest. The repository describes itself as: Search PubMed/Europe PMC, fetch articles and full text (PMC/EPMC/Unpaywall), citations, MeSH terms via MCP. STDIO or Streamable HTTP. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 5a417fb. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
API Testing loads about 8.1k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 2,611 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from cyanheads/pubmed-mcp-server at commit 5a417fb, republished under its Apache-2.0 licence (© cyanheads). 2,611 words, ~8,139 tokens.
.claude/skills/api-testing/SKILL.md (or your agent's skills folder).Tests target handler behavior directly — call handler(input, ctx), assert on the return value or thrown error. The framework's handler factory (try/catch, formatting, telemetry) is not involved. Use createMockContext from @cyanheads/mcp-ts-core/testing to construct the ctx argument.
Additional exports from /testing: createMockSession() binds a mock handler context to an HTTP session; createFetchMock() provides a strict upstream HTTP fake; runToolContract() executes a definition through schema, handler, formatting, enrichment/content, and production-shaped error-envelope checks. createMockLogger() returns a standalone MockContextLogger, createInMemoryStorage(options?) provides a real StorageService backed by InMemoryProvider, and expectInputRequired(run) returns the input_required result a multi-round-trip handler asked for (see Mock inputs).
Other testing subpaths: /testing/vitest (fixtures and conformance suites, below), /testing/fuzz (see Fuzz testing), and /testing/apps, whose renderAppTool renders an app tool's ui:// view against your server (a stdio command or an HTTP URL) in a headless MCP Apps host and reports initialization, errors, CSP violations, the view↔host messages, and screenshots. It needs the optional peers @modelcontextprotocol/client and @modelcontextprotocol/ext-apps plus a chrome-headless-shell build; the field-test skill covers installing one and reading the report.
Philosophy: Test behavior, not implementation. Refactors should not break tests. Match the repo's existing test layout: fresh scaffolds use tests/, while colocated src/**/*.test.ts files are also supported. Integration tests at I/O boundaries over unit tests of internals.
mcpTest — fixture-based Vitest testmcpTest is a test.extend-based Vitest test that provides ctx, session, fetchMock, and storage as per-test fixtures — fresh instances for every test, eliminating boilerplate and enforcing isolation automatically. fetchMock is installed as globalThis.fetch only when requested by a test and restored afterward.
import { mcpTest } from '@cyanheads/mcp-ts-core/testing/vitest';
mcpTest('echoes the message', async ({ ctx }) => {
const result = await echoTool.handler(echoTool.input.parse({ message: 'hi' }), ctx);
expect(result.message).toBe('hi');
});
mcpTest('uses storage fixture', async ({ ctx, storage }) => {
const svc = new MyService(config, storage);
const result = await svc.doWork(ctx);
expect(result).toBeDefined();
});
mcpTest('stubs an upstream HTTP boundary', async ({ fetchMock }) => {
fetchMock.route({
match: 'https://api.example.test/items/42',
respond: Response.json({ id: '42' }),
});
await expect(loadItem('42')).resolves.toMatchObject({ id: '42' });
});| Fixture | Type | Per-test? | Notes |
|---|---|---|---|
ctx | Context | Yes | Fresh createMockContext() each test |
session | MockSession | Yes | Fresh { sessionId, tenantId, ctx } from createMockSession() |
fetchMock | FetchMockHarness | Yes | Strict fetch fake installed/restored around the requesting test |
storage | StorageService | Yes | Fresh createInMemoryStorage() each test |
Override fixtures using the function form (async ({}, use) => { ... }) to preserve per-test freshness. A bare-value override shares one mutable instance across the entire file — defeating the fixture's isolation guarantee.
import { createMockContext } from '@cyanheads/mcp-ts-core/testing/vitest';
// Correct — function form gives each test a fresh context:
const tenantTest = mcpTest.extend({
ctx: async ({}, use) => { await use(createMockContext({ tenantId: 'test-tenant' })); },
});
// Wrong — bare value shares one ctx across every test in the file:
// const tenantTest = mcpTest.extend({ ctx: createMockContext({ tenantId: 'test-tenant' }) });The portable /testing helpers are re-exported from @cyanheads/mcp-ts-core/testing/vitest so fixture overrides don't need a second import.
createFetchMockUse the fetch harness at real outbound I/O boundaries. Stub the external service, not server-owned services or handlers.
import { createFetchMock } from '@cyanheads/mcp-ts-core/testing';
const http = createFetchMock([
{
method: 'GET',
match: 'https://api.example.test/items/42',
respond: Response.json({ id: '42', name: 'Example' }),
},
]);
http.install();
try {
await expect(loadItem('42')).resolves.toEqual({ id: '42', name: 'Example' });
expect(http.calls[0]?.request.url).toBe('https://api.example.test/items/42');
} finally {
http.restore();
}Routes match in registration order. match accepts an exact URL, RegExp, or request predicate; respond accepts a static Response or a response factory. A static response's body is read once, on the route's first match, and every call is served a fresh Response over those bytes with the same status, statusText, and headers — so a consumer that cancels the body, or an error-body reader like httpErrorFromResponse that stops past its cap, settles on Node as on Bun. Set once: true for one-shot behavior. Unmatched requests throw unless onUnhandled is provided.
A request predicate routes on the URL's origin, never a prefix. req.url.startsWith(BASE_URL) also matches a lookalike host (https://api.example.test.evil.com/...), which CodeQL reports as high-severity incomplete URL substring sanitization — it scans test files as readily as src/, so a suite that is green locally still fails the security check on a pull request. Parse the URL and compare origins, matching the path separately:
match: (req) => {
const url = new URL(req.url);
return url.origin === new URL(BASE_URL).origin && url.pathname.startsWith('/items/');
},toolContractSuitePoint the reusable suite at a definition plus representative success and failure inputs. It checks input/output schemas, invokes the real handler, applies formatting/enrichment/content, and validates both public error surfaces.
import { JsonRpcErrorCode } from '@cyanheads/mcp-ts-core/errors';
import { toolContractSuite } from '@cyanheads/mcp-ts-core/testing/vitest';
toolContractSuite(searchTool, {
success: [{ name: 'returns matches', input: { query: 'mcp' } }],
errors: [{
name: 'reports an empty query',
input: { query: '' },
code: JsonRpcErrorCode.InvalidParams,
reason: 'empty_query',
}],
});Use runToolContract(definition, input, { context }) from /testing when a custom test runner or an imperative assertion is a better fit. It intentionally skips transport auth and telemetry; those belong in transport/integration tests.
A declared reason thrown without a hint — a bare ctx.fail('reason') or a service throw carrying { reason } — comes back with the entry's recovery as data.recovery.hint and a Recovery: line in content[], as in production. The one production field it leaves out is data.requestId (and the request <id> term closing content[]), since there is no real request; a test asserting the factory's envelope instead expects both. Calling definition.handler(...) directly returns the McpError exactly as the throw site built it — no fill, no request id.
Arguments that fail the input schema are rejected the way the production handler factory rejects them: InvalidParams (-32602), with a message naming the tool and every failing field. That is the code a client sees on the wire, so assert it — not ValidationError (-32007), which stays the classification for a ZodError a handler throws itself. A result that breaks the tool's own output or enrichment schema is the definition's bug, so it returns InternalError (-32603) with a message naming that contract, exactly as in production.
Cancellation settles as it does in production. Pass context: { signal } and abort it: once the signal has fired, whatever the handler — or the output validation, format(), and enrichment after it — throws comes back as RequestCancelled (-32011), whether that is the signal's AbortError, its reason string, a withRetry backoff that stopped, or an McpError of the handler's own. A throw while the signal is still live keeps its own classification, and argument parsing stays outside the settle, so schema-invalid arguments on an aborted signal still return InvalidParams. A toolContractSuite error case with an aborted context.signal asserts code: JsonRpcErrorCode.RequestCancelled the same way.
Resources have no contract runner. A resource definition's handler(...) called directly returns the throw site's McpError with no recovery fill, so asserting a resource's declared errors[] hints needs a path through the resource factory. Serve the definition from createWorkerHandler({ name, title, resources: [def] }) (/worker) and fetch it a resources/read JSON-RPC request carrying the current protocol revision in both the MCP-Protocol-Version header and params._meta. The response is plain JSON or an SSE data: frame, and its error carries code, data.reason, and the filled data.recovery.hint.
createMockContext optionsimport { createMockContext } from '@cyanheads/mcp-ts-core/testing';
createMockContext() // working ctx.state on tenant 'default'
createMockContext({ tenantId: 'test-tenant' }) // explicit tenant scope for ctx.state
createMockContext({ errors: myTool.errors }) // attaches typed ctx.fail keyed by the contract reasons
createMockContext({ inputResponses: { confirm: { action: 'accept', content: { ok: true } } } }) // second round of a multi-round-trip handler
createMockContext({ requestState: 'opaque-state' }) // seeds ctx.inputs.state()
createMockContext({ clientCapabilities: { roots: {} } }) // seeds ctx.clientCapabilities and filters inputResponses to declared kinds
createMockContext({ requestId: 'my-id' }) // override request ID (default: 'test-request-id')
createMockContext({ notifyResourceListChanged: () => {} }) // with resource-list change notifier
createMockContext({ notifyResourceUpdated: (_uri) => {} }) // with resource update notifier
createMockContext({ signal: controller.signal }) // custom AbortSignal
createMockContext({ auth: { clientId: 'test', scopes: [], sub: 'test-user' } }) // with auth context
createMockContext({ uri: new URL('myscheme://item/123') }) // for resource handler testingMockContextOptions interface:
interface MockContextOptions<TErrors extends readonly ErrorContract[] | undefined> {
auth?: AuthContext;
clientCapabilities?: ClientCapabilities;
errors?: TErrors | undefined;
inputResponses?: InputResponses | Record<string, unknown>;
notifyPromptListChanged?: () => void;
notifyResourceListChanged?: () => void;
notifyResourceUpdated?: (uri: string) => void;
notifyToolListChanged?: () => void;
requestId?: string;
requestState?: unknown;
sessionId?: string;
signal?: AbortSignal;
tenantId?: string;
uri?: URL;
}| Option | Effect |
|---|---|
| (none) | Working ctx.state on tenant 'default'; ctx.inputs is empty (first round) |
auth | Sets ctx.auth for scope-checking tests |
clientCapabilities | Sets ctx.clientCapabilities (undefined when omitted) and applies the production filter to inputResponses: only the answers these capabilities cover reach ctx.inputs (elicit → elicitation, and elicitation.form when it carries content; sampling → sampling, and sampling.tools when it holds a tool_use / tool_result block; roots → roots). Omitted, every seeded response reaches ctx.inputs, so existing { inputResponses } tests are unaffected. The mock's ctx.requestInput stays ungated either way |
errors | Attaches a typed ctx.fail against the contract — same wiring the production handler factory uses. Pass myTool.errors directly; the return type narrows to HandlerContext<ReasonOf<…>>, so the context is assignable to that definition's handler parameter. |
inputResponses | Seeds ctx.inputs with the responses a retried request would carry, keyed by the identifiers the handler's ctx.requestInput(...) assigned (see below) |
notifyPromptListChanged | Assigns ctx.notifyPromptListChanged for prompt-list change notification tests |
notifyResourceListChanged | Assigns ctx.notifyResourceListChanged for resource notification tests |
notifyResourceUpdated | Assigns ctx.notifyResourceUpdated for resource update notification tests |
notifyToolListChanged | Assigns ctx.notifyToolListChanged for tool-list change notification tests |
requestId | Overrides ctx.requestId (default: 'test-request-id') |
requestState | Seeds ctx.inputs.state() — the opaque state a prior round attached |
sessionId | Sets ctx.sessionId for handlers that branch on session ID |
signal | Overrides ctx.signal — useful for cancellation testing |
tenantId | Scopes ctx.state to a specific tenant. Defaults to 'default' — the value stdio (and HTTP with MCP_AUTH_MODE=none) resolves |
uri | Sets ctx.uri for resource handler testing |
ctx.state is a real StorageService over an InMemoryProvider — the production storage path, not a Map. A test therefore sees the same rules a deployed server enforces:
^[a-zA-Z0-9_.\-/]+$ and may not contain ... Colons are rejected, so cache:v1:abc throws McpError(ValidationError) in the test exactly as it would in a deployment; use cache/v1/abc.Date reads back as its ISO string — so a test cannot pass on identity or on a Date/Map surviving storage. A value JSON cannot encode (bigint, a cyclic reference, a top-level undefined, function, or symbol) rejects with McpError(SerializationError).{ ttl: 30 } reads back as null once 30 seconds elapse — drive the clock with vi.useFakeTimers() to assert expiry.getMany / setMany / deleteMany / list validate every key and prefix, and list paginates with the same opaque cursors.ctx.signal aborts, state operations reject.const ctx = createMockContext();
await ctx.state.set('cache/v1/abc', { hits: 1 }, { ttl: 30 });
await expect(ctx.state.get('cache/v1/abc')).resolves.toEqual({ hits: 1 });
await expect(ctx.state.set('cache:v1:abc', {})).rejects.toThrow(McpError);
await ctx.state.set('seen/abc', { at: new Date('2026-01-01T00:00:00Z') });
await expect(ctx.state.get('seen/abc')).resolves.toEqual({ at: '2026-01-01T00:00:00.000Z' });Reach for createInMemoryStorage() when a service takes a StorageService directly — it builds the same pair.
ctx.requestInput is the real implementation: it throws an InputRequiredSignal the production handler factories convert into an input_required result. In a unit test the handler is called directly, so that signal surfaces as a thrown value — which is exactly how you assert the first round. The examples below drive export_report from api-context § The shape of a multi-round-trip handler, which asks for a format the caller left out:
import { isInputRequiredSignal } from '@cyanheads/mcp-ts-core';
it('asks for the format on the first round', async () => {
const ctx = createMockContext();
await expect(exportReport.handler(exportReport.input.parse({ reportId: 'r1' }), ctx))
.rejects.toSatisfy(isInputRequiredSignal);
});To assert on what was requested, use expectInputRequired from /testing. It runs the handler and returns the input_required result the handler factory would have returned; it throws when the handler returns normally, and any other error propagates untouched:
import { createMockContext, expectInputRequired } from '@cyanheads/mcp-ts-core/testing';
const asked = await expectInputRequired(() => exportReport.handler(input, createMockContext()));
expect(asked.inputRequests?.format?.method).toBe('elicitation/create');Pass asked.requestState back as createMockContext({ requestState }) when the handler reads state from the prior round. inputResponses drives the second round. ctx.inputs.accepted(key, schema) and .view(key) read it with the same helpers production uses, so a wrong response shape fails in the test:
it('exports in the format the user picked', async () => {
const ctx = createMockContext({
inputResponses: { format: { action: 'accept', content: { format: 'csv' } } },
});
await expect(exportReport.handler(input, ctx)).resolves.toMatchObject({ url: expect.any(String) });
});
it('stops when the user declines', async () => {
const ctx = createMockContext({
inputResponses: { format: { action: 'decline' } },
});
await expect(exportReport.handler(input, ctx)).rejects.toThrow(McpError);
});Seeding an answer this way is right for a handler that treats it as input. A consent gate does not: it acts only on a record it stored when it asked, so an answer seeded alone makes it ask again — see below.
ctx.inputs.dropped is always [] on a mock context — the drop only happens in the SDK's wire decoding, so cover it in an integration test rather than a unit one.
Seed clientCapabilities to test what a client without a capability gets: a pre-answered inputResponses the declared capabilities do not cover — a kind never declared, a form answer (one carrying content) from a client that declared only elicitation.url, a tool-use sampling answer without sampling.tools — never reaches ctx.inputs, so the handler asks again. A handler that falls through when a capability is missing (if (ctx.clientCapabilities?.roots) … else …) is tested by seeding both shapes.
A consent gate that redeems a ctx.state record (see api-context § Consent gates) needs the record in the second round's storage, and each mock context has its own. Copy what round one stored into the round-two context. The record carries the caller, so seed the same auth (or none) on both:
const accept = { confirm: { action: 'accept', content: { confirm: true } } };
const first = createMockContext();
const asked = await expectInputRequired(() => deletePath.handler(input, first));
const record = await first.state.get(`consent/${asked.requestState}`);
const second = createMockContext({ inputResponses: accept, requestState: asked.requestState });
await second.state.set(`consent/${asked.requestState}`, record);
await expect(deletePath.handler(input, second)).resolves.toEqual({ deleted: input.path });
// The record is spent: the same state again asks for a fresh confirmation.
await expect(deletePath.handler(input, second)).rejects.toSatisfy(isInputRequiredSignal);
// An answer with no record behind it asks as well — nothing was asked, so nothing was confirmed.
const unasked = createMockContext({ inputResponses: accept });
await expect(deletePath.handler(input, unasked)).rejects.toSatisfy(isInputRequiredSignal);A record the round-two context holds under another auth, or one written by another tool (its operation differs), asks again the same way — cover whichever of those the handler's binding is meant to catch.
ctx.log captures all log calls for inspection. Import MockContextLogger from @cyanheads/mcp-ts-core/testing and cast ctx.log to access the .calls array (the cast is necessary because createMockContext returns Context, which types log as ContextLogger):
import { createMockContext, type MockContextLogger } from '@cyanheads/mcp-ts-core/testing';
const ctx = createMockContext();
const log = ctx.log as MockContextLogger;
await myTool.handler(input, ctx);
expect(log.calls.some(c => c.level === 'info' && c.msg.includes('Processing'))).toBe(true);// tests/tools/my-tool.tool.test.ts
import { describe, expect, it } from 'vitest';
import { createMockContext } from '@cyanheads/mcp-ts-core/testing';
import { myTool } from '@/mcp-server/tools/definitions/my-tool.tool.js';
describe('myTool', () => {
it('returns expected output', async () => {
const ctx = createMockContext();
const input = myTool.input.parse({ query: 'hello' });
const result = await myTool.handler(input, ctx);
expect(result.result).toBe('Found: hello');
});
it('throws on invalid state', async () => {
const ctx = createMockContext();
const input = myTool.input.parse({ query: 'TRIGGER_ERROR' });
await expect(myTool.handler(input, ctx)).rejects.toThrow();
});
it('formats response completely', () => {
const result = { result: 'test' };
const blocks = myTool.format!(result);
expect(blocks[0].type).toBe('text');
expect((blocks[0] as { text?: string }).text).toContain('test');
});
});Parse input through myTool.input.parse(...) to validate against the Zod schema and produce the typed input the handler expects. Call myTool.handler(input, ctx) directly, not through the MCP SDK or any framework wrapper. Assert on the return value for happy paths; use .rejects.toThrow() for error paths. Test format separately if the tool defines one — it's a pure function and needs no ctx. Verify the rendered text includes the fields the LLM needs, and for projection-style tools, add a case with non-default field selections.
LLM clients only send populated fields. Form-based clients (MCP Inspector, web UIs) submit the full schema shape — optional object fields arrive with empty-string inner values instead of undefined. Both are valid MCP usage. Test that handlers handle both gracefully.
describe('form-client payloads', () => {
it('skips optional object when inner fields are empty strings', async () => {
const ctx = createMockContext();
// Form client sends the object with empty values instead of omitting it
const input = myTool.input.parse({
query: 'test',
dateRange: { minDate: '', maxDate: '' },
});
const result = await myTool.handler(input, ctx);
// Should succeed — empty dateRange is ignored, not passed downstream
expect(result.items).toBeDefined();
});
it('uses optional object when inner fields have real values', async () => {
const ctx = createMockContext();
const input = myTool.input.parse({
query: 'test',
dateRange: { minDate: '2025-01-01', maxDate: '2025-12-31' },
});
const result = await myTool.handler(input, ctx);
// Should apply the date filter
expect(result.items).toBeDefined();
});
});The pattern: parse through the schema (confirms Zod accepts the payload), call the handler, assert the empty-value case produces correct results — no errors, no corrupted downstream queries. Same applies to optional arrays: test with [] to verify the handler skips rather than passes through.
This is a different problem from form-client '' payloads. Here the upstream API omits fields entirely. The risk is either a validation failure from an over-strict schema or a quiet lie where missing data turns into a concrete fact.
describe('sparse upstream payloads', () => {
it('preserves missing upstream fields as unknown', async () => {
const upstream = {
id: 'repo-123',
name: 'Widget Repo',
// archived and star_count omitted entirely
};
const normalized = normalizeRepo(upstream);
expect(normalized).toEqual({
id: 'repo-123',
name: 'Widget Repo',
});
const output = repoSearchTool.output.parse({
repos: [normalized],
});
const blocks = repoSearchTool.format!(output);
expect((blocks[0] as { text: string }).text).toContain('Archived:** Not available');
expect((blocks[0] as { text: string }).text).not.toContain('Archived:** No');
});
});What to verify:
null or ''.format() uses explicit unknown-state fallbacks instead of inventing facts.Extend the framework's base config using mergeConfig. The base provides globals: true, pool: 'forks', isolate: true, tsconfigPaths, and a Zod SSR compatibility fix. Add only the @/ alias for your server's source:
// vitest.config.ts
import { defineConfig, mergeConfig } from 'vitest/config';
import coreConfig from '@cyanheads/mcp-ts-core/vitest.config';
export default mergeConfig(coreConfig, defineConfig({
resolve: {
alias: { '@/': new URL('./src/', import.meta.url).pathname },
},
}));mergeConfig deep-merges the framework base with your overrides. The base sets globals: true (describe, it, expect, etc. available without imports), pool: 'forks' and isolate: true (test files run in separate worker processes), and ssr: { noExternal: ['zod'] } for Zod 4 compatibility. The resolve.alias entry maps @/ to src/, matching the paths alias in tsconfig.json so imports like @/services/... resolve correctly in tests.
Construct dependencies fresh in beforeEach. Never share mutable state across tests.
import { beforeEach, describe, expect, it } from 'vitest';
import { initMyService } from '@/services/my-domain/my-service.js';
describe('myTool with service', () => {
beforeEach(() => {
// Re-initialize with a fresh instance before each test
initMyService(mockConfig, mockStorage);
});
it('calls service correctly', async () => {
const ctx = createMockContext({ tenantId: 'test-tenant' });
// ...
});
});initMyService() (or equivalent) in beforeEach when tests share a module-level singleton.createMockContext({ tenantId }) when a test needs a specific tenant; omitting it scopes state to 'default', not to a broken state surface.import { McpError, JsonRpcErrorCode } from '@cyanheads/mcp-ts-core/errors';
it('throws NotFound for missing resource', async () => {
const ctx = createMockContext();
const input = myTool.input.parse({ id: 'nonexistent' });
await expect(myTool.handler(input, ctx)).rejects.toMatchObject({
code: JsonRpcErrorCode.NotFound,
});
});Use .rejects.toThrow(McpError) to assert type only. Use .rejects.toMatchObject({ code: ... }) when the specific error code matters.
expect.schemaMatching (Vitest 4, Standard Schema) validates a value against any Zod schema — including the definition's own output. Use it to assert schema conformance without duplicating the shape in the test:
it('output conforms to the declared output schema', async () => {
const ctx = createMockContext();
const result = await myTool.handler(myTool.input.parse({ query: 'x' }), ctx);
expect(result).toEqual(expect.schemaMatching(myTool.output));
});It composes as an asymmetric matcher anywhere a value is expected — e.g. toHaveBeenCalledWith(expect.schemaMatching(schema)). Prefer exact-value assertions when the expected output is fully known; reach for schemaMatching when the output is dynamic (timestamps, generated IDs) or the schema itself is the contract under test.
errors[] (typed contract)Tools and resources that declare an errors[] contract receive a typed ctx.fail helper at runtime. Pass the definition's own errors to createMockContext and the mock wires fail the same way the production handler factory does:
import { createMockContext } from '@cyanheads/mcp-ts-core/testing';
import { JsonRpcErrorCode } from '@cyanheads/mcp-ts-core/errors';
import { fetchItems } from '@/mcp-server/tools/definitions/fetch-items.tool.js';
it('throws ctx.fail("no_match") when no items resolve', async () => {
const ctx = createMockContext({ errors: fetchItems.errors });
const input = fetchItems.input.parse({ ids: ['missing'] });
await expect(fetchItems.handler(input, ctx)).rejects.toMatchObject({
code: JsonRpcErrorCode.NotFound,
data: { reason: 'no_match' },
});
});For lower-level tests that need the raw fail helper without a full mock context (e.g. asserting the reason → code mapping), use createFail directly — see Testing the handler-side fail plumbing below.
data.reason and not just code?The contract reason is the stable machine-readable identifier — clients switch on it the same way they would on an HTTP status. A code alone (NotFound) doesn't disambiguate between contract entries that share a code ('no_match' vs 'withdrawn' both mapping to NotFound). Asserting on data.reason locks the test to the specific contract entry.
data.reason is overridable-proofThe framework spreads caller-supplied data first and writes reason last, so a handler that passes data: { reason: 'something_else' } cannot override the contract reason. Tests can rely on data.reason always equaling the contract entry's reason — write assertions that depend on it without paranoia.
fail plumbingTo verify the definition wires ctx.fail correctly without exercising the full handler factory, use the errors array directly:
import { createFail } from '@cyanheads/mcp-ts-core';
it('builds an error with the contract code and reason', () => {
const fail = createFail(myTool.errors!);
const err = fail('no_match', 'not found', { itemId: '123' });
expect(err.code).toBe(JsonRpcErrorCode.NotFound);
expect(err.data).toEqual({ reason: 'no_match', itemId: '123' });
});For schema-heavy or input-validation-critical handlers, the framework ships fuzz helpers under @cyanheads/mcp-ts-core/testing/fuzz. They generate valid + adversarial inputs from your Zod schemas via fast-check and assert handler invariants (no crashes, no prototype pollution, no stack-trace leaks).
import { fuzzTool, fuzzResource, fuzzPrompt } from '@cyanheads/mcp-ts-core/testing/fuzz';
it('survives fuzz testing', async () => {
const report = await fuzzTool(myTool, { numRuns: 100, numAdversarial: 30 });
expect(report.crashes).toHaveLength(0);
expect(report.leaks).toHaveLength(0);
expect(report.prototypePollution).toBe(false);
});| Helper | Purpose |
|---|---|
fuzzTool(def, opts) / fuzzResource(def, opts) / fuzzPrompt(def, opts) | Drive valid + adversarial inputs through the handler. Returns a FuzzReport. |
zodToArbitrary(schema) | Convert a Zod schema to a fast-check Arbitrary for custom property-based tests. |
adversarialArbitrary() / ADVERSARIAL_STRINGS | Targeted injection sets (prototype pollution probes, control characters, oversized payloads). |
FuzzOptions: numRuns (default 50), numAdversarial (default 30), seed (reproducibility), timeout (per-call ms, default 5000), ctx (MockContextOptions for stateful handlers).
report.leaks looks for a stack frame or a server-side path in what a client can observe — the code, message, and data of the thrown McpError. Strings the input itself supplied are removed before that check, so naming the offending value in error data (throw validationError(msg, { key })) never registers as a leak: the client sent those bytes and learns nothing from seeing them again.
© cyanheads, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in framework-skills/api-testing of cyanheads/pubmed-mcp-server.
Open the folder on GitHubat commit 5a417fb
API Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| API Testing this skillcyanheads/pubmed-mcp-server | 154 | — | ~8.1k | Automated safety check: Pass | Apache-2.0 | |
| MCP Testingadeze/raindrop-mcp | 188 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Storybook Testingyonatangross/orchestkit | 288 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Adding LLM MCP ToolsTriliumNext/Trilium | 38k | — | ~2.5k | Automated safety check: Pass | AGPL-3.0 | |
| Designing TestsCloudAI-X/claude-workflow-v2 | 1.4k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Web Testing with Playwright and Vitestwithkynam/vibecode-pro-max-kit | 1.1k | — | ~892 | Automated safety check: Pass | Apache-2.0 |
adeze/raindrop-mcp
MCP Testing Strategies with Vitest, Inspector, and Integration Tests
yonatangross/orchestkit
Storybook 10 testing patterns with Vitest integration, ESM-only distribution, CSF3 typesafe factories, play() interaction tests, Chromatic TurboSnap visual regression, module automocking…
TriliumNext/Trilium
A skill your agent uses when adding, changing, or reviewing an LLM/MCP tool in Trilium (the defineTools definitions under packages/trilium-core/src/services/llm/tools/ —…
CloudAI-X/claude-workflow-v2
Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.
withkynam/vibecode-pro-max-kit
Covers web testing from unit to E2E, load, visual, accessibility and security checks, with Playwright, Vitest and k6 guides plus a Playwright setup script.
DanielPodolsky/ownyourcode
Guides test strategy for unit, integration, and E2E testing.
cyanheads/pubmed-mcp-server
Scaffold an MCP App tool + UI resource pair. An agent skill from cyanheads/pubmed-mcp-server.
cyanheads/pubmed-mcp-server
Scaffold a new MCP prompt template. An agent skill from cyanheads/pubmed-mcp-server.
cyanheads/pubmed-mcp-server
Scaffold a new MCP resource definition. An agent skill from cyanheads/pubmed-mcp-server.
cyanheads/pubmed-mcp-server
Scaffold a new service integration. An agent skill from cyanheads/pubmed-mcp-server.
cyanheads/pubmed-mcp-server
Scaffold a test file for an existing tool, resource, or service.
cyanheads/pubmed-mcp-server
Authentication, authorization, and multi-tenancy patterns for @cyanheads/mcp-ts-core.
Works with
Categories
Testing patterns for MCP tool/resource handlers using createMockContext and Vitest. API Testing is an agent skill from cyanheads/pubmed-mcp-server. Testing patterns for MCP tool/resource handlers using createMockContext and Vitest.
API Testing fits situations like: tasks that involve Unit testing; tasks that involve API testing; tasks that involve MCP servers.
Run `npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a claude-code`. Or copy the skill folder (framework-skills/api-testing in cyanheads/pubmed-mcp-server) into .claude/skills/api-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a codex`. Or copy the skill folder (framework-skills/api-testing in cyanheads/pubmed-mcp-server) into .agents/skills/api-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cyanheads/pubmed-mcp-server --skill api-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/api-testing, .gemini/skills/api-testing, .github/skills/api-testing and .opencode/skills/api-testing in your project.
SKILL.md names no scripts, command-line tools or credentials: API Testing is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
API Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.1k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with API Testing: MCP Testing (adeze/raindrop-mcp, 188 stars), Storybook Testing (yonatangross/orchestkit, 288 stars), Adding LLM MCP Tools (TriliumNext/Trilium, 38k stars) and Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cyanheads (a GitHub user) maintains it in cyanheads/pubmed-mcp-server, which has 154 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 4, 2026.
Source: cyanheads/pubmed-mcp-server on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.