Agent skill

Design MCP Server

by cyanheads in cyanheads/pubmed-mcp-server

Design the tool surface, resources, and service layer for a new MCP server.

Apache-2.0Auto-check passedAgent Workflows

Install Design MCP Server

skills CLI
$ npx skills add cyanheads/pubmed-mcp-server --skill design-mcp-server -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cyanheads/pubmed-mcp-server design-mcp-server --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .claude/skills && cp -r skills-src/framework-skills/design-mcp-server .claude/skills/design-mcp-server && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
design-mcp-server
GitHub stars
158
Token cost
~22k tokens
SKILL.md length
11,367 words
Files
1
Skills in repo
30
Repo updated
First seen
Licence
Apache-2.0

At a glance

Design the tool surface, resources, and service layer for a new MCP server.

  • Works in 9 steps: Research External Dependencies → Map User Goals, Then Domain Operations → Classify into MCP Primitives → …
  • Starting a new server
  • SKILL.md covers When to Use, Inputs, Server Naming and Steps, plus 2 more sections
  • Needs MCP_REQUEST_STATE_KEY

What it does

Design MCP Server is an agent skill from cyanheads/pubmed-mcp-server. Design the tool surface, resources, and service layer for a new MCP server. Use when starting a new server, planning a major feature expansion, or when the user describes a domain/API they want to expose via MCP. Produces a design doc at docs/design.md that drives implementation.

Its SKILL.md is about 22k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering MCP servers, Architecture decision records and Design tokens. It works with Model Context Protocol. The repository describes itself as: Search PubMed/Europe PMC, fetch articles and full text (PMC/EPMC/Unpaywall), citations, MeSH terms via MCP. STDIO or Streamable HTTP. The licence is Apache-2.0.

When your agent uses it

  • Starting a new server
  • Planning a major feature expansion
  • The user describes a domain/API they want to expose via MCP

Example prompts

  • “/design-mcp-server”

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Research External Dependencies
  2. Map User Goals, Then Domain Operations
  3. Classify into MCP Primitives
  4. Design Tools
  5. Design Resources
  6. Design Prompts (if needed)
  7. Plan Services and Config
  8. Write the Design Doc
  9. Confirm and Proceed

What it can do on your machine

Read from SKILL.md and the folder at commit 79145a6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MCP_REQUEST_STATE_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Design MCP Server loads about 22k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 11,367 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cyanheads/pubmed-mcp-server at commit 79145a6, republished under its Apache-2.0 licence (© cyanheads). 11,367 words, ~22,185 tokens.

Download SKILL.mdSave it as .claude/skills/design-mcp-server/SKILL.md (or your agent's skills folder).
name
design-mcp-server
description
Design the tool surface, resources, and service layer for a new MCP server. Use when starting a new server, planning a major feature expansion, or when the user describes a domain/API they want to expose via MCP. Produces a design doc at docs/design.md that drives implementation.
metadata.author
cyanheads
metadata.version
2.32
metadata.audience
external
metadata.type
workflow

When to Use

  • User says "I want to build a ___ MCP server"
  • User has an API, database, or system they want to expose to LLMs
  • User wants to plan tools before scaffolding
  • Existing server needs a new capability area — two or more tools sharing a new noun or service, or any new upstream source (design the addition, not just a single tool)

Do NOT use for a single tool on an existing noun — use add-tool directly.

Inputs

Gather before designing. Ask the user if not obvious from context:

  1. Domain — what system, API, or capability is this server wrapping? Or is the server providing internal capability with no external dependency (computation, text/code utilities, in-memory state)?
  2. Data sources / source of truth — APIs, databases, file systems, external services? Or is the server itself the source (in-memory state, pure computation, local-only utility, embedded model)?
  3. Target users — what will the LLM (and its human) be trying to accomplish?
  4. Scope constraints — read-only? write access? admin operations? what's off-limits?
  5. Deployment — local stdio, hosted HTTP, Cloudflare Workers? The answer gates primitives: DataCanvas and MirrorService don't run on Workers, and a tool that asks the caller for input mid-call needs a stateful HTTP session (see the Client round-trip row in Step 3).

If the domain has a public API, read its docs before designing. For internal-only servers, skip API research and go straight to user goals. Don't design from vibes either way.

Server scope and audience

Before committing to a server boundary, answer: what workflow does this server serve, and who is the audience?

The unit of a server is a user workflow, not an API. A single rich API can earn its own server when the audience is large and the API surface supports a full workflow (PubMed for literature research, SEC EDGAR for financial analysis, Shodan for internet-wide device intelligence). Multiple APIs should collapse into one server when they serve the same workflow from different angles — a "threat intelligence" server that aggregates VirusTotal, AbuseIPDB, and GreyNoise is more useful than three separate servers because the user's goal is "assess this indicator," not "query VirusTotal."

Don't default to one-API-one-server. That's the right call when the API is deep enough and the audience is large enough, but it's not the starting point. The starting point is the workflow:

SignalServer boundary
Single API with rich surface, large audienceStandalone server named for the platform (pubmed-mcp-server, secedgar-mcp-server)
Multiple APIs serving the same workflowOne server named for the workflow (threat-intel-mcp-server), APIs are internal sources
Domain with distinct sub-audiencesConsider splitting — a pentester and a SOC analyst have different workflows even in the same domain
Pure computation, no external depsStandalone server named for the capability (calculator-mcp-server, pentest-mcp-server)

When multiple APIs collapse into one server, the tool surface is organized around what the user is doing, not which API gets called. The agent says "investigate this domain" and the server routes to the best available source internally. Individual APIs become service-layer implementation details, not tool-surface identities.

Server Naming

Usually settled before this skill runs; confirm it passes one test before designing against it. A name is earned when someone reading only the name — an npm result, a marketplace grid — forms no wrong expectation about scope. Name a wrapper for its source (pubmed-mcp-server, secedgar-mcp-server); a generic domain name is earned only by aggregating independent sources (threat-intel-mcp-server), and a single-jurisdiction scope goes in the name (uk-legislation-mcp-server). Spell out an acronym that reads as something else (libofcongress, not loc), or pair it with its domain (eia-energy, bls-labor). The tool prefix is a separate, stable identifier (see the Design table) — it need not move when the package is renamed, and renaming it is a breaking change for every client.

Steps

1. Research External Dependencies

Applies when: the server wraps an external API or service. Skip for internal-only servers (computation, local file ops, in-memory state, code analysis utilities) and jump to Step 2.

Before designing, verify the APIs and services the server will wrap. Read the docs, then hit the API — real requests reveal what docs omit.

Research inline by default — fetch docs, read SDK readmes, confirm assumptions before committing them to the design. For each external dependency:

  • Fetch API docs, confirm endpoint availability, auth methods, rate limits
  • Check for official SDKs or client libraries (npm packages)
  • Note any API quirks, pagination patterns, or data format considerations
  • Read the terms of use: whether storing, caching, or redistributing the data is permitted, whether AI or LLM use is, and what attribution the data carries. Note the credential model too — keyless, one operator key, or a key each user supplies. The terms decide whether the server can be hosted for others at all, whether a local mirror is allowed, and what the server instructions must credit.

When research is genuinely parallelizable (multiple independent APIs, several SDKs to evaluate), spawn background agents for the independent legs while you proceed with domain mapping. Skip the overhead for a single API — just read it yourself.

Live API probing. After reading docs, make real requests against the API to verify assumptions:

  • Response shapes — confirm actual field names, nesting, and types. Docs frequently lag or omit fields.
  • Batch/filter endpoints — look for filter.ids, bulk GET, or query-by-multiple-IDs patterns. A single batch request replaces N individual fetches and eliminates serial-request bottlenecks and rate-limit accumulation.
  • Field selection — check if the API supports fields or select parameters to request only the data you need. This reduces payload size dramatically for large objects.
  • Pagination behavior — verify token format, page size limits, and what happens when results exceed one page.
  • Error shapes — trigger real 400/404/429 responses to see the actual error format, not just what docs claim.
  • Unknown-param behavior — send one deliberately misspelled parameter. If the API silently ignores it (plausible but unfiltered results instead of an error), every typo'd or unverified param name becomes a silent-wrongness bug — the service layer then needs a strict allowlist of confirmed spellings, and new filters require a probe before they ship.
  • Omission semantics — for each major optional parameter, check what omitting it actually returns. Some APIs default to the intuitive scope; others silently widen (all historical versions, all statuses, global instead of regional). A default that changes result meaning becomes a server-side default plus an echoed output field, not something left to the agent.
  • Branch frequency, for bundled or bulk data — when the "API" is a dataset the server ships, probe not only whether each record shape exists but what fraction of records takes it: count the records where an optional field is absent, where a value is an array instead of a scalar, where an id maps to many keys rather than one, where a status is missing. A shape found in 0.1% of records is an edge case to handle; one found in 70% is the main path, and a design that specifies only the scalar case leaves the implementer to invent the common one.

Stopping condition: at minimum, probe one list/search endpoint, one single-item GET, one error case (force a 404 or 400), and one unknown-param request. For large APIs with many resource types, add one probe per major noun. Stop when the response shapes and error envelope are confirmed.

This step prevents building a service layer against assumed response shapes that don't match reality.

Probe with real data; write the design with made-up data. Every example value that reaches docs/design.md — names, emails, phone numbers, hosts, IP addresses, identifiers tied to a person — is synthetic, and no key, token, or private hostname appears in it. The doc is tracked, and git history keeps whatever its first commit carried.

2. Map User Goals, Then Domain Operations

Start with user goals, not endpoints. Enumerate the outcomes an agent (and its human) will actually try to accomplish with this server — usually 3–10, scaled to domain size. These drive the workflow tools that form the spine of the surface. Endpoint-inventory-first design produces 1:1 API mirrors; goal-first design produces tools agents reach for. For internal-only servers, goals map to capabilities rather than endpoints — e.g., "format markdown to GFM," "tokenize text by model," "compute file hash."

Example user goals for a project management server:

  • Find tasks I'm assigned to that are due soon
  • Create a task in a project, assign it, and notify the owner
  • Mark a task complete and log the outcome
  • Audit a project's overdue work

Then enumerate the underlying domain operations the system supports, grouped by noun. These are the raw material workflow tools compose and single-action tools back-fill where workflows don't cover an edge case.

NounOperations
Projectlist, get, create, archive
Tasklist (by project), get, create, update status, assign, comment
Userlist, get current

The user-goal list shapes the tool surface; the operation list fills in the gaps. Not every operation becomes a tool — an operation stays as raw material (not its own tool) when it's already fully covered by an existing tool's output, or when the only agents who'd use it are in scenarios outside this server's stated purpose.

3. Classify into MCP Primitives

Tools are the primary interface. Not all MCP clients expose resources, and those that do rarely surface one to the model without a human selecting it. Design the tool surface to be self-sufficient: an agent with only tool access should be able to do everything the server is built for. Resources add convenience for clients that support them (injectable context, stable URIs), but are not a reliable access path.

PrimitiveUse whenExamples
ToolThe default. Any operation or data access an agent needs to accomplish the server's purpose.Search, create, update, analyze, fetch-by-ID, list reference data
App ToolRare — default to a standard tool. Only when a human will actively interact with the result in real time and the target client supports MCP Apps. Most clients are tool-only and most agent workflows are read-by-LLM, not viewed-by-human. App tools add an iframe + CSP, app.ontoolresult/callServerTool plumbing, host-context wiring, and a format() text twin that still has to be content-complete (since most clients only see that). Two surfaces to keep in sync, two failure modes per change.Dense tabular state a human scrubs through; form-based human approval in an MCP Apps-capable client
ResourceAdditionally expose as a resource when the data is addressable by stable URI, read-only, and useful as injectable context.Config, schemas, status, entity-by-ID lookups
PromptReusable message template that structures how the LLM approaches a taskAnalysis framework, report template, review checklist
Client round-tripNot a registered primitive — a handler returns ctx.requestInput(...) to ask the client for what only it has, and is re-entered with the answer: a confirmation or form (inputRequired.elicit), an authorization or hosted-form URL (inputRequired.elicitUrl), the client model's judgment (inputRequired.createMessage — borrow the caller's model rather than bundling one), filesystem roots (inputRequired.listRoots). Design it into the tool that needs it; see Workflow tool safety and api-context. A server with any such tool declares createApp({ sessionMode: { default: 'stateful', require: 'stateful' } }) — under stateless HTTP a 2025-era client's round trip is refused, so the tool is unusable rather than guarded.Destructive-arm confirmation, OAuth consent, summarize-with-the-client's-model
NeitherInternal detail, admin-only, not useful to an LLMToken refresh, webhook setup, migrations

What the tool surface needs to cover depends on the server: a read-only research server has different economics than a CRUD project management server. Consider the domain, the expected agent workflows, whether it wraps one API or many, and what data relationships exist.

Common traps:

  • Data locked behind resources: If something an agent needs is only accessible via a resource, it's invisible to tool-only clients. That data might warrant its own tool, or it might already be covered by an existing tool's output — but it needs a tool path somewhere.
  • CRUD explosion: Don't map every REST endpoint to a tool. Related operations on the same noun often belong in one tool with an operation/mode parameter (see Step 4).
  • 1:1 endpoint mirroring: API endpoints are designed for programmatic consumers. LLM tools should be designed for workflows — what an agent is trying to accomplish, not what HTTP calls happen under the hood.

Irreversible operations stay in the UI. The "Neither" bucket above covers operations that aren't useful to an LLM. There's a second, sharper reason to exclude something from the tool surface: operations whose failure mode is catastrophic and unrecoverable. Examples span domains — dropping a production database table (data loss across every row), force-emptying a versioned cloud-storage bucket (no recovery once the lifecycle policy fires), revoking the workspace's last admin role (locks everyone out, recovery requires vendor support), GDPR permanent-delete on a customer profile (un-restorable by design), purging an analytics warehouse partition older than the retention window (auditable history gone), or deleting the single audience on a free-plan email platform (nukes every subscriber and historical report in one call). These are useful to an LLM in principle, but the blast radius of a mis-call is disproportionate to any agent workflow. Humans do these in the vendor UI, where confirmation dialogs and undo paths exist. Agents shouldn't have the tool at all.

This is distinct from destructiveHint — that annotation is for operations that are destructive but recoverable (deleting a task, reverting a commit) and agents should still have them. The "stays in the UI" line applies only to operations whose failure is both catastrophic and irreversible.

4. Design Tools

This is the highest-leverage step. Tool definitions — names, descriptions, parameters, output schemas — are the entire interface contract the LLM reads to decide whether and how to call a tool. Every field is context. Design accordingly.

Tool shapes you'll encounter

Most tools follow the {server}_{verb}_{noun} default — one focused responsibility, one clear verb, often (but not always) one upstream call. API-wrapping examples: pubmed_search_articles, pubmed_fetch_articles. Internal-only examples: markdown_format_text, regex_test_pattern, tokens_count_text — same naming convention, no external dep. Three variants warrant explicit design pressures of their own:

ShapePurposeTypical formExamples
WorkflowMulti-step orchestration that replaces a common agent chainN upstream calls (often parallelized); may request confirmation; may need mid-flow cleanupclinicaltrials_find_eligible (search → filter → rank)
InstructionState-aware procedural guidance — advice, not actionStatic markdown + a few live-state fetches, readOnlyHint: true, outputs nextToolSuggestions pre-filling the recommended follow-up. No writes.git_wrapup_instructions
ReferenceDecode opaque domain vocabulary — codes, enums, identifier formats, coverage windows — so agents can build valid inputs for the rest of the surfaceStatic tables or one cached fetch (often zero upstream calls); consolidate N lists under one topic enum; readOnlyHint: true, openWorldHint: false when offlinemedcode_list_systems, osv_list_ecosystems

These aren't boxes every tool must fit into — some blend shapes — but the design pressures differ enough that naming them helps avoid re-discovering the patterns per server.

Think in workflows, not endpoints

The unit of a tool is a useful action, not an API call. Ask: "What is the agent trying to accomplish?" — not "What endpoints does the API have?"

A single tool can call multiple APIs internally, apply local filtering, reshape data, and return enriched results. The LLM doesn't know or care about the underlying calls.

ts
// Workflow tool — search + local filter pipeline, not a raw API proxy
const findEligible = tool('clinicaltrials_find_eligible', {
  description: 'Match a patient profile to eligible clinical trials, filtering by age, sex, conditions, location, and healthy volunteer status. Results are ranked and carry a per-study eligibility explanation.',
  // handler: listStudies() → filter by eligibility → rank by location proximity → slice
});

Tip — mode consolidation. When a tool has several related operations on the same noun, you can consolidate them under one tool with a mode/operation enum. This affects both naming (noun-led, e.g., github_pull_request) and handler design (dispatch by mode). Use when it tightens the surface; skip when ops diverge enough to warrant separate tools. When the arms need different required fields (look up by ID vs. search by name), declare the input as z.discriminatedUnion('mode', [...]) rather than making every field optional and checking the combination by hand — each arm advertises its own required, the handler narrows on the discriminator, and mixed arguments are rejected. Two constraints: output stays a flat z.object, and a union root rules out headerParam. See add-tool § Multi-mode tools.

Multi-source tools and fallback chains

Applies when: a server aggregates multiple data sources for the same workflow, and the "best" source varies by input type, availability, or coverage. Skip for single-API servers.

When a tool's goal can be served by multiple sources, design it as a multi-source tool — the agent calls one tool, the handler routes to the best source (or fans out to several) internally. This is the difference between a "PubMed wrapper" and a "literature research server": a hypothetical literature_search_articles tries PubMed first, falls back to EuropePMC for broader coverage, then Unpaywall for open access. The agent doesn't choose which API to hit — the server makes that decision based on what works.

Two patterns:

Source fallback chains — try sources in priority order, fall through on failure or empty results. Best when sources cover the same corpus with different depth or availability. The output should indicate which source provided the data so the agent (and human) can assess provenance. When the fallback changes what is being searched — a different corpus, different identifiers, different licensing — don't chain: expose the second source as a sibling tool so the agent chooses the corpus knowingly (the shipped pubmed-mcp-server keeps pubmed_europepmc_search separate for exactly this reason).

Multi-source fan-out — query multiple sources in parallel, merge results. Best when sources provide complementary data about the same entity. Use Promise.allSettled so one unavailable source doesn't tank the whole call — but a rejected input is wrong for every source. Rethrow an input-class rejection (ValidationError, InvalidParams) from any source instead of folding it into a per-source failure: a source that ignores the bad value would otherwise answer for it, and the agent reads a caller mistake as an outage to retry.

ts
// Handler pseudocode — indicator enrichment across threat intel sources
async handler(input, ctx) {
  const settled = await Promise.allSettled([
    vtService.lookup(input.indicator),
    abuseIpService.check(input.indicator),
    greynoiseService.query(input.indicator),
  ]);
  for (const r of settled) {
    if (r.status === 'rejected' && isInputError(r.reason)) throw r.reason; // the caller's mistake, not an outage
  }
  const [vt, abuse, greynoise] = settled;
  return {
    indicator: input.indicator,
    sources: {
      // { status: 'ok', data } | { status: 'unavailable', error } — provenance per source
      virustotal: toSourceResult(vt),
      abuseipdb: toSourceResult(abuse),
      greynoise: toSourceResult(greynoise),
    },
    // Server synthesizes a verdict from the sources that answered — the agent gets a conclusion, not raw API dumps
    assessment: synthesizeVerdict(vt, abuse, greynoise),
  };
}

In both patterns, the tool surface is organized around what the user is doing. Sources are service-layer details — the agent sees threat_enrich_indicator, not virustotal_lookup + abuseipdb_check + greynoise_query. Mode-based dispatch by input type (e.g., indicator_type: 'ip' | 'domain' | 'hash') naturally routes to different source chains per mode, since different sources cover different indicator types.

Cut the surface

There is no fixed ceiling on tool count and no target either. The ceiling is workflow coverage — if the domain genuinely has 20 distinct workflows, expose 20 tools; the cut is per-tool overlap and reach. After mapping tools, review the full list critically. A tool that covers a niche use case, serves a tiny fraction of agents, or duplicates what another tool already handles is a candidate for deferral. Drop it from the design and note it as a future addition if demand warrants. Every tool in the surface is cognitive load for tool selection — a tight surface outperforms a comprehensive one.

Instruction tools

Applies when: the domain has recurring "how do I do X well given my current state" questions worth merging with static procedural content. Skip otherwise.

Some domains benefit from a tool whose output is guidance, not data — a markdown playbook tailored by live account state, with pre-filled next-step tool calls. These sit between Prompts (static templates, client-invokable) and action tools (do work, return data): they return advice, but the advice is worth more than static text because it merges procedural content with the agent's actual situation.

Characteristics:

  • Output is markdown guidance, not structured data (though the output schema still has fields — typically guidance, diagnostics, and nextToolSuggestions)
  • Merges static procedural content with live state — the value is the tailoring. "You have 12 staged files spanning 4 unrelated changes — split them into separate commits before pushing" beats a generic best-practices article. The same shape works in other domains: "Your slowest query is 2.3s on orders.customer_id — add the index before tuning the planner" (database advisor), "Error rate spiked 4× at 14:32 UTC, 4 minutes after the web@a3f9c2 deploy — roll back before chasing the upstream provider" (incident triage).
  • readOnlyHint: true; openWorldHint follows where the diagnostics come from — false when the live state is local (a repo on disk), true when it is fetched from an external API. No writes either way.
  • Outputs nextToolSuggestions — an array of recommended follow-up tool calls with arguments pre-filled from the diagnostics, not just tool names. The agent consumes the playbook, then executes steps with other tools.
  • Consolidate by topic enum — what could be N separate per-topic tools collapses into one
ts
const wrapupInstructions = tool('git_wrapup_instructions', {
  description: 'Get procedural guidance tailored to the current repo state: best-practice markdown merged with live diagnostics (staged/unstaged files, branch info, recent commits) and pre-filled follow-up tool calls. Read-only; execute the steps with other tools.',
  annotations: { readOnlyHint: true, openWorldHint: false },
  input: z.object({
    topic: z.enum(['review-changes', 'stage-and-commit', 'push-to-remote'])
      .describe('Playbook topic. Determines which static guidance is returned and which live state is fetched for tailoring.'),
  }),
  output: z.object({
    guidance: z.string()
      .describe('Markdown playbook content, tailored to current account state.'),
    diagnostics: z.record(z.string(), z.unknown())
      .describe('Live state used to tailor the guidance (e.g., staged file count, branch divergence, recent commit cadence).'),
    nextToolSuggestions: z.array(z.object({
      toolName: z.string().describe('Tool to call next.'),
      reason: z.string().describe('Why this step is recommended given current state.'),
      args: z.record(z.string(), z.unknown()).describe('Arguments pre-filled from diagnostics; {} when the tool takes none.'),
    })).describe('Recommended follow-up calls with arguments already populated.'),
  }),
});

Prior art: git_wrapup_instructions walks through staging, commit, and push with repo state inspected. If a server has recurring "how do I do X well given my state" questions, an instruction tool typically beats N topic-specific tools and duplicating guidance in tool descriptions.

One suggestion shape, every server. Each entry is exactly { toolName, reason, args } — no per-server renames (tool, rationale, suggestedArgs, arguments), which force every client to special-case every server. args is always present ({} for a tool that takes none), and the array is always present, empty when nothing is worth suggesting. A tool on another server is never an entry: this server cannot know it is installed, so name it in the guidance prose instead.

Data tools qualify on the same terms. A data tool carries nextToolSuggestions when the right next call depends on its own result and the arguments come from that result — a status sweep that finds a degraded vendor, a connect call that learns which services the upstream supports. A fixed chain (search → get by the returned ID) does not: the IDs in the result and the server instructions already carry it, and a suggestion on every call is token noise.

Suggestions are scoped to what this deployment registers. A nextToolSuggestions entry is an executable call, so it is only correct when its target is enabled under the same configuration — a tool wrapped in disabledTool() is absent from tools/list, and a suggestion naming it hands the agent a call that fails on dispatch. Build the array from the same config the registration reads, and when the target is off, drop the entry rather than the explanation: guidance can still say the capability is unavailable in this deployment and what to do instead. The audit and a worked example live under Feature-flagged tools in add-tool/SKILL.md.

Reference tools

Applies when: the domain speaks in opaque vocabulary — enum codes, classification systems, identifier formats, per-source coverage windows — that agents must supply as inputs elsewhere. Skip when inputs are self-evident (free text, ISO dates, well-known formats).

A reference tool is the surface's decoder ring: which codes exist, what they mean, what format each identifier takes, what each source covers. Consolidate what could be N per-list tools under one topic enum, and serve it from static tables or a cached fetch so it stays cheap to call. Its second job is structural: it is the standing target of recovery routing — error recovery strings, zero-hit notices, and resolver guidance across the surface end with "consult the reference tool, topic X" (see Error design), which only works when the tool exists. Implement it first — it usually has no service dependency, and it grounds field-testing for every other tool.

Workflow tool safety

Applies when: a tool performs multi-step mutations with destructive modes (send/apply/promote) that benefit from human confirmation before the irreversible step fires. Skip for read-only or idempotent workflows.

Tools that perform multi-step mutations (the Workflow shape) have two safety considerations beyond single-call tools. Both are about giving the agent — and the human behind it — a chance to catch a bad invocation before it commits.

Confirmation-gated destructive modes, with an annotation fallback. When a workflow's mode parameter switches between safe and destructive arms (draft vs send, plan vs apply), gate the destructive arm on a confirmation the handler asks for via ctx.requestInput(...), so a human approves before the irreversible step fires. The handler is re-entered with the answer on ctx.inputs; it does not await mid-call. The answer alone does not prove the prompt was shown — a client can send one unasked, and any requestState replays within its lifetime — so the gate stores the operation, the caller, the confirmed target, and a hash of what it holds in a ctx.state record keyed by a random id, sends only that id as requestState, and redeems the record before the destructive arm runs (api-context § Consent gates). Plan shared storage for that record when a 2026-07-28 retry can reach another instance (filesystem, supabase, or cloudflare-d1 — not cloudflare-kv), and set MCP_REQUEST_STATE_KEY so a retry can only carry state the server minted. Redeeming is not atomic — concurrent retries on one id can each pass the gate until ctx.state gains an atomic take (#593) — so an arm that must not run twice is idempotent per record.

The gate is always reachable — ctx.requestInput is present on every transport and both protocol revisions (2025-11-25 legacy, 2026-07-28 current) — but it is not always answerable: a client that never fulfils the input_required result simply doesn't retry, and — with the record pattern, which proceeds only on a record it redeemed — the destructive step never runs. The same holds for a 2025-11-25 HTTP client when the server runs MCP_SESSION_MODE=stateless: the legacy round-trip shim still runs, but its capability gate refuses because the serving instance never processed initialize — the destructive step never fires. That is the safe outcome, but it makes the tool unusable for those clients, so a server built around such a gate declares createApp({ sessionMode: { default: 'stateful', require: 'stateful' } }) and refuses to start stateless rather than degrading (api-context § ctx.requestInput). Keep destructiveHint: true in annotations so those clients' own approval flows still surface the risk. A decline is terminal — the handler fails the call rather than re-asking, which would loop until the round budget runs out. ctx.clientCapabilities never decides whether to ask: a gate that skips its prompt when elicitation is undeclared is the bypass the gate exists to prevent.

Safe defaults on parameters that determine blast radius. When a workflow accepts a parameter that controls how far-reaching a mutation is, default to the safer value. A bulk file-update tool defaulting mode: 'preview' (no writes) means a sloppy agent call shows a diff rather than blasting changes; an apply-plan tool defaulting dryRun: true means a misread plan previews rather than executes; an object-delete tool requiring an explicit confirmCount matching the result-set size means an unscoped query can't silently nuke a million rows. Agents that genuinely want the destructive behavior have to name it explicitly, which surfaces intent in the tool call and in logs.

Make retried writes safe. Agents re-issue a call that timed out. A write that can double-apply (create, send, charge, enqueue) takes a caller-supplied idempotency key or resolves to upsert semantics — and only then earns idempotentHint: true.

Tool descriptions

The description is the LLM's primary signal for tool selection. It must answer: what does this do, and when should I use it?

  • Imperative present tense, capability first. "Search for clinical trial studies using queries and filters" beats "Interact with studies", and beats "Searches for…", "This tool…", "Allows you to…". The tool-defs-analysis skill audits this across a finished surface.
  • Include operational guidance when it matters. If the tool has prerequisites, constraints, or gotchas the LLM needs to know, say so in the description. Don't add boilerplate workflow hints when the tool is self-explanatory.
  • Prefer a single cohesive paragraph. Pack operational guidance into prose sentences (separated by periods or em-dashes) rather than bullet lists or blank-line-separated sections. Descriptions render inline in most clients, and bullet structure reads as visual noise rather than signal. Operation-by-operation bullets also duplicate info that already lives in the operation enum's .describe().
  • Don't leak. Descriptions are for the consumer, not the author. Three categories to audit against:
    • Implementation details — endpoint paths, API call counts, internal parameter mappings, routing logic. Describe what the tool does and when to use it, not how it's wired up.
    • Meta-coaching — directives about how to use the output. "Treat X as the canonical Y", "callers should…", "the LLM should…". The description sells the tool; it doesn't coach the reader.
    • Consumer-aware phrasing — references to "LLM", "agent", "Claude", or any specific reader. The description shouldn't name who's reading it.
ts
// Good — describes a prerequisite the LLM must know
description: 'Set the session working directory for all git operations. This allows subsequent git commands to omit the path parameter.'

// Good — self-explanatory, no workflow hints needed
description: 'Show the working tree status including staged, unstaged, and untracked files.'

// Good — warns about constraints
description: 'Fetch trial results data for completed studies. Only available for studies where hasResults is true.'

Descriptions should be as long as needed — concise but complete. Don't artificially truncate, and don't pad with filler.

Parameter descriptions

Every .describe() is prompt text the LLM reads. Parameters should convey: what the value is, what it affects, and (where non-obvious) how to use it well.

  • Constrain the type. Enums and literals over free strings. Regex validation for formatted IDs. Ranges for numeric bounds.
  • The input root is already strict. tool() applies .strict() at the root and advertises additionalProperties: false, so an unknown top-level key is rejected by name instead of silently stripped; nested objects still strip unless made strict themselves. Open the root with .passthrough() (Zod 4 also spells it .loose(); tool() honors either) only on a tool that deliberately proxies arbitrary upstream parameters (a raw-query tool), and say so in its description.
  • A blank optional string is unset. Form-based clients send every optional field they display as "". Treat the blank as omitted — left off the upstream request, never forwarded as param=, which some APIs read differently from omission — and never design a .min(1) onto an optional field to catch it. add-tool has the schema pattern that keeps a validator on the field.
  • Numeric bounds earn their .max(). An output-only ceiling the content already bounds (maxSections, a character budget) takes none: a larger value means no limit. A workload cap (an ID batch, limit) keeps a hard ceiling against absurd input. Between the per-call cap and that ceiling, a read — or an operation whose contract allows partial completion — processes up to the cap and names the remainder, disclosed like a capped list (below), since a rejection for a value with one obvious reading only costs the caller a round trip. A write, destructive, or otherwise atomic batch over its cap rejects before doing anything, naming the cap, so its .max() is the cap itself: doing the first items and naming the rest splits one intended operation in two and leaves the caller to work out which items changed.
  • Use JSON-Schema-serializable types only. The MCP SDK serializes schemas to JSON Schema for tools/list. Types like z.custom(), z.date(), z.transform(), z.bigint(), z.symbol(), z.void(), z.map(), z.set() throw at runtime. Use structural equivalents (e.g., z.string().describe('ISO 8601 date') instead of z.date()).
  • Explain costs and tradeoffs when a parameter choice has meaningful consequences.
  • Name alternative approaches when a simpler path exists.
  • Include format patterns for structured values, but don't pad descriptions with redundant examples.
ts
// Good — explains cost, recommends action, names the alternative
fields: z.array(z.string()).optional()
  .describe('Specific fields to return (reduces payload size). Without this, the full study record (~70KB each) is returned. Use full data only when you need detailed eligibility criteria, locations, or results.'),

// Good — explains what the flag does AND how to override
autoExclude: z.boolean().default(true)
  .describe('Automatically exclude lock files and generated files from diff output to reduce context bloat. Set to false if you need to inspect these files.'),

// Good — names the format and gives one example
nctIds: z.union([z.string(), z.array(z.string()).max(5)])
  .describe('A single NCT ID (e.g., "NCT12345678") or an array of up to 5 NCT IDs to fetch.'),

Input-edge normalization. For every identifier, code, or enum-ish input, enumerate at design time the variants a caller will plausibly send, and decide per variant: normalize, or error. The rule is normalize what is certain, error on what is ambiguous. A variant is certain when the mapping is unambiguous, one-to-one, and preserves the submitted meaning exactly — then the call succeeds instead of returning a miss for a value the server could resolve. Everything short of that is an error naming the expected shape, with a recovery routing to the reference tool. Never fuzzy-match a value, broaden a query, swap one entity for another, or drop a filter to make a call succeed: plausible-looking rows from a guessed input are worse than a rejection the agent can act on.

ClassExampleHandling
Case or bare-leaf shorthand of a codeACS5, acs5 → acs/acs5Normalize before lookup
Domain value aliasmph → m/h, kph → km/hAlias table at the input edge
Composite identifier completed by contextpart: "52" + section: "21" → 52.21Try as-given first, retry the composed form on a miss — never rewrite unconditionally, since the bare form can be legitimate
Delimiter-joined list where an array is accepted"US,JP,KR" → ["US","JP","KR"]Split on the documented separator
Spelled-out vs. abbreviated name"Houston, Texas" → "Houston, TX"Normalize against the bundled name table

These are value-level, and the mappings are domain knowledge — settle them per input in the design doc's param table. Argument key names are not: the framework rewrites declared and case-style key aliases and drops client-added root keys before the schema sees the arguments, and, against the tool's own schema after a failed parse, repairs a JSON-stringified array or object, an integer sent for a string, a number or boolean sent as a string, and a lone string sent for a list, and deletes null sent for an optional field — so an ID field stays z.string(), never a string | number union, a list field stays z.array(), never a string | string[] union, and an optional field never needs .nullish() to absorb a client that sends null for unset. A delimiter-joined list is still yours to split: the framework never wraps a string holding a comma or a line break. Don't re-implement any of that per server — see add-tool § Three things the framework fixes before the schema sees the arguments.

This resolves one submitted value to one canonical value, and does not loosen the strict token match in MCP-side list filtering, which scores a query against many candidate names.

Output design

The output schema and format function control what the LLM reads back. Design for the agent's next decision, not for a UI or an API consumer. See the add-tool skill's Tool Response Design section for implementation-level patterns (partial success, empty results, metadata, context budget).

Principles:

  • Server reports what only the server can know; agent decides what only the agent can know. Schema, scopes, rate limits, and raw observable state belong to the server. Semantic correctness, intent-vs-effect matching, and recovery choice belong to the agent. For mutators, this means surfacing pre/post observable state rather than throwing on synthetic deltas the server can't authoritatively classify — file shrunk could be deliberate truncation or a bug; only the agent knows. See add-tool skill's Mutator response design.
  • Include IDs and references for chaining. If the agent might act on a result, return the identifiers it needs for follow-up tool calls.
  • Curate vs. pass-through depends on domain. Medical/scientific data — don't trim fields that could alter correctness. CRUD responses — return what the agent needs, not the full API payload. Match fidelity to consequence.
  • Absent upstream data stays absent. Sparse APIs omit fields; declare those output fields .optional() and render "Not available" in format() rather than coercing to false, 0, or "" — a fabricated value is worse than a gap.
  • Third-party text is data, and format() marks it. When a field carries text other people wrote — posts, reviews, bios, comments, summaries, a user-named place or project — list those fields in the design and plan how format() sets them apart: a blockquote or fence for free text, and CR/LF flattened to a space wherever a value is interpolated inline (a heading, a bolded name, a matched on "…" line), so a newline in the value can't forge structure in content[]. structuredContent keeps the value verbatim. Say in the server instructions that this content is data, never instructions; security-pass audits the result.
  • Image and audio bytes ride ctx.content, never output. ctx.content.image(data, mimeType) / .audio(...) emit a content[] block once; output keeps the metadata the agent reasons over (dimensions, duration, a reference). Base64 in a typed output field ships the bytes twice.
  • Surface what was done, not just results. After a write operation, include the post-state so the LLM can chain without an extra round trip.
  • When the effect lands after the call, wait for it by default. Some actions are fire-and-forget at the wire — send a wake packet, trigger a job, dispatch a notification, provision a resource — and their immediate result ("sent", "queued", "accepted") answers nothing the agent asked; the agent wants to know whether the machine is up, the job ran, the resource exists. Design the tool to confirm: pre-probe the observable state (cheap; if it is already in the target state, skip the action and say so), act, then poll with early return until the state is observed or a bounded window elapses. Express the window as one numeric parameter with a default (wait_for_s: 30), where 0 means act and return — not a boolean plus a timeout, which is two parameters for one decision and an awkward default. Every terminal outcome is a result, never a throw: already_<state>, <state> (with elapsed time), not_<state> within the window (with guidance naming the read-only check tool to re-poll), and unverified (window 0, or nothing to probe). Size the default to cover the common cases while staying inside client tool timeouts, and pair the action with a read-only sibling that probes the same state so an agent can re-check without re-acting.
  • Seed orientation context alongside the primary result. When a tool's call position makes the agent's next moves predictable, attaching a compact snapshot of relevant state — recent activity, tracked state, a couple of reference items — both saves round-trips and primes the LLM on the project's patterns. Surfacing recent commits teaches the commit-message style the agent should match when it later writes one; recent tags teach the versioning convention; reference records teach the naming format. Common fits: tools that open or close a session (set working dir, wrap-up), state-changing verbs where the caller wants post-action confirmation (commit, push, merge), entry points that drop the agent into a new scope (clone, checkout). Gather sub-operations in parallel with Promise.allSettled so a single failure degrades to a warning rather than tanking the outer call.
  • Communicate filtering. If the tool silently excluded content, tell the LLM what was excluded and how to get it back — an excludedFiles list whose description says "call again with autoExclude=false". The agent can't act on what it doesn't know about.
  • Empty results are a designed surface, not a fallthrough. For every search/list tool, spec the zero-hit behavior at design time: zero hits are success + an enrichment notice, never an error, and the notice is composed from condition → fragment pairs (too-narrow filter, a defaulted date window, a syntax trap), each routing to a concrete next call — relax a named filter, switch to the named sibling tool, or consult the reference tool. Relatedly, when the server applies a default that changes result semantics (an as-of date, an implicit status filter), echo the applied value in the output — an agent can't reason about a filter it can't see.
  • Capped lists disclose truncation. When a tool accepts a cap-like input (limit, per_page, page_size, max_results, max_items) and returns an array, the handler must disclose when the cap was hit — truncated: true, shown, cap in the enrichment block via ctx.enrich.truncated({ shown, cap }), and ctx.enrich.total(n) for the full count. The same applies to a display cut in format(): show the top N and say "...and X more". Silent caps leave the agent treating a partial set as complete; the capped-list-no-truncation lint rule enforces the input-cap case.
  • Continuation is a designed field. Truncation says the cap was hit; continuation says how to get the rest. Return an opaque cursor plus has_more (via extractCursor/paginateArray for local sets), and never invent page numbers over a cursor-based upstream — a page the agent can't ask for is a page it will never see.
  • Spill big analytical results to a queryable surface. When a tool's row set is something an agent would run SQL over and can exceed any reasonable context budget — paginated APIs, streamed exports, big query results — pair an inline preview with a DataCanvas table holding the full set (spillover() in api-canvas), and compute distributions or refinement hints across the full result, not the preview, so aggregate signal stays honest. The gates on when a canvas earns its keep are in Step 7.
  • Outline one large document into sections. When a single tool call returns one document-shaped record (not many rows) that can exceed context — a ~130KB FDA drug label, a big API entity dominated by a few fat fields — return a section outline (top-level keys + per-section byte size) instead of truncating, and let the agent re-call with sections: [...] to pull only what it needs. outlineOnOverflow() (@cyanheads/mcp-ts-core/utils) returns a full | outline result; pure measure + key-slice, so Cloudflare Workers-portable, unlike canvas-bound spillover(). Distinct from spillover on shape: spillover splits a row collection, this outlines one fat record. Schema shape and format() parity are in the techniques skill's outline-on-overflow reference.
  • Mirror a bulk upstream instead of paginating it live. When the server wraps a large or slow API whose corpus is queried far more than it changes, sync it once into a persistent local index and query that as the primary data path — not the live API per request. Two gates come first. The upstream's terms must permit storing the data (Step 1), and a local index must reproduce what the upstream's search returns: an API that ranks server-side — field weighting, synonym or ontology expansion — can't be mirrored for search, because the local index answers the same query with different rows and nothing flags it. Exact-key lookups and structured filters mirror faithfully; for rate-limit relief on a ranked-search API, cache responses per request with a short TTL instead. Match the backend to corpus size: below ~10⁴ rows → an in-memory index (server-level, no primitive); ~10⁴–10⁷ → the MirrorService (embedded SQLite + FTS5; declare a schema + a sync ingester via defineMirror/sqliteMirrorStore, then runSync/query, see api-mirror); above ~10⁷ → an external store. Distinct lifecycle from DataCanvas: a mirror is long-lived and cross-session, refreshed on a schedule; canvas is ephemeral and per-session.
  • Two client surfaces, both content-complete. Different MCP clients forward different surfaces to the model: some (e.g., Claude Code) read structuredContent from output, others (e.g., Claude Desktop) read content[] from format(). format() is the markdown twin of structuredContent, not a summary — a thin format() that returns only a count or title leaves content[]-only clients blind (the format-parity lint catches this). Agent-facing context that is not domain payload — empty-result notices, the query as the server parsed it, echoed defaults, totals — goes in the enrichment block via ctx.enrich(...), which reaches both surfaces automatically; hand-authored into format() text alone it reaches only one. Field-by-field rendering of that block is in the Design table's Enrichment row.
Show full SKILL.md (4,523 more words)Show less
Batch input design

Applies when: the upstream API supports batch requests (filter-by-IDs, bulk GET) OR agents commonly need multiple items per call. Skip for inherently single-target operations.

Some tools naturally operate on multiple items — fetching several entities, updating a set of records, running checks across a list. Decide during design whether a tool accepts single items, arrays, or both.

When to accept array input:

Accept arrayKeep single-itemSeparate batch tool
The upstream API supports batch requests (fetch-by-IDs, bulk update)The operation is inherently single-target (read a file, run a query)Batch has fundamentally different output shape or error semantics
Reduces N+1 round trips for a common workflowArray input adds complexity with no backend efficiency gainSingle-item tool is simple; batch version needs progress, partial failure handling
Agent commonly needs multiple items in one stepThe tool already returns a collection (search results)

If a tool accepts arrays, design for partial success. When 3 of 5 items succeed, the agent needs to know which succeeded, which failed, and why — not just a success/failure boolean. Plan the output schema to report per-item results:

ts
output: z.object({
  succeeded: z.array(ItemResultSchema).describe('Items that completed successfully.'),
  failed: z.array(z.object({
    id: z.string().describe('Item ID that failed.'),
    error: z.string().describe('What went wrong and how to resolve it.'),
  })).describe('Items that failed with per-item error details.'),
}),

Single-item tools don't need this — they either succeed or throw. The partial success question only arises when the tool can partially complete.

Convenience shortcuts for complex inputs

Applies when: a tool wraps a structured query language or filter system where the 80% case is a simple string. Skip when the primary input is already simple.

When a tool wraps a complex query language or filter system, provide a simple shortcut parameter for the 80% case alongside the full-power escape hatch. This keeps simple queries simple while preserving full expressiveness.

ts
// text_search handles the common case; query handles everything else
text_search: z.string().optional()
  .describe('Convenience shortcut: full-text search across title and abstract. For structured filters or field-specific matching, use the query parameter instead.'),
query: z.record(z.string(), z.unknown()).optional()
  .describe('Full query object for structured filters. Supports operators: _eq, _gt, _and, _or, ...'),

The pattern: name the shortcut for what it does (text_search, name_search), document what it expands to, and point to the full parameter for advanced use. Validate that at least one of the two is provided.

MCP-side list filtering

Applies when: an upstream API has no native search, the relevant set is bounded (fits one or a few fetches), and an agent needs to resolve a name → opaque ID. Skip when the API already searches, or when the set is unbounded (bills, votes, filings) — that belongs in the DataCanvas dataframe layer (*_dataframe_query), not an in-memory filter.

Two params, two behaviors — keep them named distinctly:

  • query → upstream full-text search. The API does the work; it may honor operators and ranking.
  • a local filter param → fetched-then-filtered on our side. Name it for the mechanic: filter or nameContains (the latter self-documents the local, name-keyed half of the split). Don't overload query for it — the two have different semantics and different cost.

Earns-its-keep gate — all must hold: bounded set; no native upstream search; real scan pain (opaque IDs, a large/unordered list, or a default page that hides relevant rows); and it filters the natural lookup key (name/title). When any fails, skip it — paginate, or send the agent to upstream query.

Correctness: filter the complete bounded set, not the current page. Fetch up to the cap (or page through) before filtering — filtering one page returns a misleading partial slice.

Matching: strict token match is the default. Normalize (lowercase, strip punctuation/diacritics) and require every query token to appear, so word order and missing interior words still match. That strict core is the ~90% case and needs no fuzzy library. Add a fuzzy fallback only when a caller genuinely needs typo tolerance (an LLM caller rarely does): fire it only when the strict match is empty, score against the best-matching token in each name (not the whole string) and cap the results — or one short query clears the threshold against dozens of long multi-word names — and label its hits approximate. Often a bare "no match — call the unfiltered list to browse" beats an approximate guess: it lets the model self-correct instead of committing to the wrong record. See add-tool for the param + handler implementation.

Error design

Errors are part of the tool's interface — design them during the design phase, not as an afterthought. Three aspects: the contract (which failures are public), classification (what error code), and messaging (what the LLM reads).

Declare a typed contract for domain failures. When a tool has known failure modes the agent should plan around (no_match, queue_full, vendor_down), enumerate them as errors: [{ reason, code, when, recovery, retryable? }] on the definition. recovery is required metadata — the agent's next move when this failure fires (≥ 5 words, lint-validated; the framework sends it on the wire as data.recovery.hint with any failure carrying that reason and no hint of its own). The framework types ctx.fail(reason, …) against the declared reason union (typos become TS errors) and auto-populates data.reason on the thrown error for stable observability. The error reaches clients with parity across both surfaces — structuredContent.error (Claude Code) and content[] text (Claude Desktop). Baseline codes (InternalError, ServiceUnavailable, Timeout, ValidationError, SerializationError, RequestCancelled) bubble from anywhere and don't need to be enumerated. Mark an entry the service layer throws, rather than the handler, with thrownBy: 'service' so the conformance lint doesn't report it as a reason the handler never raises. See api-errors skill for the full pattern.

Classify errors by origin. Different error sources need different codes and different recovery guidance. Map the failure modes for each tool during design:

OriginExamplesError codeAgent can recover?
Client inputBad ID format, invalid params, missing required field, out-of-range valueValidationErrorYes — fix the input and retry
Upstream API5xx, network errorServiceUnavailableMaybe — retry later, or the upstream is down
TimeoutUpstream 408/504, a fetch that ran out its timeout, an exhausted retry deadlineTimeoutMaybe — retry, or narrow the request so it finishes sooner
Rate limit429, quota exhausted, queue fullRateLimited (retryable: true; withRetry honors Retry-After)Yes — wait, then retry or reduce frequency
Not foundValid ID format but entity doesn't existNotFound (or ValidationError if ambiguous)Yes — check the ID, try a search
ConflictDuplicate key, version mismatch, concurrent modification on a writeConflictYes — re-read current state, then retry with it
Caller authInsufficient scopes, expired token, a rejected key the caller suppliedForbidden / UnauthorizedMaybe — escalate or re-auth
Server credentialThe server's own upstream key is missing, or the upstream rejects it (401/403)ConfigurationError — translated in the service, since the automatic status mapping yields Unauthorized/Forbidden, which read as the caller's credentialsNo — the operator fixes it; the recovery names the env var
Server internalParse failure, missing config, unexpected stateInternalErrorNo — server-side issue

(InvalidParams also exists — the framework's parseToolArguments emits it when input fails Zod schema validation before the handler runs. Anything the handler itself throws about inputs uses ValidationError.)

The framework auto-classifies many of these at runtime (HTTP status codes, JS error types, common patterns), but explicit classification in the handler gives better error messages. For declared contract failures, throw via ctx.fail('reason', …). For ad-hoc throws outside the contract, use error factories (notFound(), validationError(), etc.) when the code matters; plain throw new Error() when the framework's auto-classification is good enough.

Expected misses are results, not errors. When a tool's whole job is resolving one identifier — a citation, a code, a name → ID — a no-match is an expected outcome the agent must reason about, not a failure. Return { found: false, guidance } instead of throwing, and treat guidance as a first-class recovery surface: per miss outcome, say what didn't parse or resolve and route to the named tool that can recover (the broader search tool, the reference tool). Agents self-correct better from a structured miss than from a throw — a throw reads as "something broke," a miss result reads as "adjust and retry." The split with search tools: a search's empty result is a valid empty collection, so its notice rides enrichment; a resolver's miss is the primary result, so found and its guidance live in output.

Write error messages as recovery instructions. The message is the agent's only signal for what to do next — and the strongest recovery instruction ends in a named tool call, never a bare "check your input."

ts
// Bad — dead end, no recovery path
throw new Error('Not found');

// Good — names both resolution options
"No session working directory set. Please specify a 'path' or use 'git_set_working_dir' first."

// Good — structured hint in error data using the canonical `data.recovery.hint` shape.
// The framework mirrors `data.recovery.hint` into the content[] text as
// `Recovery: <hint>` so format()-only clients (Claude Desktop) see the same
// guidance structuredContent clients (Claude Code) read from `error.data.recovery.hint`.
// A hint the message already contains verbatim is dropped from the text rather
// than stated twice, and stays on structuredContent either way.
throw forbidden(
  "Cannot perform 'reset --hard' on protected branch 'main' without explicit confirmation.",
  {
    branch: 'main',
    operation: 'reset --hard',
    recovery: { hint: 'Set the confirmed parameter to true to proceed.' },
  },
);

// Good — upstream error with actionable context
throw notFound(`Paper '${id}' not found on arXiv. Verify the ID format (e.g., '2401.12345' or '2401.12345v2').`);

During design, settle the full contract for each tool — reason, code, when-clause, and the verbatim recovery string — in the tool's section of the design doc; they become the literal errors: [...] entries during scaffolding. Hold every recovery string (and zero-hit notice, and resolver guidance) to the no-dead-ends rule: it names the concrete next tool call, with the reference tool as the most common routing target. Settled at design time these stay sharp; left to implementation they degrade into "check your input." Not every failure needs a contract entry; baseline infrastructure errors (5xx, timeouts, validation) are fine to let bubble.

A routing target must be callable in the deployment doing the routing. A recovery string, notice, or guidance line that names a config-gated tool is a dead end wherever that gate is off — the agent is sent to a tool absent from tools/list, at the moment it is already recovering from a failure. Prefer routing to ungated tools (the reference tool is a good target precisely because nothing gates it). Where the target genuinely is gated, resolve the text from the same config that decides registration, and say the capability is unavailable in this deployment rather than naming a call that cannot be made. Structured follow-ups are stricter still — see Instruction tools above.

Design table

Summarize each tool:

AspectDecision
NameLowercase snake_case with a canonical server prefix. 3 segments is the strong default ({server}_{verb}_{noun} — e.g., pubmed_search_articles, clinicaltrials_find_eligible). 2 is fine only when the verb is a complete action whose object the domain implies (git_pull, git_push, git_status, git_commit — the remote, working tree, or repo is implicit); don't invent a word to pad those to 3. A verb that takes a caller-specified object — search, find, get, list, query, fetch, connect, create, update — always carries its noun (ontology_search → ontology_search_terms, geofeatures_connect → geofeatures_connect_endpoint). Smell test: if {server}_{verb} leaves "…what?" unanswered, the noun is missing. 4 is fine when the noun is inherently two words (openfda_search_device_clearances) or the prefix is multi-part. The prefix is judged on clarity, not length: the brand name or the plain well-known word for the domain both pass (pubmed_, patents_, earthquake_); an abbreviation fails only when it reads as something else out of context (loc_ → lines of code, ct_ → CT scan). The verb+noun pair should be unambiguous within the server — if two tools could plausibly share a name, the noun isn't specific enough (read_fulltext not read_text when structured metadata is a separate concept). Treat name length as a scope smell only when the extra segment is the verb overreaching (e.g., foo_create_and_send_notification → split or use modes).
GranularityScope each tool to one coherent agent action. The implementation can be a single API call (pubmed_search_articles), a multi-step workflow, or internal-only — match the unit to the work, don't constrain by call count.
DescriptionConcrete capability statement. Add operational guidance (prerequisites, constraints, gotchas) when non-obvious.
Input schema.describe() on every field. Constrained types (enums, literals, regex). Explain costs/tradeoffs of parameter choices.
Output schemaDesigned for the LLM's next action. Include chaining IDs. Communicate filtering. Post-write state where useful.
ErrorsDeclare domain failure modes as a typed contract (errors: [{ reason, code, when, recovery, retryable? }]) so ctx.fail is type-checked and capable clients can preview failures via tools/list. Every recovery string follows the no-dead-ends rule — it names the next tool call.
EnrichmentThe success-path counterpart to errors: declare the agent-facing context fields the handler populates via ctx.enrich(...) — zero-hit notice, echoed defaults, totals, truncation — with a kind-tag (notice/total/echo/delta) where one fits and an enrichmentTrailer.render for any structured field. Keys stay disjoint from output. Declare a field required only when every path writes it — a required field one branch skips fails the output parse on every call that takes another branch.
AnnotationsreadOnlyHint, destructiveHint, idempotentHint, openWorldHint. Helps clients auto-approve safely. destructiveHint defaults to true on any tool that isn't read-only, so a benign write must set destructiveHint: false explicitly; a read-only tool omits it entirely (annotation-coherence lint).
Auth scopestool:<snake_tool_name>:<verb> or resource:<kebab-resource-name>:<verb> (e.g., tool:inventory_search:read, resource:echo-app-ui:read). Domain-led <domain>:<verb> (e.g., inventory:read) is an acceptable alternative — pick one convention per server and stay consistent. Skip when no deployment will run MCP_AUTH_MODE=jwt or oauth — under none (stdio, or single-tenant HTTP) scope checks never run.
5. Design Resources

Resources are supplementary — a convenience for clients that support injectable context via stable URIs. Since many clients are tool-only, verify that any data exposed via resources is also reachable through the tool surface. This doesn't require a 1:1 resource-to-tool mapping — the data might be covered by an existing tool's output, bundled into a broader tool, or warrant its own dedicated tool, depending on the server's purpose and how agents will use it.

For each resource:

AspectDecision
URI templatescheme://{param}/path. Server domain as scheme. Keep shallow.
ParamsMinimal — typically just an identifier. Complex queries belong in tools.
PaginationNeeded if lists exceed ~50 items. Opaque cursors via extractCursor/paginateArray.
list()Provide if discoverable. Top-level categories or recent items, not exhaustive dumps.
Cache hintcacheHint: { ttlMs, cacheScope } on 2026-07-28 connections. Static reference data is public with a long TTL; per-tenant or per-session data is private or uncached.
CompletionTemplate params with a bounded vocabulary (a species list, a dataset code) get argument completion so clients can offer valid values; the same completable() wrapper applies to prompt args.
Tool coverageVerify the data is reachable via tools — either a dedicated tool, included in another tool's output, or not needed for tool-only agents.
6. Design Prompts (if needed)

Optional. Use when the server has recurring interaction patterns worth structuring:

  • Analysis frameworks, report templates, multi-step workflows

Wrap enum-like prompt args in completable() so clients can offer valid values. Skip prompts entirely for purely data/action-oriented servers.

7. Plan Services and Config

Services — one per external dependency (or per source, for multi-source servers). Init/accessor pattern. Skip if all tools are thin wrappers with no shared state. For multi-source servers, each upstream API gets its own service with its own auth, rate limits, and retry config — tools compose across services internally, agents never see the service boundary.

Server-as-service. When the server IS the source of truth (knowledge graph, in-memory task tracker, local scratchpad, embedded inference wrapper), the resilience table below doesn't apply — there's no upstream to retry. The design questions shift to state management: what's tenant-scoped vs. global, what TTLs apply, what survives a restart, what the storage backend is. Plan persistence via ctx.state for tenant-scoped KV (auto-namespaced by tenantId), or use a StorageService provider directly when data must cross tenants. Service init still happens in setup(), accessed via getMyService() at request time. Calls within the server are local and synchronous-ish — the API-efficiency table below also doesn't apply.

Analytical API servers: DataCanvas is one option. For servers that fetch analytical data — result sets an agent runs SQL over (aggregate, group, join, time-series) — and want to expose a SQL workspace, the framework's optional DataCanvas primitive (Tier 3, opt-in via CANVAS_PROVIDER_TYPE=duckdb) handles lifecycle, ID generation, eviction, and export wiring so you don't design your own. It earns its keep on shape, not size: a discovery/search surface returning categorical metadata (titles, IDs, types) — where the workflow is find-the-record-then-drill-in — does not qualify even when the result is large; resolve names over a bounded set with MCP-side list filtering instead. If you opt in, the consumer tools are mandatory: a tool that emits a canvas_id MUST be paired with a dataframe_query (and dataframe_describe) tool in the same surface — a canvas_id with no query tool is dead output the agent can't reach. Surface canvas_id as an optional input on register/query/export tools; the framework mints on omit and resolves on match. The accessor is wired once in setup() via setCanvas(core.canvas) (undefined when disabled or running on Cloudflare Workers — DuckDB has no V8-isolate build). See api-canvas for the full reference.

For services wrapping external APIs, plan the resilience layer.

ConcernDecision
Retry boundaryService method wraps full pipeline (fetch + parse), not just the network call. Use withRetry from /utils.
Backoff calibrationMatch base delay to upstream recovery time: 200–500ms (ephemeral), 1–2s (rate-limited), 2–5s (degraded).
HTTP status checkfetchWithTimeout already handles this — a non-2xx throws an McpError whose code is mapped from the status (400 → InvalidParams, 401 → Unauthorized, 403 → Forbidden, 404 → NotFound, 409 → Conflict, 422 → ValidationError, 429 → RateLimited, 408/504 → Timeout, other 5xx → ServiceUnavailable; the full table is api-errors § HTTP Response → McpError), with status and body on error.data, plus retryAfter when the upstream sent one. Plan error contracts around those codes, not a blanket ServiceUnavailable.
Parse failure classificationResponse handler detects HTML error pages and throws transient errors, not SerializationError.
Exhausted retry messagingwithRetry enriches the final error with attempt count automatically.
Total deadlineRetries multiply a per-attempt timeout: four 30s attempts plus backoff outlast a client's 60s request timeout, and the caller gets a transport timeout instead of this server's classified error. Size withRetry's deadlineMs to fit inside the client's timeout and thread attempt.signal into each fetch (api-utils).
PacingWhen the upstream mandates a rate (one request per second, N per minute, N concurrent), put a createPacer from /utils in front of that service — one pacer per upstream budget, composed as withRetry outside and the pacer inside — and note the resulting ceiling in the server instructions when it shapes how an agent should batch work.
Caller-supplied URLs or hostsRoute through fetchWithTimeout with rejectPrivateIPs: true — the SSRF guard is off by default. It blocks private, loopback, link-local, and metadata ranges, and is best-effort: DNS rebinding still gets past it, so a deployment that needs hard isolation adds egress controls. Never a bare fetch on a caller-controlled destination; security-pass audits this sink.

For API efficiency, design the service methods to minimize upstream calls:

ConcernDecision
Batch over N+1If the API supports filter-by-IDs or bulk-GET endpoints, use a single batch request instead of N individual fetches. Cross-reference the response against requested IDs to detect missing items.
Field selectionIf the API supports fields/select parameters, request only the fields the tool needs. A full study record might be 70KB; selecting 4 fields might be 5KB.
Request consolidationWhen a tool needs data from multiple related endpoints, check if a single endpoint with broader field selection can serve the same data in one round trip.
Pagination awarenessIf a batch request might exceed the API's page size, either paginate internally or assert/throw when results are truncated so callers aren't silently missing data.

Config — list env vars (API keys, base URLs). Goes in src/config/server-config.ts as a separate Zod schema.

8. Write the Design Doc

Create docs/design.md with the structure below. The MCP surface (tools, resources, prompts) goes first — it's what matters most and what the developer will reference during implementation.

markdown
# {{Server Name}} — Design

## MCP Surface

### Tools
| Name | Description | Key Inputs | Annotations |
|:-----|:------------|:-----------|:------------|

### Resources
| URI Template | Description | Pagination |
|:-------------|:------------|:-----------|

### Prompts
| Name | Description | Args |
|:-----|:------------|:-----|

## Overview

What this server does, what system it wraps, who it's for.

## Requirements

- Bullet list of capabilities and constraints
- Auth requirements, rate limits, data access scope
- Deployment targets (stdio / HTTP / Workers), session mode, upstream credential model, and the terms-of-use constraints from Step 1

## User Goals

1. Numbered outcomes agents will accomplish (from Step 2) — every tool in the surface traces back to one

## Tools — detail

One subsection per tool: a param table (param | type | maps-to | notes), the output field
list, the error contract table (reason | code | when | recovery — verbatim strings), and
zero-hit notice fragments for search tools. This is the section implementation reads
tool-by-tool.

## Resources — detail / ## Prompts — detail

Same treatment when the server has them: one subsection per resource (URI template, params,
cache hint, the tool that also covers the data) and per prompt (args, completions, the
message shape).

## Services
| Service | Wraps | Used By |
|:--------|:------|:--------|

## Config
| Env Var | Required | Description |
|:--------|:---------|:------------|

## Server Instructions

Draft `instructions` string for `createApp()` — the orientation every client sees at
initialize: canonical workflow chain, identifier semantics, rate-limit posture. A short
paragraph under 2,048 characters, essentials first — Claude Code truncates the string at that
length, mid-sentence. Tool descriptions carry the rest.

## Implementation Order

1. Config and server setup
2. Reference tool (static, no service dependency — grounds field-testing for everything else)
3. Services (external API clients)
4. Read-only tools
5. Write tools
6. Resources
7. Prompts

Each step is independently testable.

<!-- Optional sections — include when the trigger fires: -->
## Domain Mapping          <!-- nouns × operations → API endpoints; include when ≥3 nouns each with ≥3 operations -->
## Workflow Analysis        <!-- how tools chain for real tasks; include when any tool makes ≥3 upstream calls -->
## Design Decisions         <!-- rationale for consolidation, naming, tradeoffs; include when a choice would otherwise be opaque -->
## Known Limitations        <!-- inherent API/data constraints the server can't solve; include when a constraint visibly caps utility -->
## API Reference            <!-- query language, pagination, rate limits; include when worth documenting -->

Keep it concise. The design doc is a working reference, not a spec document — enough to orient a developer (or agent) implementing the server, not more. It stays true after the build: when the implementation diverges — a renamed field, a dropped method, a changed error code — the doc changes in the same commit, with the why under Design Decisions. The next agent reads it as the spec.

Workflow Analysis example. For multi-step workflow tools, document the upstream call sequence in a table — it drives several downstream decisions during implementation: the service-layer method shape, retry boundaries, where cleanup or the confirmation round belongs, and what post-action state to fetch for the response.

deploy_release (5–8 upstream calls, plus a confirmation round):

#CallPurposeMode gate
0ctx.requestInput confirmationHuman approval before promotepromote
1POST /releasesCreate release recordalways
2PUT /releases/{id}/artifactsAttach build artifactsalways
3GET /releases/{id}/preflightHealth checks, smoke testsalways
4POST /releases/{id}/canaryDeploy to 5% of trafficcanary
5POST /releases/{id}/promoteRoll out to 100%promote
6POST /releases/{id}/rollbackRestore previous versionrollback
7GET /releases/{id}Post-action state for responsealways
—DELETE /releases/{id}/canary-trafficCleanup canary if mid-flow erroron error + cleanupOnError

The table surfaces design questions early: should the confirmation round happen before or after the artifacts are attached? Does cleanup drop the canary on any failure, or only failures past the promote step? What does the response body need from the final GET — version, traffic percentage, health summary? Answering these during design is far cheaper than mid-implementation.

9. Confirm and Proceed

If the user has already authorized implementation — any message that contains both a design request and a build/implement verb in the same clause (e.g., "build me a ___ server", "design and implement a ___") — proceed directly to scaffolding using the design doc as the plan. Otherwise, present the design doc to the user for review before implementing.

After Design

Execute the plan using the scaffolding skills:

  1. add-service for each service
  2. add-tool for each standard tool
  3. add-resource for each standalone resource
  4. add-prompt for each prompt
  5. add-app-tool only if any app tools survived the design step (rare — see the App Tool row in Step 3)
  6. add-test alongside each definition, declared error contracts included
  7. devcheck after each addition

Once the surface is built, tool-defs-analysis audits the definition language and field-test exercises the tools against the live upstream.

Checklist

Items without an If …: prefix apply to every design. Conditional items only apply when the trigger fires — otherwise skip them.

  • Server scope decided — workflow identified, audience sized, boundary drawn (standalone single-API vs. multi-source aggregation vs. internal-only)
  • If multi-source: tool surface organized around user workflows, not API identity. Sources are service-layer details.
  • External APIs/dependencies researched and verified (docs fetched, SDKs identified, terms of use read — storage, redistribution, AI use, attribution, credential model)
  • If wrapping an external API: live API probed (at minimum: one list/search, one single-item GET, one error case, one unknown-param request)
  • User goals enumerated first (3–10 outcomes agents will accomplish, scaled to domain size), then domain operations mapped as raw material
  • Each operation classified as tool, resource, prompt, or excluded
  • Catastrophically irreversible operations excluded from the tool surface (stay in vendor UI) — not just destructiveHint
  • Tool surface audited — niche, overlapping, or low-value tools cut or deferred
  • Tool surface is self-sufficient — a tool-only agent can accomplish everything the server is for
  • Workflow and Instruction variants considered where they add value (single-action tools are the default)
  • Tool descriptions are imperative present tense, concrete, and include operational guidance where non-obvious
  • Parameter .describe() text explains what the value is, what it affects, and tradeoffs
  • Input schemas use constrained types (enums, literals, regex) over free strings
  • If an input is an identifier, code, or enum-ish value: the variants callers will plausibly send are enumerated per input — the unambiguous, one-to-one, meaning-preserving ones normalized before lookup; the rest rejected with a recovery naming the expected shape; any composite form tried as-given before a composed retry
  • Output schemas designed for LLM's next action — chaining IDs, post-write state, filtering communicated
  • format() renders all data the LLM needs — different clients forward different surfaces (Claude Code → structuredContent, Claude Desktop → content[]); both must carry the same data, not just a count or title
  • Error messages guide recovery — name what went wrong and the next tool call (no dead ends)
  • If a tool has known domain failure modes: typed error contract declared (errors: [{ reason, code, when, recovery, retryable? }]) with verbatim recovery strings settled in the design doc, each naming the next tool call
  • If the server has search/list tools: zero-hit notices specced (condition → fragment, each routing to a named next call); server-applied defaults that change result semantics echoed in output
  • If a tool resolves a single identifier: no-match returns { found: false, guidance } — a result, not a throw — with guidance routing per miss outcome
  • If the domain has opaque vocabulary (codes, identifier formats, coverage windows): reference tool designed (topic enum), implemented first, and used as the routing target in recovery strings and notices
  • Annotations set correctly (readOnlyHint, destructiveHint, idempotentHint, openWorldHint) — benign writes set destructiveHint: false explicitly, read-only tools omit it
  • Server-level instructions string drafted — workflow chain, identifier semantics, rate-limit posture, under 2,048 characters (ships via createApp() on every initialize)
  • If ops share a noun: related operations consolidated under one tool with a mode/operation enum — as a z.discriminatedUnion input when the arms need different required fields
  • If an upstream API has no native search but the relevant set is bounded: MCP-side list filtering considered — a distinct local filter param (filter/nameContains, not query), filtering the full set, strict token match (fuzzy only when a caller needs typo tolerance)
  • If the server has workflow tools: call-flow documented (upstream sequence + mode arms) in design doc's Workflow Analysis
  • If state-aware procedural guidance adds value: instruction tool considered with nextToolSuggestions pre-filled from diagnostics
  • If any tool is config-gated: nothing routes to it while the gate is off — recovery strings, notices, and guidance name a callable target or state the capability is unavailable, and structured follow-ups naming it are emitted only under the config that registers it
  • If workflow tools have destructive modes: destructive arm gated on a ctx.requestInput confirmation that redeems a ctx.state consent record bound to the operation, caller, and target (shared storage when retries can reach another instance, MCP_REQUEST_STATE_KEY set, an arm that must not run twice idempotent per record), with destructiveHint annotation so clients that never fulfil the round still surface the risk
  • If any tool calls ctx.requestInput: createApp() declares sessionMode with require: 'stateful'
  • If any output carries text other people wrote: those fields listed, format() quotes or fences free text and flattens CR/LF in inline slots, and the server instructions say the content is data
  • If a parameter determines blast radius: safe default set (e.g., mode: 'preview', dryRun: true, confirmCount required)
  • If an action's effect is observable only after the call (wake, trigger, dispatch, provision): confirmation on by default through one numeric window param (0 = act and return), pre-probe then poll with early return, every outcome a result rather than a throw, and a read-only sibling tool that probes the same state
  • App tools default to no. If one was proposed, verified there's a real human-in-the-loop in an MCP Apps-capable client justifying the iframe/CSP/format()-twin maintenance cost — otherwise dropped in favor of a standard tool
  • If the server exposes resources: URIs use {param} templates, pagination planned for large lists
  • If the server is itself the source of truth (no external API): state lifecycle planned — tenant-scoped vs. global, TTLs, what survives restart, storage backend chosen
  • If the server has external deps or shared state: service layer planned (or explicitly skipped with reasoning)
  • If services wrap external APIs: resilience planned (retry boundary, backoff, parse classification, a total deadline inside the client timeout, a pacer where the upstream mandates a rate, rejectPrivateIPs on caller-supplied URLs)
  • If multi-source server: each source has its own service with independent auth/retry/rate-limit config. Fallback chains or fan-out strategy documented per tool. Output includes source provenance. An input rejection from any source fails the call rather than reading as that source's outage.
  • If mirroring a bulk upstream: the terms permit storing the data, and a local index reproduces the upstream's search — no server-side ranking or query expansion on the mirrored path
  • If exposing a SQL/analytical workspace is in scope: DataCanvas considered (api-canvas skill), and it earns its keep on analytical fit (an agent would SQL it), not row count — a discovery/search surface of categorical metadata doesn't qualify. Any tool emitting a canvas_id is paired with a dataframe_query (+ dataframe_describe) tool in the same surface — a token with no query tool is dead output
  • If the server needs runtime config: env vars identified in server-config.ts
  • Design doc written to docs/design.md — every example value synthetic, no keys, tokens, or private hosts
  • Design confirmed with user (or user pre-authorized implementation)

© cyanheads, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in framework-skills/design-mcp-server of cyanheads/pubmed-mcp-server.

Open the folder on GitHubat commit 79145a6

Compare with similar skills

Design MCP Server next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Design MCP Server compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Design MCP Server this skillcyanheads/pubmed-mcp-server158—~22kAutomated safety check: PassApache-2.0
Ogham Researchogham-mcp/ogham-mcp115—~1.4kAutomated safety check: PassMIT
MCP Google Map Projectcablate/mcp-google-map469—~774Automated safety check: PassMIT
Memory686f6c61/alfred-dev117—~601Automated safety check: PassMIT
Oia Manifestruvnet/metaharness696—~910Automated safety check: PassMIT
MCP Developmentcoollabsio/coolify63k1 repos~949Automated safety check: PassMIT

Similar skills

  • Ogham Research

    ogham-mcp/ogham-mcp

    Structured memory capture for Ogham shared memory. An agent skill from ogham-mcp/ogham-mcp.

    115 GitHub stars~1.4k tokensUpdated 10 days ago
    Agent WorkflowsAuto-check passed
  • MCP Google Map Project

    cablate/mcp-google-map

    Project knowledge for developing and maintaining @cablate/mcp-google-map.

    469 GitHub stars~774 tokensUpdated 13 days ago
    Agent WorkflowsAuto-check passed
  • Memory

    686f6c61/alfred-dev

    This skill should be used when the user asks to record a design decision, search past project decisions, inspect the Alfred memory timeline, or work with the alfred-memory MCP server.

    117 GitHub stars~601 tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Oia Manifest

    ruvnet/metaharness

    Emit .harness/oia-manifest.json declaring layer alignment with the OIA v0.1 9-layer reference architecture.

    696 GitHub stars~910 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • MCP Development

    coollabsio/coolify

    A skill your agent uses for Laravel MCP development. An agent skill from coollabsio/coolify.

    63k GitHub starsUsed in 1 repo~949 tokens
    Frontend & DesignAuto-check passed
  • Repo Genome

    ruvnet/metaharness

    7-section readiness scorecard for a LOCAL repo. An agent skill from ruvnet/metaharness.

    696 GitHub stars~772 tokensUpdated today
    Product & Project ManagementAuto-check passed

More from cyanheads/pubmed-mcp-server

All 30 skills in this repo
  • Add App Tool

    cyanheads/pubmed-mcp-server

    Scaffold an MCP App tool + UI resource pair. An agent skill from cyanheads/pubmed-mcp-server.

    158 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Add Prompt

    cyanheads/pubmed-mcp-server

    Scaffold a new MCP prompt template. An agent skill from cyanheads/pubmed-mcp-server.

    158 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Add Resource

    cyanheads/pubmed-mcp-server

    Scaffold a new MCP resource definition. An agent skill from cyanheads/pubmed-mcp-server.

    158 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Add Service

    cyanheads/pubmed-mcp-server

    Scaffold a new service integration. An agent skill from cyanheads/pubmed-mcp-server.

    158 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Add Test

    cyanheads/pubmed-mcp-server

    Scaffold a test file for an existing tool, resource, or service.

    158 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • API Auth

    cyanheads/pubmed-mcp-server

    Authentication, authorization, and multi-tenancy patterns for @cyanheads/mcp-ts-core.

    158 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Design MCP Server

What does Design MCP Server do?

Design the tool surface, resources, and service layer for a new MCP server. Design MCP Server is an agent skill from cyanheads/pubmed-mcp-server. Design the tool surface, resources, and service layer for a new MCP server.

When should I use Design MCP Server?

Design MCP Server fits situations like: starting a new server; planning a major feature expansion; the user describes a domain/API they want to expose via MCP.

How do I install Design MCP Server in Claude Code?

Run `npx skills add cyanheads/pubmed-mcp-server --skill design-mcp-server -a claude-code`. Or copy the skill folder (framework-skills/design-mcp-server in cyanheads/pubmed-mcp-server) into .claude/skills/design-mcp-server in your project. Claude Code loads it when a task matches its description.

How do I install Design MCP Server in Codex?

Run `npx skills add cyanheads/pubmed-mcp-server --skill design-mcp-server -a codex`. Or copy the skill folder (framework-skills/design-mcp-server in cyanheads/pubmed-mcp-server) into .agents/skills/design-mcp-server in your project. Codex loads it when a task matches its description.

Can I use Design MCP Server in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cyanheads/pubmed-mcp-server --skill design-mcp-server -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/design-mcp-server, .gemini/skills/design-mcp-server, .github/skills/design-mcp-server and .opencode/skills/design-mcp-server in your project.

What does Design MCP Server need to run?

Going by SKILL.md and its folder, Design MCP Server needs credentials named MCP_REQUEST_STATE_KEY.

Does Design MCP Server access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Design MCP Server safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Design MCP Server use?

Design MCP Server is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Design MCP Server use?

About 22k tokens (SKILL.md is roughly 89k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Design MCP Server?

Skills that share tags, products or a category with Design MCP Server: Ogham Research (ogham-mcp/ogham-mcp, 115 stars), MCP Google Map Project (cablate/mcp-google-map, 469 stars), Memory (686f6c61/alfred-dev, 117 stars) and Oia Manifest (ruvnet/metaharness, 696 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Design MCP Server?

cyanheads (a GitHub user) maintains it in cyanheads/pubmed-mcp-server, which has 158 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 9, 2026.

Source: cyanheads/pubmed-mcp-server on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.