Setting Up Papergraph
lotchuazzz-crypto/papergraph-mcp
A skill your agent uses when a user has cloned PaperGraph MCP and asks to install, initialize, configure, set up, or start using it with an agent or MCP client.
Exercise tools, resources, and prompts against a live HTTP server via MCP JSON-RPC over curl.
$ npx skills add cyanheads/pubmed-mcp-server --skill field-test -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cyanheads/pubmed-mcp-server field-test --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .claude/skills && cp -r skills-src/framework-skills/field-test .claude/skills/field-test && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "field-test" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/field-test into .claude/skills/field-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "field-test", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/field-testType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cyanheads/pubmed-mcp-server --skill field-test -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cyanheads/pubmed-mcp-server field-test --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .agents/skills && cp -r skills-src/framework-skills/field-test .agents/skills/field-test && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "field-test" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/field-test into .agents/skills/field-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "field-test", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cyanheads/pubmed-mcp-server --skill field-test -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cyanheads/pubmed-mcp-server field-test --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/framework-skills/field-test .cursor/skills/field-test && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "field-test" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/field-test into .cursor/skills/field-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "field-test", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cyanheads/pubmed-mcp-server.git --path framework-skills/field-test--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cyanheads/pubmed-mcp-server --skill field-test -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cyanheads/pubmed-mcp-server field-test --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/framework-skills/field-test .gemini/skills/field-test && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "field-test" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/field-test into .gemini/skills/field-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "field-test", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cyanheads/pubmed-mcp-server field-testInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cyanheads/pubmed-mcp-server --skill field-test -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .github/skills && cp -r skills-src/framework-skills/field-test .github/skills/field-test && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "field-test" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/field-test into .github/skills/field-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "field-test", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cyanheads/pubmed-mcp-server --skill field-test -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cyanheads/pubmed-mcp-server field-test --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cyanheads/pubmed-mcp-server.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/framework-skills/field-test .opencode/skills/field-test && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "field-test" agent skill from https://github.com/cyanheads/pubmed-mcp-server/tree/main/framework-skills/field-test into .opencode/skills/field-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "field-test", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
field-testExercise tools, resources, and prompts against a live HTTP server via MCP JSON-RPC over curl.
Field Test is an agent skill from cyanheads/pubmed-mcp-server. Exercise tools, resources, and prompts against a live HTTP server via MCP JSON-RPC over curl. Starts the server, surfaces the catalog, runs real and adversarial inputs, measures every call (bytes, token estimate, wall-clock) and weighs the catalog, renders app tools' views in a headless MCP Apps host, and produces a tight report with concrete findings and numbered follow-up options. Use after adding or modifying definitions, or when the user asks to test, try out, or verify their MCP surface.
Its SKILL.md is about 11k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol. The repository describes itself as: Search PubMed/Europe PMC, fetch articles and full text (PMC/EPMC/Unpaywall), citations, MeSH terms via MCP. STDIO or Streamable HTTP. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5a417fb. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqbuncurlbunxnpxFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comAlso links to:
modelcontextprotocol.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Field Test loads about 11k tokens when it runs. Until then it costs about 127 tokens; SKILL.md has 3,706 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from cyanheads/pubmed-mcp-server at commit 5a417fb, republished under its Apache-2.0 licence (© cyanheads). 3,706 words, ~10,821 tokens.
.claude/skills/field-test/SKILL.md (or your agent's skills folder).Unit tests (add-test skill) verify handler logic with mocked context. Field testing exercises the real HTTP transport with real JSON-RPC: starts the server, calls initialize, surfaces the catalog, runs inputs, and checks what a client actually sees. It catches what unit tests miss — awkward input shapes, unhelpful errors, missing format output, drift between structuredContent and content[], edge-case surprises.
Actively call the tools. Don't read code and guess.
This skill drives an HTTP server because curl + JSON-RPC is the most reliable harness for shell-based agents. The same handlers run on both transports — only the framing differs — so HTTP exercises the full functional surface. Both HTTP session modes are covered: a durable Mcp-Session-Id session, and the sessionless initialization a MCP_SESSION_MODE=stateless server performs.
Stdio coverage is a boot check only — run this before Step 1. Run bun run rebuild && bun run start:stdio < /dev/null, and confirm the startup logs look clean: the Core services constructed — N tool(s) … record lists every registered tool, resource, and prompt in its tools / resources / prompts fields — the message text shows only counts — and a definition missing from them was never passed to createApp(). No errors/warnings, no missing-config gripes. The emoji startup banner prints only to a terminal, so its absence from an agent's shell is not a finding. Redirecting stdin is what ends the run: the server treats EOF as a shutdown signal, boots fully, then exits on its own, so the log also shows the graceful-shutdown path. Do not background it and reach for pkill — a pattern like pkill -f dist/index.js matches every other stdio MCP server on the machine, including the ones the calling agent's own session is connected to. Pino logs go to stderr in stdio mode (stdout is reserved for JSON-RPC), so they print straight to the terminal when you run interactively. No need to call tools over stdio — the HTTP pass already covered handler behavior.
Generate a 10-character alphanumeric ID (e.g. 9DJ73-K103L) and write the helper to /tmp/<project-name>-field-test-<ID>.sh. Use that exact path in every subsequent Bash call. Two agents in the same project tree must pick different IDs — that's what keeps their helper files, server logs, and call scratch from colliding.
The helper itself is stateless — every function takes the IDs it needs (server pid, url, port, MCP sid, server log path) as positional args. mcp_start prints them; the agent threads them through every later call. No env vars, no shared state files.
# Pick your ID — example below uses 9DJ73-K103L. Substitute your own.
# (Helper path also encodes the project name so /tmp/ stays grep-friendly.)
cat > /tmp/<project-name>-field-test-9DJ73-K103L.sh <<'HELPER_EOF'
#!/bin/bash
# Field-test helper: stateless wrappers around an MCP HTTP server + JSON-RPC
# session. Every function takes the IDs it needs as positional args — the agent
# threads pid/url/port/sid/log through each call rather than relying on a state
# file or env vars (the Bash tool wipes shell state between calls, and a
# pointer file would race the same way two agents race on shared state).
# See https://github.com/cyanheads/mcp-ts-core/issues/90, #144.
#
# Surfaces failures aggressively — field test is for finding things that fail,
# so the helper auto-tails logs and prints HTTP status/body on errors instead
# of swallowing them. It also measures: every mcp_call prints a one-line
# size/latency reading on stderr, and mcp_catalog_size weighs tools/list.
# Usage: mcp_start /path/to/server [startup-timeout-seconds] (default: 30)
# Builds, starts the HTTP server in the background, waits for the listen line,
# and prints: ready pid=<n> url=<u> port=<n> log=<path>
# Capture these — every later helper takes them as args. Raise the timeout for
# servers that build a local index at boot.
mcp_start() {
local dir="${1:-$PWD}"
local timeout="${2:-30}"
local build_log; build_log=$(mktemp /tmp/mcp-field-test-build.XXXXXX)
echo "building $dir ..." >&2
if ! (cd "$dir" && bun run rebuild) >"$build_log" 2>&1; then
echo "BUILD FAILED — last 30 lines of $build_log:" >&2
tail -30 "$build_log" >&2
return 1
fi
rm -f "$build_log"
local server_log; server_log=$(mktemp /tmp/mcp-field-test-server.XXXXXX)
echo "starting server ..." >&2
(cd "$dir" && bun run start:http) >"$server_log" 2>&1 &
local pid=$!
local line=""
local waited=0
while [ "$waited" -lt "$((timeout * 4))" ]; do
line=$(grep -Eo 'listening at http://[^" ]+/mcp' "$server_log" | head -1)
[ -n "$line" ] && break
if ! kill -0 "$pid" 2>/dev/null; then
echo "server exited during startup — last 30 lines of $server_log:" >&2
tail -30 "$server_log" >&2
rm -f "$server_log"
return 1
fi
sleep 0.25
waited=$((waited + 1))
done
if [ -z "$line" ]; then
echo "server failed to start within ${timeout}s — last 30 lines of $server_log:" >&2
tail -30 "$server_log" >&2
kill "$pid" 2>/dev/null
rm -f "$server_log"
return 1
fi
local url="${line#listening at }"
local port; port=$(echo "$url" | sed -E 's|.*:([0-9]+)/.*|\1|')
echo "ready pid=$pid url=$url port=$port log=$server_log"
}
# Internal: report a failed initialize with the raw exchange, then clean up.
_mcp_init_fail() {
local msg="$1"; local body_file="$2"; local hdr="$3"
echo "init failed — $msg" >&2
echo "--- response body ---" >&2
if [ -s "$body_file" ]; then cat "$body_file" >&2; else echo "(empty)" >&2; fi
echo "--- response headers ---" >&2
if [ -s "$hdr" ]; then cat "$hdr" >&2; else echo "(none)" >&2; fi
rm -f "$hdr" "$body_file"
return 1
}
# Usage: mcp_init <url>
# Runs `initialize`, sends `notifications/initialized`, prints:
# ready sid=<id-or-empty> protocol=<negotiated-version> requested=<want> instructions=<bytes>B (HTTP <code>)
# `instructions=` is the byte size of the server's `instructions` string — it
# loads into every client session alongside tools/list, so it is the other half
# of the per-session context tax mcp_catalog_size weighs.
# The initialize *result* is what decides success — a session ID is optional.
# A server started with MCP_SESSION_MODE=stateless mints none, and the session
# header is then omitted from every later request. Capture BOTH `sid` and
# `protocol`: mcp_call takes the protocol as its 5th arg, which is what carries
# the negotiated revision when there is no session to carry it.
# A negotiated version older than the requested one means the server capped it
# — note that in the report; you are then testing an older protocol than a
# current client would use.
mcp_init() {
local url="$1"
[ -z "$url" ] && { echo "usage: mcp_init <url>" >&2; return 1; }
local want="${MCP_FIELD_TEST_PROTOCOL:-2025-11-25}"
local hdr; hdr=$(mktemp)
local body_file; body_file=$(mktemp)
local code curl_rc
code=$(curl -sS -D "$hdr" -o "$body_file" -w '%{http_code}' -X POST "$url" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"initialize\",\"params\":{\"protocolVersion\":\"$want\",\"capabilities\":{},\"clientInfo\":{\"name\":\"field-test\",\"version\":\"1.0.0\"}}}")
curl_rc=$?
if [ "$curl_rc" -ne 0 ] || [ -z "$code" ] || [ "$code" = "000" ]; then
_mcp_init_fail "transport failure — curl exit $curl_rc, http_code '${code:-none}'; nothing listening at $url" "$body_file" "$hdr"
return 1
fi
[ "$code" -ge 400 ] && { _mcp_init_fail "HTTP $code" "$body_file" "$hdr"; return 1; }
# Unwrap SSE framing when present; a plain JSON body is used as-is.
local payload; payload=$(sed -n 's/^data: //p' "$body_file")
[ -z "$payload" ] && payload=$(cat "$body_file")
# Pick the reply frame by structure, not by substring: a server that logs to the
# client emits `notifications/message` frames first, and a `"level":"error"` or a
# log string containing `result` matches a text grep and gets read as the reply.
local reply; reply=$(printf '%s\n' "$payload" | jq -c 'select(type=="object" and (has("result") or has("error")))' 2>/dev/null | tail -1)
[ -z "$reply" ] && reply="$payload"
if printf '%s' "$reply" | grep -q '"error"'; then
_mcp_init_fail "server returned a JSON-RPC error" "$body_file" "$hdr"
return 1
fi
if ! printf '%s' "$reply" | grep -q '"result"'; then
_mcp_init_fail "HTTP $code but no JSON-RPC result in the body" "$body_file" "$hdr"
return 1
fi
local got; got=$(printf '%s' "$reply" | grep -o '"protocolVersion":"[^"]*"' | head -1 | cut -d'"' -f4)
if [ -z "$got" ]; then
_mcp_init_fail "initialize result declares no protocolVersion" "$body_file" "$hdr"
return 1
fi
local instr; instr=$(printf '%s' "$reply" | jq -r '.result.instructions // "" | utf8bytelength' 2>/dev/null || echo 0)
local sid; sid=$(grep -i '^mcp-session-id:' "$hdr" | awk '{print $2}' | tr -d '\r\n')
local init_headers=(-H "Content-Type: application/json" -H "Accept: application/json, text/event-stream" -H "MCP-Protocol-Version: $got")
[ -n "$sid" ] && init_headers+=(-H "Mcp-Session-Id: $sid")
curl -sS -X POST "$url" "${init_headers[@]}" \
-d '{"jsonrpc":"2.0","method":"notifications/initialized"}' >/dev/null
rm -f "$hdr" "$body_file"
echo "ready sid=$sid protocol=$got requested=$want instructions=${instr}B (HTTP $code)"
}
# Internal: one stderr line per call — reply bytes, the content/structured
# split, a token estimate, wall-clock. Bytes are the reply as delivered (SSE
# framing stripped). `content` is every text block's bytes, `structured` is
# structuredContent serialized. The token figure is bytes/4 — an estimate, not
# a tokenizer. Wall-clock is curl's time_total for the whole exchange.
_mcp_measure() {
local method="$1"; local params="$2"; local reply="$3"; local code="$4"; local secs="$5"
local total; total=$(printf '%s' "$reply" | wc -c | tr -d ' ')
local ms; ms=$(awk -v s="$secs" 'BEGIN { printf "%d", s * 1000 }')
local label="$method"
local split=""
case "$method" in
tools/call)
local name; name=$(printf '%s' "$params" | jq -r '.name // empty' 2>/dev/null)
[ -n "$name" ] && label="$method $name"
split=$(printf '%s' "$reply" | jq -r '
(.result // {}) as $r
| ([$r.content[]? | select(.type == "text") | .text] | join("") | utf8bytelength) as $c
| (if $r.structuredContent == null then "none" else ($r.structuredContent | tojson | utf8bytelength | tostring) end) as $s
| "content \($c) · structured \($s)"' 2>/dev/null)
;;
resources/read)
split=$(printf '%s' "$reply" | jq -r '
"text \([.result.contents[]? | .text // ""] | join("") | utf8bytelength)"' 2>/dev/null)
;;
esac
local tok; tok=$(awk -v b="$total" 'BEGIN { if (b >= 1000) printf "~%.1fk", b / 4000; else printf "~%d", b / 4 }')
echo "⏱ $label · HTTP $code · ${total} B${split:+ ($split)} · $tok tok · ${ms} ms" >&2
}
# Usage: mcp_call <url> <sid> <method> [JSON_PARAMS] [protocol]
# Prints the JSON-RPC response. SSE framing is stripped when present, and only
# the reply is emitted (a single POST can also carry progress notifications, so
# emitting every event would break `| jq .result`). A transport failure or an
# HTTP >= 400 prints the details and returns non-zero — it never returns 0 with
# empty output. Pipe to `jq`.
# Every call also prints one measurement line on stderr, e.g.
# ⏱ tools/call gbif_search_species · HTTP 200 · 18412 B (content 9100 · structured 8900) · ~4.6k tok · 812 ms
# Read it on every call — it is the size/latency evidence the report cites.
# `sid` may be empty ('') for a stateless server; the session header is then
# omitted. Pass the `protocol` mcp_init printed as the 5th arg — with no
# session carrying the negotiation, MCP-Protocol-Version is what tells the
# server which revision the request speaks.
mcp_call() {
local url="$1"; local sid="$2"; local method="$3"; local params="${4:-}"; local protocol="${5:-}"
[ -z "$url" ] || [ -z "$method" ] && { echo "usage: mcp_call <url> <sid> <method> [params] [protocol]" >&2; return 1; }
local body
if [ -z "$params" ]; then
body=$(printf '{"jsonrpc":"2.0","id":%d,"method":"%s"}' "$RANDOM" "$method")
else
body=$(printf '{"jsonrpc":"2.0","id":%d,"method":"%s","params":%s}' "$RANDOM" "$method" "$params")
fi
local resp_file; resp_file=$(mktemp)
local stats code secs curl_rc
local headers=(-H "Content-Type: application/json" -H "Accept: application/json, text/event-stream")
[ -n "$sid" ] && headers+=(-H "Mcp-Session-Id: $sid")
[ -n "$protocol" ] && headers+=(-H "MCP-Protocol-Version: $protocol")
stats=$(curl -sS -o "$resp_file" -w '%{http_code} %{time_total}' -X POST "$url" "${headers[@]}" -d "$body")
curl_rc=$?
read -r code secs <<< "$stats"
if [ "$curl_rc" -ne 0 ] || [ -z "$code" ] || [ "$code" = "000" ]; then
echo "TRANSPORT FAILURE calling $method — curl exit $curl_rc, http_code '${code:-none}'." >&2
echo "Server not reachable at $url (check it is still running: mcp_log <log>)." >&2
rm -f "$resp_file"
return 1
fi
if [ "$code" -ge 400 ]; then
echo "HTTP $code from $method — response:" >&2
cat "$resp_file" >&2
rm -f "$resp_file"
return 1
fi
local reply
local sse; sse=$(sed -n 's/^data: //p' "$resp_file")
if [ -n "$sse" ]; then
# Structural pick, same reason as in mcp_init: log-notification frames precede
# the reply and can carry the literal tokens a text grep keys on.
reply=$(printf '%s\n' "$sse" | jq -c 'select(type=="object" and (has("result") or has("error")))' 2>/dev/null | tail -1)
reply="${reply:-$sse}"
else
reply=$(cat "$resp_file")
fi
rm -f "$resp_file"
_mcp_measure "$method" "$params" "$reply" "$code" "$secs"
printf '%s\n' "$reply"
}
# Usage: mcp_catalog_size <url> <sid> [protocol]
# Weighs the catalog: the bytes of the tools/list reply — what every client
# loads into context per session before a single call — then each tool's
# serialized entry, largest first, split into description / inputSchema /
# outputSchema so the row says WHERE the weight is. A fat outputSchema costs as
# much as a fat description and is the usual surprise. Prints:
# catalog: 12 tools · 48210 B · ~12.1k tok
# <bytes> <~tok> <name> desc <b> · input <b> · output <b|none> (one row per tool)
mcp_catalog_size() {
local url="$1"; local sid="$2"; local protocol="${3:-}"
[ -z "$url" ] && { echo "usage: mcp_catalog_size <url> <sid> [protocol]" >&2; return 1; }
local reply; reply=$(mcp_call "$url" "$sid" tools/list '' "$protocol") || return 1
printf '%s' "$reply" | jq -r '
def tok: if . >= 1000 then "~\(. / 4000 * 10 | round / 10)k" else "~\(. / 4 | floor)" end;
def bytes_or_none: if . == null then "none" else (tojson | utf8bytelength | tostring) end;
(.result.tools // []) as $t
| (. | tojson | utf8bytelength) as $total
| "catalog: \($t | length) tools · \($total) B · \($total | tok) tok",
($t
| map({name, b: (tojson | utf8bytelength),
d: ((.description // "") | utf8bytelength),
i: (.inputSchema | bytes_or_none),
o: (.outputSchema | bytes_or_none)})
| sort_by(-.b) | .[]
| "\(.b)\t\(.b | tok)\t\(.name)\tdesc \(.d) · input \(.i) · output \(.o)")'
}
# Usage: mcp_log <server-log-path> [N] (default: 50 lines)
# Tail the per-server log printed by mcp_start. Useful when a call surprises
# you — pino startup banner, definition lint diagnostics, request handler
# errors, upstream calls, and rate-limit warnings all land here.
mcp_log() {
local log="$1"; local n="${2:-50}"
[ -z "$log" ] && { echo "usage: mcp_log <log-path> [n]" >&2; return 1; }
tail -n "$n" "$log"
}
# Usage: mcp_stop <pid> [server-log-path] [port]
# Kills the background server and the `bun run` child that actually holds the
# port (SIGKILL is not forwarded, so the child must be signalled directly or it
# survives as an orphaned listener). Pass the port from mcp_start to have the
# stop confirmed against the socket rather than against the wrapper PID.
# Removes the server log if a path is given.
mcp_stop() {
local pid="$1"; local log="${2:-}"; local port="${3:-}"
[ -z "$pid" ] && { echo "usage: mcp_stop <pid> [log-path] [port]" >&2; return 1; }
local kids; kids=$(pgrep -P "$pid" 2>/dev/null)
kill "$pid" $kids 2>/dev/null
for _ in $(seq 1 12); do
kill -0 "$pid" 2>/dev/null || break
sleep 0.25
done
if kill -0 "$pid" 2>/dev/null; then
echo "PID $pid didn't exit on SIGTERM — sending SIGKILL"
kill -9 "$pid" $kids 2>/dev/null
sleep 0.5
fi
local held=""
[ -n "$port" ] && held=$(lsof -ti tcp:"$port" 2>/dev/null | tr '\n' ' ')
if [ -n "$held" ]; then
echo "WARNING: port $port still held by PID(s) $held after stopping $pid — kill those before re-running"
elif kill -0 "$pid" 2>/dev/null; then
echo "WARNING: PID $pid still alive after SIGKILL"
else
echo "stopped pid=$pid${port:+ (port $port free)}"
fi
[ -n "$log" ] && rm -f "$log"
return 0
}
HELPER_EOF
. /tmp/<project-name>-field-test-9DJ73-K103L.sh
mcp_start /absolute/path/to/server # replace with the target serverCapture pid, url, port, log from the mcp_start output — every later call takes them as positional args. Two agents running concurrently in the same project tree each pick their own ID, so their helper paths, server logs, and call scratch never share a name.
Notes
MCP_HTTP_PORT is a starting port — the server auto-increments if taken. Helper parses the real URL from the log (HTTP transport listening at ...).bun run rebuild fails, stop. Don't field-test broken code — fix the build first.mcp_start /path 90) rather than reading the timeout as a real startup failure.pid is the bun run wrapper; the process that actually holds the port is its child. mcp_stop signals both — that's why it takes the port.lsof -i :<port>), confirm with the user before killing it; it may be their own session. If the user isn't available to confirm, abort the field test and surface the port conflict in your response.. /tmp/<project-name>-field-test-<ID>.sh
mcp_init <url-from-mcp_start>Runs initialize, sends notifications/initialized, prints the sid and protocol to capture for mcp_call, plus instructions= — the byte size of the server's instructions string, which every client loads per session alongside the catalog (record it with the catalog total in Step 3). Success is decided by the initialize result, so a transport failure, a non-2xx status, a JSON-RPC error, a malformed body, or a result with no protocolVersion all fail loudly with the raw exchange.
The helper requests the newest initialize-negotiated revision the SDK supports (2025-11-25). If protocol= comes back older than requested=, the server capped it — every call after that exercises an older protocol than a current client would negotiate. Note it as a bug finding and check the pinned @modelcontextprotocol/server version; don't quietly test the downgraded surface. To deliberately test an older version, set MCP_FIELD_TEST_PROTOCOL.
sid= may come back empty — that is a pass, not a failure. Under MCP_SESSION_MODE=stateless the server mints no Mcp-Session-Id, and the helper then omits the session header from every later request. Thread the empty value through positionally and pass the negotiated protocol, which is what identifies the revision when no session carries it:
mcp_call <url> '' tools/list '' <protocol-from-mcp_init>To exercise both session modes, start the server twice — once with the project's default, once with MCP_SESSION_MODE=stateless — and run the same calls against each.
Both modes exercise the 2025-era arm: initialize negotiates the revision, and the session (when there is one) carries it. The 2026-07-28 revision is a different thing from a sessionless 2025 handshake — it does not initialize at all, and is selected per request by the io.modelcontextprotocol/protocolVersion key in the request's own _meta envelope. This helper does not reach it; exercise the per-request leg from a real 2026-era client or an integration test.
. /tmp/<project-name>-field-test-<ID>.sh
mcp_call <url> <sid> tools/list | jq '.result.tools[] | {name, description, inputSchema, outputSchema}'
mcp_call <url> <sid> resources/list | jq '.result.resources[] | {uri, name, mimeType}'
mcp_call <url> <sid> prompts/list | jq '.result.prompts[] | {name, description, arguments}'
mcp_catalog_size <url> <sid> <protocol>Weigh the catalog. mcp_catalog_size prints the tools/list bytes — the context every client loads per session before a single call — and each tool's entry, largest first, split into description / inputSchema / outputSchema. Record the total alongside the instructions= bytes from Step 2; together they are the per-session tax. The split says where a heavy tool's weight lives: an outputSchema narrating every field of a 60-field record is the common surprise, an over-long description the obvious one. Hand the outliers to tool-defs-analysis (its length-outliers pass) rather than trimming blind.
Present a compact catalog to the user: each definition's name + 1-line description. Flag vague or missing descriptions as you go — those feed into the report. Use this to build the test plan.
Audit every description for leaks — tool description, every parameter .describe() in inputSchema, and every field .describe() in outputSchema (the outputSchema projection above is what surfaces these; don't skim past it). Three categories:
Treat any hit as a ux finding in the report. The authoring rule lives under Tool descriptions in design-mcp-server/SKILL.md — same categories, applied at review time.
Budget. Don't run every category against every definition — the cross-product is infeasible. Apply the universal battery to everything; apply situational categories only when the definition triggers them.
Universal battery — run on every tool
| Category | What to verify |
|---|---|
| Happy path | One realistic input. Output shape matches schema. content[] text reads clearly to a human. |
structuredContent ↔ content[] parity | Dump the whole array (jq '.result.content') and check every structuredContent field is surfaced somewhere in it — enrichment lands in its own trailing block, not in content[0]. Parity gap = client-specific blindness. |
| Input error | One invalid input (wrong type or missing required). Error text says what, why, how to fix. |
| Size & latency | Read the ⏱ line mcp_call prints on every call. A happy-path response over 24,000 B (the framework's DEFAULT_OUTLINE_BUDGET_BYTES — the line at which it would outline a document itself) with no truncation disclosure and no retrieval path (cursor, offset, sections, canvas handle) is a ux finding: the agent pays the whole payload with no way to ask for less. content ≈ structured with the text starting { means the JSON is on the wire twice — a missing format(). A call over ~5 s on a happy-path input is worth a mcp_log look before calling it upstream latency. |
Situational — add only when triggered
Trigger (look in input schema or annotations) | Add category |
|---|---|
include / fields / expand / view / projection parameter | Field selection: non-default value renders requested fields |
Array return with query / filter inputs | Empty result: does response explain why (echo criteria, suggest broadening)? |
| Identifier, code, or enum-ish input (an ID format, a classification code, a unit, a place name, a list the docs say may be comma-joined) | Value-variant tolerance: re-send the happy-path call with each obvious variant of that value — lowercase, the bare leaf of a hierarchical code, a common domain alias, a delimiter-joined list where an array is accepted, the spelled-out form of an abbreviated name. Pass is either outcome: the call succeeds, or it fails with an error naming the expected shape. A miss or a bare validation failure on a variant that maps one-to-one onto a valid value is a ux finding. Probe values — variants of the argument key name, and a JSON-stringified array or object or an integer sent for a string as a value, are handled by the framework, not the server. |
| Batch / bulk input (arrays of IDs, multi-item ops) | Partial success: mix valid + invalid items |
annotations.readOnlyHint: true | Confirm no mutation happened |
annotations.idempotentHint: true | Call twice with same input — safe? |
| Hits external API / live upstream | One call that exercises upstream; note rate-limit / timeout / transient-failure behavior |
| Chained with other tools (search → detail → act) | Run one representative chain end-to-end; does each step return the IDs/cursors the next needs? |
cursor / offset / limit params | Pagination: second page, end-of-list |
Output can be truncated, capped, or spilled (maxLength-style caps, outline-on-overflow, canvas/dataframe spill, "showing N of M") | Truncation retrievability: force a response that truncates, then confirm the response both discloses the truncation and hands back the means to reach the rest — a cursor, an offset, a document/section selector, a canvas handle. Truncated data with no retrieval path is a bug, not a nit. |
Tool declared an errors: [...] contract | Error contract (tool): trigger ≥1 declared failure mode. Verify result.structuredContent.error.code matches the contract entry, result.structuredContent.error.data.reason is the declared reason (only present when the handler threw an McpError — ctx.fail always does, plain throw new Error(...) does not), and content[0].text is actionable. Reasons declared but unreachable from any input are dead contract entries. |
Resource declared an errors: [...] contract | Error contract (resource): trigger ≥1 declared failure mode by reading a URI that exercises it. Resources re-throw errors at the JSON-RPC level — verify error.code matches the contract entry and error.data.reason is the declared reason. (Resources don't use the result.isError envelope — they fail the request itself.) |
Mutator (write/update/delete/append/patch verbs, or destructiveHint: true) | Mutator response observability: run an intentionally-ambiguous input (typo path, wrong ID, already-deleted target). Confirm the response carries enough state (pre/post values, state-change discriminator) for the agent to detect intent-effect divergence without re-fetching. |
_meta.ui.resourceUri on the tool (an app tool) | View rendering: Step 6. The curl calls see only the format() text and the view's raw HTML; whether the view runs is visible only when it is rendered. |
Resources. Happy path, not-found URI (use a syntactically valid but non-existent ID — e.g., substitute a fake ID into the URI template), list if defined, pagination if used.
Prompts. Happy path, defaults omitted, skim message quality.
Sampling for large servers. If more than 15 tools, run the universal battery on all, but pick roughly 30–40% for situational testing. Weight toward: write-shaped tools, complex schemas, external deps. List which ones you skipped in the report.
Auth & external state.
skipped — requires $VAR and move on. Don't fabricate inputs.Use TaskCreate — one task per definition. Mark complete as you go. Don't batch.
For each call, capture: input sent, the ⏱ line (bytes, split, ms), response (trim huge payloads to files), whether isError: true appeared, anything surprising (slow response, parity drift, unhelpful text, crash).
When a call surprises you — slow, hangs, returns terse output, surfaces an unhelpful error — run . /tmp/<project-name>-field-test-<ID>.sh && mcp_log <log> to tail the server log. The pino startup banner, request handler errors, upstream API call traces, and rate-limit warnings all land in the per-server log (read via mcp_log) rather than coming back through mcp_call. Don't guess at runtime behavior from response text alone. To count upstream requests from the log, start the server with MCP_LOG_RATE_LIMIT_THRESHOLD=0: by default the logger drops a message repeated past its per-window threshold and reports it later as a Suppressed N line, so a count read mid-run comes up short.
Interpreting responses
content[] is an array of blocks — read all of them, never just content[0]. A success result is assembled as [...ctx.content media blocks, ...the format()/JSON domain render, ...the enrichment trailer]. Everything the handler put on ctx.enrich — empty-result notices, totals, query echoes, truncation disclosure — renders in that trailer, a separate trailing block, not inside the format() block. Quoting content[0].text and reporting those fields as absent from content[] is a false parity gap; the suggested fix (render them in format() too) would double-render them. Dump .result.content in full before claiming drift.{result: {content: [...], isError: true}} — they live in result, not error. Check isError, not the JSON-RPC error field.result.structuredContent.error.{code, message, data?.reason} — inspect that, not just the text. data carries what the handler threw as an McpError (or a ZodError's issues) plus the framework's data.requestId; plain throw new Error(...) won't populate data.reason. Use ctx.fail-thrown errors when the contract reason matters — a declared reason arrives with its contract recovery as data.recovery.hint even when the throw site passed none. The text in result.content[0].text mirrors the message, adds Recovery: <hint> when data.recovery.hint says something the message does not already say, and closes with (reason <reason> · not retryable · request <id>) for whichever of data.reason / data.retryable / data.requestId is present — the numeric code stays JSON-only. The request id matches the requestId on that call's server log records.error.{code, data.reason} field, not inside result. Resource handlers re-throw rather than producing an isError envelope.error only appears for protocol issues (bad session, malformed envelope, unknown method).mcp_call already strips SSE framing. Pipe to jq for readability.Skip this step when no tool carries _meta.ui.resourceUri. List the app tools:
. /tmp/<project-name>-field-test-<ID>.sh
mcp_call <url> <sid> tools/list '' <protocol> | jq -r '.result.tools[] | select(._meta.ui.resourceUri) | .name'Render each one with mcp-ts-core app-render against the Step 1 server. It connects as an MCP Apps client, calls the tool, loads the ui:// view into headless chrome-headless-shell inside the double-iframe sandbox and CSP the MCP Apps spec prescribes, plays the host half of the protocol, and writes report.json plus screenshots under --out. No window opens.
bunx @cyanheads/mcp-ts-core app-render --url <url> --tool <app_tool> --args '<happy-path JSON>' \
--click '<selector>' --out /tmp/<project-name>-field-test-<ID>-apps/<app_tool> > /dev/null; echo "exit=$?"
jq '{initialized, failure, toolError, errors, cspViolations, size, steps, text}' \
/tmp/<project-name>-field-test-<ID>-apps/<app_tool>/report.jsonThe report also goes to stdout; read the file instead, since messages grows with every exchange. Pass one --click per control whose handler matters (a screenshot follows each), run a second pass with --theme dark when the view applies host theming, and add --stream-input when it renders partial input. Fill, wait-for, and evaluate steps need renderAppTool from @cyanheads/mcp-ts-core/testing/apps in a script; it returns the same report.
A setup failure exits 1 with app-render: <reason> on stderr and writes no report:
| Reason names | Meaning |
|---|---|
| A missing optional peer | @modelcontextprotocol/client or @modelcontextprotocol/ext-apps is not installed. With the user's go-ahead, bun add -d it; otherwise record app views as skipped — requires <package>. |
| No browser, or a browser path | No usable chrome-headless-shell. With the user's go-ahead, npx @puppeteer/browsers install chrome-headless-shell@stable --path ~/.cache/puppeteer (without --path it installs into the working directory, where the host never looks), or pass an existing build with --browser; otherwise record skipped — requires chrome-headless-shell. Never point --browser at a desktop browser. |
| The server is unreachable | Wrong URL, or the Step 1 server died — mcp_log <log>. |
No such tool, no UI resource, an unreadable UI resource, or a _meta.ui.csp entry | A server finding (bug): the tool is unregistered, its resourceUri is missing, resources/read on the ui:// URI fails, or a CSP domain entry is not a plain origin. |
A view failure exits 0: the run reached the browser, and the report says what the view did. Read every field:
initialized — the view completed ui/initialize. false, with failure naming the timeout, means the view never connected: a script threw before app.connect() (see errors), the CSP blocked its SDK import (see cspViolations), or the HTML does not use the ext-apps App class. bug.toolError — the tool call failed at the protocol level, so the view received tool-cancelled instead of a result.errors — uncaught exceptions and console errors, each tagged with its frame. A view entry is the server's own UI code: bug. sandbox, host, and unknown entries come from the host pages, not the server; report them apart from the server's findings.cspViolations — loads the view's CSP blocked, as directive + blockedURI. An origin the view needs but its resource's _meta.ui.csp omits (resourceDomains for scripts, styles, images, fonts, and media; connectDomains for fetch and WebSocket; frameDomains for iframes) is a bug, because a spec-conformant host blocks the same load. blockedURI: "eval" means the view calls eval or new Function, which the policy never allows. An eval entry whose sourceURL points into the ext-apps bundle (app-with-deps.js) is Zod's guarded capability probe, which catches the refusal, not the view's own code, so it is no finding.messages — every message between view and host, in order, with its direction. A healthy run opens view-to-host ui/initialize → ui/notifications/initialized, then host-to-view ui/notifications/tool-input → ui/notifications/tool-result. A click that should reach the server shows a view-to-host tools/call and its host-to-view response; an error there is the server's rejection, or the host refusing a tool whose _meta.ui.visibility excludes app. Requests that need a user or a model (ui/message, ui/open-link, ui/update-model-context, ui/request-display-mode) are recorded and acknowledged — check their params carry what the control meant to send.steps — one entry per click and screenshot; ok: false with No element matches <selector> means the control is missing from the rendered view.text and the screenshots — the rendered text after the steps, and the PNGs (final.png, one per click). Open them: a blank, unstyled, clipped, or unreadable-in-dark view is a ux finding even when every field above is clean, and so is data in the tool result's structuredContent that the view never shows.size — the view's last ui/notifications/size-changed. Absent means the view never reports its size, so a host cannot fit its frame to it: ux.The harness stops the browser and deletes its profile itself; the --out directories are removed in Step 7.
. /tmp/<project-name>-field-test-<ID>.sh
mcp_stop <pid> <log> <port>
rm -rf /tmp/<project-name>-field-test-<ID>-apps
rm -f /tmp/<project-name>-field-test-<ID>.shKills the background server and its port-holding child, removes the server log and any app-render output, then removes the helper script itself. Do this before writing the report so nothing leaks into the next session. Pass the port — it's what turns "the wrapper PID is gone" into "the socket is actually free." If mcp_stop warns the port is still held or the PID survived SIGKILL, note it in the report and proceed — don't block on a zombie process, but do say which PID to kill.
Four sections. Tight. The user should be able to skim the summary, scan the numbers, read details only for what matters, and act on numbered options.
One paragraph. How many definitions exercised, how many passed clean, how many have issues, and the single most important finding. No tables, no lists.
Per-session tax on its own line (instructions bytes + catalog bytes, with the heaviest tool named), then one row per tool exercised, sorted by happy-path bytes descending: tool · bytes · ~tok · ms. Over 15 tools, keep every row over 24,000 B plus the three slowest and fold the rest into one line ("N more under budget, median X B"). Numbers only — what they mean goes in Findings.
Only include definitions with issues. Group by severity. Each finding is 2–4 lines unless it genuinely needs more. A parity finding cites the full content[] dump as its evidence — a quote from one index doesn't establish drift.
| Severity | Meaning |
|---|---|
| bug | Broken: crash, wrong output, isError: true on valid input, data loss, schema violation |
| ux | Works but degrades the user/LLM experience: vague description, leaky description (implementation details, meta-coaching, consumer-aware phrasing), unhelpful error text, missing format(), parity drift, annotation mismatches behavior |
| nit | Polish: phrasing, inconsistent tone, minor doc gaps |
Format:
**<tool_name> — <bug|ux|nit>**
Input: `<short input>` → <what happened>
Expected: <what should happen>
Fix: <one sentence>Numbered, actionable, cherry-pickable. Each item maps to a concrete change.
1. Fix empty-result message in `pubmed_search_articles` — echo criteria (finding #2)
2. Add `format()` to `pubmed_lookup_mesh` — currently returns raw JSON (finding #5)
3. Tighten `ids` description in `pubmed_fetch_articles` — silent on PMID vs DOI (finding #8)End with:
Pick by number (e.g. "do 1, 3, 5" or "expand on 2").
bun run rebuild && bun run start:stdio < /dev/null shows clean startup (every expected definition listed in the Core services constructed record's tools / resources / prompts fields, no errors) and a graceful shutdown on EOFsid — still a pass); notifications/initialized sent; negotiated protocol version matches the requested one (a downgrade is a finding)mcp_catalog_size); total + instructions= bytes recorded for the report⏱ line read; any happy-path response over 24,000 B with no disclosure + retrieval path filed as uxcontent[] array, input error)errors: [...] contract: ≥1 declared failure mode triggered; result.structuredContent.error.code and data.reason verified against the contract entryerrors: [...] contract: ≥1 declared failure mode triggered; top-level JSON-RPC error.code and error.data.reason verified against the contract entry_meta.ui.resourceUri: each view rendered with app-render (Step 6); initialized, errors, cspViolations, the message log, and the screenshots checked; a missing peer or browser recorded as skipped, not as a server finding© cyanheads, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in framework-skills/field-test of cyanheads/pubmed-mcp-server.
Open the folder on GitHubat commit 5a417fb
Field Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Field Test this skillcyanheads/pubmed-mcp-server | 156 | — | ~11k | Automated safety check: Pass | Apache-2.0 | |
| Setting Up Papergraphlotchuazzz-crypto/papergraph-mcp | 285 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Just PRs MCPClawBio/ClawBio | 1.2k | — | ~3.5k | Automated safety check: Pass | MIT | |
| Patsnap Current Awarenesspatsnap/mcp | 113 | — | ~671 | Automated safety check: Pass | Apache-2.0 | |
| Patsnap Scientific Translational Evidencepatsnap/mcp | 113 | — | ~728 | Automated safety check: Pass | Apache-2.0 | |
| Peer Review Loophashgraph-online/awesome-codex-plugins | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 |
lotchuazzz-crypto/papergraph-mcp
A skill your agent uses when a user has cloned PaperGraph MCP and asks to install, initialize, configure, set up, or start using it with an agent or MCP client.
ClawBio/ClawBio
Compute evidence-aware polygenic risk scores from a local VCF or WGS file through the validated just-prs engine and a pinned local just-prs MCP server.
patsnap/mcp
Patsnap Current Awareness MCP for AI agents. An agent skill from patsnap/mcp.
patsnap/mcp
Patsnap Scientific & Translational Evidence MCP for AI agents.
hashgraph-online/awesome-codex-plugins
Peer Review Ralph Loop — combines Cavekit kits with a Ralph Loop and true cross-model peer review using Codex (OpenAI).
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
cyanheads/pubmed-mcp-server
Scaffold an MCP App tool + UI resource pair. An agent skill from cyanheads/pubmed-mcp-server.
cyanheads/pubmed-mcp-server
Scaffold a new MCP prompt template. An agent skill from cyanheads/pubmed-mcp-server.
cyanheads/pubmed-mcp-server
Scaffold a new MCP resource definition. An agent skill from cyanheads/pubmed-mcp-server.
cyanheads/pubmed-mcp-server
Scaffold a new service integration. An agent skill from cyanheads/pubmed-mcp-server.
cyanheads/pubmed-mcp-server
Scaffold a test file for an existing tool, resource, or service.
cyanheads/pubmed-mcp-server
Authentication, authorization, and multi-tenancy patterns for @cyanheads/mcp-ts-core.
Works with
Categories
Exercise tools, resources, and prompts against a live HTTP server via MCP JSON-RPC over curl. Field Test is an agent skill from cyanheads/pubmed-mcp-server. Exercise tools, resources, and prompts against a live HTTP server via MCP JSON-RPC over curl.
Field Test fits situations like: verify their MCP surface; tasks that involve MCP servers.
Run `npx skills add cyanheads/pubmed-mcp-server --skill field-test -a claude-code`. Or copy the skill folder (framework-skills/field-test in cyanheads/pubmed-mcp-server) into .claude/skills/field-test in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cyanheads/pubmed-mcp-server --skill field-test -a codex`. Or copy the skill folder (framework-skills/field-test in cyanheads/pubmed-mcp-server) into .agents/skills/field-test in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cyanheads/pubmed-mcp-server --skill field-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/field-test, .gemini/skills/field-test, .github/skills/field-test and .opencode/skills/field-test in your project.
Going by SKILL.md and its folder, Field Test needs the command-line tools its instructions call (jq, bun, curl, bunx and npx).
SKILL.md names 2 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: modelcontextprotocol.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Field Test is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 11k tokens (SKILL.md is roughly 43k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Field Test: Setting Up Papergraph (lotchuazzz-crypto/papergraph-mcp, 285 stars), Just PRs MCP (ClawBio/ClawBio, 1.2k stars), Patsnap Current Awareness (patsnap/mcp, 113 stars) and Patsnap Scientific Translational Evidence (patsnap/mcp, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cyanheads (a GitHub user) maintains it in cyanheads/pubmed-mcp-server, which has 156 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 4, 2026.
Source: cyanheads/pubmed-mcp-server on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.