MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions.
$ npx skills add pproenca/dot-skills --skill eval-mcp -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pproenca/dot-skills eval-mcp --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/eval-mcp .claude/skills/eval-mcp && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-mcp" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/eval-mcp into .claude/skills/eval-mcp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-mcp", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/eval-mcpType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pproenca/dot-skills --skill eval-mcp -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pproenca/dot-skills eval-mcp --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.experimental/eval-mcp .agents/skills/eval-mcp && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-mcp" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/eval-mcp into .agents/skills/eval-mcp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-mcp", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill eval-mcp -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pproenca/dot-skills eval-mcp --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.experimental/eval-mcp .cursor/skills/eval-mcp && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-mcp" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/eval-mcp into .cursor/skills/eval-mcp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-mcp", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pproenca/dot-skills.git --path skills/.experimental/eval-mcp--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pproenca/dot-skills --skill eval-mcp -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pproenca/dot-skills eval-mcp --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.experimental/eval-mcp .gemini/skills/eval-mcp && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-mcp" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/eval-mcp into .gemini/skills/eval-mcp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-mcp", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pproenca/dot-skills eval-mcpInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pproenca/dot-skills --skill eval-mcp -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.experimental/eval-mcp .github/skills/eval-mcp && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-mcp" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/eval-mcp into .github/skills/eval-mcp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-mcp", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill eval-mcp -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pproenca/dot-skills eval-mcp --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.experimental/eval-mcp .opencode/skills/eval-mcp && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-mcp" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/eval-mcp into .opencode/skills/eval-mcp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-mcp", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-mcpMeasures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions.
Eval MCP is an agent skill from pproenca/dot-skills. Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions. Triggers when the user asks to "evaluate MCP tools", "test tool selection", "improve tool descriptions", "check MCP schema quality", or "eval my MCP server". Companion to build-mcp-server.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `gotchas.md`, `metadata.json` and `references/eval-patterns.md`).
It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
bashnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval MCP loads about 2.4k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 755 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 755 words, ~2,389 tokens.
.claude/skills/eval-mcp/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Tool descriptions are prompt engineering — they land directly in Claude's context window and determine whether Claude picks the right tool with the right arguments. This skill makes tool quality measurable and improvable instead of guesswork.
Three levels of testing, each building on the last:
build-mcp-server and wants to validate qualityPhase 1: Connect → Phase 2: Static Analysis → Phase 3: Selection Testing → Phase 4: Optimize
↑__________________________|Phase 4 loops back: apply rewrites → refetch schemas → retest → compare accuracy.
npx)tools/list. Use build-mcp-server/scripts/test-server.sh to verify connectivity first.Connect to the user's MCP server and fetch the tool schemas.
Ask the user how to reach their server:
http://localhost:3000/mcp)node dist/server.js)bash scripts/fetch-tools.sh <url-or-command> <transport> <workspace>/tools.jsonThis calls tools/list via the MCP Inspector CLI and saves the schemas.
Show a summary table:
| # | Tool | Description (preview) | Params | Annotations |
|---|------|-----------------------|--------|-------------|
| 1 | search_issues | Search issues by keyword... | 3 | readOnlyHint |
| 2 | create_issue | Create a new issue... | 4 | — |Flag tool count: 1-15 optimal, 15-30 warning, 30+ excessive (consider search+execute pattern).
Create workspace at {server-name}-eval/ adjacent to the skill directory or in the user's project:
{server-name}-eval/
├── tools.json
├── evals/
│ └── evals.json
└── iteration-N/Run deterministic quality checks — no Claude calls needed. This gives immediate feedback during development.
bash scripts/analyze-schemas.sh <workspace>/tools.json <workspace>/iteration-N/static-analysis.jsonShow per-tool quality scores. Read references/quality-checklist.md for the criteria being checked.
| Tool | Desc | Params | Schema | Annotations | Overall | Issues |
|------|------|--------|--------|-------------|---------|--------|
| search_issues | 3/3 | 3/3 | 2/3 | 2/3 | 2.5 | No negation |
| create_issue | 1/3 | 1/3 | 0/3 | 0/3 | 0.5 | 4 issues |If the analysis found tools with high description overlap, highlight them as confusion risks:
### Sibling Pairs (confusion risk)
| Tool A | Tool B | Overlap | Risk |
|--------|--------|---------|------|
| search_issues | list_issues | 52% | HIGH |If critical issues exist (missing descriptions, zero annotations), recommend fixing them before Phase 3. Static issues create noise in selection testing — fix the obvious problems first, then measure the subtle ones.
If all tools score well, proceed to Phase 3.
Test whether Claude picks the right tool for each user intent. This is the core eval.
Read references/eval-patterns.md for intent generation patterns.
For each tool, generate:
For each sibling pair flagged in Phase 2:
Present all intents to the user for review. Ask if any should be added, removed, or modified.
Save to {workspace}/evals/evals.json:
{
"server_name": "my-server",
"generated_from": "tools.json",
"intents": [
{
"id": 1,
"intent": "Are there any open bugs related to checkout?",
"expected_tool": "search_issues",
"type": "should_trigger",
"target_tool": "search_issues",
"notes": "Implicit intent — doesn't name the action"
}
]
}For each intent, spawn a subagent that receives:
The subagent prompt:
You have access to the following MCP tools:
{tool schemas as JSON}
A user sends this message:
"{intent text}"
Which tool would you call? Respond with JSON:
{
"selected_tool": "tool_name" or null,
"arguments": { ... } or {},
"reasoning": "One sentence explaining your choice"
}
If no tool fits the user's request, set selected_tool to null.
Select exactly ONE tool. Do not suggest calling multiple tools.Save each result to {workspace}/iteration-N/selection/intent-{ID}/result.json.
Launch all selection tests in parallel for efficiency.
bash scripts/grade-selection.sh \
<workspace>/iteration-N/selection \
<workspace>/evals/evals.json \
<workspace>/iteration-N/benchmark.json## Selection Results — Iteration N
**Accuracy:** 82% (41/50 correct)
| Metric | Count |
|--------|-------|
| Correct | 41 |
| Wrong tool | 5 |
| False accept | 2 |
| False reject | 2 |
### Per-Tool Accuracy
| Tool | Precision | Recall |
|------|-----------|--------|
| search_issues | 0.90 | 0.85 |
| create_issue | 1.00 | 1.00 |
### Worst Confusions
| Expected | Selected Instead | Times |
|----------|-----------------|-------|
| list_issues | search_issues | 3 |
| get_user | find_user_by_email | 2 |Analyze confusion patterns and suggest description improvements. Read references/optimization.md for rewrite patterns.
For each confused pair (from worst_confusions):
## Suggested Improvements
### search_issues ↔ list_issues (confused 3 times)
**search_issues — Before:**
> Search issues by keyword.
**search_issues — After:**
> Search issues by keyword across title and body. Returns up to `limit` results ranked by relevance. Does NOT filter by status, assignee, or date — use list_issues for structured filtering.
**Reason:** Adding scope boundary and cross-reference to disambiguate from list_issues.Save to {workspace}/iteration-N/suggestions.json (format defined in optimization.md).
After the user applies the rewrites to their server code:
iteration-N+1 using the same evals.json## Iteration Comparison
| Metric | Iteration 1 | Iteration 2 | Delta |
|--------|------------|------------|-------|
| Accuracy | 82% | 94% | +12% |
| search↔list confusion | 3 | 0 | -3 |Read these when you reach the relevant phase — not upfront:
references/quality-checklist.md — Testable quality criteria for tool schemas (Phase 2)references/eval-patterns.md — How to write tool selection test intents (Phase 3)references/optimization.md — How to improve descriptions from eval results (Phase 4)build-mcp-server — Design and scaffold MCP servers (run this first, then eval-mcp to validate)build-mcp-app — MCP servers with interactive UI widgets© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references) in skills/.experimental/eval-mcp of pproenca/dot-skills.
Open the folder on GitHubat commit cf93c57
Eval MCP next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval MCP this skillpproenca/dot-skills | 215 | — | ~2.4k | Automated safety check: Pass | MIT | |
| MCP Server Builderanthropics/skills | 180k | 64 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| MCP Server BuildershareAI-lab/learn-claude-code | 78k | 5 repos | ~1.2k | Automated safety check: Pass | MIT | |
| MCP Integration for Pluginsanthropics/claude-plugins-official | 38k | 11 repos | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Fastmcp Client CLIPrefectHQ/fastmcp | 28k | 1 repos | ~823 | Automated safety check: Pass | Apache-2.0 | |
| Crush Configurationcharmbracelet/crush | 29k | — | ~3.7k | Automated safety check: Pass | Custom licence |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
shareAI-lab/learn-claude-code
Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.
anthropics/claude-plugins-official
Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.
PrefectHQ/fastmcp
Query and invoke tools on MCP servers using fastmcp list and fastmcp call.
charmbracelet/crush
Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.
mksglu/context-mode
Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.
pproenca/dot-skills
Audio forensics and voice recovery guidelines for CSI-level audio analysis.
pproenca/dot-skills
Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.
pproenca/dot-skills
Create well-structured RFCs and technical proposals for software projects.
pproenca/dot-skills
Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.
pproenca/dot-skills
Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…
pproenca/dot-skills
Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…
Works with
Categories
Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions. Eval MCP is an agent skill from pproenca/dot-skills. Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions.
Eval MCP fits situations like: the user asks to evaluate MCP tools; test tool selection; improve tool descriptions; check MCP schema quality.
Run `npx skills add pproenca/dot-skills --skill eval-mcp -a claude-code`. Or copy the skill folder (skills/.experimental/eval-mcp in pproenca/dot-skills) into .claude/skills/eval-mcp in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pproenca/dot-skills --skill eval-mcp -a codex`. Or copy the skill folder (skills/.experimental/eval-mcp in pproenca/dot-skills) into .agents/skills/eval-mcp in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill eval-mcp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-mcp, .gemini/skills/eval-mcp, .github/skills/eval-mcp and .opencode/skills/eval-mcp in your project.
Going by SKILL.md and its folder, Eval MCP needs a shell for the scripts in its folder and the command-line tools its instructions call (bash and node). Our summary lists: Node.js; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Eval MCP is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Eval MCP: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars) and Fastmcp Client CLI (PrefectHQ/fastmcp, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 215 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on August 15, 2026.
Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.