Agent skill

Eval MCP

by pproenca in pproenca/dot-skills

Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions.

MITAuto-check passedAgent Workflows

Install Eval MCP

skills CLI
$ npx skills add pproenca/dot-skills --skill eval-mcp -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pproenca/dot-skills eval-mcp --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/eval-mcp .claude/skills/eval-mcp && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-mcp
GitHub stars
215
Token cost
~2.4k tokens
SKILL.md length
755 words
Files
9 (incl. scripts, references)
Skills in repo
41
Repo updated
First seen
Licence
MIT

At a glance

Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions.

  • Works in 4 steps: Connect & Inventory → Static Analysis → Selection Testing → …
  • The user asks to evaluate MCP tools
  • SKILL.md covers When to Apply, Workflow Overview, Prerequisites and Phase 1 — Connect & Inventory, plus 5 more sections
  • Runs Shell scripts from its folder; calls bash and node

What it does

Eval MCP is an agent skill from pproenca/dot-skills. Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions. Triggers when the user asks to "evaluate MCP tools", "test tool selection", "improve tool descriptions", "check MCP schema quality", or "eval my MCP server". Companion to build-mcp-server.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `gotchas.md`, `metadata.json` and `references/eval-patterns.md`).

It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.

When your agent uses it

  • The user asks to evaluate MCP tools
  • Test tool selection
  • Improve tool descriptions
  • Check MCP schema quality

Example prompts

  • “evaluate MCP tools”
  • “test tool selection”
  • “improve tool descriptions”
  • “/eval-mcp”

Requirements

  • Node.js
  • A Bash shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Connect & Inventory
  2. Static Analysis
  3. Selection Testing
  4. Optimize & Iterate

What it can do on your machine

Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval MCP loads about 2.4k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 755 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 755 words, ~2,389 tokens.

Download SKILL.mdSave it as .claude/skills/eval-mcp/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
eval-mcp
description
Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions. Triggers when the user asks to "evaluate MCP tools", "test tool selection", "improve tool descriptions", "check MCP schema quality", or "eval my MCP server". Companion to build-mcp-server.

Evaluate MCP Tools

Tool descriptions are prompt engineering — they land directly in Claude's context window and determine whether Claude picks the right tool with the right arguments. This skill makes tool quality measurable and improvable instead of guesswork.

Three levels of testing, each building on the last:

  1. Static Analysis — deterministic schema quality checks (no Claude calls)
  2. Selection Testing — does Claude pick the right tool for each intent?
  3. Description Optimization — iterative improvement based on confusion patterns

When to Apply

  • User wants to check if their MCP tool schemas are well-designed
  • User wants to test whether Claude selects the right tools for user intents
  • User is debugging tool confusion (Claude picks the wrong tool)
  • User wants to optimize tool descriptions for better selection accuracy
  • User has finished scaffolding with build-mcp-server and wants to validate quality

Workflow Overview

Phase 1: Connect → Phase 2: Static Analysis → Phase 3: Selection Testing → Phase 4: Optimize
                                                            ↑__________________________|

Phase 4 loops back: apply rewrites → refetch schemas → retest → compare accuracy.

Prerequisites

  • Node.js >= 18 — required for the MCP Inspector CLI (npx)
  • jq — required for schema analysis scripts
  • A running MCP server — the server must respond to tools/list. Use build-mcp-server/scripts/test-server.sh to verify connectivity first.

Phase 1 — Connect & Inventory

Connect to the user's MCP server and fetch the tool schemas.

1a: Get connection details

Ask the user how to reach their server:

  • HTTP/SSE: URL (e.g., http://localhost:3000/mcp)
  • stdio: spawn command (e.g., node dist/server.js)
1b: Fetch tool schemas
bash
bash scripts/fetch-tools.sh <url-or-command> <transport> <workspace>/tools.json

This calls tools/list via the MCP Inspector CLI and saves the schemas.

1c: Display inventory

Show a summary table:

markdown
| # | Tool | Description (preview) | Params | Annotations |
|---|------|-----------------------|--------|-------------|
| 1 | search_issues | Search issues by keyword... | 3 | readOnlyHint |
| 2 | create_issue | Create a new issue... | 4 | — |

Flag tool count: 1-15 optimal, 15-30 warning, 30+ excessive (consider search+execute pattern).

1d: Create workspace

Create workspace at {server-name}-eval/ adjacent to the skill directory or in the user's project:

{server-name}-eval/
├── tools.json
├── evals/
│   └── evals.json
└── iteration-N/

Phase 2 — Static Analysis

Run deterministic quality checks — no Claude calls needed. This gives immediate feedback during development.

2a: Run analysis
bash
bash scripts/analyze-schemas.sh <workspace>/tools.json <workspace>/iteration-N/static-analysis.json
2b: Display results

Show per-tool quality scores. Read references/quality-checklist.md for the criteria being checked.

markdown
| Tool | Desc | Params | Schema | Annotations | Overall | Issues |
|------|------|--------|--------|-------------|---------|--------|
| search_issues | 3/3 | 3/3 | 2/3 | 2/3 | 2.5 | No negation |
| create_issue | 1/3 | 1/3 | 0/3 | 0/3 | 0.5 | 4 issues |
2c: Flag sibling pairs

If the analysis found tools with high description overlap, highlight them as confusion risks:

markdown
### Sibling Pairs (confusion risk)
| Tool A | Tool B | Overlap | Risk |
|--------|--------|---------|------|
| search_issues | list_issues | 52% | HIGH |
2d: Decision point

If critical issues exist (missing descriptions, zero annotations), recommend fixing them before Phase 3. Static issues create noise in selection testing — fix the obvious problems first, then measure the subtle ones.

If all tools score well, proceed to Phase 3.


Phase 3 — Selection Testing

Test whether Claude picks the right tool for each user intent. This is the core eval.

3a: Generate test intents

Read references/eval-patterns.md for intent generation patterns.

For each tool, generate:

  • 3 should-trigger intents — direct, implicit, and casual phrasings
  • 2 should-not-trigger intents — near-miss and keyword overlap

For each sibling pair flagged in Phase 2:

  • 1 disambiguation intent per tool — tests whether Claude picks the RIGHT sibling

Present all intents to the user for review. Ask if any should be added, removed, or modified.

Show full SKILL.md (303 more words)Show less
3b: Save intents

Save to {workspace}/evals/evals.json:

json
{
  "server_name": "my-server",
  "generated_from": "tools.json",
  "intents": [
    {
      "id": 1,
      "intent": "Are there any open bugs related to checkout?",
      "expected_tool": "search_issues",
      "type": "should_trigger",
      "target_tool": "search_issues",
      "notes": "Implicit intent — doesn't name the action"
    }
  ]
}
3c: Run selection tests

For each intent, spawn a subagent that receives:

  1. The full tool schemas from tools.json (formatted as they'd appear in Claude's context)
  2. The user intent text
  3. Instructions to select exactly one tool and provide arguments, or decline if no tool fits

The subagent prompt:

You have access to the following MCP tools:

{tool schemas as JSON}

A user sends this message:
"{intent text}"

Which tool would you call? Respond with JSON:
{
  "selected_tool": "tool_name" or null,
  "arguments": { ... } or {},
  "reasoning": "One sentence explaining your choice"
}

If no tool fits the user's request, set selected_tool to null.
Select exactly ONE tool. Do not suggest calling multiple tools.

Save each result to {workspace}/iteration-N/selection/intent-{ID}/result.json.

Launch all selection tests in parallel for efficiency.

3d: Grade results
bash
bash scripts/grade-selection.sh \
  <workspace>/iteration-N/selection \
  <workspace>/evals/evals.json \
  <workspace>/iteration-N/benchmark.json
3e: Display results
markdown
## Selection Results — Iteration N

**Accuracy:** 82% (41/50 correct)

| Metric | Count |
|--------|-------|
| Correct | 41 |
| Wrong tool | 5 |
| False accept | 2 |
| False reject | 2 |

### Per-Tool Accuracy
| Tool | Precision | Recall |
|------|-----------|--------|
| search_issues | 0.90 | 0.85 |
| create_issue | 1.00 | 1.00 |

### Worst Confusions
| Expected | Selected Instead | Times |
|----------|-----------------|-------|
| list_issues | search_issues | 3 |
| get_user | find_user_by_email | 2 |

Phase 4 — Optimize & Iterate

Analyze confusion patterns and suggest description improvements. Read references/optimization.md for rewrite patterns.

4a: Analyze confusions

For each confused pair (from worst_confusions):

  1. Read both tools' current descriptions
  2. Identify why they're confusing (missing negation, overlapping scope, no cross-reference)
  3. Draft a specific rewrite following the disambiguation patterns in optimization.md
4b: Present suggestions
markdown
## Suggested Improvements

### search_issues ↔ list_issues (confused 3 times)

**search_issues — Before:**
> Search issues by keyword.

**search_issues — After:**
> Search issues by keyword across title and body. Returns up to `limit` results ranked by relevance. Does NOT filter by status, assignee, or date — use list_issues for structured filtering.

**Reason:** Adding scope boundary and cross-reference to disambiguate from list_issues.

Save to {workspace}/iteration-N/suggestions.json (format defined in optimization.md).

4c: Apply and retest

After the user applies the rewrites to their server code:

  1. Restart the server
  2. Re-run Phase 1 to refetch tools.json (descriptions may have changed)
  3. Re-run Phase 2 for updated static analysis
  4. Re-run Phase 3 into iteration-N+1 using the same evals.json
  5. Compare accuracy:
markdown
## Iteration Comparison

| Metric | Iteration 1 | Iteration 2 | Delta |
|--------|------------|------------|-------|
| Accuracy | 82% | 94% | +12% |
| search↔list confusion | 3 | 0 | -3 |
4d: Iteration guidance
  • Change one sibling pair per iteration so you can attribute improvements
  • If accuracy plateaus, the remaining confusions may need architectural changes (merging tools, renaming, or restructuring the tool surface)
  • Stop when accuracy exceeds 90% or when remaining confusions are in ambiguous edge cases that humans would also struggle with

Reference Files

Read these when you reach the relevant phase — not upfront:

  • build-mcp-server — Design and scaffold MCP servers (run this first, then eval-mcp to validate)
  • build-mcp-app — MCP servers with interactive UI widgets

© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/.experimental/eval-mcp of pproenca/dot-skills.

  • SKILL.md
  • gotchas.md
  • metadata.json
  • references/eval-patterns.md
  • references/optimization.md
  • references/quality-checklist.md
  • scripts/analyze-schemas.sh
  • scripts/fetch-tools.sh
  • scripts/grade-selection.sh

Open the folder on GitHubat commit cf93c57

Compare with similar skills

Eval MCP next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval MCP compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval MCP this skillpproenca/dot-skills215—~2.4kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
MCP Integration for Pluginsanthropics/claude-plugins-official38k11 repos~3.1kAutomated safety check: PassApache-2.0
Fastmcp Client CLIPrefectHQ/fastmcp28k1 repos~823Automated safety check: PassApache-2.0
Crush Configurationcharmbracelet/crush29k—~3.7kAutomated safety check: PassCustom licence

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • MCP Integration for Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.

    38k GitHub starsUsed in 11 repos~3.1k tokens
    Agent WorkflowsAuto-check passed
  • Fastmcp Client CLI

    PrefectHQ/fastmcp

    Query and invoke tools on MCP servers using fastmcp list and fastmcp call.

    28k GitHub starsUsed in 1 repo~823 tokens
    Agent WorkflowsAuto-check passed
  • Crush Configuration

    charmbracelet/crush

    Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.

    29k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Context Mode Output Sandbox

    mksglu/context-mode

    Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.

    26k GitHub stars~4.1k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from pproenca/dot-skills

All 41 skills in this repo
  • Audio Voice Recovery

    pproenca/dot-skills

    Audio forensics and voice recovery guidelines for CSI-level audio analysis.

    215 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Codemod React Pipeline

    pproenca/dot-skills

    Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.

    215 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dev Rfc

    pproenca/dot-skills

    Create well-structured RFCs and technical proposals for software projects.

    215 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Dx Harness

    pproenca/dot-skills

    Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.

    215 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Language Spec Author

    pproenca/dot-skills

    Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…

    215 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    215 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Eval MCP

What does Eval MCP do?

Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions. Eval MCP is an agent skill from pproenca/dot-skills. Measures whether Claude uses an MCP server's tools correctly — tests tool selection accuracy, analyzes schema quality, and iteratively optimizes descriptions.

When should I use Eval MCP?

Eval MCP fits situations like: the user asks to evaluate MCP tools; test tool selection; improve tool descriptions; check MCP schema quality.

How do I install Eval MCP in Claude Code?

Run `npx skills add pproenca/dot-skills --skill eval-mcp -a claude-code`. Or copy the skill folder (skills/.experimental/eval-mcp in pproenca/dot-skills) into .claude/skills/eval-mcp in your project. Claude Code loads it when a task matches its description.

How do I install Eval MCP in Codex?

Run `npx skills add pproenca/dot-skills --skill eval-mcp -a codex`. Or copy the skill folder (skills/.experimental/eval-mcp in pproenca/dot-skills) into .agents/skills/eval-mcp in your project. Codex loads it when a task matches its description.

Can I use Eval MCP in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill eval-mcp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-mcp, .gemini/skills/eval-mcp, .github/skills/eval-mcp and .opencode/skills/eval-mcp in your project.

What does Eval MCP need to run?

Going by SKILL.md and its folder, Eval MCP needs a shell for the scripts in its folder and the command-line tools its instructions call (bash and node). Our summary lists: Node.js; A Bash shell.

Does Eval MCP access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval MCP safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eval MCP use?

Eval MCP is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval MCP use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Eval MCP?

Skills that share tags, products or a category with Eval MCP: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars) and Fastmcp Client CLI (PrefectHQ/fastmcp, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval MCP?

pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 215 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on August 15, 2026.

Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.