Generate a search quality benchmark for the AI Registry. An agent skill from agentic-community/mcp-gateway-registry.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Search Benchmark

skills CLI
$ npx skills add agentic-community/mcp-gateway-registry --skill search-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentic-community/mcp-gateway-registry search-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentic-community/mcp-gateway-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/search-benchmark .claude/skills/search-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
search-benchmark
GitHub stars
962
Token cost
~1.9k tokens
SKILL.md length
746 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generate a search quality benchmark for the AI Registry. An agent skill from agentic-community/mcp-gateway-registry.

  • Works in 4 steps: Check for Existing Ground Truth → Run Benchmark → Review Report → …
  • You want to measure search quality after changes to the scoring algorithm
  • SKILL.md covers Prerequisites, Input, Workflow and Output Format, plus 4 more sections
  • Calls curl, uv and python3; reaches d2xl2zfuhgc4l0.cloudfront.net

What it does

Search Benchmark is an agent skill from agentic-community/mcp-gateway-registry. Generate a search quality benchmark for the AI Registry. Generates ground truth from the registry's assets, runs 100+ queries against the semantic search API, evaluates results using NDCG@10/MRR/Recall, and produces a markdown report. Use when you want to measure search quality after changes to the scoring algorithm, embedding model, or indexed content.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Embeddings. The repository describes itself as: Enterprise-ready MCP Gateway & Registry that centralizes AI development tools with secure OAuth authentication, dynamic tool discovery, and unified access for both autonomous AI… The licence is Apache-2.0.

When your agent uses it

  • You want to measure search quality after changes to the scoring algorithm
  • Embedding model
  • Indexed content

Example prompts

  • “/search-benchmark”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Check for Existing Ground Truth
  2. Run Benchmark
  3. Review Report
  4. Compare Before/After (Optional)

What it can do on your machine

Read from SKILL.md and the folder at commit ec3a197. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • uv
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • d2xl2zfuhgc4l0.cloudfront.net

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Search Benchmark loads about 1.9k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 746 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentic-community/mcp-gateway-registry at commit ec3a197, republished under its Apache-2.0 licence (© agentic-community). 746 words, ~1,914 tokens.

Download SKILL.mdSave it as .claude/skills/search-benchmark/SKILL.md (or your agent's skills folder).
name
search-benchmark
description
Generate a search quality benchmark for the AI Registry. Generates ground truth from the registry's assets, runs 100+ queries against the semantic search API, evaluates results using NDCG@10/MRR/Recall, and produces a markdown report. Use when you want to measure search quality after changes to the scoring algorithm, embedding model, or indexed content.
license
Apache-2.0
metadata.author
mcp-gateway-registry
metadata.version
1.0

Search Benchmark Skill

Measure semantic search quality against a deployed AI Registry. Generates a ground truth dataset from the registry's own assets, runs queries against the live API, evaluates results using standard information retrieval metrics (NDCG@10, MRR, Recall@10), and produces a markdown report.

Prerequisites

  1. Registry URL - The base URL of the deployed registry (e.g., https://d2xl2zfuhgc4l0.cloudfront.net)
  2. JWT Token - A valid admin token in .token file (get from "Get JWT Token" button in registry UI)
  3. Registry must have assets indexed - At least some servers, agents, or skills registered

The .token file supports both raw JWT format and the full JSON response from the registry UI.

Input

/search-benchmark [REGISTRY_URL] [TOKEN_FILE]
  • REGISTRY_URL - Base URL of the registry to benchmark (default: reads from user or uses http://localhost)
  • TOKEN_FILE - Path to the token file (default: .token)

Workflow

Step 1: Check for Existing Ground Truth

Check if a ground truth dataset already exists:

bash
ls tests/fixtures/search_dataset/ground_truth.json 2>/dev/null

If the file exists, report how many queries it contains and ask the user: "A ground truth dataset already exists (N queries). Do you want to use it or generate a new one from this registry?"

  • If use existing: skip to Step 2
  • If generate new: proceed to Step 1b
Step 1b: Generate Expert Ground Truth

This is NOT a simple programmatic generation. You must deeply analyze the registry's assets and craft queries like a search expert. Follow this process:

1b.1: Fetch all assets as JSON

bash
TOKEN=$(python3 -c "
import json
with open('{TOKEN_FILE}') as f:
    raw = f.read().strip()
if raw.startswith('{'):
    data = json.loads(raw)
    print(data.get('tokens',{}).get('access_token') or data.get('access_token',''))
else:
    print(raw.replace('Bearer ',''))
")

curl -s -H "Authorization: Bearer $TOKEN" "{REGISTRY_URL}/api/servers?limit=2000" > /tmp/servers.json
curl -s -H "Authorization: Bearer $TOKEN" "{REGISTRY_URL}/api/agents?limit=2000" > /tmp/agents.json
curl -s -H "Authorization: Bearer $TOKEN" "{REGISTRY_URL}/api/skills?limit=2000" > /tmp/skills.json

1b.2: Analyze the assets deeply

Read the JSON dumps. Filter out stress-test and security-pending assets. For each real asset, understand:

  • What it does (from description)
  • What tools it has (tool names and descriptions)
  • What tags it uses
  • How users would naturally search for it

1b.3: Craft 100 queries across these categories

CategoryCountStrategy
exact-name10Product/server/agent names as queries
semantic15Natural language paraphrases with NO keyword overlap
tool-precision10Exact tool names as queries
tag-based10Tag vocabulary combinations
multi-entity10Queries where correct answers span servers + agents + skills
conflict10Ambiguous terms where vector and keyword disagree
no-answer10Queries completely outside the dataset (no match exists)
agent-focused10Queries specifically targeting agents
skill-focused10Queries specifically targeting skills
tricky15Edge cases, short queries, non-English, adversarial

1b.4: For each query, define expected results

json
{
  "query": "the search query",
  "category": "one of the categories above",
  "description": "what this query tests and why",
  "expected": [
    {"path": "/exact-path-from-registry", "grade": 3, "reason": "why this should match"},
    {"path": "/another-path", "grade": 2, "reason": "why this is relevant but not perfect"}
  ]
}

Grade scale: 3 = perfect match, 2 = highly relevant, 1 = somewhat relevant. For no-answer queries, expected should be an empty list.

1b.5: Validate all paths exist

Every path in expected results must exist in the fetched assets. Verify before saving.

1b.6: Save

Write the ground truth to tests/fixtures/search_dataset/ground_truth.json.

Tell the user how many queries were created per category and that they should review it.

Step 2: Run Benchmark

Run all queries against the live semantic search API and generate a report:

bash
uv run python scripts/benchmark_search.py \
    --url {REGISTRY_URL} \
    --token-file {TOKEN_FILE} \
    --queries tests/fixtures/search_dataset/generated_ground_truth.json

Output:

  • tests/fixtures/search_dataset/benchmark_results.json (raw results)
  • tests/fixtures/search_dataset/benchmark_results.md (markdown report)
Show full SKILL.md (290 more words)Show less
Step 3: Review Report

Show the user the key metrics from the report:

  1. Quality Metrics - NDCG@10 > 0.7 is good, > 0.8 is excellent
  2. Score Health - Saturated scores at 1.0 should be < 15% (if higher, scoring formula may have issues)
  3. Quality by Category - Identify weak areas (e.g., agent-focused queries underperforming)
  4. Per-query results - Check specific queries that scored 0.0 (complete miss) or low NDCG

Open the report in the editor for the user to review.

Step 4: Compare Before/After (Optional)

If asked to compare two runs (e.g., before and after a scoring algorithm change):

bash
uv run python scripts/benchmark_search.py \
    --compare tests/fixtures/search_dataset/benchmark_results.json other_results.json

Output Format

The report includes:

  • Registry metadata: version, database backend, server/agent/skill counts
  • Quality Metrics: NDCG@10, MRR, Recall@10 averaged across all queries
  • Score Health: saturation analysis (unique scores, % at 1.0, range)
  • Quality by Category: breakdown per query type
  • Per-query results: top 5 results with scores and ground truth comparison

Interpreting Metrics

MetricGoodExcellentPoor
NDCG@10> 0.65> 0.80< 0.50
MRR> 0.70> 0.85< 0.50
Recall@10> 0.75> 0.90< 0.60
Score saturation< 15%< 5%> 30%

Example

$ /search-benchmark https://d2xl2zfuhgc4l0.cloudfront.net .token

This will:

  1. Generate ~90 queries from the registry's assets
  2. Run each query against the semantic search API
  3. Produce a report showing search quality metrics
  4. Open the report for review

Troubleshooting

  • 401 errors: Token expired. Get a fresh one from the registry UI "Get JWT Token" button.
  • 0 assets found: Registry has no servers/agents/skills registered, or token lacks permissions.
  • Low NDCG on "description-derived" category: Expected for programmatically generated queries. Add hand-curated queries for better evaluation.
  • High saturation (>30% at 1.0): Scoring formula may be using the legacy additive method. Set SEARCH_FUSION_METHOD=rrf on the registry.

© agentic-community, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/search-benchmark of agentic-community/mcp-gateway-registry.

Open the folder on GitHubat commit ec3a197

Compare with similar skills

Search Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Search Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Search Benchmark this skillagentic-community/mcp-gateway-registry962—~1.9kAutomated safety check: PassApache-2.0
Hybrid Search Implementationwshobson/agents40k9 repos~497Automated safety check: PassMIT
Gemini Live APIgoogle/skills21k—~2.5kAutomated safety check: NotesApache-2.0
Turnstile Spinswyxio/skills172—~987Automated safety check: PassMIT
Multimodal Embedding Serving Devopen-edge-platform/edge-ai-libraries168—~1.2kAutomated safety check: PassApache-2.0
Storing And Querying Vectorsaws/agent-toolkit-for-aws2.8k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Shows how to run vector and keyword search side by side and merge their results, so retrieval catches both meaning and exact terms in RAG and search systems.

    40k GitHub starsUsed in 9 repos~497 tokens
    AI & LLM EngineeringAuto-check passed
  • Gemini Live API

    google/skills

    Official

    Generates a Gemini LiveAPI client service class in the user's chosen programming language.

    21k GitHub stars~2.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Turnstile Spin

    swyxio/skills

    Add or repair Cloudflare Turnstile on an existing web form by creating or reusing a widget, embedding it, wiring mandatory server-side Siteverify in the existing backend, and validating the result.

    172 GitHub stars~987 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Multimodal Embedding Serving Dev

    open-edge-platform/edge-ai-libraries

    Develop the Multimodal Embedding Serving microservice itself — Poetry install, run the existing tests, navigate the wrapper/registry/handler architecture, add a new model family, and build the image…

    168 GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Storing And Querying Vectors

    aws/agent-toolkit-for-aws

    Official

    Store and query vector embeddings using Amazon S3 Vectors, a cost-effective long-term vector storage service with its own API namespace (s3vectors).

    2.8k GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Debugging Signals Pipeline

    PostHog/posthog-foss

    Official

    Debug the signals pipeline locally end-to-end. An agent skill from PostHog/posthog-foss.

    721 GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from agentic-community/mcp-gateway-registry

All 17 skills in this repo
  • Explainer

    agentic-community/mcp-gateway-registry

    Explain a GitHub issue or pull request at 100, 200, and 300 level.

    962 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Debug

    agentic-community/mcp-gateway-registry

    Debug issues in the MCP Gateway Registry using first-principles thinking.

    962 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check: notes
  • Infra Sync

    agentic-community/mcp-gateway-registry

    Keep Terraform and CDK infrastructure in sync. An agent skill from agentic-community/mcp-gateway-registry.

    962 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Writing

    agentic-community/mcp-gateway-registry

    Write prose people will actually read. An agent skill from agentic-community/mcp-gateway-registry.

    962 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Agentcore Register

    agentic-community/mcp-gateway-registry

    Given an MCP server URL, probe the server via curl to discover its metadata and tools, then generate a markdown file with copy-pasteable content for each field in the Amazon Bedrock AgentCore…

    962 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Benchmark Report

    agentic-community/mcp-gateway-registry

    Generate a benchmark report from stress test results (registration, API performance, search concurrency).

    962 GitHub stars~590 tokensUpdated yesterday
    Auto-check passed

Questions about Search Benchmark

What does Search Benchmark do?

Generate a search quality benchmark for the AI Registry. An agent skill from agentic-community/mcp-gateway-registry. Search Benchmark is an agent skill from agentic-community/mcp-gateway-registry. Generate a search quality benchmark for the AI Registry.

When should I use Search Benchmark?

Search Benchmark fits situations like: you want to measure search quality after changes to the scoring algorithm; embedding model; indexed content.

How do I install Search Benchmark in Claude Code?

Run `npx skills add agentic-community/mcp-gateway-registry --skill search-benchmark -a claude-code`. Or copy the skill folder (.claude/skills/search-benchmark in agentic-community/mcp-gateway-registry) into .claude/skills/search-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Search Benchmark in Codex?

Run `npx skills add agentic-community/mcp-gateway-registry --skill search-benchmark -a codex`. Or copy the skill folder (.claude/skills/search-benchmark in agentic-community/mcp-gateway-registry) into .agents/skills/search-benchmark in your project. Codex loads it when a task matches its description.

Can I use Search Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentic-community/mcp-gateway-registry --skill search-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/search-benchmark, .gemini/skills/search-benchmark, .github/skills/search-benchmark and .opencode/skills/search-benchmark in your project.

What does Search Benchmark need to run?

Going by SKILL.md and its folder, Search Benchmark needs the command-line tools its instructions call (curl, uv and python3). Our summary lists: Python 3.

Does Search Benchmark access the network?

SKILL.md names 1 domain. In commands or code: d2xl2zfuhgc4l0.cloudfront.net; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Search Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Search Benchmark use?

Search Benchmark is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Search Benchmark use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Search Benchmark?

Skills that share tags, products or a category with Search Benchmark: Hybrid Search Implementation (wshobson/agents, 40k stars), Gemini Live API (google/skills, 21k stars), Turnstile Spin (swyxio/skills, 172 stars) and Multimodal Embedding Serving Dev (open-edge-platform/edge-ai-libraries, 168 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Search Benchmark?

agentic-community (a GitHub organization) maintains it in agentic-community/mcp-gateway-registry, which has 962 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 6, 2026.

Source: agentic-community/mcp-gateway-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.