MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
为 OCE 检索评测构建或扩展基准集——也就是 bench/datasets/.jsonl 里的「问题 + 标准答案(expectedfiles)」。当被要求「出评测题/加题/扩充 benchmark/把题量提到 N 道/增加刁钻角度/为新仓库做基准/校准难度/复核 expectedfiles」时使用。涵盖题目 schema、评分语义、10…
$ npx skills add oce-ai/oce --skill build-eval-benchmark -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install oce-ai/oce build-eval-benchmark --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/oce-ai/oce.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/build-eval-benchmark .claude/skills/build-eval-benchmark && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "build-eval-benchmark" agent skill from https://github.com/oce-ai/oce/tree/master/.claude/skills/build-eval-benchmark into .claude/skills/build-eval-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-eval-benchmark", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/oce-ai/oce/tree/master/.claude/skills/build-eval-benchmarkType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add oce-ai/oce --skill build-eval-benchmark -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install oce-ai/oce build-eval-benchmark --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oce-ai/oce.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/build-eval-benchmark .agents/skills/build-eval-benchmark && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "build-eval-benchmark" agent skill from https://github.com/oce-ai/oce/tree/master/.claude/skills/build-eval-benchmark into .agents/skills/build-eval-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-eval-benchmark", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add oce-ai/oce --skill build-eval-benchmark -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install oce-ai/oce build-eval-benchmark --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oce-ai/oce.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/build-eval-benchmark .cursor/skills/build-eval-benchmark && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "build-eval-benchmark" agent skill from https://github.com/oce-ai/oce/tree/master/.claude/skills/build-eval-benchmark into .cursor/skills/build-eval-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-eval-benchmark", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/oce-ai/oce.git --path .claude/skills/build-eval-benchmark--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add oce-ai/oce --skill build-eval-benchmark -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install oce-ai/oce build-eval-benchmark --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oce-ai/oce.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/build-eval-benchmark .gemini/skills/build-eval-benchmark && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "build-eval-benchmark" agent skill from https://github.com/oce-ai/oce/tree/master/.claude/skills/build-eval-benchmark into .gemini/skills/build-eval-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-eval-benchmark", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install oce-ai/oce build-eval-benchmarkInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add oce-ai/oce --skill build-eval-benchmark -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/oce-ai/oce.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/build-eval-benchmark .github/skills/build-eval-benchmark && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "build-eval-benchmark" agent skill from https://github.com/oce-ai/oce/tree/master/.claude/skills/build-eval-benchmark into .github/skills/build-eval-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-eval-benchmark", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add oce-ai/oce --skill build-eval-benchmark -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install oce-ai/oce build-eval-benchmark --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oce-ai/oce.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/build-eval-benchmark .opencode/skills/build-eval-benchmark && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "build-eval-benchmark" agent skill from https://github.com/oce-ai/oce/tree/master/.claude/skills/build-eval-benchmark into .opencode/skills/build-eval-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-eval-benchmark", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
build-eval-benchmark为 OCE 检索评测构建或扩展基准集——也就是 bench/datasets/.jsonl 里的「问题 + 标准答案(expectedfiles)」。当被要求「出评测题/加题/扩充 benchmark/把题量提到 N 道/增加刁钻角度/为新仓库做基准/校准难度/复核 expectedfiles」时使用。涵盖题目 schema、评分语义、10…
Build Eval Benchmark is an agent skill from oce-ai/oce. 为 OCE 检索评测构建或扩展基准集——也就是 bench/datasets/.jsonl 里的「问题 + 标准答案(expectedfiles)」。当被要求「出评测题/加题/扩充 benchmark/把题量提到 N 道/增加刁钻角度/为新仓库做基准/校准难度/复核 expectedfiles」时使用。涵盖题目 schema、评分语义、10 类题型、难度梯度、刁钻角度清单、答案正确性校验与防过拟合规则。
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows. The repository describes itself as: OpenContextEngine is a self-hosted, ACE-compatible code retrieval service. It indexes source files with cAST-aware chunking, stores metadata in PostgreSQL or SQLite, performs… The licence is Apache-2.0.
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d1f5493. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythongitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Build Eval Benchmark loads about 1.7k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 359 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from oce-ai/oce at commit d1f5493, republished under its Apache-2.0 licence (© oce-ai). 359 words, ~1,656 tokens.
.claude/skills/build-eval-benchmark/SKILL.md (or your agent's skills folder).这个技能教你为一个目标仓库生成高质量的检索评测题:每行一个问题,连同它的标准答案
(expected_files)。题目最终喂给 oce bench(评测 harness 在 src/oce/bench/,评分口径在
src/oce/bench/scoring.py),对一个运行中的 OCE 服务打分。答案错误 = 评测失效,所以正确性
优先于数量。
expected_files 是否真的正确读题之前必须先懂评分器怎么用你的答案,否则会写出无法被正确评分的题。
评分逻辑见 src/oce/bench/scoring.py(ndcg_at_k / score_query / relevance_grades):
expected_files 中任意一项。expected_files 是按相关性从高到低书写的:expected_files 支持 glob:含 *?[ 任一字符即按 fnmatch 匹配,例如
src-tauri/src/commands/*.rs。用于「一族文件里命中任一即可」的题。\ 归一成 /;**一律用正斜杠的仓库相对路径**,不带盘符、不带前导 /。推论:想让一道题「难」,不是把
expected_files写得更深,而是让正确的首项文件与 查询的字面词重叠更少、且仓库里存在强干扰项会把排序带偏。
每行一个 JSON 对象(UTF-8,可带 BOM):
{"id": "Q31", "category": "symbol_location", "difficulty": 2,
"query": "`add_provider` 函数在哪个 Rust 文件定义?",
"expected_files": ["src-tauri/src/commands/provider.rs"],
"must_contain": ["add_provider"]}| 字段 | 必填 | 说明 |
|---|---|---|
id | 是 | Q01…QNN,全文件唯一(harness 也接受 query_id,但统一用 id) |
category | 是 | 见第 3 节十类之一,用于报告分组 |
difficulty | 是 | 1/2/3,见第 4 节 |
query | 是 | 自然语言问题。默认中文(与现有基准一致);评测跨语言/英文检索时可英文 |
expected_files | 是 | 标准答案,按相关性降序,仓库相对正斜杠路径或 glob |
must_contain | 否 | 人工复核用的关键词提示。评分器当前不读取它——别依赖它判分,只当注记 |
README、目录树、入口文件、构建配置,搞清技术栈与模块边界。
先用 git ls-files / Glob 拿到真实文件清单——这是后面校验答案存在性的依据。expected_files 里的位置(首项 = rel=2 的真答案)。绝不凭文件名猜测答案。| category | 考什么 | 典型答案形态 |
|---|---|---|
file_exact_match | 按文件名/类型精确定位 | 单个唯一文件(package.json) |
configuration_lookup | 找某项配置所在文件 | 配置文件 |
symbol_location | 找函数/类/常量定义处 | 定义所在文件 |
component_location | 找 UI 组件/模块所在 | 组件文件 |
api_usage | 找某 API/接口被调用的地方 | 调用点(注意区别于定义) |
semantic_feature | 按功能语义找实现(词面常不重叠) | 实现文件,常跨前后端 |
call_chain | 追一条调用的完整路径 | 有序多文件:UI→api 层→命令层 |
architecture_understanding | 理解某机制/流程落在哪 | 核心机制文件 |
cross_language | 同一概念跨语言两端的对应 | 多文件,每种语言一项 |
error_handling | 找错误处理/边界/日志逻辑 | 错误处理文件 |
call_chain 和 cross_language 的 expected_files 天然多项;首项放调用链起点或最核心
的那一端。
Cargo.toml 在哪")。lib.rs/mod.rs/index.ts/App 这类 barrel 文件压制)、需要多跳追踪、
或存在强干扰项(同名不同义、近义文件)。检索系统稳定答错的题多在此档。给难度 2/3 的题挑 1 个角度施加。每个角度都对应一种真实检索失败模式:
init_status.rs)mod.rs/lib.rs/index.ts——除非题目就是要那个聚合入口。provider.rs / providers.ts / ProviderForm.tsx /
provider/mod.rs 多个高度相似路径,题目精确指向其中一个。难度来自答案与查询之间的语义距离 + 干扰强度,不是来自把路径写得又长又偏。
把候选写到 bench/datasets/<name>.jsonl 后,用目标仓库根目录跑:
python - "$REPO_ROOT" bench/datasets/<name>.jsonl <<'PY'
import json, sys, glob as g, os
repo, path = sys.argv[1], sys.argv[2]
ids, errs = set(), []
for n, line in enumerate(open(path, encoding="utf-8-sig"), 1):
if not line.strip(): continue
r = json.loads(line)
qid = r.get("id", f"line{n}")
if qid in ids: errs.append(f"{qid}: 重复 id")
ids.add(qid)
for k in ("id","category","difficulty","query","expected_files"):
if k not in r: errs.append(f"{qid}: 缺字段 {k}")
if r.get("difficulty") not in (1,2,3): errs.append(f"{qid}: difficulty 非 1/2/3")
ef = r.get("expected_files") or []
if not ef: errs.append(f"{qid}: expected_files 为空"); continue
for e in ef:
if "\\" in e or e.startswith("/"): errs.append(f"{qid}: 路径不规范 {e}")
# glob 或精确路径都必须能在仓库里命中至少一个真实文件
hit = g.glob(os.path.join(repo, e), recursive=True) if any(c in e for c in "*?[") \
else [os.path.join(repo, e)]
if not any(os.path.isfile(h) for h in hit):
errs.append(f"{qid}: 答案不存在 {e}")
print("\n".join(errs) if errs else f"OK: {len(ids)} 题,schema 与答案路径全部通过")
PY脚本只校验机械正确性(存在、schema、路径规范、id 唯一)。语义正确性(这个文件 是否真的回答这个问题、首项是否真该排第一)必须你亲自打开文件确认——机器查不出来。
一条经验教训:把「从失败案例提取的词典 / 单题调参」写进检索代码会让分数虚高,去过拟合后 明显回落,且泛化未经验证。出题时同理:
<name>.metadata.json(schema 见 src/oce/bench/datasets.py 文档串):
钉住被测仓库的 repository.commit / describe,用于 RunRecord 溯源与 lock 状态判断。docs/evaluation-guide.md,新旧分数不可直接比。id/category/difficulty/query/expected_files)id 全文件唯一,difficulty ∈ {1,2,3}expected_files 路径在目标仓库真实存在(glob 至少命中一个文件)/bench/datasets/(配套 <name>.metadata.json),必要时更新
docs/evaluation-guide.md 的题数/类别分布表© oce-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/build-eval-benchmark of oce-ai/oce.
Open the folder on GitHubat commit d1f5493
Build Eval Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Build Eval Benchmark this skilloce-ai/oce | 121 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| MCP Server Builderanthropics/skills | 180k | 63 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Hook Development for Claude Code Pluginsanthropics/claude-plugins-official | 38k | 10 repos | ~4.1k | Automated safety check: Notes | Apache-2.0 | |
| Using Superpowersfarm-fe/farm | 5.6k | 35 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Executing Plans Inlineobra/superpowers | 297k | 2 repos | ~5.1k | Automated safety check: Pass | MIT | |
| Skill CreatorAzure/azqr | 795 | 89 repos | ~8.2k | Automated safety check: Pass | Apache-2.0 |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
anthropics/claude-plugins-official
Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.
farm-fe/farm
A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions
obra/superpowers
Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.
Azure/azqr
Create new skills, modify and improve existing skills, and measure skill performance.
anthropics/claude-plugins-official
Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.
oce-ai/oce
OpenContextEngine (oce) 仓库的版本发布运营技能。使用 scripts/ 下的 bumpversion.py 更新版本号(pyproject.toml 与 src/oce/init.py 同步)、generatechangelog.py 从 git 提交生成 CHANGELOG、release.py 执行 bump→changelog→build→commit→tag…
Categories
为 OCE 检索评测构建或扩展基准集——也就是 bench/datasets/.jsonl 里的「问题 + 标准答案(expectedfiles)」。当被要求「出评测题/加题/扩充 benchmark/把题量提到 N 道/增加刁钻角度/为新仓库做基准/校准难度/复核 expectedfiles」时使用。涵盖题目 schema、评分语义、10…. Build Eval Benchmark is an agent skill from oce-ai/oce.
Build Eval Benchmark fits situations like: agent Workflows work in your project.
Run `npx skills add oce-ai/oce --skill build-eval-benchmark -a claude-code`. Or copy the skill folder (.claude/skills/build-eval-benchmark in oce-ai/oce) into .claude/skills/build-eval-benchmark in your project. Claude Code loads it when a task matches its description.
Run `npx skills add oce-ai/oce --skill build-eval-benchmark -a codex`. Or copy the skill folder (.claude/skills/build-eval-benchmark in oce-ai/oce) into .agents/skills/build-eval-benchmark in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oce-ai/oce --skill build-eval-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-eval-benchmark, .gemini/skills/build-eval-benchmark, .github/skills/build-eval-benchmark and .opencode/skills/build-eval-benchmark in your project.
Going by SKILL.md and its folder, Build Eval Benchmark needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Build Eval Benchmark is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Build Eval Benchmark: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
oce-ai (a GitHub organization) maintains it in oce-ai/oce, which has 121 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 8, 2026.
Source: oce-ai/oce on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.