MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
Define metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-3-metrics --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-3-metrics .claude/skills/kayba-stage-3-metrics && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "kayba-stage-3-metrics" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-3-metrics into .claude/skills/kayba-stage-3-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-3-metrics", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-3-metricsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-3-metrics --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-3-metrics .agents/skills/kayba-stage-3-metrics && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "kayba-stage-3-metrics" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-3-metrics into .agents/skills/kayba-stage-3-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-3-metrics", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-3-metrics --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-3-metrics .cursor/skills/kayba-stage-3-metrics && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "kayba-stage-3-metrics" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-3-metrics into .cursor/skills/kayba-stage-3-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-3-metrics", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/kayba-ai/agentic-context-engine.git --path .claude/skills/kayba-pipeline/stage-3-metrics--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-3-metrics --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-3-metrics .gemini/skills/kayba-stage-3-metrics && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "kayba-stage-3-metrics" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-3-metrics into .gemini/skills/kayba-stage-3-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-3-metrics", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-3-metricsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-3-metrics .github/skills/kayba-stage-3-metrics && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "kayba-stage-3-metrics" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-3-metrics into .github/skills/kayba-stage-3-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-3-metrics", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-3-metrics --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-3-metrics .opencode/skills/kayba-stage-3-metrics && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "kayba-stage-3-metrics" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-3-metrics into .opencode/skills/kayba-stage-3-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-3-metrics", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
kayba-stage-3-metricsDefine metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful.
Kayba Stage 3 Metrics is an agent skill from kayba-ai/agentic-context-engine. Define metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful. Trigger when the user says "run stage 3", "define metrics", "build metrics", "compute baselines", or when invoked by the kayba-pipeline orchestrator. Requires eval/stage1insightssummary.md and eval/stage2domaincontext.md to exist.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows. It works with Python. The repository describes itself as: 🧠 Make your agents learn from experience. Now available as a hosted solution at kayba.ai. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 3a31983. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Kayba Stage 3 Metrics loads about 3.4k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 1,633 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from kayba-ai/agentic-context-engine at commit 3a31983, republished under its Apache-2.0 licence (© kayba-ai). 1,633 words, ~3,426 tokens.
.claude/skills/kayba-stage-3-metrics/SKILL.md (or your agent's skills folder).Define metrics from insights, implement as code, run, review, iterate.
TRACES_FOLDER — path to directory containing trace JSON fileseval/stage1_insights_summary.md — output from Stage 1eval/stage2_domain_context.md — output from Stage 2Read both input files before starting.
This stage is iterative. You cycle through define → implement → run → review, with a hard cap of 3 iterations. A metric set is "clean" when ALL of the following hold:
"confidence": "directional-only" and excluded from priority sorting."extreme_justification" field.|events_A ∩ events_B| / min(|events_A|, |events_B|). If > 0.70, merge or drop one.If after 3 iterations the set is not fully clean, ship what you have and log remaining issues in eval/baseline_metrics.json under a top-level "warnings" key.
For each insight from the Kayba analysis, use the evidence fields to identify observable signals in the traces:
Recovery detectors — consecutive calls to the same function where first has error, next succeeds
def has_recovery(calls, function_name):
for i in range(len(calls) - 1):
if calls[i]['name'] == function_name and is_error(calls[i]['output']):
if calls[i+1]['name'] == function_name and is_success(calls[i+1]['output']):
return True
return FalseLoop detectors — N+ consecutive calls to the same function (stuck agent)
Give-up detectors — regex match agent output for abandonment phrases ("I'm unable to", "cannot complete", "beyond my capabilities")
Error classifiers — match function outputs against domain-specific error patterns. Build a pattern table:
ERROR_PATTERNS = {
'pattern_name': r'regex matching the error',
# one entry per distinct error type
}Over-exploration detectors — ratio of explore vs action calls. Use the tool categories from Stage 2. If explore ratio exceeds threshold AND task didn't complete → analysis paralysis
Ground-truth comparison detectors — agent claims a value (dollar amount, flight number, policy rule) in natural language, and the preceding tool response contains the actual value. Extract candidate values from agent text via regex, then compare against structured fields in the tool response JSON. Examples:
# Extract dollar amounts from agent text
DOLLAR_PATTERN = r'\$\s?([\d,]+(?:\.\d{2})?)'
# Extract flight numbers (3 letters + 3 digits)
FLIGHT_PATTERN = r'\b([A-Z]{2,3}\d{3,4})\b'
def check_agent_claims_against_tool(agent_text, preceding_tool_response):
"""Compare values the agent states against the tool response ground truth."""
claimed_amounts = re.findall(DOLLAR_PATTERN, agent_text)
actual_amounts = extract_amounts_from_json(preceding_tool_response)
# A claim is fabricated if it doesn't match any actual value
fabricated = [c for c in claimed_amounts if not any(matches(c, a) for a in actual_amounts)]
return len(fabricated) == 0, fabricatedThis pattern covers data accuracy (fabricated prices/flights), post-action verification (quoted vs actual cost), and policy accuracy (claimed restrictions vs policy text). These are NOT qualitative-only — regex + JSON comparison is noisy but produces a real signal. Build the detector even if it's imperfect; a noisy metric that produces a fix is better than a clean classification that produces nothing.
Ordering/sequencing detectors — agent performs actions in the wrong order (e.g., searches for flights before checking if the reservation is even modifiable). Check whether tool call A appears before tool call B when B should come first.
Clean success — threads where all tasks completed with no errors and no other tags
Write eval/compute_baselines.py with:
--traces-dir (required), --output (default: eval/baseline_metrics.json)load_traces(traces_dir) — loads all JSON trace filestag_thread(thread) — combines all detectors, returns list of tagsnumerator / denominatorcompute_all_baselines(traces_dir) — runs all metrics, returns dictRun it:
python eval/compute_baselines.py --traces-dir {TRACES_FOLDER} --output eval/baseline_metrics.jsonRun these checks in order after every run. Each check either passes or produces a concrete fix action.
Check A — Script health. Did the script error or produce null values? → fix and re-run. This is iteration 0-cost; don't count it toward the 3-iteration cap.
Check B — Small-sample guard. For each metric, examine the denominator:
"confidence": "directional-only" in the output JSON. The metric stays in the report but is excluded from priority sorting in Stage 4. Do NOT drop it — small-sample metrics can still inform qualitative analysis."confidence": "not-observed" and move on).Check C — Extreme-value triage. For any metric at exactly 0% or 100%:
"extreme_justification" in the output. Example: M5=0% is correct because both cancellations in the dataset were on ineligible reservations."at_ceiling": true (or "at_floor": true) to its entry in the output JSON. This signals to Stage 4 (direction setting) and Stage 5 (action planning) that the metric is already optimal and should NOT be listed as needing improvement. Stage 4 must set its direction to "↑ maintain" or "— already optimal", never bare "↑".Check D — Correlation / overlap audit. For every pair of metrics, compute event overlap: |denom_A ∩ denom_B| / min(|denom_A|, |denom_B|). If > 0.70:
Check E — Coverage (strict). For EVERY Stage 1 insight, verify it has a corresponding metric. If an insight has no metric:
After checks, if any produced a fix action: apply fixes and re-run (counts as one iteration). If all checks pass → the metric set is clean. Stop iterating.
Target one metric per insight. Every insight should have a metric unless it is genuinely unmeasurable (see above). If you end up with fewer metrics than insights, you are being too conservative. Directional-only metrics (denominator < 5) still count — they produce fixes in Stage 5. Only apply the redundancy check (Check D) to merge metrics that truly overlap; do not use the metric count as a reason to skip building detectors.
Express every metric as a ratio or percentage. Absolute counts aren't comparable across trace sets.
Prefer per-event denominators over per-thread. "% of EditScript calls with errors" is sharper than "% of threads with any EditScript error." Per-thread denominators compress information — a thread with 10 violations and a thread with 1 both count the same.
One metric per behavioral change. If two would always move together, keep only the sharper one. Use Check D (overlap audit) to enforce this mechanically, not just by intuition.
Build a metric for EVERY insight. "Unmeasurable" is a last resort, not a default. Before classifying an insight as unmeasurable, you MUST attempt to build a programmatic detector. The bar for "unmeasurable" is: you tried a concrete approach, it fundamentally cannot work (not just "it's noisy"), and you can explain why in one sentence.
Specifically:
"confidence": "directional-only". Do NOT skip building the metric. A directional-only metric still produces a fix in Stage 5.If after genuine effort an insight truly cannot be measured programmatically, classify it as:
"qualitative-only" — requires semantic understanding that regex/JSON comparison cannot approximate. Must explain what specific semantic judgment is needed and why pattern matching fails."insufficient-data" — detector exists but denominator is 0 (not just small — literally zero applicable events). Note what scenarios would need to appear in traces."needs-ground-truth" — requires task-specific expected outcomes that aren't in the trace format.Record any remaining unmeasurable insights in the output JSON under a "unmeasurable" key. The goal is for this list to be as short as possible — ideally empty.
eval/compute_baselines.py — runnable script with --traces-dir and --output CLI argseval/baseline_metrics.json — computed baseline values, structured as:{
"M1": {
"name": "single_tool_call_compliance",
"value": 0.414,
"numerator": 12,
"denominator": 29,
"confidence": "full"
},
"M5": {
"name": "cancellation_policy_compliance",
"value": 0.0,
"numerator": 0,
"denominator": 2,
"confidence": "directional-only",
"extreme_justification": "0% correct: both cancellations in dataset were on ineligible reservations"
},
"warnings": ["M5 and M6 have denominator < 5; excluded from priority ranking"],
"unmeasurable": [
{
"insight_id": "d7494740",
"name": "Cabin Change Constraints",
"classification": "insufficient-data",
"reason": "Only 1 update_reservation_flights call in dataset"
}
]
}© kayba-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/kayba-pipeline/stage-3-metrics of kayba-ai/agentic-context-engine.
Open the folder on GitHubat commit 3a31983
Kayba Stage 3 Metrics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Kayba Stage 3 Metrics this skillkayba-ai/agentic-context-engine | 2.6k | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | |
| MCP Server Builderanthropics/skills | 180k | 63 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| MCP Server BuildershareAI-lab/learn-claude-code | 78k | 4 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Mem0 CLI Memory Commandsmem0ai/mem0 | 67k | — | ~2k | Automated safety check: Notes | Apache-2.0 | |
| MemPalace Setup and OperationMemPalace/mempalace | 60k | — | ~2.2k | Automated safety check: Pass | MIT | |
| Google Antigravity SDKgoogle-antigravity/antigravity-sdk-python | 3.7k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
shareAI-lab/learn-claude-code
Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.
mem0ai/mem0
Adds, searches, lists, updates and deletes memories on the Mem0 platform from the terminal with the mem0 command, including a JSON mode built for agents.
MemPalace/mempalace
Installs and configures MemPalace as a private local palace, a shared-brain hub or a client of an existing hub, including MCP registration and version-correct initialization.
google-antigravity/antigravity-sdk-python
Design, implement, and debug autonomous AI agents and multi-agent systems using the Google Antigravity (AGY) SDK.
microsoft/SkillOpt
Runs a nightly or on-demand sleep cycle for a local Codex agent: review past sessions, replay recurring tasks and stage validated skill and memory edits for adoption.
kayba-ai/agentic-context-engine
End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.
kayba-ai/agentic-context-engine
Fetch pre-computed insights from the Kayba API and build a structured summary.
kayba-ai/agentic-context-engine
Gather domain context about the repository and agent — system prompt, tool definitions, domain docs, and behavior patterns from traces.
kayba-ai/agentic-context-engine
Organize computed metrics into a tiered evaluation rubric with leading, lagging, and quality indicators.
kayba-ai/agentic-context-engine
Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations.
kayba-ai/agentic-context-engine
Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.
Works with
Categories
Define metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful. Kayba Stage 3 Metrics is an agent skill from kayba-ai/agentic-context-engine. Define metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful.
Kayba Stage 3 Metrics fits situations like: the user says run stage 3; compute baselines; invoked by the kayba-pipeline orchestrator.
Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a claude-code`. Or copy the skill folder (.claude/skills/kayba-pipeline/stage-3-metrics in kayba-ai/agentic-context-engine) into .claude/skills/kayba-stage-3-metrics in your project. Claude Code loads it when a task matches its description.
Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a codex`. Or copy the skill folder (.claude/skills/kayba-pipeline/stage-3-metrics in kayba-ai/agentic-context-engine) into .agents/skills/kayba-stage-3-metrics in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-3-metrics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kayba-stage-3-metrics, .gemini/skills/kayba-stage-3-metrics, .github/skills/kayba-stage-3-metrics and .opencode/skills/kayba-stage-3-metrics in your project.
Going by SKILL.md and its folder, Kayba Stage 3 Metrics needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Kayba Stage 3 Metrics is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Kayba Stage 3 Metrics: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), Mem0 CLI Memory Commands (mem0ai/mem0, 67k stars) and MemPalace Setup and Operation (MemPalace/mempalace, 60k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
kayba-ai (a GitHub organization) maintains it in kayba-ai/agentic-context-engine, which has 2,590 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 24, 2026.
Source: kayba-ai/agentic-context-engine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.