MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments.
$ npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install AgentToolkit/altk-evolve agent-wiki-compare-outcomes --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/AgentToolkit/altk-evolve.git skills-src && mkdir -p .claude/skills && cp -r skills-src/explorations/agent-wiki/skills/agent-wiki-compare-outcomes .claude/skills/agent-wiki-compare-outcomes && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-wiki-compare-outcomes" agent skill from https://github.com/AgentToolkit/altk-evolve/tree/main/explorations/agent-wiki/skills/agent-wiki-compare-outcomes into .claude/skills/agent-wiki-compare-outcomes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-wiki-compare-outcomes", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/AgentToolkit/altk-evolve/tree/main/explorations/agent-wiki/skills/agent-wiki-compare-outcomesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install AgentToolkit/altk-evolve agent-wiki-compare-outcomes --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AgentToolkit/altk-evolve.git skills-src && mkdir -p .agents/skills && cp -r skills-src/explorations/agent-wiki/skills/agent-wiki-compare-outcomes .agents/skills/agent-wiki-compare-outcomes && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-wiki-compare-outcomes" agent skill from https://github.com/AgentToolkit/altk-evolve/tree/main/explorations/agent-wiki/skills/agent-wiki-compare-outcomes into .agents/skills/agent-wiki-compare-outcomes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-wiki-compare-outcomes", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install AgentToolkit/altk-evolve agent-wiki-compare-outcomes --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AgentToolkit/altk-evolve.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/explorations/agent-wiki/skills/agent-wiki-compare-outcomes .cursor/skills/agent-wiki-compare-outcomes && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-wiki-compare-outcomes" agent skill from https://github.com/AgentToolkit/altk-evolve/tree/main/explorations/agent-wiki/skills/agent-wiki-compare-outcomes into .cursor/skills/agent-wiki-compare-outcomes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-wiki-compare-outcomes", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/AgentToolkit/altk-evolve.git --path explorations/agent-wiki/skills/agent-wiki-compare-outcomes--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install AgentToolkit/altk-evolve agent-wiki-compare-outcomes --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AgentToolkit/altk-evolve.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/explorations/agent-wiki/skills/agent-wiki-compare-outcomes .gemini/skills/agent-wiki-compare-outcomes && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-wiki-compare-outcomes" agent skill from https://github.com/AgentToolkit/altk-evolve/tree/main/explorations/agent-wiki/skills/agent-wiki-compare-outcomes into .gemini/skills/agent-wiki-compare-outcomes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-wiki-compare-outcomes", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install AgentToolkit/altk-evolve agent-wiki-compare-outcomesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/AgentToolkit/altk-evolve.git skills-src && mkdir -p .github/skills && cp -r skills-src/explorations/agent-wiki/skills/agent-wiki-compare-outcomes .github/skills/agent-wiki-compare-outcomes && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-wiki-compare-outcomes" agent skill from https://github.com/AgentToolkit/altk-evolve/tree/main/explorations/agent-wiki/skills/agent-wiki-compare-outcomes into .github/skills/agent-wiki-compare-outcomes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-wiki-compare-outcomes", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install AgentToolkit/altk-evolve agent-wiki-compare-outcomes --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AgentToolkit/altk-evolve.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/explorations/agent-wiki/skills/agent-wiki-compare-outcomes .opencode/skills/agent-wiki-compare-outcomes && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-wiki-compare-outcomes" agent skill from https://github.com/AgentToolkit/altk-evolve/tree/main/explorations/agent-wiki/skills/agent-wiki-compare-outcomes into .opencode/skills/agent-wiki-compare-outcomes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-wiki-compare-outcomes", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-wiki-compare-outcomesContrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments.
Used after the summarize, extract and synthesize passes, this step looks at several normalized trajectories of the same or similar task and sets those that succeeded against those that failed. A bundled Python script, compare_outcomes.py, groups runs by task id, falling back to the normalized task request, and builds an evidence pack: request text, stored outcomes and failure snippets, observed tool and API calls, and tool descriptions shown in the transcript.
Outcomes can come from stored labels or from an LLM judging the transcript, selected with a --judge-outcomes flag set to never, missing or always. The always setting is preferred when stored labels come from a benchmark-specific evaluator. A rule is written only when the trajectories contain evidence for it, so guidelines are never filled in from prior knowledge of the benchmark or application.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 9e5bb56. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Wiki Compare Outcomes loads about 1.7k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 612 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from AgentToolkit/altk-evolve at commit 9e5bb56, republished under its Apache-2.0 licence (© AgentToolkit). 612 words, ~1,654 tokens.
.claude/skills/agent-wiki-compare-outcomes/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use this pass after summarize/extract/synthesize when there are multiple trajectories that can be judged as successful or failed. It can judge outcomes with an LLM from the normalized transcript, so it does not need to depend on benchmark-specific success/failure labels. It derives contrastive guidelines: rules that are supported by a failed path, a successful path, and concrete evidence from task wording, tool/API documentation, tool/API calls, transcript evidence, optional failure snippets, and optionally an LLM success/failure judgment.
This pass exists to avoid hand-authored domain knowledge. Do not write a rule just because you know the benchmark or application. Write a rule only when the input trajectories contain the evidence.
Run the bundled script over normalized trajectory JSON files:
uv run python explorations/agent-wiki/skills/agent-wiki-compare-outcomes/scripts/compare_outcomes.py \
--input <normalized-dir-or-json> \
--out-json <analysis.json> \
--out-md <analysis.md> \
--judge-outcomes alwaysPass --input multiple times to compare several experiment arms.
The script groups traces by metadata.task_id when present; otherwise it uses
a normalized task request. For each group it compares successful and failed
runs, then extracts:
--judge-outcomes missing or
--judge-outcomes always is set;stats.top_tools, code snippets, and source
api_calls.jsonl when available;Judging modes:
--judge-outcomes never: use only stored outcome.success.--judge-outcomes missing: judge only traces without stored outcomes.--judge-outcomes always: ignore stored success labels and use the LLM
judgment for all traces.Prefer --judge-outcomes always when the available stored labels come from a
benchmark evaluator or another dataset-specific schema. Use stored outcomes
only when they are trusted, dataset-neutral annotations you are comfortable
using as ground truth.
Use --judge-include-failures when generic failure reports or evaluator
snippets are available and you want the LLM to interpret them. This does not
require benchmark-specific code; the snippets are passed as opaque evidence.
Without failure snippets or ground truth, an LLM can still identify obvious
tool errors, step-limit failures, missing finalization, or apparent success,
but it may not detect silent semantic mismatches.
Read the generated Markdown. A candidate is promotable only if it has:
If the evidence is incomplete, keep it as a hypothesis. Hypotheses are useful for evaluation notes but should not be promoted into future-agent instructions.
When a candidate is strong, render it as a guideline with provenance:
{
"entities": [
{
"type": "guideline",
"title": "Choose record source from task wording",
"content": "Apply this rule only when the live choice is between the observed successful and failed APIs, or between APIs with the same documented meanings. Prefer the successful source when the request matches its observed documentation. Do not apply this rule when the request explicitly uses failed-side terms; inspect the failed-side source instead. Do not generalize this rule to other record families or unrelated APIs unless a separate contrast includes those APIs.",
"rationale": "In the contrasted trajectories, failed runs used a feed endpoint for a task about the user's own transactions, while the successful run used the documented account-owned transaction endpoint.",
"trigger": "Use only when choosing between the observed successful and failed APIs and the task wording aligns with the successful-side documentation; skip when the task explicitly mentions failed-side terms or asks about a different record family.",
"session_id": "<comparison-id>",
"agent": "agent-wiki-compare-outcomes",
"tags": ["contrastive", "tool-selection", "data-source-routing"],
"normalized_path": "<analysis.json>"
}
]
}Pipe through the normal helper:
cat /tmp/contrastive-guideline.json | uv run python explorations/agent-wiki/skills/scripts/build_agent_wiki.py --wiki-root <wiki-root> render-guidelines
uv run python explorations/agent-wiki/skills/scripts/build_agent_wiki.py --wiki-root <wiki-root> catalog© AgentToolkit, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in explorations/agent-wiki/skills/agent-wiki-compare-outcomes of AgentToolkit/altk-evolve.
Open the folder on GitHubat commit 9e5bb56
Agent Wiki Compare Outcomes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Wiki Compare Outcomes this skillAgentToolkit/altk-evolve | 122 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| MCP Server Builderanthropics/skills | 180k | 62 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Mem0 CLI Memory Commandsmem0ai/mem0 | 67k | — | ~1.4k | Automated safety check: Notes | Apache-2.0 | |
| MemPalace Setup and OperationMemPalace/mempalace | 59k | — | ~2.2k | Automated safety check: Pass | MIT | |
| SkillOpt-Sleep Self-Improvement Cyclemicrosoft/SkillOpt | 18k | — | ~2.3k | Automated safety check: Pass | MIT | |
| Agent Memorytigerless-labs/agent-memory | 2.4k | — | ~1.2k | Automated safety check: Pass | MIT |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
mem0ai/mem0
Adds, searches, lists, updates and deletes memories on the Mem0 platform from the terminal with the mem0 command, including a JSON mode built for agents.
MemPalace/mempalace
Installs and configures MemPalace as a private local palace, a shared-brain hub or a client of an existing hub, including MCP registration and version-correct initialization.
microsoft/SkillOpt
Runs a nightly or on-demand sleep cycle for a local Codex agent: review past sessions, replay recurring tasks and stage validated skill and memory edits for adoption.
tigerless-labs/agent-memory
Read and write the shared long-term memory store. An agent skill from tigerless-labs/agent-memory.
mindscale-noah/MindMemOS
Give an AI agent persistent, cross-session long-term memory through MindMemOS.
AgentToolkit/altk-evolve
Turns a saved agent trajectory into a reusable skill with a SKILL.md and supporting scripts, so later sessions can call the workflow instead of rediscovering it.
AgentToolkit/altk-evolve
Copies a memory the agent just saved into the shared evolve store with a retrieval trigger, so it can be shared across a team and audited like other entities.
AgentToolkit/altk-evolve
Has the agent read a project wiki's AGENTS.md and pull the guidelines that fit the task it is about to start, instead of loading the wiki at session start.
AgentToolkit/altk-evolve
Reads a normalized Claude Code trajectory JSON and turns its errors and solutions into reusable guideline pages in wiki-twobatch/guidelines.
AgentToolkit/altk-evolve
Ingest one or more agent trajectories (raw bob/claude traces or normalized JSON) into an agent-wiki end-to-end — convert, summarize, extract guidelines, synthesize skills, optionally compare…
AgentToolkit/altk-evolve
Read a normalized Claude Code trajectory JSON and write an episodic summary page to wiki-twobatch/summaries/.
Works with
Categories
Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments. Used after the summarize, extract and synthesize passes, this step looks at several normalized trajectories of the same or similar task and sets those that succeeded against those that failed.py, groups runs by task id, falling back to the normalized task request, and builds an evidence pack: request text, stored outcomes and failure snippets, observed tool and API calls, and tool descriptions shown in the transcript.
Agent Wiki Compare Outcomes fits situations like: learning rules from multiple successful and failed runs of the same task; turning benchmark trajectories into evidence-backed agent guidelines; judging run outcomes with an LLM when evaluator labels are missing or benchmark-specific.
Run `npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a claude-code`. Or copy the skill folder (explorations/agent-wiki/skills/agent-wiki-compare-outcomes in AgentToolkit/altk-evolve) into .claude/skills/agent-wiki-compare-outcomes in your project. Claude Code loads it when a task matches its description.
Run `npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a codex`. Or copy the skill folder (explorations/agent-wiki/skills/agent-wiki-compare-outcomes in AgentToolkit/altk-evolve) into .agents/skills/agent-wiki-compare-outcomes in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AgentToolkit/altk-evolve --skill agent-wiki-compare-outcomes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-wiki-compare-outcomes, .gemini/skills/agent-wiki-compare-outcomes, .github/skills/agent-wiki-compare-outcomes and .opencode/skills/agent-wiki-compare-outcomes in your project.
Going by SKILL.md and its folder, Agent Wiki Compare Outcomes needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: uv and Python to run compare_outcomes.py; Normalized trajectory JSON files.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Agent Wiki Compare Outcomes is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agent Wiki Compare Outcomes: MCP Server Builder (anthropics/skills, 180k stars), Mem0 CLI Memory Commands (mem0ai/mem0, 67k stars), MemPalace Setup and Operation (MemPalace/mempalace, 59k stars) and SkillOpt-Sleep Self-Improvement Cycle (microsoft/SkillOpt, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
AgentToolkit (a GitHub organization) maintains it in AgentToolkit/altk-evolve, which has 122 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 7, 2026.
Source: AgentToolkit/altk-evolve on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.