MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote.
$ npx skills add Q00/ouroboros --skill evaluate -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Q00/ouroboros evaluate --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluate .claude/skills/evaluate && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluate" agent skill from https://github.com/Q00/ouroboros/tree/main/skills/evaluate into .claude/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Q00/ouroboros/tree/main/skills/evaluateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Q00/ouroboros --skill evaluate -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Q00/ouroboros evaluate --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/evaluate .agents/skills/evaluate && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluate" agent skill from https://github.com/Q00/ouroboros/tree/main/skills/evaluate into .agents/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Q00/ouroboros --skill evaluate -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Q00/ouroboros evaluate --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/evaluate .cursor/skills/evaluate && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluate" agent skill from https://github.com/Q00/ouroboros/tree/main/skills/evaluate into .cursor/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Q00/ouroboros.git --path skills/evaluate--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Q00/ouroboros --skill evaluate -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Q00/ouroboros evaluate --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/evaluate .gemini/skills/evaluate && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluate" agent skill from https://github.com/Q00/ouroboros/tree/main/skills/evaluate into .gemini/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Q00/ouroboros evaluateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Q00/ouroboros --skill evaluate -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/evaluate .github/skills/evaluate && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluate" agent skill from https://github.com/Q00/ouroboros/tree/main/skills/evaluate into .github/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Q00/ouroboros --skill evaluate -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Q00/ouroboros evaluate --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/evaluate .opencode/skills/evaluate && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluate" agent skill from https://github.com/Q00/ouroboros/tree/main/skills/evaluate into .opencode/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluateScores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote.
Stage 1, mechanical verification, costs nothing and runs lint, build, tests, static analysis and coverage; it is the only stage that can grant approval, needing at least one configured check to run and all of them to pass. Stage 2, semantic evaluation, scores acceptance-criteria compliance, goal alignment and drift and explains its reasoning, but can only withhold approval, never grant it. Stage 3, multi-model consensus, is optional and fires only on uncertainty or a manual request; a rejection withholds approval while an approval cannot grant it or lift a Stage 2 block. Without an executed Stage 1 check the result is always unverified.
Invoking the skill starts by loading the Ouroboros MCP tools, which the skill says are commonly registered as deferred tools that must be discovered before they can be called, and it insists on running that discovery step even if the tool does not already appear in the current tool list, since an empty discovery result for an already-exposed tool is expected rather than a failure. Only after discovery still finds nothing does the skill fall back to treating the tool as genuinely absent. A deferred-schema guard in the instructions addresses an invalid-parameters failure that can occur after a fresh conversation turn.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0df5b98. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ouroboros Three-Stage Evaluate loads about 2.2k tokens when it runs. Until then it costs about 17 tokens; SKILL.md has 1,043 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Q00/ouroboros at commit 0df5b98, republished under its MIT licence (© Q00). 1,043 words, ~2,216 tokens.
.claude/skills/evaluate/SKILL.md (or your agent's skills folder).Evaluate an execution session using the three-stage verification pipeline.
/ouroboros:evaluate <session_id> [artifact]Trigger keywords: "evaluate this", "3-stage check"
The evaluation pipeline runs three progressive stages:
Stage 1: Mechanical Verification ($0 cost)
Stage 2: Semantic Evaluation (Standard tier, advisory)
Stage 3: Multi-Model Consensus (Frontier tier, optional, advisory)
Without an executed Stage 1 check the outcome is acceptance_state: unverified
(not approved), with the model review attached as feedback.
When the user invokes this skill:
The Ouroboros MCP tools are often registered as deferred tools that must be explicitly loaded before use. You MUST perform this step before proceeding.
tool discovery query: "+ouroboros evaluate"mcp__plugin_ouroboros_ouroboros__ouroboros_start_evaluate (with a plugin prefix). After runtime tool discovery returns, the tool becomes callable.IMPORTANT: Do NOT skip this step. Do NOT assume MCP tools are unavailable just because they don't appear in your immediate tool list. They are almost always available as deferred tools that need to be loaded first.
CRITICAL — deferred-schema guard (prevents "Invalid tool parameters"):
This skill can call ouroboros_start_evaluate after a fresh turn. A deferred tool's
schema loaded on one turn is NOT guaranteed to still be loaded on the next. If
you call it while its schema is not loaded in the current turn, the runtime
rejects the call with "Invalid tool parameters" before it reaches the server.
Therefore: immediately before EVERY ouroboros_start_evaluate call in this skill,
re-run tool discovery query: "+ouroboros evaluate" (idempotent — a no-op when
already loaded). If the load returns no matching tool (and the tool is not already callable — an empty load for an already-exposed tool is an expected no-op, not absence), switch to the documented
fallback instead of retrying the failing call.
Determine what to evaluate:
session_id provided: Use it directlyGather the artifact to evaluate:
2.5. Acting verification — reproduce and OBSERVE (do not skip for behaviour-bearing work):
Stage 1 already runs mechanical checks (build/test). Go further when the runtime
exposes acting tools — computer-use / browser, Bash/shell, file reads: don't
just reason over the diff, run the result and observe the real effect (the
command's output, the endpoint's response, the rendered UI via a screenshot). Do
it via a dedicated verification sub-agent to keep the main session lean — or
inline in the main session where the runtime restricts sub-agent spawning (the
observation is what matters; the delegation is only an optimization). Probe the
acceptance criteria against the ACTUAL observable behaviour and the adversarial
classes (misleading_output, hung_command, stale_state, dirty_worktree, …).
Feed the captured evidence (commands, outputs, artifact paths) into the evaluate
call as part of the artifact. If acting tools are unavailable, note that
behaviour was not observed and evaluate on the text alone.
Call the background ouroboros_start_evaluate MCP tool so rejected verdicts
can continue through the configured Ralph convergence chain:
Tool: ouroboros_start_evaluate
Arguments:
session_id: <session ID>
artifact: <the code/output to evaluate, plus observed-behaviour evidence from 2.5>
seed_content: <original seed YAML, if available>
acceptance_criterion: <specific AC to check, optional>
artifact_type: "code" (or "docs", "config")
working_dir: <absolute project root, recommended>
trigger_consensus: false (true if user requests Stage 3)
auto_evolve: <optional override; omit to use execution.auto_evolve>working_dir controls both Stage 1 command execution and Stage 2 source-file visibility. Pass the absolute project root whenever available; if omitted, the MCP handler falls back to the registered brownfield default, seed project metadata, then the MCP server cwd.
Observe the returned evaluation job. If its terminal result contains
chained_ralph_job_id, follow that Ralph job to terminal before presenting
the convergence outcome. A missing Seed produces
chained_ralph_skipped: seed_unavailable; preserve the rejected verdict and
explain that automatic continuation was safely skipped. In OpenCode plugin
mode, auto_evolve=true intentionally returns this pollable parent-owned job;
with automatic evolution disabled, the plugin child remains the terminal
surface and job_id is None.
Present results clearly:
◆ Evaluation approved → next: accept, or ooo evolve to iteratively refinecode_changes_detected: true): ◆ Current state → next: Fix the build/test failures above, then ooo evaluate — or ooo ralph for automated fix loopcode_changes_detected: false): ◆ Current state → next: Run ooo run first to produce code, then ooo evaluate◆ Current state → next: ooo run to re-execute with fixes — or ooo evolve for iterative refinement◆ Current state → next: ooo interview to re-examine requirements — or ooo unstuck to challenge assumptionsacceptance_state: unverified, no executed check): ◆ Current state → next: add executable checks to .ouroboros/mechanical.toml (or run ouroboros detect), then ooo evaluate; the semantic review above is feedback, not a verdictIf the MCP server is not available, use the ouroboros:evaluator agent to perform a prompt-based evaluation:
ouroboros:evaluator agentUser: /ouroboros:evaluate sess-abc-123
Evaluation Results
============================================================
Final Approval: APPROVED
Highest Stage Completed: 2
Stage 1: Mechanical Verification
[PASS] lint: No issues found
[PASS] build: Build successful
[PASS] test: 12/12 tests passing
Stage 2: Semantic Evaluation
Score: 0.85
AC Compliance: YES
Goal Alignment: 0.90
Drift Score: 0.08
◆ Evaluation approved → next: accept, or `ooo evolve` to iteratively refineYour final response MUST end with exactly one breadcrumb footer line:
◆ <current state> → next: <recommended action>Derive <current state> from live session state via ouroboros_session_status when that MCP projection is available; otherwise derive it from this skill's actual outcome. Never use a linear Step N of M footer because Ouroboros is an evolutionary loop. When the next action is genuinely a choice, list 2-3 honest options in the next: clause. The breadcrumb line must be the last line of the response.
© Q00, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/evaluate of Q00/ouroboros.
Open the folder on GitHubat commit 0df5b98
Ouroboros Three-Stage Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ouroboros Three-Stage Evaluate this skillQ00/ouroboros | 6.2k | — | ~2.2k | Automated safety check: Pass | MIT | |
| MCP Server Builderanthropics/skills | 180k | 62 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Autocontext for Hermesgreyhaven-ai/autocontext | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Skill ForgeAgriciDaniel/skill-forge | 177 | — | ~1.9k | Automated safety check: Notes | MIT | |
| Eval Answermalloydata/publisher | 116 | — | ~4.3k | Automated safety check: Pass | MIT | |
| Waza Interactivemicrosoft/waza | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
greyhaven-ai/autocontext
Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.
AgriciDaniel/skill-forge
Ultimate Claude Code skill creator and architect. An agent skill from AgriciDaniel/skill-forge.
malloydata/publisher
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
microsoft/waza
Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.
bgauryy/octocode
Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.
Q00/ouroboros
Triages and works through GitHub issues and pull requests in the Q00/ouroboros repo as a maintainer, within a stated review boundary and clear limits on what it may change.
Q00/ouroboros
Runs a guided product-manager interview that classifies each question automatically and produces a Product Requirements Document.
Q00/ouroboros
Scans a directory for existing git repositories and worktrees, then registers and manages which ones serve as default context during interviews.
Q00/ouroboros
Starts, monitors or rewinds an evolutionary development loop that refines an ontology and acceptance criteria generation by generation until it converges, using the Ouroboros MCP tools.
Q00/ouroboros
Opens or drives the Ouroboros settings GUI, picking a browser, TUI or chat-based approach depending on whether the user can reach a browser window.
Q00/ouroboros
Reference guide to the Ouroboros commands and agents, covering interviews, seed specs, evaluation, lateral-thinking personas and the evolutionary loop.
Works with
Categories
Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote. Stage 1, mechanical verification, costs nothing and runs lint, build, tests, static analysis and coverage; it is the only stage that can grant approval, needing at least one configured check to run and all of them to pass. Stage 2, semantic evaluation, scores acceptance-criteria compliance, goal alignment and drift and explains its reasoning, but can only withhold approval, never grant it.
Ouroboros Three-Stage Evaluate fits situations like: checking whether a finished execution session passed lint, build and tests; getting an advisory assessment of how well work matches its acceptance criteria; escalating an uncertain result to a multi-model consensus vote; evaluating a session with the three-stage pipeline by its session id.
Run `npx skills add Q00/ouroboros --skill evaluate -a claude-code`. Or copy the skill folder (skills/evaluate in Q00/ouroboros) into .claude/skills/evaluate in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Q00/ouroboros --skill evaluate -a codex`. Or copy the skill folder (skills/evaluate in Q00/ouroboros) into .agents/skills/evaluate in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Q00/ouroboros --skill evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate, .gemini/skills/evaluate, .github/skills/evaluate and .opencode/skills/evaluate in your project.
SKILL.md names no scripts, command-line tools or credentials: Ouroboros Three-Stage Evaluate is instructions for the agent only. Our summary lists: The Ouroboros MCP server.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ouroboros Three-Stage Evaluate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ouroboros Three-Stage Evaluate: MCP Server Builder (anthropics/skills, 180k stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars), Skill Forge (AgriciDaniel/skill-forge, 177 stars) and Eval Answer (malloydata/publisher, 116 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Q00 (a GitHub user) maintains it in Q00/ouroboros, which has 6,189 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 6, 2026.
Source: Q00/ouroboros on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.