Agent skill

Ouroboros Three-Stage Evaluate

by Q00 in Q00/ouroboros

Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote.

MITAuto-check passedAgent Workflows

Install Ouroboros Three-Stage Evaluate

skills CLI
$ npx skills add Q00/ouroboros --skill evaluate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Q00/ouroboros evaluate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluate .claude/skills/evaluate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate
GitHub stars
6.2k
Token cost
~2.2k tokens
SKILL.md length
1,043 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote.

  • Works in 3 steps: Stage 1: Mechanical Verification ($0 cost) → Stage 2: Semantic Evaluation (Standard… → Stage 3: Multi-Model Consensus (Frontier…
  • Checking whether a finished execution session passed lint, build and tests
  • SKILL.md covers Usage, How It Works, Instructions and Fallback (No MCP Server), plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stage 1, mechanical verification, costs nothing and runs lint, build, tests, static analysis and coverage; it is the only stage that can grant approval, needing at least one configured check to run and all of them to pass. Stage 2, semantic evaluation, scores acceptance-criteria compliance, goal alignment and drift and explains its reasoning, but can only withhold approval, never grant it. Stage 3, multi-model consensus, is optional and fires only on uncertainty or a manual request; a rejection withholds approval while an approval cannot grant it or lift a Stage 2 block. Without an executed Stage 1 check the result is always unverified.

Invoking the skill starts by loading the Ouroboros MCP tools, which the skill says are commonly registered as deferred tools that must be discovered before they can be called, and it insists on running that discovery step even if the tool does not already appear in the current tool list, since an empty discovery result for an already-exposed tool is expected rather than a failure. Only after discovery still finds nothing does the skill fall back to treating the tool as genuinely absent. A deferred-schema guard in the instructions addresses an invalid-parameters failure that can occur after a fresh conversation turn.

When your agent uses it

  • Checking whether a finished execution session passed lint, build and tests
  • Getting an advisory assessment of how well work matches its acceptance criteria
  • Escalating an uncertain result to a multi-model consensus vote
  • Evaluating a session with the three-stage pipeline by its session id

Example prompts

  • “Evaluate session abc123 with the three-stage check.”
  • “Run the evaluation on the latest session and show me the Stage 2 feedback.”
  • “This result looks borderline, trigger the Stage 3 consensus vote.”

Requirements

  • The Ouroboros MCP server

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Stage 1: Mechanical Verification ($0 cost)
  2. Stage 2: Semantic Evaluation (Standard tier, advisory)
  3. Stage 3: Multi-Model Consensus (Frontier tier, optional, advisory)

What it can do on your machine

Read from SKILL.md and the folder at commit 0df5b98. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ouroboros Three-Stage Evaluate loads about 2.2k tokens when it runs. Until then it costs about 17 tokens; SKILL.md has 1,043 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~17
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Q00/ouroboros at commit 0df5b98, republished under its MIT licence (© Q00). 1,043 words, ~2,216 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate/SKILL.md (or your agent's skills folder).
name
evaluate
description
Evaluate execution with three-stage verification pipeline
aliases
eval

/ouroboros:evaluate

Evaluate an execution session using the three-stage verification pipeline.

Usage

/ouroboros:evaluate <session_id> [artifact]

Trigger keywords: "evaluate this", "3-stage check"

How It Works

The evaluation pipeline runs three progressive stages:

  1. Stage 1: Mechanical Verification ($0 cost)

    • Lint checks, build validation, test execution
    • Static analysis, coverage measurement
    • Fails fast if mechanical checks don't pass
    • The only stage that can grant approval: at least one configured check must run and all must pass
  2. Stage 2: Semantic Evaluation (Standard tier, advisory)

    • AC compliance assessment
    • Goal alignment scoring
    • Drift measurement
    • Reasoning explanation
    • Can withhold approval and supply feedback; cannot grant it
  3. Stage 3: Multi-Model Consensus (Frontier tier, optional, advisory)

    • Multiple models vote on approval
    • Only triggered by uncertainty or manual request
    • A rejection withholds approval; an approval cannot grant it or lift a Stage 2 block

Without an executed Stage 1 check the outcome is acceptance_state: unverified (not approved), with the model review attached as feedback.

Instructions

When the user invokes this skill:

Load MCP Tools (Required first)

The Ouroboros MCP tools are often registered as deferred tools that must be explicitly loaded before use. You MUST perform this step before proceeding.

  1. Use the active runtime's tool-discovery capability to find and load the evaluate MCP tools:
    tool discovery query: "+ouroboros evaluate"
  2. The tool will typically be named mcp__plugin_ouroboros_ouroboros__ouroboros_start_evaluate (with a plugin prefix). After runtime tool discovery returns, the tool becomes callable.
  3. If the tool is callable — already exposed, or loaded by discovery — proceed with the MCP-based evaluation below. An empty discovery result for an already-exposed tool is expected, not a failure. Skip to the Fallback section only if the tool is genuinely absent (no Ouroboros MCP server).

IMPORTANT: Do NOT skip this step. Do NOT assume MCP tools are unavailable just because they don't appear in your immediate tool list. They are almost always available as deferred tools that need to be loaded first.

CRITICAL — deferred-schema guard (prevents "Invalid tool parameters"): This skill can call ouroboros_start_evaluate after a fresh turn. A deferred tool's schema loaded on one turn is NOT guaranteed to still be loaded on the next. If you call it while its schema is not loaded in the current turn, the runtime rejects the call with "Invalid tool parameters" before it reaches the server. Therefore: immediately before EVERY ouroboros_start_evaluate call in this skill, re-run tool discovery query: "+ouroboros evaluate" (idempotent — a no-op when already loaded). If the load returns no matching tool (and the tool is not already callable — an empty load for an already-exposed tool is an expected no-op, not absence), switch to the documented fallback instead of retrying the failing call.

Evaluation Steps
  1. Determine what to evaluate:

    • If session_id provided: Use it directly
    • If no session_id: Check conversation for recent execution session IDs
  2. Gather the artifact to evaluate:

    • If user specifies a file: Read it with Read tool
    • If recent execution output exists in conversation: Use that
    • Ask user if unclear what to evaluate

2.5. Acting verification — reproduce and OBSERVE (do not skip for behaviour-bearing work): Stage 1 already runs mechanical checks (build/test). Go further when the runtime exposes acting tools — computer-use / browser, Bash/shell, file reads: don't just reason over the diff, run the result and observe the real effect (the command's output, the endpoint's response, the rendered UI via a screenshot). Do it via a dedicated verification sub-agent to keep the main session lean — or inline in the main session where the runtime restricts sub-agent spawning (the observation is what matters; the delegation is only an optimization). Probe the acceptance criteria against the ACTUAL observable behaviour and the adversarial classes (misleading_output, hung_command, stale_state, dirty_worktree, …). Feed the captured evidence (commands, outputs, artifact paths) into the evaluate call as part of the artifact. If acting tools are unavailable, note that behaviour was not observed and evaluate on the text alone.

  1. Call the background ouroboros_start_evaluate MCP tool so rejected verdicts can continue through the configured Ralph convergence chain:

    Tool: ouroboros_start_evaluate
    Arguments:
      session_id: <session ID>
      artifact: <the code/output to evaluate, plus observed-behaviour evidence from 2.5>
      seed_content: <original seed YAML, if available>
      acceptance_criterion: <specific AC to check, optional>
      artifact_type: "code"  (or "docs", "config")
      working_dir: <absolute project root, recommended>
      trigger_consensus: false  (true if user requests Stage 3)
      auto_evolve: <optional override; omit to use execution.auto_evolve>

    working_dir controls both Stage 1 command execution and Stage 2 source-file visibility. Pass the absolute project root whenever available; if omitted, the MCP handler falls back to the registered brownfield default, seed project metadata, then the MCP server cwd.

  2. Observe the returned evaluation job. If its terminal result contains chained_ralph_job_id, follow that Ralph job to terminal before presenting the convergence outcome. A missing Seed produces chained_ralph_skipped: seed_unavailable; preserve the rejected verdict and explain that automatic continuation was safely skipped. In OpenCode plugin mode, auto_evolve=true intentionally returns this pollable parent-owned job; with automatic evolution disabled, the plugin child remains the terminal surface and job_id is None.

  3. Present results clearly:

    • Show each stage's pass/fail status
    • Highlight the final approval decision
    • If rejected, explain the failure reason
    • Suggest fixes if evaluation fails
    • Always end with a state breadcrumb based on the outcome:
      • APPROVED: ◆ Evaluation approved → next: accept, or ooo evolve to iteratively refine
      • REJECTED at Stage 1 (mechanical, code_changes_detected: true): ◆ Current state → next: Fix the build/test failures above, then ooo evaluate — or ooo ralph for automated fix loop
      • REJECTED at Stage 1 (mechanical, code_changes_detected: false): ◆ Current state → next: Run ooo run first to produce code, then ooo evaluate
      • REJECTED at Stage 2 (semantic): ◆ Current state → next: ooo run to re-execute with fixes — or ooo evolve for iterative refinement
      • REJECTED at Stage 3 (consensus): ◆ Current state → next: ooo interview to re-examine requirements — or ooo unstuck to challenge assumptions
      • NOT APPROVED (unverified) (acceptance_state: unverified, no executed check): ◆ Current state → next: add executable checks to .ouroboros/mechanical.toml (or run ouroboros detect), then ooo evaluate; the semantic review above is feedback, not a verdict
Show full SKILL.md (128 more words)Show less

Fallback (No MCP Server)

If the MCP server is not available, use the ouroboros:evaluator agent to perform a prompt-based evaluation:

  1. Delegate to ouroboros:evaluator agent
  2. The agent performs qualitative evaluation based on the seed spec
  3. Results are advisory (no numerical scoring without Python core)

Example

User: /ouroboros:evaluate sess-abc-123

Evaluation Results
============================================================
Final Approval: APPROVED
Highest Stage Completed: 2

Stage 1: Mechanical Verification
  [PASS] lint: No issues found
  [PASS] build: Build successful
  [PASS] test: 12/12 tests passing

Stage 2: Semantic Evaluation
  Score: 0.85
  AC Compliance: YES
  Goal Alignment: 0.90
  Drift Score: 0.08

◆ Evaluation approved → next: accept, or `ooo evolve` to iteratively refine

Your final response MUST end with exactly one breadcrumb footer line:

◆ <current state> → next: <recommended action>

Derive <current state> from live session state via ouroboros_session_status when that MCP projection is available; otherwise derive it from this skill's actual outcome. Never use a linear Step N of M footer because Ouroboros is an evolutionary loop. When the next action is genuinely a choice, list 2-3 honest options in the next: clause. The breadcrumb line must be the last line of the response.

© Q00, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/evaluate of Q00/ouroboros.

Open the folder on GitHubat commit 0df5b98

Compare with similar skills

Ouroboros Three-Stage Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ouroboros Three-Stage Evaluate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ouroboros Three-Stage Evaluate this skillQ00/ouroboros6.2k—~2.2kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.0
Autocontext for Hermesgreyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0
Skill ForgeAgriciDaniel/skill-forge177—~1.9kAutomated safety check: NotesMIT
Eval Answermalloydata/publisher116—~4.3kAutomated safety check: PassMIT
Waza Interactivemicrosoft/waza1.4k—~1.3kAutomated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Skill Forge

    AgriciDaniel/skill-forge

    Ultimate Claude Code skill creator and architect. An agent skill from AgriciDaniel/skill-forge.

    177 GitHub stars~1.9k tokensUpdated 6 mo ago
    Agent WorkflowsAuto-check: notes
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Waza Interactive

    microsoft/waza

    Official

    Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

    1.4k GitHub stars~1.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Octocode Graph Eval Loop

    bgauryy/octocode

    Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

    946 GitHub stars~1.6k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed

More from Q00/ouroboros

All 23 skills in this repo
  • Triages and works through GitHub issues and pull requests in the Q00/ouroboros repo as a maintainer, within a stated review boundary and clear limits on what it may change.

    6.2k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Runs a guided product-manager interview that classifies each question automatically and produces a Product Requirements Document.

    6.2k GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Scans a directory for existing git repositories and worktrees, then registers and manages which ones serve as default context during interviews.

    6.2k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Starts, monitors or rewinds an evolutionary development loop that refines an ontology and acceptance criteria generation by generation until it converges, using the Ouroboros MCP tools.

    6.2k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Opens or drives the Ouroboros settings GUI, picking a browser, TUI or chat-based approach depending on whether the user can reach a browser window.

    6.2k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Reference guide to the Ouroboros commands and agents, covering interviews, seed specs, evaluation, lateral-thinking personas and the evolutionary loop.

    6.2k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Ouroboros Three-Stage Evaluate

What does Ouroboros Three-Stage Evaluate do?

Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote. Stage 1, mechanical verification, costs nothing and runs lint, build, tests, static analysis and coverage; it is the only stage that can grant approval, needing at least one configured check to run and all of them to pass. Stage 2, semantic evaluation, scores acceptance-criteria compliance, goal alignment and drift and explains its reasoning, but can only withhold approval, never grant it.

When should I use Ouroboros Three-Stage Evaluate?

Ouroboros Three-Stage Evaluate fits situations like: checking whether a finished execution session passed lint, build and tests; getting an advisory assessment of how well work matches its acceptance criteria; escalating an uncertain result to a multi-model consensus vote; evaluating a session with the three-stage pipeline by its session id.

How do I install Ouroboros Three-Stage Evaluate in Claude Code?

Run `npx skills add Q00/ouroboros --skill evaluate -a claude-code`. Or copy the skill folder (skills/evaluate in Q00/ouroboros) into .claude/skills/evaluate in your project. Claude Code loads it when a task matches its description.

How do I install Ouroboros Three-Stage Evaluate in Codex?

Run `npx skills add Q00/ouroboros --skill evaluate -a codex`. Or copy the skill folder (skills/evaluate in Q00/ouroboros) into .agents/skills/evaluate in your project. Codex loads it when a task matches its description.

Can I use Ouroboros Three-Stage Evaluate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Q00/ouroboros --skill evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate, .gemini/skills/evaluate, .github/skills/evaluate and .opencode/skills/evaluate in your project.

What does Ouroboros Three-Stage Evaluate need to run?

SKILL.md names no scripts, command-line tools or credentials: Ouroboros Three-Stage Evaluate is instructions for the agent only. Our summary lists: The Ouroboros MCP server.

Does Ouroboros Three-Stage Evaluate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ouroboros Three-Stage Evaluate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ouroboros Three-Stage Evaluate use?

Ouroboros Three-Stage Evaluate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ouroboros Three-Stage Evaluate use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ouroboros Three-Stage Evaluate?

Skills that share tags, products or a category with Ouroboros Three-Stage Evaluate: MCP Server Builder (anthropics/skills, 180k stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars), Skill Forge (AgriciDaniel/skill-forge, 177 stars) and Eval Answer (malloydata/publisher, 116 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ouroboros Three-Stage Evaluate?

Q00 (a GitHub user) maintains it in Q00/ouroboros, which has 6,189 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 6, 2026.

Source: Q00/ouroboros on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.