Agent skill

Ouroboros Evolve Loop

by Q00 in Q00/ouroboros

Starts, monitors or rewinds an evolutionary development loop that refines an ontology and acceptance criteria generation by generation until it converges, using the Ouroboros MCP tools.

MITAuto-check passedAgent Workflows

Install Ouroboros Evolve Loop

skills CLI
$ npx skills add Q00/ouroboros --skill evolve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Q00/ouroboros evolve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evolve .claude/skills/evolve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evolve
GitHub stars
6.2k
Token cost
~3.2k tokens
SKILL.md length
1,564 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Starts, monitors or rewinds an evolutionary development loop that refines an ontology and acceptance criteria generation by generation until it converges, using the Ouroboros MCP tools.

  • Works in 3 steps: Use the active runtime's tool-discovery… → The tools will typically be named with… → If the tools are callable — already…
  • Starting an iterative build loop that refines requirements across generations
  • SKILL.md covers Description, Flow, Usage and Instructions, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The loop works in generations. The first seeds an ontology, executes every node, judges the results and applies a gate. The second asks about the failed nodes, reflects on the active ones, executes only those and gates everything. Later generations repeat with a smaller active set, and nodes that already passed are frozen and re-verified instead of regenerated. You start it with `ooo evolve` and a goal, use a fast `--no-execute` mode that builds the ontology only, check a lineage with `--status`, or go back with `--rewind` to a chosen generation.

Before choosing between the MCP path and the fallback path, the agent must load the Ouroboros tools through the runtime's tool discovery, since they are often deferred and not listed at first. Only if they are genuinely absent does it fall back. The skill also describes a benchmark control run, which calls the evolve step with a control flag on an existing lineage from a clean Git directory and a separate clean worktree, and notes that normal calls never start a control arm.

When your agent uses it

  • Starting an iterative build loop that refines requirements across generations
  • Checking the status of an existing evolution lineage
  • Rewinding to an earlier generation after a bad round

Example prompts

  • “Run ooo evolve on "build a task management CLI".”
  • “Evolve the idea in fast mode with no execution so I can review the ontology first.”
  • “Show the status of the lineage we started yesterday and rewind it to generation two.”

Requirements

  • The Ouroboros MCP server and its tools available in the runtime

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Use the active runtime's tool-discovery capability to find and load the evolve MCP tools
  2. The tools will typically be named with prefix mcpplugin_ouroboros_ouroboros (e.g., ouroboros_evolve_step, ouroboros_interview…
  3. If the tools are callable — already exposed, or loaded by discovery — proceed to Path A. An empty discovery result for already-exposed…

What it can do on your machine

Read from SKILL.md and the folder at commit 0df5b98. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ouroboros Evolve Loop loads about 3.2k tokens when it runs. Until then it costs about 14 tokens; SKILL.md has 1,564 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~14
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Q00/ouroboros at commit 0df5b98, republished under its MIT licence (© Q00). 1,564 words, ~3,190 tokens.

Download SKILL.mdSave it as .claude/skills/evolve/SKILL.md (or your agent's skills folder).
name
evolve
description
Start or monitor an evolutionary development loop

ooo evolve - Evolutionary Loop

Description

Start, monitor, or rewind an evolutionary development loop. The loop iteratively refines the ontology and acceptance criteria across generations until convergence.

Flow

Gen 1: Seed(O₁) → Execute(all nodes) → Judge → Gate
Gen 2: Wonder(failed nodes) → Reflect(active nodes only) → Execute(active) → Gate(all)
Gen 3: Repeat with a smaller active set; frozen PASS nodes are reverified, not regenerated
...until the outcome gate passes, the working set stagnates, ontology converges,
or max 30 generations is reached

Usage

Start a new evolutionary loop
ooo evolve "build a task management CLI"
Fast mode (ontology-only, no execution)
ooo evolve "build a task management CLI" --no-execute
Check lineage status
ooo evolve --status <lineage_id>
Record an isolated full-graph benchmark control

Call ouroboros_evolve_step for an existing Gen 2+ lineage with benchmark_control: true, execute: true, and an explicit clean Git project_dir. Use a distinct clean worktree at the treatment commit. Normal evolve calls keep benchmark_control false and never launch a control arm.

Rewind to a previous generation
ooo evolve --rewind <lineage_id> <generation_number>

Instructions

Load MCP Tools (Required before Path A/B decision)

The Ouroboros MCP tools are often registered as deferred tools that must be explicitly loaded before use. You MUST perform this step before deciding between Path A and Path B.

  1. Use the active runtime's tool-discovery capability to find and load the evolve MCP tools:
    tool discovery query: "+ouroboros evolve"
  2. The tools will typically be named with prefix mcp__plugin_ouroboros_ouroboros__ (e.g., ouroboros_evolve_step, ouroboros_interview, ouroboros_generate_seed). After runtime tool discovery returns, the tools become callable.
  3. If the tools are callable — already exposed, or loaded by discovery — proceed to Path A. An empty discovery result for already-exposed tools is expected, not a failure. Proceed to Path B only if they are genuinely absent (no Ouroboros MCP server).

IMPORTANT: Do NOT skip this step. Do NOT assume MCP tools are unavailable just because they don't appear in your immediate tool list. They are almost always available as deferred tools that need to be loaded first.

CRITICAL — deferred-schema guard (prevents "Invalid tool parameters"): This skill makes ouroboros_* MCP calls across multiple turns, and each turn runs in a fresh tool context. A deferred tool's schema loaded on one turn is NOT guaranteed to still be loaded on the next. If you call any ouroboros_* MCP tool while its schema is not loaded in the current turn, the runtime rejects the call with "Invalid tool parameters" before it ever reaches the server. Therefore: immediately before EVERY ouroboros_* MCP call in this skill, re-run the tool-discovery load query for the specific MCP tool or documented tool family you are about to call. Use "+ouroboros evolve" for ouroboros_evolve_step, ouroboros_lineage_status, and the evolve flow's documented tool family; use "+ouroboros interview" before ouroboros_interview, "+ouroboros seed" before ouroboros_generate_seed, and "+ouroboros lateral" before ouroboros_lateral_think. If a load returns no matching tool (and the tool is not already callable — an empty load for an already-exposed tool is an expected no-op, not absence), switch to the documented fallback / Path B instead of retrying the failing call.

Path A: MCP Available (loaded via runtime tool discovery above)

Starting a new evolutionary loop:

  1. Parse the user's input as initial_context
  2. Run the interview: call ouroboros_interview with initial_context
  3. Complete the interview (3+ rounds until ambiguity ≤ 0.2)
  4. Generate seed: call ouroboros_generate_seed with the session_id
  5. Call ouroboros_evolve_step with:
    • lineage_id: new unique ID (e.g., lin_<seed_id>)
    • seed_content: the generated seed YAML
    • execute: true (default) for full Execute→Evaluate pipeline, false for fast ontology-only evolution (no seed execution)
    • benchmark_control: false (default). Set true only for a deliberate Gen 2+ full-graph control in an explicit clean Git project/worktree.
  6. Check the action in the response:
    • continue → Inspect active_ac_indices, then call ouroboros_evolve_step again with just lineage_id. Only active failed/reopened nodes evolve; frozen PASS nodes stay immutable and are boundary-reverified.
    • ontology_stable → This is not success. Call ouroboros_evolve_step again for the same lineage_id with execute: true so the stable Seed goes through Execute→Evaluate. Do not call standalone evaluate or report convergence before that step returns converged.
    • converged → Evolution complete! Display final ontology
    • stagnated → Ontology unchanged for 3+ gens. Consider ouroboros_lateral_think
    • exhausted → Max 30 generations reached. Display best result
    • failed → Check error, possibly retry. If the error reports an expired lineage owner, first confirm that the prior owner process is dead, then make one explicit recovery call with recover_expired_claim: true. Never set this flag for a merely slow or still-running owner.
  7. Repeat step 6 while action is continue. Treat ontology_stable as the explicit transition above, then process the execute: true response using step 6. Stop only on converged, stagnated, exhausted, or failed.
  8. When the loop terminates, display a result summary with next step:
    • converged: ◆ Current state → next: Ontology converged! Run ooo evaluate for formal verification
    • ontology_stable: ◆ Current state → next: Run the same lineage with execute=true for Execute→Evaluate; this is not verified convergence yet
    • stagnated: ◆ Current state → next: ooo unstuck to break through, then ooo evolve --status <lineage_id> to resume
    • exhausted: ◆ Current state → next: ooo evaluate to check best result — or ooo unstuck to try a new approach
    • failed: ◆ Current state → next: Check the error above. ooo status to inspect session, or ooo unstuck if blocked

Checking status:

  1. Call ouroboros_lineage_status with the lineage_id
  2. Display: generation count, ontology evolution, convergence progress

Rewinding:

  1. Call ouroboros_evolve_step with:
    • lineage_id: the lineage to continue from a rewind point
    • seed_content: the seed YAML from the target generation (Future: dedicated ouroboros_evolve_rewind tool)
Path B: Plugin-only (no MCP tools available)

If MCP tools are not available, explain the evolutionary loop concept and suggest installing the Ouroboros MCP server. See Getting Started for install options, then run:

ouroboros mcp serve --runtime claude-cli

Then add to your runtime's MCP configuration (e.g., ~/.claude/mcp.json for Claude Code).

Show full SKILL.md (707 more words)Show less

Key Concepts

  • Wonder: "What do we still not know?" - examines evaluation results to identify ontological gaps and hidden assumptions
  • Reflect: "How should the ontology evolve?" - proposes specific mutations to fields, acceptance criteria, and constraints
  • Convergence: Verified success requires the independently evaluated outcome to pass. In ontology-only mode, similarity ≥ 0.95 returns the non-success ontology_stable handoff and the same lineage must run once with execute: true. Judge-score plateau and the 30-generation cap are also non-success stops.
  • Focused evolution: Gen 1 establishes the baseline. Gen 2+ derives an active node set from failed/regressed verifier results. Only those nodes are open to Wonder, Reflect, and execution; PASS nodes are frozen. The active output should shrink toward zero rather than feeding the full graph back.
  • Frugality proof: A smaller active-node count is a working-set observation, not a savings claim. Gen 2+ records measured total-generation runtime tokens (Wonder, all Reflect attempts, validator/evaluator providers, executor AC attempts, dependency analysis, decomposition policy/attestation/repair, coordinator review, and shadow replay when enabled), wall time, calls, retries, and quality evidence. A PASS requires a paired full-graph control and focused treatment from distinct clean worktrees at the same Git commit, at least 10% fewer total-generation runtime tokens, and no final-gate, evaluation score/stage, drift, reward-hacking, per-AC verdict/score, lineage regression, or TraceGuard degradation. The comparison key covers the complete semantic Seed and previous evaluation. Every expected primary attempt needs exactly one runtime-token receipt and TraceGuard verdict bound to the exact session, primary dispatch, and root identity; every active root must be dispatched; every other generation provider call needs runtime usage; normalized backend/model/tier/mode/effort/permission and provider request configurations must be complete and match the same root AC or auxiliary phase role and preserve per-unit call multiplicity across the paired arms. Shared-unit sequences must be exactly equal; only whole control-only units may be removed. Blank/unknown phase roles are incomplete evidence. Full-graph controls must execute every Seed root; treatment active/frozen sets must exactly partition the Seed. Auxiliary tools, system-prompt identity, request kwargs, and fresh/scoped session mode are configuration-bound. Completion profiles are resolved once into a sealed dispatch config; only registered exact-key, single-attempt, secret-safe adapter attestations whose in-memory endpoint/credential authority reaches the actual call boundary unchanged are proof-eligible. Global or post-attestation mutable routing is ineligible. Malformed attestations invalidate the proof, not evolution. True resumed contexts without semantic identity and unknown/conflicting effective models or unreconciled usage counters are incomplete. Every current Seed AC needs exactly one final verdict against the same final semantic Seed contract. Missing, duplicate, malformed, partial, or opaque evidence is insufficient_data. Evidence reads and provider-call capture are capped; oversized or non-finite individual/subtotal/combined usage, cap overflow, and Wonder-only early stops also produce non-PASS durable receipts. Evolve never launches the control automatically because doing so would consume the production savings being measured.
  • Judge/Gate split: Evaluation records score and evidence; deterministic convergence logic decides accept/continue/stagnate. A passing Gen 1 may end immediately—minimum-generation churn is not required after the gate passes.
  • No-drift stop: If rejected evaluation scores move by less than 0.01 for the 3-generation window, stop as stagnated and hand off to ooo unstuck instead of spending the 30-generation cap.
  • Rewind: Each generation is a snapshot. You can rewind to any generation and branch evolution from there
  • evolve_step: Runs exactly ONE generation per call. Designed for Ralph integration — state is fully reconstructed from events between calls
  • execute flag: true (default) runs full Execute→Evaluate each generation. false skips execution for fast ontology exploration. Previous generation's execution output is fed into Wonder/Reflect for informed evolution. An ontology_stable response from execute: false must be followed by the same lineage with execute: true; ontology stability alone is never convergence.
  • QA verdict: Each generation's response includes a QA Verdict section (when execute=true and skip_qa is not set). Use the QA score to track quality progression across generations. Pass skip_qa: true to disable

Your final response MUST end with exactly one breadcrumb footer line:

◆ <current state> → next: <recommended action>

Derive <current state> from live session state via ouroboros_session_status when that MCP projection is available; otherwise derive it from this skill's actual outcome. Never use a linear Step N of M footer because Ouroboros is an evolutionary loop. When the next action is genuinely a choice, list 2-3 honest options in the next: clause. The breadcrumb line must be the last line of the response.

© Q00, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/evolve of Q00/ouroboros.

Open the folder on GitHubat commit 0df5b98

Compare with similar skills

Ouroboros Evolve Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ouroboros Evolve Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ouroboros Evolve Loop this skillQ00/ouroboros6.2k—~3.2kAutomated safety check: PassMIT
Hermes Mission ControlTh0rgal/sandboxed.sh515—~8.6kAutomated safety check: PassNone
Agent Native Architecturesandgardenhq/sgai137—~2kAutomated safety check: PassCustom licence
Peer Review Loophashgraph-online/awesome-codex-plugins1.2k—~2.3kAutomated safety check: PassApache-2.0
Improving MCP ToolsPostHog/posthog-foss721—~1.5kAutomated safety check: PassMIT
Strandsstrands-agents/harness-sdk8.7k—~1kAutomated safety check: PassApache-2.0

Similar skills

  • Hermes Mission Control

    Th0rgal/sandboxed.sh

    Teaches Hermes to monitor and steer long-running sandboxed.sh missions: spot where a model is stuck, switch backends or models between turns, and send targeted hints.

    515 GitHub stars~8.6k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Agent Native Architecture

    sandgardenhq/sgai

    Build AI agents using prompt-native architecture where features are defined in prompts, not code.

    137 GitHub stars~2k tokensUpdated 16 days ago
    Agent WorkflowsAuto-check passed
  • Peer Review Loop

    hashgraph-online/awesome-codex-plugins

    Peer Review Ralph Loop — combines Cavekit kits with a Ralph Loop and true cross-model peer review using Codex (OpenAI).

    1.2k GitHub stars~2.3k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Improving MCP Tools

    PostHog/posthog-foss

    Official

    Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…

    721 GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Strands

    strands-agents/harness-sdk

    Build, extend, evaluate, or migrate applications with Strands Agents in Python or TypeScript.

    8.7k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Peer Review

    hashgraph-online/awesome-codex-plugins

    Patterns for using a second AI agent or model to challenge the primary builder agent's work.

    1.2k GitHub stars~5.8k tokensUpdated today
    Research & ScienceAuto-check passed

More from Q00/ouroboros

All 23 skills in this repo
  • Triages and works through GitHub issues and pull requests in the Q00/ouroboros repo as a maintainer, within a stated review boundary and clear limits on what it may change.

    6.2k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Runs a guided product-manager interview that classifies each question automatically and produces a Product Requirements Document.

    6.2k GitHub stars~5.7k tokensUpdated today
    Auto-check passed
  • Scans a directory for existing git repositories and worktrees, then registers and manages which ones serve as default context during interviews.

    6.2k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote.

    6.2k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Opens or drives the Ouroboros settings GUI, picking a browser, TUI or chat-based approach depending on whether the user can reach a browser window.

    6.2k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Reference guide to the Ouroboros commands and agents, covering interviews, seed specs, evaluation, lateral-thinking personas and the evolutionary loop.

    6.2k GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Ouroboros Evolve Loop

What does Ouroboros Evolve Loop do?

Starts, monitors or rewinds an evolutionary development loop that refines an ontology and acceptance criteria generation by generation until it converges, using the Ouroboros MCP tools. The loop works in generations. The first seeds an ontology, executes every node, judges the results and applies a gate.

When should I use Ouroboros Evolve Loop?

Ouroboros Evolve Loop fits situations like: starting an iterative build loop that refines requirements across generations; checking the status of an existing evolution lineage; rewinding to an earlier generation after a bad round.

How do I install Ouroboros Evolve Loop in Claude Code?

Run `npx skills add Q00/ouroboros --skill evolve -a claude-code`. Or copy the skill folder (skills/evolve in Q00/ouroboros) into .claude/skills/evolve in your project. Claude Code loads it when a task matches its description.

How do I install Ouroboros Evolve Loop in Codex?

Run `npx skills add Q00/ouroboros --skill evolve -a codex`. Or copy the skill folder (skills/evolve in Q00/ouroboros) into .agents/skills/evolve in your project. Codex loads it when a task matches its description.

Can I use Ouroboros Evolve Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Q00/ouroboros --skill evolve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evolve, .gemini/skills/evolve, .github/skills/evolve and .opencode/skills/evolve in your project.

What does Ouroboros Evolve Loop need to run?

SKILL.md names no scripts, command-line tools or credentials: Ouroboros Evolve Loop is instructions for the agent only. Our summary lists: The Ouroboros MCP server and its tools available in the runtime.

Does Ouroboros Evolve Loop access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ouroboros Evolve Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ouroboros Evolve Loop use?

Ouroboros Evolve Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ouroboros Evolve Loop use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ouroboros Evolve Loop?

Skills that share tags, products or a category with Ouroboros Evolve Loop: Hermes Mission Control (Th0rgal/sandboxed.sh, 515 stars), Agent Native Architecture (sandgardenhq/sgai, 137 stars), Peer Review Loop (hashgraph-online/awesome-codex-plugins, 1.2k stars) and Improving MCP Tools (PostHog/posthog-foss, 721 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ouroboros Evolve Loop?

Q00 (a GitHub user) maintains it in Q00/ouroboros, which has 6,189 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 6, 2026.

Source: Q00/ouroboros on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.