Agent skill

Evolve

by SethGammon in SethGammon/Citadel

Research-driven multi-cycle improvement director. An agent skill from SethGammon/Citadel.

MITAuto-check passedEducation

Install Evolve

skills CLI
$ npx skills add SethGammon/Citadel --skill evolve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SethGammon/Citadel evolve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evolve .claude/skills/evolve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evolve
GitHub stars
922
Token cost
~3.7k tokens
SKILL.md length
1,579 words
Files
3
Skills in repo
48
Repo updated
First seen
Licence
MIT

At a glance

Research-driven multi-cycle improvement director. An agent skill from SethGammon/Citadel.

  • Works in 8 steps: Survey → Hypothesize → Scout → …
  • Education work in your project
  • SKILL.md covers Orientation, Invocation, Campaign Artifacts and Director Cycle Protocol, plus 5 more sections
  • Calls node and git

What it does

Evolve is an agent skill from SethGammon/Citadel. Research-driven multi-cycle improvement director. Forms causal hypotheses about why scores are low, validates them with scout agents before attacking, dispatches axis-parallel fleet attacks, extracts transferable patterns, and runs indefinitely within a budget envelope. Accumulates a persistent belief model and pattern library across sessions.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `__benchmarks__/happy-path-status.md` and `__benchmarks__/no-rubric.md`).

It sits in Education. The repository describes itself as: The operating layer for Claude Code + OpenAI Codex: persistent project memory, intent routing, safety hooks, cost telemetry, and parallel agent fleets. The licence is MIT.

When your agent uses it

  • Education work in your project

Example prompts

  • “/evolve”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Survey
  2. Hypothesize
  3. Scout
  4. Prioritize
  5. Fleet Attack
  6. Synthesize
  7. Cross-Pollinate
  8. Loop or Halt

What it can do on your machine

Read from SKILL.md and the folder at commit e41ff1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evolve loads about 3.7k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 1,579 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from SethGammon/Citadel at commit e41ff1d, republished under its MIT licence (© SethGammon). 1,579 words, ~3,689 tokens.

Download SKILL.mdSave it as .claude/skills/evolve/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evolve
description
Research-driven multi-cycle improvement director. Forms causal hypotheses about why scores are low, validates them with scout agents before attacking, dispatches axis-parallel fleet attacks, extracts transferable patterns, and runs indefinitely within a budget envelope. Accumulates a persistent belief model and pattern library across sessions.
license
MIT
user-invocable
true
auto-trigger
false
trigger_keywords
evolve, sustained improve, improvement director, research-driven improve, multi-cycle improve, run until done, improve until ceiling, keep improving…
last-updated
2026-05-03

/evolve — Improvement Director

Orientation

Use when: You want sustained autonomous quality advancement — the director forms hypotheses, scouts before attacking, and builds a belief model that compounds across cycles. Runs until a natural ceiling, budget exhaustion, or you say stop.

Don't use when: You want a single scored loop (/improve), a known axis attacked directly (/improve --axis), or a one-time audit (/improve --score-only).

Key difference from /improve: /improve follows the rubric mechanically. /evolve asks why scores are where they are, validates those theories before spending fleet budget, and extracts cross-skill patterns that propagate to skills never directly attacked.

Invocation

/evolve {target}                  # run until ceiling, velocity drop, or budget
/evolve {target} --n={N}          # exactly N director cycles then stop
/evolve {target} --budget=${X}    # run until cumulative spend reaches $X
/evolve {target} --continue       # resume from saved director state
/evolve {target} --status         # show belief model, velocity, spend — no attack
/evolve {target} --axis={name}    # focus director on one axis (scout + attack only)

target maps to .planning/rubrics/{target}.md. If no rubric exists, run /improve {target} Phase 0 first — /evolve requires an approved rubric and will not auto-generate one.

Campaign Artifacts

All findings are externalized incrementally — written after every phase, not only at cycle end. A crashed or compacted session resumes with full context.

ArtifactPathContents
Director state.planning/evolve/{target}/director-state.jsoncycle count, spend, velocity history, current phase, halt status
Belief model.planning/evolve/{target}/belief-model.jsonlone record per (axis, skill) per cycle: score, hypothesis, evidence, confidence
Experiment log.planning/evolve/{target}/experiment-log.jsonlevery experiment: hypothesis → prediction → actual delta → mechanism confirmed
Pattern library.planning/evolve/{target}/pattern-library.mdtransferable patterns: what change to what axis class caused what delta in which skills
Cycle digest.planning/evolve/{target}/cycle-{n}-digest.mdhuman-readable per-cycle summary for review
Global patterns.planning/research/patterns.mdcross-target patterns written outside campaign scope; available to future sessions and other targets
Knowledge wiki.planning/wiki/compiled wiki pages from /learn; integrates evolve discoveries across sessions

Create .planning/evolve/{target}/ on first invocation. Create .planning/research/ if absent.

Cycle digest contents: scores table (axis, prior, this cycle, delta), hypotheses table (id, axis, hypothesis, scout result, confidence), what was attacked (axis, skill, delta, mechanism confirmed), patterns discovered this cycle, belief model updates, and the spend/velocity line. Full template: docs/QUALITY_LOOPS.md#cycle-digest-format.

Director Cycle Protocol

Phase 1: Survey

Run /improve {target} --score-only. Record scores to belief model with delta from prior cycle (empty on cycle 1). Flag any axis that dropped since last cycle as regression-watch — these are checked first in Phase 2.

Phase 2: Hypothesize

For every axis below 8.0, generate one primary hypothesis in this form:

HYPOTHESIS: {axis} scores {n}/10 because {specific mechanism},
            not because {common misread}.
PREDICTION: Fixing {mechanism} will raise score ≥ {delta} across {N} skills.
FALSIFICATION: If we apply {change} and score does not rise > 0.5, hypothesis rejected.

Draw hypotheses from: evaluator justifications in Phase 1, prior evidence in the belief model, and programmatic check failures. Do not hypothesize from score alone — the number is the symptom.

Write each hypothesis to the experiment log as { id, status: "pending", ... }.

Skip hypothesis generation for an axis if the belief model already has a confidence >= 0.8 confirmed hypothesis for it that has not yet been attacked.

Phase 3: Scout

For axes below 7.0, or axes with unconfirmed hypotheses: dispatch one scout agent per hypothesis. Scouts read — they do not modify files.

Each scout returns { "hypothesis_id", "confirmed", "evidence", "confidence" } (schema example: docs/QUALITY_LOOPS.md#scout-result-schema).

Scout confidence protocol: Scouts read relevant files only — no edits, no test runs. Assign confidence:

  • 0.9+: mechanism is directly observable (explicit absence, missing section, wrong value in file)
  • 0.7–0.89: strong indirect evidence from 2+ corroborating observations
  • 0.4–0.69: single observation that supports the hypothesis; alternative explanations plausible
  • < 0.4: no direct evidence found; hypothesis is speculative from this file set

Run scouts in parallel. Update experiment log:

  • confidence >= 0.7 → confirmed
  • confidence 0.4–0.69 → needs-evidence (do not attack; add to next cycle)
  • confidence < 0.4 → rejected

Skip Phase 3 for any hypothesis already confirmed at confidence >= 0.8 in the belief model from a prior cycle.

Phase 4: Prioritize

For each confirmed hypothesis compute:

EV = (delta_estimate × axis_weight × confidence) / (effort_tier × collision_multiplier)
  • effort_tier: low=1.0, medium=1.5, high=2.5
  • collision_multiplier: 2.0 if axis shares primary files with another attack in this cycle

Select top K axes where K = min(confirmed count, 4). Document selection rationale in cycle digest. If --axis was set, skip ranking — attack only that axis.

Phase 5: Fleet Attack

Dispatch one agent per selected axis in an isolated worktree (Agent tool, isolation: "worktree"). Each agent receives:

  • The confirmed hypothesis and its falsification criterion
  • The specific files to modify
  • Verification oracle: node scripts/run-with-timeout.js 300 node scripts/test-all.js

Each agent returns a structured result: { "axis", "skill", "delta", "mechanism_confirmed", "files_changed", "approach" } (schema example: docs/QUALITY_LOOPS.md#fleet-agent-result-schema).

Merge rules:

  • Non-conflicting worktrees: merge all
  • Conflicting worktrees (same file): keep higher delta, discard lower
  • Regression on any previously passing programmatic check: abort that worktree, do not merge
  • mechanism_confirmed: false (score improved but not via predicted mechanism): record as incidental_improvement, mark hypothesis as needs-revision

Commit each merged worktree with a message citing the hypothesis ID.

After committing changes to any SKILL.md (here or in Phase 7), run /reload-skills if the running Claude Code version supports it so the change is live this session; otherwise note that a fresh session is required before the updated skill takes effect.

Phase 6: Synthesize

For each result:

  1. Update belief model — append evidence record for (axis, skill)
  2. Update experiment log — mark verified / refuted / incidental
  3. Identify transferable patterns:
PATTERN: {axis_class} | Mechanism: {what caused improvement} | Delta: {avg} across {N} instances | Applies to: {skill list} | Confidence: high/medium/low

Write patterns to .planning/evolve/{target}/pattern-library.md.

Compile into wiki: After writing to the pattern library, call /learn --from-evolve {target} --cycle {n}. This compiles cycle discoveries into .planning/wiki/ — integrating with findings from prior cycles and campaigns rather than siloing them in the evolve directory. Skip if /learn is not available in this session (log the skip, do not block the cycle).

Phase 7: Cross-Pollinate

For each confidence: high pattern, or any pattern confirmed in 2+ skills: apply to all other applicable skills as targeted single-file edits — without running a full attack cycle.

Run verification oracle per cross-pollinated skill. Commit only if all programmatic checks pass and no axis drops > 0.3. Revert on regression; mark pattern as context-dependent.

Write patterns that apply beyond this target to .planning/research/patterns.md.

Show full SKILL.md (680 more words)Show less
Phase 8: Loop or Halt

Compute learning velocity:

velocity = Σ(delta across all attacked axes this cycle) / axes_attacked

Append to director-state.json velocity history.

Halt conditions (check in order):

  1. --n cycles completed
  2. --budget reached (cumulative cost ≥ limit)
  3. All axes ≥ 9.0 across all scored skills
  4. velocity < 0.2 for 3 consecutive cycles AND no needs-evidence hypotheses remain
  5. Level-up triggered (see below)
  6. User says stop

On velocity drop, before halting: attempt one axis-class switch — attack the highest-EV axis from a category not touched in the last 2 cycles. If velocity is still < 0.2 after that cycle, halt.

On level-up trigger (no axis improved > 0.5 for 2 loops, ≥ 3 loops run, no programmatic failures): write level-up proposals to .planning/rubrics/{target}-proposals.md, set status: level-up-pending in director state, halt. The campaign resumes only after the human approves and edits the live rubric.

On normal loop: increment cycle, compress prior cycle findings to continuation context, return to Phase 1.

Unlimited Mode

No --n and no --budget = unlimited. Declare before starting: target, exit conditions (all axes ≥ 9.0 OR velocity < 0.2 for 3 cycles), estimated cost ($12–18/cycle), spend so far ($0), and how to halt (type /stop or press Escape to stop after the current cycle). At the end of every cycle report cycle spend, cumulative spend, and velocity. Literal declaration and report templates: docs/QUALITY_LOOPS.md#unlimited-mode-templates.

When context approaches compression territory (session duration > 30 min or /compact recommended): write continuation checkpoint to director state, surface the --continue command. The next session picks up exactly where this one stopped.

For overnight / unattended runs: combine with /daemon. The director is daemon-compatible — daemon calls /evolve {target} --continue each session. Set --budget to cap total spend.

Fringe Cases

  • .planning/ does not exist: error — run /do setup first to initialize the harness state directory, then retry.
  • No rubric: error — run /improve {target} Phase 0 first. List available targets in .planning/rubrics/ as hint.
  • No prior scores in belief model: proceed from cycle 1; all deltas empty on first survey. Expected.
  • All scouts return needs-evidence: attack the top-EV axis anyway under low-confidence flag; record as exploratory. Mark result regardless.
  • Scout agent hangs or times out (dispatched scout never returns): After 10 minutes without a response, log the scout as status: timed-out in the experiment log with confidence: 0. Proceed with the remaining returned scouts. Never let a hung scout block the cycle — if all scouts time out, treat as "all scouts return needs-evidence" and attack the top-EV axis under low-confidence flag.
  • All axes collide (every axis shares files): serialize top 2 axes; parallelize remainder. Log collision.
  • Cross-pollination causes regression: revert that skill, mark pattern context-dependent, do not propagate further.
  • Level-up mid-campaign: pause, write proposals, set level-up-pending. /evolve --continue after human approval resumes cycle numbering from where it stopped.
  • Budget overrun risk: if projected spend for current cycle would exceed --budget by > 20%, warn and confirm before dispatching fleet.
  • --continue with no director state: error — no campaign to resume. Suggest /evolve {target} to start fresh.
  • Pattern library > 50 entries: consolidate — group by axis class, merge similar patterns, keep highest-confidence instance of each class. Log consolidation.
  • Zero skills match target rubric: error with message listing all .planning/rubrics/*.md targets.

Quality Gates

  • Every hypothesis must have an explicit falsification criterion before Phase 3
  • Scouts must run before fleet dispatch on any unconfirmed hypothesis
  • Belief model written after every phase, not only end of cycle
  • Cross-pollination requires passing verification oracle before commit
  • Regression on any previously-passing axis aborts that worktree commit
  • Pattern library and global patterns updated at every cycle end, even zero-improvement cycles
  • Cycle digest written even on abort or no-change cycles

Contextual Gates

Disclosure:

  • State mode (unlimited / fixed / budget) and exit conditions before first cycle
  • Estimate $12–18 per full cycle before starting
  • Report spend and velocity at end of every cycle
  • Confirm before continuing if cumulative spend exceeds $50

Reversibility: Red. Cross-pollination modifies many files across the repo; level-up rewrites rubric anchors permanently. Each commit is individually revertable; high volume. Range: git revert {first}^..{last}.

Trust gates:

  • Novice (0-4 sessions): --status and --n=1 only; unlimited blocked
  • Familiar (5-19): up to --n=5; unlimited requires explicit --budget cap
  • Trusted (20+): no cap; confirm if projected total > $100

Exit Protocol

---HANDOFF---
- Target: {target} | Cycles: {n} | Spend: ${total} | Mode: {unlimited/n/budget}
- Axes improved: {list with deltas}
- Belief model: .planning/evolve/{target}/belief-model.jsonl ({N} confirmed, {M} rejected)
- Pattern library: .planning/evolve/{target}/pattern-library.md ({N} patterns)
- Global patterns: .planning/research/patterns.md
- Knowledge wiki: .planning/wiki/index.md (compiled via /learn --from-evolve after each cycle)
- Cycle digests: .planning/evolve/{target}/cycle-*-digest.md
- Halt reason: {ceiling/velocity/budget/n-complete/user-stop/level-up-pending}
- Level-up proposals: {path or N/A}
- Reversibility: red — {N} commits across {M} files; revert range: git revert {range}
- Recommended next: {level-up and re-run / new target / done}
---

© SethGammon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/evolve of SethGammon/Citadel.

  • SKILL.md
  • __benchmarks__/happy-path-status.md
  • __benchmarks__/no-rubric.md

Open the folder on GitHubat commit e41ff1d

Compare with similar skills

Evolve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evolve compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evolve this skillSethGammon/Citadel922—~3.7kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.3kAutomated safety check: PassApache-2.0
Zhang Xuefeng Perspectivealchaincyf/zhangxuefeng-skill10k1 repos~2.6kAutomated safety check: PassMIT
Deep Reading Analystginobefun/deep-reading-analyst-skill3535 repos~3.6kAutomated safety check: PassMIT
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch65k—~2kAutomated safety check: PassMIT
OpenMAIC Setup and ExtensionTHU-MAIC/OpenMAIC40k—~1.7kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.3k tokensUpdated 3 days ago
    EducationAuto-check passed
  • Zhang Xuefeng Perspective

    alchaincyf/zhangxuefeng-skill

    Answers education and career questions in the voice of Zhang Xuefeng, looking up current employment and admissions data before giving a direct verdict.

    10k GitHub starsUsed in 1 repo~2.6k tokens
    EducationAuto-check passed
  • Deep Reading Analyst

    ginobefun/deep-reading-analyst-skill

    Comprehensive framework for deep analysis of articles, papers, and long-form content using 10+ thinking models (SCQA, 5W2H, critical thinking, inversion, mental models, first principles, systems…

    353 GitHub starsUsed in 5 repos~3.6k tokens
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    65k GitHub stars~2k tokensUpdated yesterday
    EducationAuto-check passed
  • Guides setup, classroom generation and secondary development for OpenMAIC, the multi-agent interactive classroom, one confirmed phase at a time.

    40k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check: notes
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed

More from SethGammon/Citadel

All 48 skills in this repo
  • Create Skill

    SethGammon/Citadel

    Creates new skills from the user's repeating patterns. An agent skill from SethGammon/Citadel.

    922 GitHub stars~1.9k tokensUpdated 6 days ago
    Auto-check passed
  • Houseclean

    SethGammon/Citadel

    Cross-drive storage audit and cleanup. An agent skill from SethGammon/Citadel.

    922 GitHub stars~2.2k tokensUpdated 6 days ago
    Auto-check passed
  • Loop

    SethGammon/Citadel

    Bounded foreground repetition for the current session. An agent skill from SethGammon/Citadel.

    922 GitHub stars~1.4k tokensUpdated 6 days ago
    Auto-check passed
  • Triage

    SethGammon/Citadel

    GitHub issue and PR investigator. An agent skill from SethGammon/Citadel.

    922 GitHub stars~2.7k tokensUpdated 6 days ago
    Auto-check passed
  • Watch

    SethGammon/Citadel

    File sentinel that monitors the working directory for changes and marker comments, then auto-triggers appropriate skills.

    922 GitHub stars~2.9k tokensUpdated 6 days ago
    Auto-check passed
  • Archon

    SethGammon/Citadel

    Autonomous multi-session campaign agent. An agent skill from SethGammon/Citadel.

    922 GitHub stars~5.4k tokensUpdated 6 days ago
    Auto-check passed

Categories

Questions about Evolve

What does Evolve do?

Research-driven multi-cycle improvement director. An agent skill from SethGammon/Citadel. Evolve is an agent skill from SethGammon/Citadel. Research-driven multi-cycle improvement director.

When should I use Evolve?

Evolve fits situations like: education work in your project.

How do I install Evolve in Claude Code?

Run `npx skills add SethGammon/Citadel --skill evolve -a claude-code`. Or copy the skill folder (skills/evolve in SethGammon/Citadel) into .claude/skills/evolve in your project. Claude Code loads it when a task matches its description.

How do I install Evolve in Codex?

Run `npx skills add SethGammon/Citadel --skill evolve -a codex`. Or copy the skill folder (skills/evolve in SethGammon/Citadel) into .agents/skills/evolve in your project. Codex loads it when a task matches its description.

Can I use Evolve in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SethGammon/Citadel --skill evolve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evolve, .gemini/skills/evolve, .github/skills/evolve and .opencode/skills/evolve in your project.

What does Evolve need to run?

Going by SKILL.md and its folder, Evolve needs the command-line tools its instructions call (node and git).

Does Evolve access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evolve safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evolve use?

Evolve is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evolve use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evolve?

Skills that share tags, products or a category with Evolve: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), Zhang Xuefeng Perspective (alchaincyf/zhangxuefeng-skill, 10k stars), Deep Reading Analyst (ginobefun/deep-reading-analyst-skill, 353 stars) and AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 65k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evolve?

SethGammon (a GitHub user) maintains it in SethGammon/Citadel, which has 922 GitHub stars. The repository holds 48 skills in this directory. The repository was last updated on October 1, 2026.

Source: SethGammon/Citadel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.