Agent skill

Evolve Agent

by simple-agent-lab in simple-agent-lab/RSIHub

Run evidence-driven evolution of agents, prompts, skills, and agent harnesses.

Apache-2.0Auto-check passed

Install Evolve Agent

skills CLI
$ npx skills add simple-agent-lab/RSIHub --skill evolve-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install simple-agent-lab/RSIHub evolve-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/simple-agent-lab/RSIHub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evolve-agent .claude/skills/evolve-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evolve-agent
GitHub stars
180
Token cost
~2.3k tokens
SKILL.md length
1,076 words
Files
10 (incl. references)
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run evidence-driven evolution of agents, prompts, skills, and agent harnesses.

  • Works in 6 steps: Establish the contract → Choose a method and control path → Author reusable operators in a source… → …
  • Asked to initialize
  • SKILL.md covers 1. Establish the contract, 2. Choose a method and control…, 3. Author reusable operators… and 4. Prefer capabilities over…, plus 4 more sections
  • Calls uv

What it does

Evolve Agent is an agent skill from simple-agent-lab/RSIHub. Run evidence-driven evolution of agents, prompts, skills, and agent harnesses. Use when asked to initialize or operate an evolution workspace, choose Hill Climb, A-Evolve, GEPA, AHE, or HyperAgents, let an outer Agent adapt the evolution process, invoke operators directly, improve a candidate through repeated evaluation, recover interrupted evolution, or report an evidence-backed champion.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `agents/openai.yaml`, `references/a-evolve.md` and `references/agent-driven.md`).

It works with Python. The repository describes itself as: A research framework for principled agent self-improvement under frozen evaluators and declared mutation boundaries, recording verifiable lineage to make it reproducible and… The licence is Apache-2.0.

When your agent uses it

  • Asked to initialize
  • Operate an evolution workspace
  • Choose Hill Climb
  • Let an outer Agent adapt the evolution process

Example prompts

  • “/evolve-agent”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Establish the contract
  2. Choose a method and control path
  3. Author reusable operators in a source checkout
  4. Prefer capabilities over source
  5. Close the loop
  6. Report only what the chain proves

What it can do on your machine

Read from SKILL.md and the folder at commit 64491da. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evolve Agent loads about 2.3k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,076 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from simple-agent-lab/RSIHub at commit 64491da, republished under its Apache-2.0 licence (© simple-agent-lab). 1,076 words, ~2,302 tokens.

Download SKILL.mdSave it as .claude/skills/evolve-agent/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
evolve-agent
description
Run evidence-driven evolution of agents, prompts, skills, and agent harnesses. Use when asked to initialize or operate an evolution workspace, choose Hill Climb, A-Evolve, GEPA, AHE, or HyperAgents, let an outer Agent adapt the evolution process, invoke operators directly, improve a candidate through repeated evaluation, recover interrupted evolution, or report an evidence-backed champion.

Build an evidence chain

Treat evolution as an evidence chain:

text
contract → baseline → evidence → hypothesis → candidate → evaluation → lineage

A higher score alone is insufficient. Link every candidate to the evidence that motivated it, its exact snapshot, frozen evaluation, and lineage decision.

1. Establish the contract

Detect whether the current directory is an initialized evolution workspace.

  • In a workspace, read AGENTS.md, evolve.yaml, program.md, then run ./evolve status . and ./evolve verify ..
  • For a new experiment, identify the target, mutable surface, evaluator, data partitions, budget, and execution boundary before initialization.
  • Read the workspace contract before creating, operating, recovering, or interpreting a workspace. Its "Create a workspace" section gives the initialization and baseline-certification commands and the preconditions they enforce.

Completion check: Name the target, mutable surface, frozen evaluator (the evaluator/ contract that scores every candidate), data partitions, candidate budget, and execution boundary. In an existing workspace, also identify the current champion, next generation, and interrupted state.

2. Choose a method and control path

The initialized operators define the starting method. The control path may be driver-led, where evolve run fixes the lifecycle, or agent-led, where the outer Agent chooses which direct capabilities to invoke and may change the active process when the surface permits it. For a new experiment, GEPA is the default; choose Hill Climb when the experiment needs the simplest attributable control. Match the method and control path to the research question, available evidence, and mutable surface. Read only the relevant method card.

Observable conditionMethodRead
A minimal attributable control is enoughHill Climbhill-climb.md
Behavioral traces or generated-artifact rubrics should guide prompt or skill mutationA-Evolvea-evolve.md
The evaluator returns per-task results and the target splits into componentsGEPAgepa.md
Failures are execution-shaped and justify harness changesAHEahe.md
The evolution process itself may also changeHyperAgentshyperagents.md

When the outer Agent, rather than a configured mutate stage, should decide the sequence of investigation, operator calls, edits, retries, and stopping, read agent-driven control. This is an experimental control path over existing workspace capabilities, not a new operator method.

Read scientific foundations only when defining or changing evaluator semantics, partitions, acceptance rules, or research claims.

Completion check: State why the method and control path match the evidence and declared mutable surface. If they do not, choose again before running.

For artifact-producing Skills, prefer replaying a selected parent's certified artifacts over executing that parent again. Re-execute the parent only when the current task set, evaluator identity, runtime identity, or required artifacts do not match the retained evidence. Execute every child freshly.

3. Author reusable operators in a source checkout

Use the library when creating a reusable policy, not an edit to one already initialized workspace. Discover available entries, then scaffold and verify one operator before selecting it from a recipe:

bash
uv run --frozen evolve operator list [stage]
uv run --frozen evolve operator new mutate <name>
uv run --frozen evolve operator describe mutate/<name>
uv run --frozen evolve operator check mutate/<name> --config '{"attempts": 3}'
uv run --frozen evolve recipe check <recipe-path>

new writes exactly one entry at library/mutate/<name>.py. Implement the generated MutateOperator, keep validate_config, and use sdk.main(..., validate_config=validate_config). A recipe selects it with an operator: value and nested config: mapping:

yaml
operators:
  mutate:
    operator: critic_editor
    timeout_s: 3600
    config:
      attempts: 3

Do not put a reusable implementation beside a recipe or alter a library entry to change a running workspace. Run recipe check before initialization; a new workspace freezes the selected source. Existing workspaces retain their own frozen active operators.

Completion check: The operator is in the central library, its configuration passes operator check, the recipe passes recipe check, and the source change is separated from any initialized workspace it does not retroactively alter.

Show full SKILL.md (524 more words)Show less

4. Prefer capabilities over source

For agent-led evolution, start from the stable workspace interface:

bash
./evolve operator active . --json
./evolve operator run . <stage> --genid <id> [stage arguments]

Treat operator active --json as the live authority for which stages are configured and whether their access is direct, driver, or finalize. Invoke configured direct operators, read their retained artifacts under runs/gen-<id>/, and make the candidate change yourself.

Escalate progressively:

  1. Tune one call with --config when the capability is right but its bounds are wrong.
  2. Read PROTOCOL.md, operator guidance, or operators/README.md when an input or artifact is unclear.
  3. Read the active operators/<stage>.py only to diagnose behavior or change the active evolution process.
  4. Read library/<stage>/ only to compare or adapt another implementation.

Do not read implementation source merely to invoke a working operator. Do not edit library/ and assume runtime behavior changed; active code lives under operators/.

Use the configured driver when its mutation stage should own the edit and an unattended run is desired:

bash
./evolve run . --max-generations 1

Driver and agent-led paths share the same evaluation and lineage mechanism. Do not run them concurrently. Ordinary agent-led work should close one generation through the stable commands below. For an explicitly Agent Driven experiment, the outer Agent may adapt its action sequence under the Agent Driven control reference; it must still use the mechanism for candidate identity, evaluation, and finalization.

Completion check: Choose exactly one control path for the active work. For agent-led evolution, name the available direct operators, hard budget, and the evidence supporting the next action; source inspection must have a concrete reason.

5. Close the loop

  1. Establish and inspect the certified baseline.
  2. Select a parent and retain the method's required evidence.
  3. State one evidence-linked hypothesis and predicted effect.
  4. Produce one candidate inside the declared mutable surface.
  5. Run every configured admission check against the final candidate snapshot.
  6. Evaluate and finalize through the workspace mechanism.
  7. Verify lineage before beginning another generation.

Completion check: The candidate has an exact lineage identity; required admission decisions and evaluator-stamped results exist; lineage verification passes; accepted and rejected outcomes remain auditable.

6. Report only what the chain proves

Start from ./evolve report ., which writes the experiment report and research-claim checklist from stamped records. Around it, report the baseline, champion, parent-child changes, accepted and rejected mutations, evaluation scope, retained evidence, and limitations. Tie every quality claim to evaluator-stamped artifacts from the run.

Completion check: Every score and champion identity is derivable from trusted lineage records, and every generalization claim names its data partition.

Guard the chain

  • Keep one evaluator and runtime identity within an experiment. Start a new experiment when the evaluator changes.
  • Take scores and champion state only from mechanism-owned stamped records.
  • Keep optimization, gate, and sealed task identities disjoint.
  • Change only the declared mutable surface.
  • Treat linked worktrees outside runs/worktrees/ as user-owned. Report them; never remove or modify them without explicit authorization.
  • Match the execution boundary to candidate trust.
  • Keep credentials out of prompts, artifacts, and reports.
  • Spend live evaluation budget only when the request authorizes execution.

Historical-workspace note

Older initialized workspaces retain the stage files and configuration frozen at creation time. Treat those as historical metadata only; start a new workspace to use the current operator model.

© simple-agent-lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in skills/evolve-agent of simple-agent-lab/RSIHub.

  • SKILL.md
  • agents/openai.yaml
  • references/a-evolve.md
  • references/agent-driven.md
  • references/ahe.md
  • references/gepa.md
  • references/hill-climb.md
  • references/hyperagents.md
  • references/scientific-foundations.md
  • references/workspace-contract.md

Open the folder on GitHubat commit 64491da

Compare with similar skills

Evolve Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evolve Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evolve Agent this skillsimple-agent-lab/RSIHub180—~2.3kAutomated safety check: PassApache-2.0
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
NotebookLM Research AssistantPleasePrompto/notebooklm-skill7.8k14 repos~2.4kAutomated safety check: NotesMIT
Manim Video Productionbrowser-use/video-use28k6 repos~3kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • NotebookLM Research Assistant

    PleasePrompto/notebooklm-skill

    Lets Claude Code ask questions of your Google NotebookLM notebooks through browser automation and return answers grounded in your uploaded sources.

    7.8k GitHub starsUsed in 14 repos~2.4k tokens
    Knowledge ManagementAuto-check: notes
  • Manim Video Production

    browser-use/video-use

    Produces math and technical explainer videos with Manim Community Edition: concept animations, equation derivations, algorithm walkthroughs and data stories.

    28k GitHub starsUsed in 6 repos~3k tokens
    Media & CreativeAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • PPT Master

    hugohe3/ppt-master

    Generates editable PowerPoint decks, rebuilds slides from images, fills .pptx templates and polishes existing presentations through routed workflows.

    58k GitHub starsUsed in 1 repo~2.5k tokens
    Documents & OfficeAuto-check passed

More from simple-agent-lab/RSIHub

  • Task Execution

    simple-agent-lab/RSIHub

    Solve repository tasks by inspecting first, making focused edits, and verifying the result.

    180 GitHub stars~212 tokensUpdated 24 days ago
    Auto-check passed
  • Task Execution

    simple-agent-lab/RSIHub

    Baseline procedure for completing a terminal task reliably. An agent skill from simple-agent-lab/RSIHub.

    180 GitHub stars~147 tokensUpdated 24 days ago
    Auto-check passed

Works with

Questions about Evolve Agent

What does Evolve Agent do?

Run evidence-driven evolution of agents, prompts, skills, and agent harnesses. Evolve Agent is an agent skill from simple-agent-lab/RSIHub. Run evidence-driven evolution of agents, prompts, skills, and agent harnesses.

When should I use Evolve Agent?

Evolve Agent fits situations like: asked to initialize; operate an evolution workspace; choose Hill Climb; let an outer Agent adapt the evolution process.

How do I install Evolve Agent in Claude Code?

Run `npx skills add simple-agent-lab/RSIHub --skill evolve-agent -a claude-code`. Or copy the skill folder (skills/evolve-agent in simple-agent-lab/RSIHub) into .claude/skills/evolve-agent in your project. Claude Code loads it when a task matches its description.

How do I install Evolve Agent in Codex?

Run `npx skills add simple-agent-lab/RSIHub --skill evolve-agent -a codex`. Or copy the skill folder (skills/evolve-agent in simple-agent-lab/RSIHub) into .agents/skills/evolve-agent in your project. Codex loads it when a task matches its description.

Can I use Evolve Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add simple-agent-lab/RSIHub --skill evolve-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evolve-agent, .gemini/skills/evolve-agent, .github/skills/evolve-agent and .opencode/skills/evolve-agent in your project.

What does Evolve Agent need to run?

Going by SKILL.md and its folder, Evolve Agent needs the command-line tools its instructions call (uv).

Does Evolve Agent access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evolve Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evolve Agent use?

Evolve Agent is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evolve Agent use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.2k tokens, read only when the agent opens those files.

What are the alternatives to Evolve Agent?

Skills that share tags, products or a category with Evolve Agent: MCP Server Builder (anthropics/skills, 180k stars), PDF Processing (anthropics/skills, 180k stars), NotebookLM Research Assistant (PleasePrompto/notebooklm-skill, 7.8k stars) and Manim Video Production (browser-use/video-use, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evolve Agent?

simple-agent-lab (a GitHub organization) maintains it in simple-agent-lab/RSIHub, which has 180 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 13, 2026.

Source: simple-agent-lab/RSIHub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.