Agent skill

Harness Evolve

by ruvnet in ruvnet/ruflo

Run @metaharness/darwin evolve <repo to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and…

MITAuto-check: notesDevelopment

Install Harness Evolve

skills CLI
$ npx skills add ruvnet/ruflo --skill harness-evolve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo harness-evolve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-metaharness/skills/harness-evolve .claude/skills/harness-evolve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-evolve
GitHub stars
74k
Token cost
~1.6k tokens
SKILL.md length
601 words
Files
1
Skills in repo
264
Repo updated
First seen
Licence
MIT

At a glance

Run @metaharness/darwin evolve <repo to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and…

  • Works in 6 steps: Validate args (--repo exists, caps on… → Without --confirm: print plan + exit 0… → With --confirm: run metaharness-darwin… → …
  • Tasks that involve Architecture decision records
  • SKILL.md covers When to use, When NOT to use, Algorithm and The seven mutation surfaces, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Harness Evolve is an agent skill from ruvnet/ruflo. Run @metaharness/darwin evolve <repo to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Architecture decision records and Bioinformatics. The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

When your agent uses it

  • Tasks that involve Architecture decision records
  • Tasks that involve Bioinformatics

Example prompts

  • “/harness-evolve”

Requirements

  • Pre-approved tools (allowed-tools): Bash

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Validate args (--repo exists, caps on --generations ≤ 50, --children
  2. Without --confirm: print plan + exit 0 (mirrors harness-mint safety
  3. With --confirm: run metaharness-darwin evolve ... from the installed
  4. Compute timeout from generations × children × per-variant (per-variant
  5. Honor upstream exit code 99 — propagate as "safety-disqualified", do not
  6. Optional --alert-on-no-improvement: exit 1 when champion ≤ parent.

What it can do on your machine

Read from SKILL.md and the folder at commit de590e1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Evolve loads about 1.6k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 601 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit de590e1, republished under its MIT licence (© ruvnet). 601 words, ~1,578 tokens.

Download SKILL.mdSave it as .claude/skills/harness-evolve/SKILL.md (or your agent's skills folder).
name
harness-evolve
description
Run `@metaharness/darwin evolve <repo>` to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).
allowed-tools
Bash
argument-hint
--repo <path> [--generations 3] [--children 3] [--concurrency 2] [--sandbox real|mock|agent] [--selection pareto|quality-diversity|...] [--mutator…

Surfaces the upstream metaharness-darwin evolve CLI as a ruflo skill. The write layer that pairs with ADR-150's read layer (score / genome / mcp-scan / threat-model / oia-audit). Use when you have a harness whose readiness scores are flat and you want to discover which surface mutation moves them — without retraining the foundation model.

When to use

  • A harness-score result is below target and you don't know which policy surface is responsible.
  • You're seeding a harness for a new vertical and want to find a good starting configuration empirically rather than hand-tuning.
  • You're comparing your hand-tuned harness against an evolved baseline (treat darwin's champion as the strawman).

When NOT to use

  • For continuous background optimization. Darwin Mode is human-initiated. Wire it into CI for one-shot exploration, not for autonomous self-modification.
  • For ruflo itself in CI. ADR-153 §5 explicitly rejects auto-evolving ruflo — the CI gate verifies graceful degradation, not convergence.

Algorithm

Implementation: scripts/evolve.mjs.

  1. Validate args (--repo exists, caps on --generations ≤ 50, --children ≤ 20, --concurrency ≤ 8, sandbox/selection/mutator are known values).
  2. Without --confirm: print plan + exit 0 (mirrors harness-mint safety convention; defense in depth over the upstream safety.ts checks).
  3. With --confirm: run metaharness-darwin evolve <repo> ... from the installed @metaharness/darwin (pin in _darwin.mjs; one-time ~/.ruflo/darwin-cache-<pin> install only if no installed copy qualifies — never npx) via the shared _darwin.mjs async helper. Per-generation progress is forwarded to stderr; final champion JSON is captured from stdout.
  4. Compute timeout from generations × children × per-variant (per-variant ≈ 60s real, ≈ 2s mock). Caller may override with --timeout-ms.
  5. Honor upstream exit code 99 — propagate as "safety-disqualified", do not remap. This is a designed-in tripwire (a variant tripped inspectVariant for secrets / shell-out / network / dynamic-eval). See ADR-153 §"Safety model".
  6. Optional --alert-on-no-improvement: exit 1 when champion ≤ parent.

The seven mutation surfaces

SurfaceWhat it owns
plannertask decomposition / step ordering
contextBuilderwhat gets fed into the prompt
reviewerself-critique / output verification
retryPolicywhen + how to retry on failure
toolPolicywhich tools the agent may use, under which conditions
memoryPolicywhat to persist, recall, forget
scorePolicyhow the agent grades its own output

One mutation per variant. Multi-surface mutations are not allowed (causal attribution stays clean).

Show full SKILL.md (250 more words)Show less

Output

Reports land under <repo>/.metaharness/:

.metaharness/
  archive.json         # full lineage tree (sampling next gen draws from this)
  lineage.json         # parent→child edges only
  variants/<id>/       # per-variant code (kept for audit)
  runs/<id>/           # per-variant sandbox test output
  reports/winner.json  # final champion + score delta vs parent

Skill stdout = JSON {success, data: {champion, plan, durationMs, improved}} (plus data.diagnosis when --diagnose is passed — see below).

Failure diagnosis (--diagnose)

GEPA's key trick is natural-language failure diagnosis from execution traces feeding the next mutation — not just scalar fitness. --diagnose adds a modest slice of that: after the evolution completes, the losing / failed variants' transcripts are run through darwin's GEPA library ops (analyzeTranscript + classifyFailure, via the shared importGepa resolver in scripts/_darwin.mjs) and a diagnosis section is appended to the emitted JSON:

json
"diagnosis": {
  "available": true,
  "scope": "losing-variants",
  "variants": [
    { "id": "g1_v0", "transcripts": 2,
      "failureClasses": { "exploration-loop": 1, "edit-mechanics": 1 },
      "dominantClass": "exploration-loop" }
  ],
  "totals": { "exploration-loop": 1, "edit-mechanics": 1 }
}

Upstream shape caveats (verified against @metaharness/darwin@0.8.0):

  • metaharness-darwin evolve --json prints a TEXT leaderboard — the stdout carries no JSON and no transcripts. Per-variant run records live at <repo>/.metaharness/runs/<id>.json.
  • Those run records hold sandbox exec traces ({taskId, exitCode, stdout, stderr}), which are NOT GEPA {actionRaw, obs} transcripts. Diagnosis therefore uses GEPA-shaped transcripts when a run record embeds them (agent sandbox / future upstream), falls back to the champion's transcript, and otherwise emits diagnosis: {available: false, reason, traceSummary} where traceSummary is a mechanical per-variant tally (tasks / failed / timedOut / blockedActions).
  • --diagnose NEVER fails the run — any internal error degrades to {available: false, reason: "diagnosis-failed: ..."}.

Exit codes

CodeMeaning
0Evolved OK, or dry-run, or degraded (Darwin absent)
1--alert-on-no-improvement and champion did not beat parent
2Config error or evolution infrastructure failure
99Upstream "safety-disqualified" (PROPAGATED, not remapped)

Graceful degradation (ADR-150 constraint 3 + ADR-153)

When @metaharness/darwin is not installed, the script emits {degraded: true, reason: 'metaharness-darwin-not-available', hint: ...} and exits 0. ruflo continues to function. CI's no-metaharness-smoke.yml-style job asserts this path.

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-metaharness/skills/harness-evolve of ruvnet/ruflo.

Open the folder on GitHubat commit de590e1

Compare with similar skills

Harness Evolve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Evolve compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Evolve this skillruvnet/ruflo74k—~1.6kAutomated safety check: NotesMIT
Repo Genomeruvnet/metaharness688—~772Automated safety check: PassMIT
PR Design DocOpenHands/OpenHands90k—~2.4kAutomated safety check: PassMIT
Cto AdvisorIbrahim-3d/orchestrator-supaconductor3804 repos~2.4kAutomated safety check: PassMIT
Improve Codebase Architectureywwynm/EverythingDone14415 repos~1.3kAutomated safety check: PassGPL-3.0
Domain Modelingbrim-borium/spotify_sdk1665 repos~806Automated safety check: PassApache-2.0

Similar skills

  • Repo Genome

    ruvnet/metaharness

    7-section readiness scorecard for a LOCAL repo. An agent skill from ruvnet/metaharness.

    688 GitHub stars~772 tokensUpdated yesterday
    Product & Project ManagementAuto-check passed
  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Cto Advisor

    Ibrahim-3d/orchestrator-supaconductor

    Technical leadership guidance for engineering teams, architecture decisions, and technology strategy.

    380 GitHub starsUsed in 4 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Improve Codebase Architecture

    ywwynm/EverythingDone

    Find deepening opportunities in a codebase, informed by the domain language in CONTEXT.md and the decisions in docs/adr/.

    144 GitHub starsUsed in 15 repos~1.3k tokens
    DevelopmentAuto-check passed
  • Domain Modeling

    brim-borium/spotify_sdk

    Build and sharpen a project's domain model. An agent skill from brim-borium/spotify_sdk.

    166 GitHub starsUsed in 5 repos~806 tokens
    DevelopmentAuto-check passed
  • Design Doc Mermaid

    SpillwaveSolutions/design-doc-mermaid

    Create Mermaid diagrams (flowchart, sequence, class, ER, state, C4, architecture) from text or source code.

    175 GitHub starsUsed in 1 repo~5.6k tokens
    DevelopmentAuto-check passed

More from ruvnet/ruflo

All 264 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 2 repos~830 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 2 repos~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 2 repos~779 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Agent Coordination

    ruvnet/ruflo

    Reference for spawning, listing, monitoring and stopping agents with claude-flow commands, with agent type families, routing codes and coordination tips.

    74k GitHub starsUsed in 2 repos~519 tokens
    Auto-check passed

Categories

Questions about Harness Evolve

What does Harness Evolve do?

Run @metaharness/darwin evolve <repo to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and…. Harness Evolve is an agent skill from ruvnet/ruflo. Run @metaharness/darwin evolve <repo to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins.

When should I use Harness Evolve?

Harness Evolve fits situations like: tasks that involve Architecture decision records; tasks that involve Bioinformatics.

How do I install Harness Evolve in Claude Code?

Run `npx skills add ruvnet/ruflo --skill harness-evolve -a claude-code`. Or copy the skill folder (plugins/ruflo-metaharness/skills/harness-evolve in ruvnet/ruflo) into .claude/skills/harness-evolve in your project. Claude Code loads it when a task matches its description.

How do I install Harness Evolve in Codex?

Run `npx skills add ruvnet/ruflo --skill harness-evolve -a codex`. Or copy the skill folder (plugins/ruflo-metaharness/skills/harness-evolve in ruvnet/ruflo) into .agents/skills/harness-evolve in your project. Codex loads it when a task matches its description.

Can I use Harness Evolve in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill harness-evolve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-evolve, .gemini/skills/harness-evolve, .github/skills/harness-evolve and .opencode/skills/harness-evolve in your project.

What does Harness Evolve need to run?

SKILL.md names no scripts, command-line tools or credentials: Harness Evolve is instructions for the agent only. Its frontmatter pre-approves these tools: Bash.

Does Harness Evolve access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harness Evolve safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Harness Evolve use?

Harness Evolve is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Evolve use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Harness Evolve?

Skills that share tags, products or a category with Harness Evolve: Repo Genome (ruvnet/metaharness, 688 stars), PR Design Doc (OpenHands/OpenHands, 90k stars), Cto Advisor (Ibrahim-3d/orchestrator-supaconductor, 380 stars) and Improve Codebase Architecture (ywwynm/EverythingDone, 144 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Evolve?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,012 GitHub stars. The repository holds 264 skills in this directory. The repository was last updated on October 7, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.