Agent skill

Harness Bench

by ruvnet in ruvnet/ruflo

Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness.

MITAuto-check: notesDevelopment

Install Harness Bench

skills CLI
$ npx skills add ruvnet/ruflo --skill harness-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo harness-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-metaharness/skills/harness-bench .claude/skills/harness-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-bench
GitHub stars
74k
Used in
1 other repo
Token cost
~586 tokens
SKILL.md length
217 words
Files
1
Skills in repo
264
Repo updated
First seen
Licence
MIT

At a glance

Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness.

  • Works in 4 steps: Resolve --repo path; reject if missing. → Shell to metaharness-darwin bench create… → Default output path:… → …
  • Tasks that involve Project scaffolding
  • SKILL.md covers When to use, Algorithm, Output shape and Exit codes, plus 1 more section
  • Calls npm

What it does

Harness Bench is an agent skill from ruvnet/ruflo. Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness. Bench suites are the fixed evaluation corpora that harness-evolve --bench <suite.json scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent.

Its SKILL.md is about 590 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Project scaffolding. The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

When your agent uses it

  • Tasks that involve Project scaffolding

Example prompts

  • “/harness-bench”

Requirements

  • Pre-approved tools (allowed-tools): Bash

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Resolve --repo path; reject if missing.
  2. Shell to metaharness-darwin bench create [--out ].
  3. Default output path: /.metaharness/bench/suite.json (chosen by upstream).
  4. Suite shape (per upstream): array of { input, expectedOutput, weight } tasks

What it can do on your machine

Read from SKILL.md and the folder at commit 6051f67. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Bench loads about 586 tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 217 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~586

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit 6051f67, republished under its MIT licence (© ruvnet). 217 words, ~586 tokens.

Download SKILL.mdSave it as .claude/skills/harness-bench/SKILL.md (or your agent's skills folder).
name
harness-bench
description
Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent.
allowed-tools
Bash
argument-hint
--op create --repo <path> [--out <path>] | --op verify --suite <path>

Surfaces metaharness-darwin bench <create|verify> — the supporting verb for harness-evolve --bench. Use when you want evolution scored against a fixed corpus (independent of npm test) so champion fitness is comparable across commits or across forks of the same harness.

When to use

  • Setting up a new evolution pipeline for a repo whose npm test is flaky, slow, or undersized — scaffold a deterministic bench suite once, then evolve against it repeatedly.
  • CI: bench verify the checked-in suite on every PR that touches it (cheap; ~5s).
  • Forking a harness to a new domain: copy and edit the suite to retarget the evaluation without losing comparability to the parent.

Algorithm

Implementation: scripts/bench.mjs.

--op create
  1. Resolve --repo path; reject if missing.
  2. Shell to metaharness-darwin bench create <repo> [--out <suite.json>].
  3. Default output path: <repo>/.metaharness/bench/suite.json (chosen by upstream).
  4. Suite shape (per upstream): array of { input, expectedOutput, weight } tasks derived from existing test cases.
--op verify
  1. Resolve --suite path; reject if missing.
  2. Shell to metaharness-darwin bench verify <suite.json>.
  3. Exit 1 if any task malformed (upstream's signal).

Output shape

json
{
  "success": true,
  "data": {
    "op": "verify",
    "taskCount": 42,
    "wellFormed": true,
    "durationMs": 870
  }
}

Exit codes

CodeMeaning
0OK (or degraded — Darwin absent)
1--op verify and suite malformed
2Config error or upstream invocation failure

Graceful degradation

When @metaharness/darwin is absent, emits the standard {degraded: true, reason: 'metaharness-darwin-not-available'} payload and exits 0.

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-metaharness/skills/harness-bench of ruvnet/ruflo.

Open the folder on GitHubat commit 6051f67

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ruvnet/ruflo, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Harness Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Bench this skillruvnet/ruflo74k1 repos~586Automated safety check: NotesMIT
Nx Generatenomcopter/react-mosaic4.8k7 repos~1.9kAutomated safety check: PassCustom licence
PonytailDavidObando/gsharp5658 repos~1.7kAutomated safety check: PassMIT
Run Nx Generatornrwl/nx29k2 repos~592Automated safety check: NotesMIT
Conductor Setupgemini-cli-extensions/conductor3.8k—~4.2kAutomated safety check: PassApache-2.0
Mirage VFS Adapter Authoringstrukto-ai/mirage3.7k—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Nx Generate

    nomcopter/react-mosaic

    Generate code using nx generators. An agent skill from nomcopter/react-mosaic.

    4.8k GitHub starsUsed in 7 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Ponytail

    DavidObando/gsharp

    Forces the laziest solution that actually works, simplest, shortest, most minimal.

    565 GitHub starsUsed in 8 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Run Nx generators with prioritization for workspace-plugin generators.

    29k GitHub starsUsed in 2 repos~592 tokens
    DevelopmentAuto-check: notes
  • Conductor Setup

    gemini-cli-extensions/conductor

    Scaffolds the project and sets up the Conductor environment.

    3.8k GitHub stars~4.2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Builds or extends a custom Mirage virtual filesystem adapter for an API, database, object store or app data, with a working mount configuration and filesystem tests.

    3.7k GitHub stars~2.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Enforces this repository's TypeScript backend module architecture under server/: feature folders, barrel exports, and where shared types and utilities belong.

    14k GitHub stars~1.2k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from ruvnet/ruflo

All 264 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 3 repos~830 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 2 repos~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 2 repos~779 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Agent Coordination

    ruvnet/ruflo

    Reference for spawning, listing, monitoring and stopping agents with claude-flow commands, with agent type families, routing codes and coordination tips.

    74k GitHub starsUsed in 2 repos~519 tokens
    Auto-check passed

Categories

Questions about Harness Bench

What does Harness Bench do?

Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness. Harness Bench is an agent skill from ruvnet/ruflo.json checks suite well-formedness.

When should I use Harness Bench?

Harness Bench fits situations like: tasks that involve Project scaffolding.

How do I install Harness Bench in Claude Code?

Run `npx skills add ruvnet/ruflo --skill harness-bench -a claude-code`. Or copy the skill folder (plugins/ruflo-metaharness/skills/harness-bench in ruvnet/ruflo) into .claude/skills/harness-bench in your project. Claude Code loads it when a task matches its description.

How do I install Harness Bench in Codex?

Run `npx skills add ruvnet/ruflo --skill harness-bench -a codex`. Or copy the skill folder (plugins/ruflo-metaharness/skills/harness-bench in ruvnet/ruflo) into .agents/skills/harness-bench in your project. Codex loads it when a task matches its description.

Can I use Harness Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill harness-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-bench, .gemini/skills/harness-bench, .github/skills/harness-bench and .opencode/skills/harness-bench in your project.

What does Harness Bench need to run?

Going by SKILL.md and its folder, Harness Bench needs the command-line tools its instructions call (npm). Its frontmatter pre-approves these tools: Bash.

Does Harness Bench access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Harness Bench safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Harness Bench use?

Harness Bench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Bench use?

About 586 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Harness Bench?

Skills that share tags, products or a category with Harness Bench: Nx Generate (nomcopter/react-mosaic, 4.8k stars), Ponytail (DavidObando/gsharp, 565 stars), Run Nx Generator (nrwl/nx, 29k stars) and Conductor Setup (gemini-cli-extensions/conductor, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Bench?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,089 GitHub stars. The repository holds 264 skills in this directory. The repository was last updated on October 8, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.