Agent skill

Harness Security Bench

by ruvnet in ruvnet/ruflo

Run @metaharness/darwin security bench (upstream "Darwin Shield" / ADR-155) — evolves a champion security-detection harness against a 10-vuln / 9-decoy corpus and grades it on…

MITAuto-check: notesDevelopment

Install Harness Security Bench

skills CLI
$ npx skills add ruvnet/ruflo --skill harness-security-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo harness-security-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-metaharness/skills/harness-security-bench .claude/skills/harness-security-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-security-bench
GitHub stars
74k
Token cost
~1.1k tokens
SKILL.md length
357 words
Files
1
Skills in repo
264
Repo updated
First seen
Licence
MIT

At a glance

Run @metaharness/darwin security bench (upstream "Darwin Shield" / ADR-155) — evolves a champion security-detection harness against a 10-vuln / 9-decoy corpus and grades it on…

  • Works in 5 steps: Run metaharness-darwin security bench… → Default timeout = 3s × 19 evaluations ×… → Parse the markdown report — overall… → …
  • Tasks that involve Architecture decision records
  • SKILL.md covers Why this matters for ruflo's…, Algorithm, Output shape and Wiring into ADR-155 nightly…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Harness Security Bench is an agent skill from ruvnet/ruflo. Run @metaharness/darwin security bench (upstream "Darwin Shield" / ADR-155) — evolves a champion security-detection harness against a 10-vuln / 9-decoy corpus and grades it on TPR/FPR/patch-pass/repro/unsafe vs four baselines (B0 static, B1 LLM-single-pass, B2 fixed-agent, B3 Darwin-champion). Closest reference implementation for ruflo's own ADR-155 nightly self-learning security harness (PR

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Architecture decision records. The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

When your agent uses it

  • Tasks that involve Architecture decision records

Example prompts

  • “Darwin Shield”
  • “/harness-security-bench”

Requirements

  • Pre-approved tools (allowed-tools): Bash

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Run metaharness-darwin security bench --population N --cycles N [--seed S] from the
  2. Default timeout = 3s × 19 evaluations × population × cycles + 30s overhead.
  3. Parse the markdown report — overall PASS/FAIL plus per-gate
  4. Parse the baselines-vs-champion table (4 rows: fitness/TPR/FPR/patchPass/
  5. Emit structured JSON. With --alert-on-fail, exit 1 when overall = FAIL.

What it can do on your machine

Read from SKILL.md and the folder at commit 58e0ae7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json and jsonc).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Security Bench loads about 1.1k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 357 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit 58e0ae7, republished under its MIT licence (© ruvnet). 357 words, ~1,133 tokens.

Download SKILL.mdSave it as .claude/skills/harness-security-bench/SKILL.md (or your agent's skills folder).
name
harness-security-bench
description
Run `@metaharness/darwin security bench` (upstream "Darwin Shield" / ADR-155) — evolves a champion security-detection harness against a 10-vuln / 9-decoy corpus and grades it on TPR/FPR/patch-pass/repro/unsafe vs four baselines (B0 static, B1 LLM-single-pass, B2 fixed-agent, B3 Darwin-champion). Closest reference implementation for ruflo's own ADR-155 nightly self-learning security harness (PR
allowed-tools
Bash
argument-hint
[--population 2] [--cycles 1] [--seed N] [--alert-on-fail]

Surfaces the upstream metaharness-darwin security bench command. This is the upstream's own ADR-155 — Darwin Shield — and is the closest reference implementation for ruflo's nightly self-learning security harness (#2417).

Why this matters for ruflo's ADR-155

ruflo's ADR-155 proposes three learning loops (per-dimension confidence, severity calibration, auto-fix bid). Loop A trains on accumulated (finding, dimension, human_outcome) tuples — but the gradient signal is only sound if the underlying detection mechanism converges on a known-good corpus. Darwin Shield evolves exactly that mechanism on a 10-vuln/9-decoy ground-truth set. Running this nightly gives us:

  • Empirical floor: if Darwin Shield's champion can't reach TPR=1/FPR=0 on the bench corpus, our Loop A's reward signal is noise.
  • Drift detection: week-over-week champion fitness deltas surface when the security landscape (or our mutator policy) shifts.
  • Baseline diversity: the 4 baselines (B0–B3) give us 4 anchor points to weight per-dimension confidence against.

Algorithm

Implementation: scripts/security-bench.mjs.

  1. Run metaharness-darwin security bench --population N --cycles N [--seed S] from the installed @metaharness/darwin (pin in _darwin.mjs; one-time cache install only if no installed copy qualifies — never npx).
  2. Default timeout = 3s × 19 evaluations × population × cycles + 30s overhead. At default --population 2 --cycles 1 ≈ 144s; at --population 4 --cycles 3 ≈ 12 min.
  3. Parse the markdown report — overall PASS/FAIL plus per-gate pass/fail rows (gate examples: "TPR improvement ≥ 25% vs fixed", "FPR reduction ≥ 40%", "Patch-test pass rate ≥ 80%", "Reproduction success ≥ 90%", "Unsafe outputs = 0", "Cost increase ≤ 2× fixed", "Beyond SOTA: champion statistically beats previous champion", "Compounding: false-positive repeat-rate drop ≥ 35%").
  4. Parse the baselines-vs-champion table (4 rows: fitness/TPR/FPR/patchPass/ repro/unsafe/cost per harness).
  5. Emit structured JSON. With --alert-on-fail, exit 1 when overall = FAIL.
Show full SKILL.md (92 more words)Show less

Output shape

json
{
  "success": true,
  "data": {
    "overall": { "ok": true, "icon": "✅" },
    "gates": {
      "total": 11,
      "passed": 11,
      "failed": 0,
      "details": [{ "ok": true, "criterion": "TPR improvement ≥ 25% vs fixed harness", "measured": "+150% (B2 0.4 → B3 1)" }, ...]
    },
    "baselines": [
      { "harness": "static-only", "fitness": 0.5665, "tpr": 0.3, "fpr": 1, "unsafe": 0, ... },
      { "harness": "LLM single-pass", "fitness": 0.1365, ... },
      { "harness": "fixed agent", "fitness": 0.598, ... },
      { "harness": "Darwin champion", "fitness": 0.93275, "tpr": 1, "fpr": 0, ... }
    ],
    "rawMarkdown": "...",
    "shape": { "population": 2, "cycles": 1, "seed": null },
    "durationMs": 142000
  }
}

Wiring into ADR-155 nightly harness

The ADR-155 nightly workflow (per #2418 task W1.5) will spawn this as one of the active-pentest dimension's calls — its results become a trajectory record:

jsonc
{
  "dimension": "mcp-pentest",
  "subdimension": "darwin-shield-bench",
  "champion_fitness": 0.93275,
  "champion_tpr": 1, "champion_fpr": 0,
  "gates_passed": 11, "gates_failed": 0,
  "shape": { "population": 4, "cycles": 3 }
}

Loop A learns: if darwin-shield-bench consistently passes on the seeded corpus, weight findings caught only by mcp-pentest higher.

Exit codes

CodeMeaning
0Bench ran (overall PASS or FAIL — distinguish via JSON overall.ok), or degraded
1--alert-on-fail and overall.ok === false
2Config error or upstream infrastructure failure

Graceful degradation

When @metaharness/darwin is absent, emits {degraded: true, reason: 'metaharness-darwin-not-available'} and exits 0.

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-metaharness/skills/harness-security-bench of ruvnet/ruflo.

Open the folder on GitHubat commit 58e0ae7

Compare with similar skills

Harness Security Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Security Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Security Bench this skillruvnet/ruflo74k—~1.1kAutomated safety check: NotesMIT
PR Design DocOpenHands/OpenHands90k—~2.4kAutomated safety check: PassMIT
Cto AdvisorIbrahim-3d/orchestrator-supaconductor3814 repos~2.4kAutomated safety check: PassMIT
Architecture DecisionDonchitos/Claude-Code-Game-Studios26k—~1.7kAutomated safety check: PassMIT
Improve Codebase Architectureywwynm/EverythingDone14415 repos~1.3kAutomated safety check: PassGPL-3.0
Domain Modelingbrim-borium/spotify_sdk1665 repos~806Automated safety check: PassApache-2.0

Similar skills

  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Cto Advisor

    Ibrahim-3d/orchestrator-supaconductor

    Technical leadership guidance for engineering teams, architecture decisions, and technology strategy.

    381 GitHub starsUsed in 4 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Architecture Decision

    Donchitos/Claude-Code-Game-Studios

    Create an ADR documenting a technical decision: context, alternatives considered, consequences.

    26k GitHub stars~1.7k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Improve Codebase Architecture

    ywwynm/EverythingDone

    Find deepening opportunities in a codebase, informed by the domain language in CONTEXT.md and the decisions in docs/adr/.

    144 GitHub starsUsed in 15 repos~1.3k tokens
    DevelopmentAuto-check passed
  • Domain Modeling

    brim-borium/spotify_sdk

    Build and sharpen a project's domain model. An agent skill from brim-borium/spotify_sdk.

    166 GitHub starsUsed in 5 repos~806 tokens
    DevelopmentAuto-check passed
  • Design Doc Mermaid

    SpillwaveSolutions/design-doc-mermaid

    Create Mermaid diagrams (flowchart, sequence, class, ER, state, C4, architecture) from text or source code.

    176 GitHub starsUsed in 1 repo~5.6k tokens
    DevelopmentAuto-check passed

More from ruvnet/ruflo

All 264 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 2 repos~830 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 2 repos~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 2 repos~779 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Agent Coordination

    ruvnet/ruflo

    Reference for spawning, listing, monitoring and stopping agents with claude-flow commands, with agent type families, routing codes and coordination tips.

    74k GitHub starsUsed in 2 repos~519 tokens
    Auto-check passed

Categories

Questions about Harness Security Bench

What does Harness Security Bench do?

Run @metaharness/darwin security bench (upstream "Darwin Shield" / ADR-155) — evolves a champion security-detection harness against a 10-vuln / 9-decoy corpus and grades it on…. Harness Security Bench is an agent skill from ruvnet/ruflo. Run @metaharness/darwin security bench (upstream "Darwin Shield" / ADR-155) — evolves a champion security-detection harness against a 10-vuln / 9-decoy corpus and grades it on TPR/FPR/patch-pass/repro/unsafe vs four baselines (B0 static, B1 LLM-single-pass, B2 fixed-agent, B3 Darwin-champion).

When should I use Harness Security Bench?

Harness Security Bench fits situations like: tasks that involve Architecture decision records.

How do I install Harness Security Bench in Claude Code?

Run `npx skills add ruvnet/ruflo --skill harness-security-bench -a claude-code`. Or copy the skill folder (plugins/ruflo-metaharness/skills/harness-security-bench in ruvnet/ruflo) into .claude/skills/harness-security-bench in your project. Claude Code loads it when a task matches its description.

How do I install Harness Security Bench in Codex?

Run `npx skills add ruvnet/ruflo --skill harness-security-bench -a codex`. Or copy the skill folder (plugins/ruflo-metaharness/skills/harness-security-bench in ruvnet/ruflo) into .agents/skills/harness-security-bench in your project. Codex loads it when a task matches its description.

Can I use Harness Security Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill harness-security-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-security-bench, .gemini/skills/harness-security-bench, .github/skills/harness-security-bench and .opencode/skills/harness-security-bench in your project.

What does Harness Security Bench need to run?

SKILL.md names no scripts, command-line tools or credentials: Harness Security Bench is instructions for the agent only. Its frontmatter pre-approves these tools: Bash.

Does Harness Security Bench access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harness Security Bench safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Harness Security Bench use?

Harness Security Bench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Security Bench use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Harness Security Bench?

Skills that share tags, products or a category with Harness Security Bench: PR Design Doc (OpenHands/OpenHands, 90k stars), Cto Advisor (Ibrahim-3d/orchestrator-supaconductor, 381 stars), Architecture Decision (Donchitos/Claude-Code-Game-Studios, 26k stars) and Improve Codebase Architecture (ywwynm/EverythingDone, 144 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Security Bench?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,159 GitHub stars. The repository holds 264 skills in this directory. The repository was last updated on October 9, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.