Harness Bench
ruvnet/ruflo
Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness.
Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis.
$ npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Gentleman-Programming/gentle-ai gentle-ai-bench --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Gentleman-Programming/gentle-ai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gentle-ai-bench .claude/skills/gentle-ai-bench && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gentle-ai-bench" agent skill from https://github.com/Gentleman-Programming/gentle-ai/tree/main/skills/gentle-ai-bench into .claude/skills/gentle-ai-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gentle-ai-bench", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Gentleman-Programming/gentle-ai/tree/main/skills/gentle-ai-benchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Gentleman-Programming/gentle-ai gentle-ai-bench --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Gentleman-Programming/gentle-ai.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gentle-ai-bench .agents/skills/gentle-ai-bench && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gentle-ai-bench" agent skill from https://github.com/Gentleman-Programming/gentle-ai/tree/main/skills/gentle-ai-bench into .agents/skills/gentle-ai-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gentle-ai-bench", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Gentleman-Programming/gentle-ai gentle-ai-bench --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Gentleman-Programming/gentle-ai.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gentle-ai-bench .cursor/skills/gentle-ai-bench && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gentle-ai-bench" agent skill from https://github.com/Gentleman-Programming/gentle-ai/tree/main/skills/gentle-ai-bench into .cursor/skills/gentle-ai-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gentle-ai-bench", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Gentleman-Programming/gentle-ai.git --path skills/gentle-ai-bench--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Gentleman-Programming/gentle-ai gentle-ai-bench --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Gentleman-Programming/gentle-ai.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gentle-ai-bench .gemini/skills/gentle-ai-bench && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gentle-ai-bench" agent skill from https://github.com/Gentleman-Programming/gentle-ai/tree/main/skills/gentle-ai-bench into .gemini/skills/gentle-ai-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gentle-ai-bench", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Gentleman-Programming/gentle-ai gentle-ai-benchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Gentleman-Programming/gentle-ai.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gentle-ai-bench .github/skills/gentle-ai-bench && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gentle-ai-bench" agent skill from https://github.com/Gentleman-Programming/gentle-ai/tree/main/skills/gentle-ai-bench into .github/skills/gentle-ai-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gentle-ai-bench", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Gentleman-Programming/gentle-ai gentle-ai-bench --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Gentleman-Programming/gentle-ai.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gentle-ai-bench .opencode/skills/gentle-ai-bench && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gentle-ai-bench" agent skill from https://github.com/Gentleman-Programming/gentle-ai/tree/main/skills/gentle-ai-bench into .opencode/skills/gentle-ai-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gentle-ai-bench", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gentle-ai-benchTrigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis.
Gentle AI Bench is an agent skill from Gentleman-Programming/gentle-ai. Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Author and verify gentle-ai bench journeys; go test ./bench never proves driven execution.
Its SKILL.md is about 800 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Gentle-AI configures the AI coding agents you already use: Claude Code, Cursor, OpenCode, Codex, Pi, and more. Choose persistent memory, Organic-Driven Development, curated… The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 417c7a2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
goFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gentle AI Bench loads about 795 tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 442 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Gentleman-Programming/gentle-ai at commit 417c7a2, republished under its Apache-2.0 licence (© Gentleman-Programming). 442 words, ~795 tokens.
.claude/skills/gentle-ai-bench/SKILL.md (or your agent's skills folder).Load when touching bench/ in gentle-ai, adding or changing a journey, changing a product semantic a journey might pin, or diagnosing a bench failure in CI's Unit Tests job.
go test ./bench validates corpus declarations only. It does NOT execute journeys. The only driven proof is building the harness and the product binary and running the harness against it; a green go test ./bench claims nothing about execution..github/workflows/ci.yml and copy its exact build and gentle-ai-bench run --binary ... commands. Use --only <journey-id> to drive one journey.journeys_*.go file. The collision guard fails loudly naming both files; pick an unused ID by reading the corpus, never reuse a retired one.Review: — reviewOptedIn (the runner enables receipt-driven development globally before the first step, uncounted, and fails the journey if the switch does not come on) or reviewUntouched (its subject IS the switch, or it has nothing to do with reviews). The declaration is mandatory; validateCorpus fails the run without it. Lifecycle journeys must not depend on the product default. Reviews default to ON; reviewUntouched does not imply OFF. Journeys requiring OFF must explicitly disable it, while default-mode journeys must assert ON/default with unset sources.execute transition must carry a runnable command; the dead-execute guard fails the run otherwise.dead_end prints n/a unless the run actually measured one. Never fabricate a value to move the column.by_design exemption costs a shape from the closed vocabulary plus a verified quote of the product's own next-action text. If the quote no longer tells the operator what to do, it is a defect wearing an exemption.journeys_*.go file when the shared ones are owned by open PRs; bump the core journey-count pin in the same change.go test ./... in bench/ for declarations, THEN the driven harness for execution; both results go in the PR body.PR evidence includes the driven-mode summary line (completed / unsupported / failed counts) from a locally built binary, not only go test output.
© Gentleman-Programming, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/gentle-ai-bench of Gentleman-Programming/gentle-ai.
Open the folder on GitHubat commit 417c7a2
Gentle AI Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gentle AI Bench this skillGentleman-Programming/gentle-ai | 7.6k | — | ~795 | Automated safety check: Pass | Apache-2.0 | |
| Harness Benchruvnet/ruflo | 74k | — | ~586 | Automated safety check: Notes | MIT | |
| Sandbox Benchvercel/next.js | 143k | — | ~4.1k | Automated safety check: Pass | MIT | |
| Customer Journey Mapphuryn/pm-skills | 27k | — | ~816 | Automated safety check: Pass | MIT | |
| Bench Readgithub/awesome-copilot | 40k | — | ~747 | Automated safety check: Pass | MIT | |
| Benchddalcu/mlx-serve | 1.8k | 1 repos | ~1.1k | Automated safety check: Pass | Custom licence |
ruvnet/ruflo
Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness.
vercel/next.js
Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…
phuryn/pm-skills
Create an end-to-end customer journey map with stages, touchpoints, emotions, pain points, and opportunities.
github/awesome-copilot
Read artifacts from the shared bench — the workspace where desks leave findings, verdicts, and work products for each other and the operator.
ddalcu/mlx-serve
mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules (same-methodology cells only, spec-decode variance, thermal lies, engine naming), perf-claim etiquette.
paperclipai/paperclip
Run one Terminal-Bench task through a bounded Paperclip smoke/diagnosis/fix loop.
Gentleman-Programming/gentle-ai
Drafts, creates, comments on and approves GitHub issues under strict rules: YAML Issue Forms, a duplicate search first, and guarded protected labels.
Gentleman-Programming/gentle-ai
Guides writing a new agent skill: when one is warranted, the required folder layout and frontmatter, section order and size limits, with a bundled style guide.
Gentleman-Programming/gentle-ai
Prepares pull requests for the Gentle AI project under an issue-first rule: a linked approved issue, one type label, a valid branch name and confirmed required CI.
Gentleman-Programming/gentle-ai
Plans commits and PRs as reviewable work units, keeping tests and docs with the code they cover and splitting large changes into chained PRs.
Gentleman-Programming/gentle-ai
Trigger: Go tests, go test coverage, Bubbletea teatest, golden files.
Gentleman-Programming/gentle-ai
Trigger: a mapping need, parallel units, context backstop, high-risk verify, fresh review, or multi-step debug.
Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Gentle AI Bench is an agent skill from Gentleman-Programming/gentle-ai. Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis.
Run `npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a claude-code`. Or copy the skill folder (skills/gentle-ai-bench in Gentleman-Programming/gentle-ai) into .claude/skills/gentle-ai-bench in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a codex`. Or copy the skill folder (skills/gentle-ai-bench in Gentleman-Programming/gentle-ai) into .agents/skills/gentle-ai-bench in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gentle-ai-bench, .gemini/skills/gentle-ai-bench, .github/skills/gentle-ai-bench and .opencode/skills/gentle-ai-bench in your project.
Going by SKILL.md and its folder, Gentle AI Bench needs the command-line tools its instructions call (go).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gentle AI Bench is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 795 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gentle AI Bench: Harness Bench (ruvnet/ruflo, 74k stars), Sandbox Bench (vercel/next.js, 143k stars), Customer Journey Map (phuryn/pm-skills, 27k stars) and Bench Read (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Gentleman-Programming (a GitHub organization) maintains it in Gentleman-Programming/gentle-ai, which has 7,622 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.
Source: Gentleman-Programming/gentle-ai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.