Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis.

Apache-2.0Auto-check passed

Install Gentle AI Bench

skills CLI
$ npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Gentleman-Programming/gentle-ai gentle-ai-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Gentleman-Programming/gentle-ai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gentle-ai-bench .claude/skills/gentle-ai-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gentle-ai-bench
GitHub stars
7.6k
Token cost
~795 tokens
SKILL.md length
442 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
Apache-2.0

At a glance

Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis.

  • Works in 4 steps: Read the corpus area you touch and the… → Author or adapt the journey; update its… → Run go test ./... in bench/ for… → …
  • SKILL.md covers Activation Contract, Hard Rules, Execution Steps and Output Contract
  • Calls go

What it does

Gentle AI Bench is an agent skill from Gentleman-Programming/gentle-ai. Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Author and verify gentle-ai bench journeys; go test ./bench never proves driven execution.

Its SKILL.md is about 800 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Gentle-AI configures the AI coding agents you already use: Claude Code, Cursor, OpenCode, Codex, Pi, and more. Choose persistent memory, Organic-Driven Development, curated… The licence is Apache-2.0.

Example prompts

  • “/gentle-ai-bench”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read the corpus area you touch and the CI invocation before writing.
  2. Author or adapt the journey; update its title, step names, and comment to say WHY the expectation holds (cite the issue or ratified…
  3. Run go test ./... in bench/ for declarations, THEN the driven harness for execution; both results go in the PR body.
  4. On semantic changes, list the journeys you checked for stale pins.

What it can do on your machine

Read from SKILL.md and the folder at commit 417c7a2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • go

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gentle AI Bench loads about 795 tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 442 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~795

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Gentleman-Programming/gentle-ai at commit 417c7a2, republished under its Apache-2.0 licence (© Gentleman-Programming). 442 words, ~795 tokens.

Download SKILL.mdSave it as .claude/skills/gentle-ai-bench/SKILL.md (or your agent's skills folder).
name
gentle-ai-bench
description
Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Author and verify gentle-ai bench journeys; go test ./bench never proves driven execution.
license
Apache-2.0
metadata.author
Gentleman-Programming
metadata.version
1.0

Activation Contract

Load when touching bench/ in gentle-ai, adding or changing a journey, changing a product semantic a journey might pin, or diagnosing a bench failure in CI's Unit Tests job.

Hard Rules

  • go test ./bench validates corpus declarations only. It does NOT execute journeys. The only driven proof is building the harness and the product binary and running the harness against it; a green go test ./bench claims nothing about execution.
  • Reproduce CI, do not guess invocations: read the Unit Tests step in .github/workflows/ci.yml and copy its exact build and gentle-ai-bench run --binary ... commands. Use --only <journey-id> to drive one journey.
  • Journey IDs are unique across every journeys_*.go file. The collision guard fails loudly naming both files; pick an unused ID by reading the corpus, never reuse a retired one.
  • Every journey declares Review: — reviewOptedIn (the runner enables receipt-driven development globally before the first step, uncounted, and fails the journey if the switch does not come on) or reviewUntouched (its subject IS the switch, or it has nothing to do with reviews). The declaration is mandatory; validateCorpus fails the run without it. Lifecycle journeys must not depend on the product default. Reviews default to ON; reviewUntouched does not imply OFF. Journeys requiring OFF must explicitly disable it, while default-mode journeys must assert ON/default with unset sources.
  • Every execute transition must carry a runnable command; the dead-execute guard fails the run otherwise.
  • When a ratified product semantic changes, grep the corpus for journeys pinning the OLD behavior before shipping. The corpus is a second test surface beyond unit tests; a journey asserting the defect keeps the defect green.
  • dead_end prints n/a unless the run actually measured one. Never fabricate a value to move the column.
  • A by_design exemption costs a shape from the closed vocabulary plus a verified quote of the product's own next-action text. If the quote no longer tells the operator what to do, it is a defect wearing an exemption.
  • Prefer a NEW journeys_*.go file when the shared ones are owned by open PRs; bump the core journey-count pin in the same change.
Show full SKILL.md (96 more words)Show less

Execution Steps

  1. Read the corpus area you touch and the CI invocation before writing.
  2. Author or adapt the journey; update its title, step names, and comment to say WHY the expectation holds (cite the issue or ratified decision).
  3. Run go test ./... in bench/ for declarations, THEN the driven harness for execution; both results go in the PR body.
  4. On semantic changes, list the journeys you checked for stale pins.

Output Contract

PR evidence includes the driven-mode summary line (completed / unsupported / failed counts) from a locally built binary, not only go test output.

© Gentleman-Programming, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gentle-ai-bench of Gentleman-Programming/gentle-ai.

Open the folder on GitHubat commit 417c7a2

Compare with similar skills

Gentle AI Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gentle AI Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gentle AI Bench this skillGentleman-Programming/gentle-ai7.6k—~795Automated safety check: PassApache-2.0
Harness Benchruvnet/ruflo74k—~586Automated safety check: NotesMIT
Sandbox Benchvercel/next.js143k—~4.1kAutomated safety check: PassMIT
Customer Journey Mapphuryn/pm-skills27k—~816Automated safety check: PassMIT
Bench Readgithub/awesome-copilot40k—~747Automated safety check: PassMIT
Benchddalcu/mlx-serve1.8k1 repos~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Harness Bench

    ruvnet/ruflo

    Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness.

    74k GitHub stars~586 tokensUpdated today
    DevelopmentAuto-check: notes
  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Customer Journey Map

    phuryn/pm-skills

    Create an end-to-end customer journey map with stages, touchpoints, emotions, pain points, and opportunities.

    27k GitHub stars~816 tokensUpdated 25 days ago
    Product & Project ManagementAuto-check passed
  • Bench Read

    github/awesome-copilot

    Official

    Read artifacts from the shared bench — the workspace where desks leave findings, verdicts, and work products for each other and the operator.

    40k GitHub stars~747 tokensUpdated today
    Auto-check passed
  • Bench

    ddalcu/mlx-serve

    mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules (same-methodology cells only, spec-decode variance, thermal lies, engine naming), perf-claim etiquette.

    1.8k GitHub starsUsed in 1 repo~1.1k tokens
    Media & CreativeAuto-check passed
  • Terminal Bench Loop

    paperclipai/paperclip

    Run one Terminal-Bench task through a bounded Paperclip smoke/diagnosis/fix loop.

    99k GitHub stars~6.3k tokensUpdated today
    Auto-check passed

More from Gentleman-Programming/gentle-ai

All 15 skills in this repo
  • GitHub Issue Creation

    Gentleman-Programming/gentle-ai

    Drafts, creates, comments on and approves GitHub issues under strict rules: YAML Issue Forms, a duplicate search first, and guarded protected labels.

    7.6k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Skill Creator

    Gentleman-Programming/gentle-ai

    Guides writing a new agent skill: when one is warranted, the required folder layout and frontmatter, section order and size limits, with a bundled style guide.

    7.6k GitHub stars~998 tokensUpdated today
    Auto-check passed
  • Gentle AI Pull Requests

    Gentleman-Programming/gentle-ai

    Prepares pull requests for the Gentle AI project under an issue-first rule: a linked approved issue, one type label, a valid branch name and confirmed required CI.

    7.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Work-Unit Commits

    Gentleman-Programming/gentle-ai

    Plans commits and PRs as reviewable work units, keeping tests and docs with the code they cover and splitting large changes into chained PRs.

    7.6k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Go Testing

    Gentleman-Programming/gentle-ai

    Trigger: Go tests, go test coverage, Bubbletea teatest, golden files.

    7.6k GitHub stars~550 tokensUpdated today
    Auto-check passed
  • Hermes Ephemeral Delegation

    Gentleman-Programming/gentle-ai

    Trigger: a mapping need, parallel units, context backstop, high-risk verify, fresh review, or multi-step debug.

    7.6k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Gentle AI Bench

What does Gentle AI Bench do?

Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Gentle AI Bench is an agent skill from Gentleman-Programming/gentle-ai. Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis.

How do I install Gentle AI Bench in Claude Code?

Run `npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a claude-code`. Or copy the skill folder (skills/gentle-ai-bench in Gentleman-Programming/gentle-ai) into .claude/skills/gentle-ai-bench in your project. Claude Code loads it when a task matches its description.

How do I install Gentle AI Bench in Codex?

Run `npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a codex`. Or copy the skill folder (skills/gentle-ai-bench in Gentleman-Programming/gentle-ai) into .agents/skills/gentle-ai-bench in your project. Codex loads it when a task matches its description.

Can I use Gentle AI Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gentle-ai-bench, .gemini/skills/gentle-ai-bench, .github/skills/gentle-ai-bench and .opencode/skills/gentle-ai-bench in your project.

What does Gentle AI Bench need to run?

Going by SKILL.md and its folder, Gentle AI Bench needs the command-line tools its instructions call (go).

Does Gentle AI Bench access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gentle AI Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gentle AI Bench use?

Gentle AI Bench is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gentle AI Bench use?

About 795 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gentle AI Bench?

Skills that share tags, products or a category with Gentle AI Bench: Harness Bench (ruvnet/ruflo, 74k stars), Sandbox Bench (vercel/next.js, 143k stars), Customer Journey Map (phuryn/pm-skills, 27k stars) and Bench Read (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gentle AI Bench?

Gentleman-Programming (a GitHub organization) maintains it in Gentleman-Programming/gentle-ai, which has 7,622 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.

Source: Gentleman-Programming/gentle-ai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.