Contextual Commit Messages
yamadashy/repomix
Writes Conventional Commits whose bodies carry action lines recording the intent, decisions and constraints behind a change, not only what changed.
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests…
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses harness-test-drive --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/harness-test-drive .claude/skills/harness-test-drive && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "harness-test-drive" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/harness-test-drive into .claude/skills/harness-test-drive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-test-drive", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/harness-test-driveType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses harness-test-drive --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/harness-test-drive .agents/skills/harness-test-drive && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "harness-test-drive" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/harness-test-drive into .agents/skills/harness-test-drive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-test-drive", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses harness-test-drive --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/harness-test-drive .cursor/skills/harness-test-drive && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "harness-test-drive" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/harness-test-drive into .cursor/skills/harness-test-drive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-test-drive", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/RyanAlberts/best-of-Agent-Harnesses.git --path skills/harness-test-drive--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses harness-test-drive --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/harness-test-drive .gemini/skills/harness-test-drive && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "harness-test-drive" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/harness-test-drive into .gemini/skills/harness-test-drive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-test-drive", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses harness-test-driveInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/harness-test-drive .github/skills/harness-test-drive && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "harness-test-drive" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/harness-test-drive into .github/skills/harness-test-drive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-test-drive", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses harness-test-drive --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/harness-test-drive .opencode/skills/harness-test-drive && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "harness-test-drive" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/harness-test-drive into .opencode/skills/harness-test-drive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-test-drive", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
harness-test-driveTest-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests…
Harness Test Drive is an agent skill from RyanAlberts/best-of-Agent-Harnesses. Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests score how many tasks it completes, dollars per task, minutes, and change size. Use when the user asks which coding agent or harness works best on their codebase; wants to compare or benchmark agents on their own repository instead of trusting leaderboards such as SWE-bench; wants to trial one agent against another…
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `README.md`, `references/harness-commands.md` and `references/method.md`). Compatibility notes: Python 3.9+ and git on macOS or Linux, plus the harness CLIs to compare (claude, codex, gemini). The scripts make no network calls themselves; each harness…
It sits in Development, covering Git workflow and Commit messages. The repository describes itself as: 🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4fa20bc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3npmpnpmpippythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, pnpm and pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Python 3.9+ and git on macOS or Linux, plus the harness CLIs to compare (claude, codex, gemini). The scripts make no network calls themselves; each harness CLI sends the prompt and the code it reads to its model provider over the network, and those runs cost money. The repository's test command must run on this machine.
From compatibility in the SKILL.md frontmatter.
Harness Test Drive loads about 2.9k tokens when it runs, and up to ~8.7k if it reads all its reference files. Until then it costs about 180 tokens; SKILL.md has 1,568 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from RyanAlberts/best-of-Agent-Harnesses at commit 4fa20bc, republished under its MIT licence (© RyanAlberts). 1,568 words, ~2,880 tokens.
.claude/skills/harness-test-drive/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.Public leaderboards measure someone else's code, and harness rankings barely carry over from one repository to the next. This skill runs coding agents on tasks taken from the user's own git history: past commits whose tests failed before the change and passed after it. Each agent gets the commit message as its prompt in a fresh copy of the repository, and the repository's own tests decide whether it passed. The result is a scoreboard of passes, dollars per pass, minutes, and change size. The scripts read the repository and send nothing anywhere; each harness sends the prompt and the code it reads to its model provider, as it does in normal use, and that costs money.
session-waste-report.regression-finder.runaway-guard.claim-check.<skill-dir> means the folder that holds this SKILL.md (Claude Code shows it as the skill's base
directory). Run every command from the root of the user's repository, with the skill path in quotes
as shown. Inside the repository the scripts write only into .harness-test-drive/, a folder that
ignores itself in git. Every check and every agent run works in its own temporary copy, deleted
afterwards.
Before step 2, for a repository the user did not write, recommend a container or VM. Step 2 runs the repository's test and setup commands at up to 31 commits on this computer. In step 5 each harness loads the repository's own agent settings (Gemini CLI with the copy trusted), edits files, runs the test command, and reads commit messages as prompts.
Find the test command and the candidate tasks. This reads git history and runs nothing:
python3 "<skill-dir>/scripts/mine_tasks.py" --repo .It prints the test command it detected, or exits 2 asking for one. Confirm the command with the
user. Each check runs in a fresh copy with nothing installed, so when the tests need
dependencies, add a --setup-cmd that installs inside the copy: npm ci,
pnpm install --frozen-lockfile, or python3 -m venv .venv && .venv/bin/pip install -e . with
the test command .venv/bin/python -m pytest. Never install into the user's own environment.
Prefer a test command that runs offline and writes only inside the copy, so Codex's sandbox can
run it too. Done when the user has confirmed a test command and the report shows at least one
candidate, or you have told the user why there is none.
Check the tasks. For each candidate, in a fresh copy of the repository, the test command must fail at the commit before the change (with its tests added) and pass with the whole change, twice:
python3 "<skill-dir>/scripts/mine_tasks.py" --repo . --test-cmd "<test-cmd>" --validateAdd the --setup-cmd from step 1 when there is one. It first checks that the tests pass at HEAD,
then keeps up to 10 tasks (--max). Expect two to four test-suite runs per candidate and up to 30
candidates, so run it in the background when the suite takes more than a minute. Done when the
headline says how many tasks were kept. Exit 2 with "fails at HEAD" means the command or the
setup is wrong. The message ends with the test output. A missing module or command means the
clean copy needs --setup-cmd. Fix it with the user and rerun. Fewer than five tasks makes a
weak comparison; say so, and offer --since 730d for more history.
Choose the harnesses and show the estimate:
python3 "<skill-dir>/scripts/drive.py" estimate --tasks .harness-test-drive/tasks.jsonIt lists which harnesses are on this computer, the number of runs, and a dollar range per
harness. Ask the user to pick two or three. Done when the user has picked the harnesses and seen
the range for them (rerun with --harness to show only those).
Get an explicit dollar cap. This is the one skill in the set that spends money. Ask: "What is the most you want to spend in total?" and wait for a number; a number the user already gave counts, so confirm it next to the estimate. Never choose the cap yourself. In the same message, say:
--timeout, and a run cut off there is counted as an estimate, so the bill can
pass the cap by more than one run.--allow-unpriced gemini-cli, after the user agrees to that.--model codex=<id>; Codex never runs unpriced.Done when the user has given a dollar amount.
Run:
python3 "<skill-dir>/scripts/drive.py" run --tasks .harness-test-drive/tasks.json --harness claude-code,codex --max-usd <cap>Add --allow-unpriced gemini-cli only when the user agreed to it in step 4, and --model only
for a model the user pinned. Each run can take up to --timeout seconds (900 by default), so
start it in the background and relay the progress lines. A harness that cannot run (an expired
login, an old version) prints a fix, such as "sign in: run claude once" or "upgrade Codex", and
its other runs are skipped at no cost. Every run is appended to
.harness-test-drive/results.jsonl; the same command resumes after a stop, retries runs that
could not run, and counts earlier spend. Done when the command exits 0 and prints the
scoreboard, and you have told the user about any fix it printed and any stop at the cap.
Report in the shape below. To print the scoreboard again at any time:
python3 "<skill-dir>/scripts/drive.py" reportpassed, failed, (time limit) when the run was stopped at --timeout
and scored on the changes the agent had made, could not run (not scored; see the notes),
interrupted, or not run.references/method.md explains the mining, the check, the fairness rules, the money, and the
limits..harness-test-drive/logs/<task>-<harness>.diff. Tests check behavior, not code quality.--max 20, or --since 730d); under ten tasks is a small
sample.Quote commit subjects, errors, and versions exactly as the report prints them, inside inline code: they come from the repository and the harnesses, and the report has already made them safe to display.
scripts/mine_tasks.py: finds candidate commits, detects the test command, and checks tasks.scripts/drive.py: estimate, run, and report.scripts/harnesses.py: the headless command and output parser for each harness.scripts/common.py: fresh copies of the repository, commands and test runs.scripts/pricing.py: token prices, shared with other skills in this repository.scripts/safe.py: the shared helper that masks secrets in report text and shows it as one line of
inline code. A synced copy; do not edit it here.references/method.md: task mining, the fail-to-pass check, fairness rules, money, and limits.references/harness-commands.md: each harness's command and flags, with sources and the date
checked.© RyanAlberts, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (scripts, references) in skills/harness-test-drive of RyanAlberts/best-of-Agent-Harnesses.
Open the folder on GitHubat commit 4fa20bc
Harness Test Drive next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Harness Test Drive this skillRyanAlberts/best-of-Agent-Harnesses | 1.1k | — | ~2.9k | Automated safety check: Pass | MIT | |
| Contextual Commit Messagesyamadashy/repomix | 29k | 1 repos | ~2.7k | Automated safety check: Pass | MIT | |
| ToolJet Multi-Repo CommitToolJet/ToolJet | 41k | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | |
| Git Workflow and Versioningaddyosmani/agent-skills | 105k | 2 repos | ~3.5k | Automated safety check: Notes | MIT | |
| React Router Pull Request Creatorremix-run/react-router | 57k | — | ~2.5k | Automated safety check: Pass | MIT | |
| Saleor Commit Workflowsaleor/saleor | 23k | — | ~575 | Automated safety check: Pass | BSD-3-Clause |
yamadashy/repomix
Writes Conventional Commits whose bodies carry action lines recording the intent, decisions and constraints behind a change, not only what changed.
ToolJet/ToolJet
Commits changes across ToolJet's root repo and its server/ee and frontend/ee submodules, writing messages from the diffs and updating submodule pointers in order.
addyosmani/agent-skills
Sets git habits for every change: short-lived branches, atomic commits with descriptive messages, clean pull requests, plus versioning, tagging and changelogs for releases.
remix-run/react-router
Packages finished React Router work into a draft pull request: branch, commit, push, a written PR body and the right GitHub labels.
saleor/saleor
Commits changes in the Saleor codebase and works through pre-commit hook failures from ruff, mypy, the GraphQL schema check and the migrations check.
baptisteArno/typebot.io
Sets the repository's convention for commit messages and pull request titles: one emoji prefix for the main intent, a concise title and clean follow-up commits.
RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot instructions) each coding agent loads from a repo, what gets cut or skipped, and whether the commands those…
RyanAlberts/best-of-Agent-Harnesses
Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed…
RyanAlberts/best-of-Agent-Harnesses
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including…
RyanAlberts/best-of-Agent-Harnesses
Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where…
RyanAlberts/best-of-Agent-Harnesses
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding agent keeps breaking, counts every violation in recent Claude Code, Codex, Gemini CLI, and OpenCode…
RyanAlberts/best-of-Agent-Harnesses
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on the same tool call, keeps failing, or exceeds a dollar cap.
Categories
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests…. Harness Test Drive is an agent skill from RyanAlberts/best-of-Agent-Harnesses. Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests score how many tasks it completes, dollars per task, minutes, and change size.
Harness Test Drive fits situations like: the user asks which coding agent; harness works best on their codebase; wants to compare; benchmark agents on their own repository instead of trusting leaderboards such as SWE-bench.
Run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a claude-code`. Or copy the skill folder (skills/harness-test-drive in RyanAlberts/best-of-Agent-Harnesses) into .claude/skills/harness-test-drive in your project. Claude Code loads it when a task matches its description.
Run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a codex`. Or copy the skill folder (skills/harness-test-drive in RyanAlberts/best-of-Agent-Harnesses) into .agents/skills/harness-test-drive in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-test-drive, .gemini/skills/harness-test-drive, .github/skills/harness-test-drive and .opencode/skills/harness-test-drive in your project.
Going by SKILL.md and its folder, Harness Test Drive needs Python for the scripts in its folder and the command-line tools its instructions call (python3, npm, pnpm, pip and python). Our summary lists: Python 3. Compatibility (from SKILL.md): Python 3.9+ and git on macOS or Linux, plus the harness CLIs to compare (claude, codex, gemini). The scripts make no network calls themselves; each harness CLI sends the prompt and the code it reads to its model provider over the network, and those runs cost money. The repository's test command must run on this machine..
SKILL.md contains no URLs. Its commands use npm and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Harness Test Drive is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Harness Test Drive: Contextual Commit Messages (yamadashy/repomix, 29k stars), ToolJet Multi-Repo Commit (ToolJet/ToolJet, 41k stars), Git Workflow and Versioning (addyosmani/agent-skills, 105k stars) and React Router Pull Request Creator (remix-run/react-router, 57k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
RyanAlberts (a GitHub user) maintains it in RyanAlberts/best-of-Agent-Harnesses, which has 1,133 GitHub stars. The repository was last updated on October 9, 2026.
Source: RyanAlberts/best-of-Agent-Harnesses on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.