LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.
$ npx skills add borghei/Claude-Skills --skill agent-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install borghei/Claude-Skills agent-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/agent-harness .claude/skills/agent-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-harness" agent skill from https://github.com/borghei/Claude-Skills/tree/main/engineering/agent-harness into .claude/skills/agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/borghei/Claude-Skills/tree/main/engineering/agent-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add borghei/Claude-Skills --skill agent-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install borghei/Claude-Skills agent-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/engineering/agent-harness .agents/skills/agent-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-harness" agent skill from https://github.com/borghei/Claude-Skills/tree/main/engineering/agent-harness into .agents/skills/agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add borghei/Claude-Skills --skill agent-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install borghei/Claude-Skills agent-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/engineering/agent-harness .cursor/skills/agent-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-harness" agent skill from https://github.com/borghei/Claude-Skills/tree/main/engineering/agent-harness into .cursor/skills/agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/borghei/Claude-Skills.git --path engineering/agent-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add borghei/Claude-Skills --skill agent-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install borghei/Claude-Skills agent-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/engineering/agent-harness .gemini/skills/agent-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-harness" agent skill from https://github.com/borghei/Claude-Skills/tree/main/engineering/agent-harness into .gemini/skills/agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install borghei/Claude-Skills agent-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add borghei/Claude-Skills --skill agent-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/engineering/agent-harness .github/skills/agent-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-harness" agent skill from https://github.com/borghei/Claude-Skills/tree/main/engineering/agent-harness into .github/skills/agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add borghei/Claude-Skills --skill agent-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install borghei/Claude-Skills agent-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/engineering/agent-harness .opencode/skills/agent-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-harness" agent skill from https://github.com/borghei/Claude-Skills/tree/main/engineering/agent-harness into .opencode/skills/agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-harnessTest and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.
Agent Harness is an agent skill from borghei/Claude-Skills. Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets. Use when agent quality is vibe-checked, before shipping a prompt or model change, or when evals drift.
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts, reference files and assets (for example `assets/eval_report_template.md`, `assets/sample_baseline_report.json` and `assets/sample_candidate_report.json`).
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Harness loads about 3.1k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 1,507 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 1,507 words, ~3,068 tokens.
.claude/skills/agent-harness/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Most agents ship on vibes: someone tries eight prompts, the output looks good, it goes to production, and the next prompt tweak silently breaks a refusal nobody re-tested. This skill builds the harness around an agent so its behaviour becomes measurable — scenario suites with structural assertions, deterministic replay of recorded tool calls, paired regression diffing across prompt and model changes, and per-scenario cost and latency budgets. The tools here score an agent; they never invoke one, so they run offline on every commit.
Before building the harness, confirm these inputs. If any is unknown or vague, ASK — do not assume:
tool_not_called assertions at critical severity, and what the release gate blocks onStop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
assets/scenario_authoring_checklist.md. Structural assertions first — tool
called / not called / order / arguments — text assertions only on domain tokens.defaults for latency, cost, and turn ceilings so every
scenario is budgeted without repeating yourself.model and prompt_sha.python3 engineering/agent-harness/scripts/scenario_runner.py \
--suite engineering/agent-harness/assets/sample_suite.json \
--transcripts engineering/agent-harness/assets/sample_transcripts_baseline.json \
--strict-criticalassets/eval_report_template.md and promote the
accepted candidate report to the new baseline.python3 engineering/agent-harness/scripts/scenario_runner.py \
--suite engineering/agent-harness/assets/sample_suite.json \
--transcripts engineering/agent-harness/assets/sample_transcripts_candidate.json \
--format json > /tmp/candidate.report.json
python3 engineering/agent-harness/scripts/eval_diff.py \
--baseline engineering/agent-harness/assets/sample_baseline_report.json \
--candidate /tmp/candidate.report.json \
--fail-on-regression --drift-threshold 0.15The shipped sample data demonstrates the core lesson: both runs score 83.3%, and the candidate contains a critical prompt-injection regression. A gate on pass rate ships it; the paired diff catches it.
defaults, overriding
only where a scenario is legitimately expensive.minor assertions, so
they report without blocking.python3 engineering/agent-harness/scripts/eval_diff.py \
--baseline engineering/agent-harness/assets/sample_baseline_report.json \
--candidate engineering/agent-harness/assets/sample_candidate_report.json \
--drift-threshold 0.10 --format json| Need | Use | Durability |
|---|---|---|
| The agent must take an action | tool_called, tool_call_order | [PROVEN] Exact; survives rewording |
| The agent must NOT take an action | tool_not_called | [PROVEN] The single highest-value assertion in any agent suite |
| The action must use the right data | tool_arg_equals | [PROVEN] Catches the right tool with wrong arguments |
| Structured output correctness | json_field_equals | [PROVEN] Exact when the agent has a JSON mode |
| A required domain fact appears | output_contains on an ID, number, or policy name | [RECOMMENDED] Stable if you never quote sentences |
| A forbidden phrase must not appear | output_not_contains | [RECOMMENDED] Good for injection and leak checks |
| Tone, helpfulness, faithfulness | Model-graded rubric (outside this harness) | [EXPERIMENTAL] Noisy and drifts with the judge; calibrate against human labels first, and never gate on it alone |
| Severity | Covers | Gate |
|---|---|---|
critical | Safety, money movement, data loss, refusals that must hold | Blocks on a single failure (--strict-critical) |
major | Task correctness — the user did not get what they asked for | Blocks below the pass-rate floor (--fail-under) |
minor | Budgets, verbosity, style | Reported; never blocks |
| Discordant scenarios (flipped either way) | Read it as |
|---|---|
| 0 | No behavioural change detected at this suite's resolution |
| 1-5 | Read the individual scenarios; the p-value has no power here |
| 6-24 | Exact McNemar p is meaningful; eval_diff.py reports it |
| 25+ | Both the p-value and the aggregate rate movement are informative |
A single critical regression is actionable at n = 1. Significance testing is
for aggregate movement, never for safety failures.
Mistake: The release check is "pass rate ≥ 90%," and everything else is advisory. Why it happens: One number is easy to put in a dashboard and easy to explain to leadership, and it genuinely looks like the summary statistic. Instead: Gate on critical-severity failures and on the paired per-scenario diff. The pass rate is the last number you read, always with its confidence interval — at 30 scenarios that interval is ±13 points, which cannot resolve the regressions you care about. The sample data here shows two runs at an identical 83.3% where one refunds money on an injected instruction.
Mistake: output_contains: "I've issued your refund of $49.00 and it should arrive in 3-5 business days".
Why it happens: It is the fastest thing to do — copy the good output into the assertion and move on.
Instead: Assert on the tool call (issue_refund with order_id=A-10041) and on a domain token in the text ("refund", the order ID). Structural assertions do not break when the model rewords, so the suite keeps signal across model upgrades instead of generating a wall of false failures that trains the team to ignore it.
Mistake: Every scenario is a happy path; the suite has no tool_not_called assertions.
Why it happens: Suites get written from the product spec, and specs describe intended behaviour, not forbidden behaviour.
Instead: For every irreversible action the agent can take, write a scenario where taking it is wrong. Refusal and adversarial scenarios are where prompt changes actually regress, because a change that makes an agent more capable usually makes it more eager. Target roughly 35% of the suite across refusal and adversarial buckets.
Mistake: Iterating on the prompt with the full suite visible until every scenario passes. Why it happens: It feels like the tight feedback loop that good engineering is supposed to have. Instead: Hold out 20% of scenarios and never look at them while iterating; run them only at the gate. Thirty scenarios is a small enough surface to overfit in an afternoon, producing an agent that passes the suite and fails users.
Mistake: Four scenarios flip after a prompt edit, so the team spends two days finding the cause. Why it happens: Nobody ever ran the identical configuration twice, so run-to-run variance is unmeasured and every flip looks causal. Instead: Before trusting any diff, score the same configuration twice and diff it against itself. That flip count is your noise floor. Then reduce it — temperature 0 where the product allows, replayed tool results rather than live backends, and re-runs of flipped scenarios to separate flaky from real.
| File | Purpose |
|---|---|
scripts/scenario_runner.py | Runs a JSON scenario suite against recorded transcripts; reports pass/fail per assertion with severity, budget checks, and CI exit codes |
scripts/eval_diff.py | Diffs two runs into regressed/fixed/stable, with Wilson intervals, exact McNemar on discordant pairs, and cost/latency drift |
references/scenario-and-fixture-design.md | The six scenario buckets, replay modes, fixture recording rules, assertion tiers, suite sizing |
references/eval-methodology-and-budgets.md | Scoring layers, small-sample statistics, budget setting, CI wiring, methodology anti-patterns |
assets/sample_suite.json | Six-scenario support-agent suite covering all assertion types |
assets/sample_transcripts_baseline.json | Recorded baseline run |
assets/sample_transcripts_candidate.json | Recorded candidate run containing a critical regression at an unchanged pass rate |
assets/sample_baseline_report.json | Scored baseline report — input for eval_diff.py |
assets/sample_candidate_report.json | Scored candidate report — input for eval_diff.py |
assets/eval_report_template.md | Release-decision report template |
assets/scenario_authoring_checklist.md | Pre-merge checklist for any scenario joining a gating suite |
© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (scripts, references, assets) in engineering/agent-harness of borghei/Claude-Skills.
Open the folder on GitHubat commit 4a698e8
Agent Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Harness this skillborghei/Claude-Skills | 891 | — | ~3.1k | Automated safety check: Pass | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
borghei/Claude-Skills
Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.
borghei/Claude-Skills
Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.
borghei/Claude-Skills
Idea to AI-generated prototype to customer validation to engineering handoff.
borghei/Claude-Skills
Analytics engineering across data modeling, dbt, transformation, and semantic layers.
borghei/Claude-Skills
Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.
borghei/Claude-Skills
OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.
Categories
Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets. Agent Harness is an agent skill from borghei/Claude-Skills. Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.
Agent Harness fits situations like: agent quality is vibe-checked; before shipping a prompt.
Run `npx skills add borghei/Claude-Skills --skill agent-harness -a claude-code`. Or copy the skill folder (engineering/agent-harness in borghei/Claude-Skills) into .claude/skills/agent-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add borghei/Claude-Skills --skill agent-harness -a codex`. Or copy the skill folder (engineering/agent-harness in borghei/Claude-Skills) into .agents/skills/agent-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill agent-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-harness, .gemini/skills/agent-harness, .github/skills/agent-harness and .opencode/skills/agent-harness in your project.
Going by SKILL.md and its folder, Agent Harness needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Agent Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agent Harness: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 891 GitHub stars. The repository holds 354 skills in this directory. The repository was last updated on October 7, 2026.
Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.