Agent skill

Agent Harness

by borghei in borghei/Claude-Skills

Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.

MITAuto-check passedAI & LLM Engineering

Install Agent Harness

skills CLI
$ npx skills add borghei/Claude-Skills --skill agent-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills agent-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/agent-harness .claude/skills/agent-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-harness
GitHub stars
891
Token cost
~3.1k tokens
SKILL.md length
1,507 words
Files
12 (incl. scripts, references, assets)
Skills in repo
354
Repo updated
First seen
Licence
MIT

At a glance

Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.

  • Works in 5 steps: Enumerate the agent's irreversible… → Draft 20-30 scenarios across all six… → Declare suite-wide defaults for latency,… → …
  • Agent quality is vibe-checked
  • SKILL.md covers When to use this skill, Inputs the skill expects, Clarify First and Workflows, plus 3 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Agent Harness is an agent skill from borghei/Claude-Skills. Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets. Use when agent quality is vibe-checked, before shipping a prompt or model change, or when evals drift.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts, reference files and assets (for example `assets/eval_report_template.md`, `assets/sample_baseline_report.json` and `assets/sample_candidate_report.json`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Agent quality is vibe-checked
  • Before shipping a prompt

Example prompts

  • “/agent-harness”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Enumerate the agent's irreversible actions; each one gets a refusal scenario.
  2. Draft 20-30 scenarios across all six buckets (happy, boundary, refusal,
  3. Declare suite-wide defaults for latency, cost, and turn ceilings so every
  4. Record one transcript per scenario, scrubbing PII at record time, and stamp
  5. Score the run and read critical failures before the pass rate.

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Harness loads about 3.1k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 1,507 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 1,507 words, ~3,068 tokens.

Download SKILL.mdSave it as .claude/skills/agent-harness/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
agent-harness
description
Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets. Use when agent quality is vibe-checked, before shipping a prompt or model change, or when evals drift.
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
agent-evaluation
metadata.updated
2026-07-21
metadata.tags
agent-eval, regression-testing, harness, replay, llm-testing

Agent Harness

Most agents ship on vibes: someone tries eight prompts, the output looks good, it goes to production, and the next prompt tweak silently breaks a refusal nobody re-tested. This skill builds the harness around an agent so its behaviour becomes measurable — scenario suites with structural assertions, deterministic replay of recorded tool calls, paired regression diffing across prompt and model changes, and per-scenario cost and latency budgets. The tools here score an agent; they never invoke one, so they run offline on every commit.

When to use this skill

  • An agent is going to production and the only quality evidence is manual spot-checking
  • A prompt, tool schema, or model version is changing and you need to know what broke
  • Two model or configuration options need a defensible comparison, not a demo
  • An incident happened and you need the behaviour encoded as a permanent regression test
  • Agent cost or latency is climbing across releases and nobody can point to when
  • An existing eval suite reports a healthy pass rate that nobody trusts

Inputs the skill expects

  • The agent's tool inventory — names, arguments, and which tools are irreversible
  • Recorded transcripts per scenario: tool calls, final output, turns, latency, cost, error state
  • The behavioural rules the agent must hold (refusals, escalation triggers, policy boundaries)
  • Known failure history — past incidents, customer complaints, internal bug reports
  • Current cost and latency expectations per interaction
  • The release gate that consumes the result (CI job, review checklist, launch review)

Clarify First

Before building the harness, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Which agent actions are irreversible — determines which scenarios need tool_not_called assertions at critical severity, and what the release gate blocks on
  • Whether transcripts are already recorded — decides whether workflow 1 starts from replay or from an instrumentation task first
  • What the suite gates — a CI blocking check, a nightly report, or a one-off comparison; changes suite size, runtime budget, and severity strictness
  • The known failure modes — past incidents seed the adversarial and refusal buckets, which is where regressions actually hide

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Workflows

Workflow 1 — Stand up a scenario suite and score a run
  1. Enumerate the agent's irreversible actions; each one gets a refusal scenario.
  2. Draft 20-30 scenarios across all six buckets (happy, boundary, refusal, adversarial, failure-recovery, ambiguity) using assets/scenario_authoring_checklist.md. Structural assertions first — tool called / not called / order / arguments — text assertions only on domain tokens.
  3. Declare suite-wide defaults for latency, cost, and turn ceilings so every scenario is budgeted without repeating yourself.
  4. Record one transcript per scenario, scrubbing PII at record time, and stamp the run with model and prompt_sha.
  5. Score the run and read critical failures before the pass rate.
bash
python3 engineering/agent-harness/scripts/scenario_runner.py \
  --suite engineering/agent-harness/assets/sample_suite.json \
  --transcripts engineering/agent-harness/assets/sample_transcripts_baseline.json \
  --strict-critical
Workflow 2 — Gate a prompt or model change on a paired regression diff
  1. Score the baseline and the candidate with the same suite file, saving both as JSON reports.
  2. Diff them. Read regressions and budget drift before the aggregate rate.
  3. Triage every regression: intended trade, real defect, or flaky scenario (re-run the flipped scenario five times to tell the last two apart).
  4. Record the decision in assets/eval_report_template.md and promote the accepted candidate report to the new baseline.
bash
python3 engineering/agent-harness/scripts/scenario_runner.py \
  --suite engineering/agent-harness/assets/sample_suite.json \
  --transcripts engineering/agent-harness/assets/sample_transcripts_candidate.json \
  --format json > /tmp/candidate.report.json

python3 engineering/agent-harness/scripts/eval_diff.py \
  --baseline engineering/agent-harness/assets/sample_baseline_report.json \
  --candidate /tmp/candidate.report.json \
  --fail-on-regression --drift-threshold 0.15

The shipped sample data demonstrates the core lesson: both runs score 83.3%, and the candidate contains a critical prompt-injection regression. A gate on pass rate ships it; the paired diff catches it.

Workflow 3 — Establish cost and latency budgets, then track drift
  1. Take the last release's accepted run as the reference.
  2. Set per-scenario latency at p95 × 1.3, cost at median × 1.5, and the turn ceiling at observed max + 2. Put them in the suite defaults, overriding only where a scenario is legitimately expensive.
  3. Score the current run; budget breaches surface as minor assertions, so they report without blocking.
  4. Diff against the reference with a tight drift threshold to catch the slow bleed that stays inside budget.
bash
python3 engineering/agent-harness/scripts/eval_diff.py \
  --baseline engineering/agent-harness/assets/sample_baseline_report.json \
  --candidate engineering/agent-harness/assets/sample_candidate_report.json \
  --drift-threshold 0.10 --format json

Decision frameworks

Which assertion type to reach for
NeedUseDurability
The agent must take an actiontool_called, tool_call_order[PROVEN] Exact; survives rewording
The agent must NOT take an actiontool_not_called[PROVEN] The single highest-value assertion in any agent suite
The action must use the right datatool_arg_equals[PROVEN] Catches the right tool with wrong arguments
Structured output correctnessjson_field_equals[PROVEN] Exact when the agent has a JSON mode
A required domain fact appearsoutput_contains on an ID, number, or policy name[RECOMMENDED] Stable if you never quote sentences
A forbidden phrase must not appearoutput_not_contains[RECOMMENDED] Good for injection and leak checks
Tone, helpfulness, faithfulnessModel-graded rubric (outside this harness)[EXPERIMENTAL] Noisy and drifts with the judge; calibrate against human labels first, and never gate on it alone
Severity, and what each one gates
SeverityCoversGate
criticalSafety, money movement, data loss, refusals that must holdBlocks on a single failure (--strict-critical)
majorTask correctness — the user did not get what they asked forBlocks below the pass-rate floor (--fail-under)
minorBudgets, verbosity, styleReported; never blocks
Can I trust this diff?
Discordant scenarios (flipped either way)Read it as
0No behavioural change detected at this suite's resolution
1-5Read the individual scenarios; the p-value has no power here
6-24Exact McNemar p is meaningful; eval_diff.py reports it
25+Both the p-value and the aggregate rate movement are informative

A single critical regression is actionable at n = 1. Significance testing is for aggregate movement, never for safety failures.

Show full SKILL.md (581 more words)Show less

Anti-Patterns

Gating on the aggregate pass rate

Mistake: The release check is "pass rate ≥ 90%," and everything else is advisory. Why it happens: One number is easy to put in a dashboard and easy to explain to leadership, and it genuinely looks like the summary statistic. Instead: Gate on critical-severity failures and on the paired per-scenario diff. The pass rate is the last number you read, always with its confidence interval — at 30 scenarios that interval is ±13 points, which cannot resolve the regressions you care about. The sample data here shows two runs at an identical 83.3% where one refunds money on an injected instruction.

Asserting on sentences instead of structure

Mistake: output_contains: "I've issued your refund of $49.00 and it should arrive in 3-5 business days". Why it happens: It is the fastest thing to do — copy the good output into the assertion and move on. Instead: Assert on the tool call (issue_refund with order_id=A-10041) and on a domain token in the text ("refund", the order ID). Structural assertions do not break when the model rewords, so the suite keeps signal across model upgrades instead of generating a wall of false failures that trains the team to ignore it.

Only testing what the agent should do

Mistake: Every scenario is a happy path; the suite has no tool_not_called assertions. Why it happens: Suites get written from the product spec, and specs describe intended behaviour, not forbidden behaviour. Instead: For every irreversible action the agent can take, write a scenario where taking it is wrong. Refusal and adversarial scenarios are where prompt changes actually regress, because a change that makes an agent more capable usually makes it more eager. Target roughly 35% of the suite across refusal and adversarial buckets.

Tuning the prompt until the suite goes green

Mistake: Iterating on the prompt with the full suite visible until every scenario passes. Why it happens: It feels like the tight feedback loop that good engineering is supposed to have. Instead: Hold out 20% of scenarios and never look at them while iterating; run them only at the gate. Thirty scenarios is a small enough surface to overfit in an afternoon, producing an agent that passes the suite and fails users.

Chasing regressions without a noise floor

Mistake: Four scenarios flip after a prompt edit, so the team spends two days finding the cause. Why it happens: Nobody ever ran the identical configuration twice, so run-to-run variance is unmeasured and every flip looks causal. Instead: Before trusting any diff, score the same configuration twice and diff it against itself. That flip count is your noise floor. Then reduce it — temperature 0 where the product allows, replayed tool results rather than live backends, and re-runs of flipped scenarios to separate flaky from real.

Files

FilePurpose
scripts/scenario_runner.pyRuns a JSON scenario suite against recorded transcripts; reports pass/fail per assertion with severity, budget checks, and CI exit codes
scripts/eval_diff.pyDiffs two runs into regressed/fixed/stable, with Wilson intervals, exact McNemar on discordant pairs, and cost/latency drift
references/scenario-and-fixture-design.mdThe six scenario buckets, replay modes, fixture recording rules, assertion tiers, suite sizing
references/eval-methodology-and-budgets.mdScoring layers, small-sample statistics, budget setting, CI wiring, methodology anti-patterns
assets/sample_suite.jsonSix-scenario support-agent suite covering all assertion types
assets/sample_transcripts_baseline.jsonRecorded baseline run
assets/sample_transcripts_candidate.jsonRecorded candidate run containing a critical regression at an unchanged pass rate
assets/sample_baseline_report.jsonScored baseline report — input for eval_diff.py
assets/sample_candidate_report.jsonScored candidate report — input for eval_diff.py
assets/eval_report_template.mdRelease-decision report template
assets/scenario_authoring_checklist.mdPre-merge checklist for any scenario joining a gating suite

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references, assets) in engineering/agent-harness of borghei/Claude-Skills.

  • SKILL.md
  • assets/eval_report_template.md
  • assets/sample_baseline_report.json
  • assets/sample_candidate_report.json
  • assets/sample_suite.json
  • assets/sample_transcripts_baseline.json
  • assets/sample_transcripts_candidate.json
  • assets/scenario_authoring_checklist.md
  • references/eval-methodology-and-budgets.md
  • references/scenario-and-fixture-design.md
  • scripts/eval_diff.py
  • scripts/scenario_runner.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Agent Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Harness this skillborghei/Claude-Skills891—~3.1kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from borghei/Claude-Skills

All 354 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    891 GitHub stars~4.2k tokensUpdated 3 days ago
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    891 GitHub stars~3.4k tokensUpdated 3 days ago
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    891 GitHub stars~3.6k tokensUpdated 3 days ago
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    891 GitHub stars~3.4k tokensUpdated 3 days ago
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    891 GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    891 GitHub stars~1.4k tokensUpdated 3 days ago
    Auto-check passed

Questions about Agent Harness

What does Agent Harness do?

Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets. Agent Harness is an agent skill from borghei/Claude-Skills. Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.

When should I use Agent Harness?

Agent Harness fits situations like: agent quality is vibe-checked; before shipping a prompt.

How do I install Agent Harness in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill agent-harness -a claude-code`. Or copy the skill folder (engineering/agent-harness in borghei/Claude-Skills) into .claude/skills/agent-harness in your project. Claude Code loads it when a task matches its description.

How do I install Agent Harness in Codex?

Run `npx skills add borghei/Claude-Skills --skill agent-harness -a codex`. Or copy the skill folder (engineering/agent-harness in borghei/Claude-Skills) into .agents/skills/agent-harness in your project. Codex loads it when a task matches its description.

Can I use Agent Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill agent-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-harness, .gemini/skills/agent-harness, .github/skills/agent-harness and .opencode/skills/agent-harness in your project.

What does Agent Harness need to run?

Going by SKILL.md and its folder, Agent Harness needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Agent Harness access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Harness use?

Agent Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Harness use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.

What are the alternatives to Agent Harness?

Skills that share tags, products or a category with Agent Harness: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Harness?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 891 GitHub stars. The repository holds 354 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.