Agent skill

OpenHarness End-to-End Evals

by HKUDS in HKUDS/OpenHarness

Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

MITAuto-check: notesTesting & QA

Install OpenHarness End-to-End Evals

skills CLI
$ npx skills add HKUDS/OpenHarness --skill harness-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenHarness harness-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenHarness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/harness-eval .claude/skills/harness-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-eval
GitHub stars
16k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
851 words
Files
3 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

  • Works in 7 steps: Prepare Workspace → Configure Environment → Prepare Real Sandbox Runtime When Relevant → …
  • Validating OpenHarness features with real model calls
  • SKILL.md covers Core Principles, Workflow, Feature Coverage Checklist and Common Pitfalls, plus 1 more section
  • Calls python, apt-get and git; reaches github.com and api.moonshot.cn; needs ANTHROPIC_API_KEY

What it does

Every test exercises the full stack from API client through model, tool calls and execution to result. Five principles apply: test on an unfamiliar project (never OpenHarness itself, since the agent would modify its own code), use real API calls with no mocks, hold multi-turn conversations of at least two turns that depend on earlier context, combine features such as hooks, skills and the agent loop, and verify tool execution by inspecting tool call lists and output files rather than only model text.

The workflow clones a real project into a temporary workspace, sets the Anthropic-style environment variables for a real endpoint, and leaves `max_turns` at the product default of 200 for long runs. When the target is sandbox behavior, it installs and checks the real sandbox runtime first and runs a smoke test through OpenHarness's own bash tool, treating missing dependencies as a setup failure rather than a regression. Tests follow a pattern of building an engine, submitting messages and collecting text, tools, turns and tokens, with code templates and helpers in a test-patterns reference and a feature matrix alongside.

When your agent uses it

  • Validating OpenHarness features with real model calls
  • Running end-to-end agent loop tests against a cloned project
  • Checking that sandbox behavior works with the real runtime

Example prompts

  • “Run the harness eval against a freshly cloned project and check the tool calls it made.”
  • “Test hooks and skills together in a multi-turn agent loop with the real API.”
  • “Verify the sandbox feature end to end and tell me whether any failure is a setup problem.”

Requirements

  • A real LLM endpoint and API key set through environment variables
  • Git, to clone an unfamiliar project as the workspace
  • For sandbox tests: the sandbox runtime package, bubblewrap and ripgrep

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Prepare Workspace
  2. Configure Environment
  3. Prepare Real Sandbox Runtime When Relevant
  4. Design Tests
  5. Prefer Long-Horizon, Real Agent Loops
  6. Run Tests
  7. Interpret Results

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2efd7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • apt-get
    • git
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • api.moonshot.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OpenHarness End-to-End Evals loads about 2.1k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 851 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:45
    sudo apt-get update
  • NoteRuns commands with sudoSKILL.md:46
    sudo apt-get install -y bubblewrap ripgrep

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenHarness at commit 9b2efd7, republished under its MIT licence (© HKUDS). 851 words, ~2,112 tokens.

Download SKILL.mdSave it as .claude/skills/harness-eval/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
harness-eval
description
This skill should be used when the user asks to "test the harness", "run integration tests", "validate features with real API", "test with real model calls", "run agent loop tests", "verify end-to-end", or needs to verify OpenHarness features on a real codebase with actual LLM calls.
version
0.2.0

Harness Eval — End-to-End Feature Validation

Validate OpenHarness features by running real agent loops against an unfamiliar codebase with actual LLM API calls. Every test exercises the full stack: API client → model → tool calls → execution → result.

Core Principles

  1. Test on an unfamiliar project — never test on OpenHarness itself (the agent modifies its own code). Clone a real project as the workspace.
  2. Use real API calls — no mocks. Configure a real LLM endpoint.
  3. Multi-turn conversations — always test 2+ turns where the model needs prior context.
  4. Combine features — test hooks+skills+agent loop together, not in isolation.
  5. Verify tool execution — inspect tool call lists and output files, not just model text.

Workflow

1. Prepare Workspace

Clone an unfamiliar project (do not use OpenHarness):

bash
git clone https://github.com/HKUDS/AutoAgent /tmp/eval-workspace
2. Configure Environment
bash
export ANTHROPIC_API_KEY=sk-xxx
export ANTHROPIC_BASE_URL=https://api.moonshot.cn/anthropic  # or any provider
export ANTHROPIC_MODEL=kimi-k2.5

For long-running real evals, do not artificially lower max_turns. Use the product default (200) unless the user explicitly wants a tighter bound.

3. Prepare Real Sandbox Runtime When Relevant

If the task is validating sandbox behavior, install and verify the actual runtime before running agent loops:

bash
npm install -g @anthropic-ai/sandbox-runtime
sudo apt-get update
sudo apt-get install -y bubblewrap ripgrep
which srt
which bwrap
which rg
srt --version

Then run a minimal smoke check through OpenHarness, not just raw srt, so you verify the real adapter path:

python
from pathlib import Path
from openharness.config.settings import Settings, SandboxSettings, save_settings
from openharness.tools.bash_tool import BashTool

cfg = Path("/tmp/openharness-sandbox-settings.json")
save_settings(Settings(sandbox=SandboxSettings(enabled=True, fail_if_unavailable=True)), cfg)
# Point config loader at this file, then run BashTool on a tiny command such as `pwd`.

If sandbox dependencies are missing, treat that as an environment/setup failure, not a feature regression.

4. Design Tests

Each test follows this pattern:

python
engine = make_engine(system_prompt="...", cwd=UNFAMILIAR_PROJECT)
evs1 = [ev async for ev in engine.submit_message("Read X, analyze Y")]
r1 = collect(evs1)  # text, tools, turns, tokens
evs2 = [ev async for ev in engine.submit_message("Based on what you found...")]
r2 = collect(evs2)
assert "grep" in r1["tools"]  # verify tools ran

For detailed code templates and the make_engine/collect helpers, consult references/test-patterns.md.

5. Prefer Long-Horizon, Real Agent Loops

For meaningful end-to-end validation, prefer unfamiliar-repo tasks that force multiple turns, context reuse, and mixed tool usage.

Recommended pattern:

  • Use a real external workspace such as AutoAgent
  • Use real provider credentials and the actual target model
  • Keep max_turns=200
  • Use per-prompt timeouts large enough for real exploration, such as 240-600s
  • Require at least 2 turns per scenario
  • Verify both text quality and tool traces
  • Keep polling long-running sessions until they finish; do not abandon a run after the first long pause

Recommended long-horizon scenarios:

  • architecture_multiturn

    • Turn 1: map architecture, shell/subprocess surfaces, and test entrypoints
    • Turn 2: identify top risks and propose refactors
    • Turn 3: condense into onboarding or remediation actions
    • Success: bash, glob, grep, read_file all appear; no timeout; no MaxTurnsExceeded
  • hook_block_and_recover

    • Force the model to try bash
    • Block it with a real pre-tool hook
    • Verify the model adapts with glob/grep/read_file
  • sandbox_multiturn

    • Enable real sandbox settings with fail_if_unavailable=true
    • First prompt must start with exactly one shell command such as pwd && ls -la
    • Second prompt must explicitly reuse the prior shell findings
    • Success: bash executes via sandbox, non-shell tools continue the task, and the agent recovers from incidental repo errors

When a scenario fails, classify it before changing code:

  • MaxTurnsExceeded: likely eval harness misconfiguration if max_turns was manually lowered
  • timeout: task is too broad or per-prompt timeout is too small
  • sandbox unavailable: environment missing srt, bwrap, or rg
  • tool error with task still completed: feature may still be healthy; inspect recovery behavior
6. Run Tests
bash
python tests/test_merged_prs_on_autoagent.py   # PR feature tests
python tests/test_real_large_tasks.py           # large multi-step tasks
python tests/test_hooks_skills_plugins_real.py  # hooks/skills/plugins
python -m pytest tests/ -q -k "not autoagent"  # unit tests (no API)

For ad hoc long-horizon validation, it is acceptable to run a temporary Python driver script as long as it:

  • uses real OpenHarness engine/tool objects
  • targets an unfamiliar repository
  • prints per-scenario JSON summaries
  • records tools, errors, turns, and token usage
  • stays attached until completion
Show full SKILL.md (335 more words)Show less
7. Interpret Results
ResultMeaningAction
PASS with tool callsFeature works end-to-endDone
PASS without tool callsModel answered from knowledgeRewrite prompt to force tool use
FAIL with exceptionCode bugRead traceback
FAIL with wrong outputModel behavior issueCheck system prompt and tool schemas
TimeoutTask too complexIncrease max_turns or simplify prompt

For long-running real evals, refine the timeout guidance:

  • First check whether max_turns was manually set too low
  • If max_turns=200 and the run still fails, the next suspect is wall-clock timeout, not turn count
  • Distinguish environment failures from product failures
    • Example: missing dependency in the unfamiliar target repo is not automatically an OpenHarness regression
    • Example: missing srt/bwrap/rg is an eval environment issue

Feature Coverage Checklist

  • Engine: multi-turn memory, tool chaining, parallel tools, error recovery, auto-compaction
  • Swarm: InProcessBackend lifecycle, concurrent teammates, coordinator+notifications
  • Hooks: pre_tool_use blocking → model adapts, post_tool_use firing
  • Skills: skill tool invocation → model follows instructions
  • Plugins: plugin-provided skill loaded and used in agent loop
  • Memory: YAML frontmatter parsing, body content search, context injection
  • Session: save → load → resume with context preserved
  • Providers: Anthropic client, OpenAI client (with reasoning_content), multi-turn
  • Cost: token accumulation across turns

Common Pitfalls

  • Testing on OpenHarness itself — agent modifies its own running code
  • Using mocks — misses serialization and API compatibility bugs
  • Single-turn only — misses context accumulation and compaction bugs
  • Artificially lowering max_turns during real evals — can create false failures that do not reflect product defaults
  • Not checking tool call list — model may claim tool use without calling it
  • Hardcoding paths — use WORKSPACE variable, skip in CI with pytest.mark.skipif
  • Declaring sandbox “tested” after only checking raw srt — verify the OpenHarness adapter path too
  • Abandoning long tasks too early — some real tasks pause for minutes before the next event arrives

Additional Resources

Reference Files
  • references/test-patterns.md — Complete code templates for make_engine, collect, and each feature category
  • references/feature-matrix.md — Detailed test cases for every OpenHarness module
Existing Test Files

Working test suites in the repo:

  • tests/test_merged_prs_on_autoagent.py — PR feature validation
  • tests/test_real_large_tasks.py — Large multi-step tasks
  • tests/test_hooks_skills_plugins_real.py — Hooks/skills/plugins in agent loops
  • tests/test_untested_features.py — Module-level integration tests

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .claude/skills/harness-eval of HKUDS/OpenHarness.

  • SKILL.md
  • references/feature-matrix.md
  • references/test-patterns.md

Open the folder on GitHubat commit 9b2efd7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in HKUDS/OpenHarness, which our catalogue first saw on October 7, 2026.

Compare with similar skills

OpenHarness End-to-End Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OpenHarness End-to-End Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OpenHarness End-to-End Evals this skillHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
CodeGraph Agent Evalcolbymchenry/codegraph73k—~950Automated safety check: PassMIT
Godot E2ERandallLiuXin/GodotMaker5491 repos~3.9kAutomated safety check: PassCustom licence
Interactive CLI Testing With tui-testslopus/happy24k—~603Automated safety check: PassMIT
Test Pyramidkubernetes-sigs/agent-sandbox4.2k—~1.6kAutomated safety check: PassApache-2.0
Blockless Extension E2EFreakStudioCN/mpy-hardware-extension132—~1.2kAutomated safety check: NotesCustom licence

Similar skills

  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    73k GitHub stars~950 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Godot E2E

    RandallLiuXin/GodotMaker

    Write and run E2E (end-to-end) game tests using the godot-e2e framework.

    549 GitHub starsUsed in 1 repo~3.9k tokens
    Testing & QAAuto-check passed
  • Tests interactive CLI and TUI programs with Microsoft's tui-test, driving prompts, arrow keys and screen output in a real pseudo-terminal.

    24k GitHub stars~603 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Test Pyramid

    kubernetes-sigs/agent-sandbox

    Official

    Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need…

    4.2k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Blockless Extension E2E

    FreakStudioCN/mpy-hardware-extension

    Run and debug the Blockless VS Code extension release gate: CI-equivalent API and extension tests, V0 protocol smoke, live DeepSeek full-stack e2e, VSIX packaging, local reinstall, direct…

    132 GitHub stars~1.2k tokensUpdated 10 days ago
    Testing & QAAuto-check: notes
  • Fly E2E Test

    cmpnd-ai/dspy-cli

    Deploy and test dspy-cli on Fly.io using local changes via temp git branch.

    137 GitHub stars~2.3k tokensUpdated 7 mo ago
    Testing & QAAuto-check: notes

More from HKUDS/OpenHarness

  • Merges external GitHub pull requests while keeping the original author credited, and fixes conflicts after the merge instead of rewriting the contribution.

    16k GitHub starsUsed in 1 repo~847 tokens
    Auto-check passed

Works with

Questions about OpenHarness End-to-End Evals

What does OpenHarness End-to-End Evals do?

Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution. Every test exercises the full stack from API client through model, tool calls and execution to result. Five principles apply: test on an unfamiliar project (never OpenHarness itself, since the agent would modify its own code), use real API calls with no mocks, hold multi-turn conversations of at least two turns that depend on earlier context, combine features such as hooks, skills and the agent loop, and verify tool execution by inspecting tool call lists and output files rather than only model text.

When should I use OpenHarness End-to-End Evals?

OpenHarness End-to-End Evals fits situations like: validating OpenHarness features with real model calls; running end-to-end agent loop tests against a cloned project; checking that sandbox behavior works with the real runtime.

How do I install OpenHarness End-to-End Evals in Claude Code?

Run `npx skills add HKUDS/OpenHarness --skill harness-eval -a claude-code`. Or copy the skill folder (.claude/skills/harness-eval in HKUDS/OpenHarness) into .claude/skills/harness-eval in your project. Claude Code loads it when a task matches its description.

How do I install OpenHarness End-to-End Evals in Codex?

Run `npx skills add HKUDS/OpenHarness --skill harness-eval -a codex`. Or copy the skill folder (.claude/skills/harness-eval in HKUDS/OpenHarness) into .agents/skills/harness-eval in your project. Codex loads it when a task matches its description.

Can I use OpenHarness End-to-End Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenHarness --skill harness-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-eval, .gemini/skills/harness-eval, .github/skills/harness-eval and .opencode/skills/harness-eval in your project.

What does OpenHarness End-to-End Evals need to run?

Going by SKILL.md and its folder, OpenHarness End-to-End Evals needs the command-line tools its instructions call (python, apt-get, git and npm) and credentials named ANTHROPIC_API_KEY. Our summary lists: A real LLM endpoint and API key set through environment variables; Git, to clone an unfamiliar project as the workspace; For sandbox tests: the sandbox runtime package, bubblewrap and ripgrep.

Does OpenHarness End-to-End Evals access the network?

SKILL.md names 2 domains. In commands or code: github.com and api.moonshot.cn; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is OpenHarness End-to-End Evals safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does OpenHarness End-to-End Evals use?

OpenHarness End-to-End Evals is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does OpenHarness End-to-End Evals use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to OpenHarness End-to-End Evals?

Skills that share tags, products or a category with OpenHarness End-to-End Evals: CodeGraph Agent Eval (colbymchenry/codegraph, 73k stars), Godot E2E (RandallLiuXin/GodotMaker, 549 stars), Interactive CLI Testing With tui-test (slopus/happy, 24k stars) and Test Pyramid (kubernetes-sigs/agent-sandbox, 4.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OpenHarness End-to-End Evals?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenHarness, which has 15,919 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on June 4, 2026.

Source: HKUDS/OpenHarness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.