CodeGraph Agent Eval
colbymchenry/codegraph
Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
$ npx skills add HKUDS/OpenHarness --skill harness-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install HKUDS/OpenHarness harness-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/HKUDS/OpenHarness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/harness-eval .claude/skills/harness-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "harness-eval" agent skill from https://github.com/HKUDS/OpenHarness/tree/main/.claude/skills/harness-eval into .claude/skills/harness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/HKUDS/OpenHarness/tree/main/.claude/skills/harness-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add HKUDS/OpenHarness --skill harness-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install HKUDS/OpenHarness harness-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenHarness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/harness-eval .agents/skills/harness-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "harness-eval" agent skill from https://github.com/HKUDS/OpenHarness/tree/main/.claude/skills/harness-eval into .agents/skills/harness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add HKUDS/OpenHarness --skill harness-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install HKUDS/OpenHarness harness-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenHarness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/harness-eval .cursor/skills/harness-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "harness-eval" agent skill from https://github.com/HKUDS/OpenHarness/tree/main/.claude/skills/harness-eval into .cursor/skills/harness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/HKUDS/OpenHarness.git --path .claude/skills/harness-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add HKUDS/OpenHarness --skill harness-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install HKUDS/OpenHarness harness-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenHarness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/harness-eval .gemini/skills/harness-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "harness-eval" agent skill from https://github.com/HKUDS/OpenHarness/tree/main/.claude/skills/harness-eval into .gemini/skills/harness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install HKUDS/OpenHarness harness-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add HKUDS/OpenHarness --skill harness-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/HKUDS/OpenHarness.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/harness-eval .github/skills/harness-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "harness-eval" agent skill from https://github.com/HKUDS/OpenHarness/tree/main/.claude/skills/harness-eval into .github/skills/harness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add HKUDS/OpenHarness --skill harness-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install HKUDS/OpenHarness harness-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenHarness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/harness-eval .opencode/skills/harness-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "harness-eval" agent skill from https://github.com/HKUDS/OpenHarness/tree/main/.claude/skills/harness-eval into .opencode/skills/harness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
harness-evalValidates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
Every test exercises the full stack from API client through model, tool calls and execution to result. Five principles apply: test on an unfamiliar project (never OpenHarness itself, since the agent would modify its own code), use real API calls with no mocks, hold multi-turn conversations of at least two turns that depend on earlier context, combine features such as hooks, skills and the agent loop, and verify tool execution by inspecting tool call lists and output files rather than only model text.
The workflow clones a real project into a temporary workspace, sets the Anthropic-style environment variables for a real endpoint, and leaves `max_turns` at the product default of 200 for long runs. When the target is sandbox behavior, it installs and checks the real sandbox runtime first and runs a smoke test through OpenHarness's own bash tool, treating missing dependencies as a setup failure rather than a regression. Tests follow a pattern of building an engine, submitting messages and collecting text, tools, turns and tokens, with code templates and helpers in a test-patterns reference and a feature matrix alongside.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 9b2efd7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonapt-getgitnpmFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comapi.moonshot.cnFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
ANTHROPIC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
OpenHarness End-to-End Evals loads about 2.1k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 851 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
sudo apt-get updatesudo apt-get install -y bubblewrap ripgrepAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from HKUDS/OpenHarness at commit 9b2efd7, republished under its MIT licence (© HKUDS). 851 words, ~2,112 tokens.
.claude/skills/harness-eval/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Validate OpenHarness features by running real agent loops against an unfamiliar codebase with actual LLM API calls. Every test exercises the full stack: API client → model → tool calls → execution → result.
Clone an unfamiliar project (do not use OpenHarness):
git clone https://github.com/HKUDS/AutoAgent /tmp/eval-workspaceexport ANTHROPIC_API_KEY=sk-xxx
export ANTHROPIC_BASE_URL=https://api.moonshot.cn/anthropic # or any provider
export ANTHROPIC_MODEL=kimi-k2.5For long-running real evals, do not artificially lower max_turns. Use the product default (200) unless the user explicitly wants a tighter bound.
If the task is validating sandbox behavior, install and verify the actual runtime before running agent loops:
npm install -g @anthropic-ai/sandbox-runtime
sudo apt-get update
sudo apt-get install -y bubblewrap ripgrep
which srt
which bwrap
which rg
srt --versionThen run a minimal smoke check through OpenHarness, not just raw srt, so you verify the real adapter path:
from pathlib import Path
from openharness.config.settings import Settings, SandboxSettings, save_settings
from openharness.tools.bash_tool import BashTool
cfg = Path("/tmp/openharness-sandbox-settings.json")
save_settings(Settings(sandbox=SandboxSettings(enabled=True, fail_if_unavailable=True)), cfg)
# Point config loader at this file, then run BashTool on a tiny command such as `pwd`.If sandbox dependencies are missing, treat that as an environment/setup failure, not a feature regression.
Each test follows this pattern:
engine = make_engine(system_prompt="...", cwd=UNFAMILIAR_PROJECT)
evs1 = [ev async for ev in engine.submit_message("Read X, analyze Y")]
r1 = collect(evs1) # text, tools, turns, tokens
evs2 = [ev async for ev in engine.submit_message("Based on what you found...")]
r2 = collect(evs2)
assert "grep" in r1["tools"] # verify tools ranFor detailed code templates and the make_engine/collect helpers, consult references/test-patterns.md.
For meaningful end-to-end validation, prefer unfamiliar-repo tasks that force multiple turns, context reuse, and mixed tool usage.
Recommended pattern:
AutoAgentmax_turns=200240-600sRecommended long-horizon scenarios:
architecture_multiturn
bash, glob, grep, read_file all appear; no timeout; no MaxTurnsExceededhook_block_and_recover
bashglob/grep/read_filesandbox_multiturn
fail_if_unavailable=truepwd && ls -labash executes via sandbox, non-shell tools continue the task, and the agent recovers from incidental repo errorsWhen a scenario fails, classify it before changing code:
MaxTurnsExceeded: likely eval harness misconfiguration if max_turns was manually loweredtimeout: task is too broad or per-prompt timeout is too smallsrt, bwrap, or rgpython tests/test_merged_prs_on_autoagent.py # PR feature tests
python tests/test_real_large_tasks.py # large multi-step tasks
python tests/test_hooks_skills_plugins_real.py # hooks/skills/plugins
python -m pytest tests/ -q -k "not autoagent" # unit tests (no API)For ad hoc long-horizon validation, it is acceptable to run a temporary Python driver script as long as it:
| Result | Meaning | Action |
|---|---|---|
| PASS with tool calls | Feature works end-to-end | Done |
| PASS without tool calls | Model answered from knowledge | Rewrite prompt to force tool use |
| FAIL with exception | Code bug | Read traceback |
| FAIL with wrong output | Model behavior issue | Check system prompt and tool schemas |
| Timeout | Task too complex | Increase max_turns or simplify prompt |
For long-running real evals, refine the timeout guidance:
max_turns was manually set too lowmax_turns=200 and the run still fails, the next suspect is wall-clock timeout, not turn countsrt/bwrap/rg is an eval environment issuemax_turns during real evals — can create false failures that do not reflect product defaultsWORKSPACE variable, skip in CI with pytest.mark.skipifsrt — verify the OpenHarness adapter path tooreferences/test-patterns.md — Complete code templates for make_engine, collect, and each feature categoryreferences/feature-matrix.md — Detailed test cases for every OpenHarness moduleWorking test suites in the repo:
tests/test_merged_prs_on_autoagent.py — PR feature validationtests/test_real_large_tasks.py — Large multi-step taskstests/test_hooks_skills_plugins_real.py — Hooks/skills/plugins in agent loopstests/test_untested_features.py — Module-level integration tests© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in .claude/skills/harness-eval of HKUDS/OpenHarness.
Open the folder on GitHubat commit 9b2efd7
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in HKUDS/OpenHarness, which our catalogue first saw on October 7, 2026.
OpenHarness End-to-End Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| OpenHarness End-to-End Evals this skillHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| CodeGraph Agent Evalcolbymchenry/codegraph | 73k | — | ~950 | Automated safety check: Pass | MIT | |
| Godot E2ERandallLiuXin/GodotMaker | 549 | 1 repos | ~3.9k | Automated safety check: Pass | Custom licence | |
| Interactive CLI Testing With tui-testslopus/happy | 24k | — | ~603 | Automated safety check: Pass | MIT | |
| Test Pyramidkubernetes-sigs/agent-sandbox | 4.2k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Blockless Extension E2EFreakStudioCN/mpy-hardware-extension | 132 | — | ~1.2k | Automated safety check: Notes | Custom licence |
colbymchenry/codegraph
Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.
RandallLiuXin/GodotMaker
Write and run E2E (end-to-end) game tests using the godot-e2e framework.
slopus/happy
Tests interactive CLI and TUI programs with Microsoft's tui-test, driving prompts, arrow keys and screen output in a real pseudo-terminal.
kubernetes-sigs/agent-sandbox
Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need…
FreakStudioCN/mpy-hardware-extension
Run and debug the Blockless VS Code extension release gate: CI-equivalent API and extension tests, V0 protocol smoke, live DeepSeek full-stack e2e, VSIX packaging, local reinstall, direct…
cmpnd-ai/dspy-cli
Deploy and test dspy-cli on Fly.io using local changes via temp git branch.
HKUDS/OpenHarness
Merges external GitHub pull requests while keeping the original author credited, and fixes conflicts after the merge instead of rewriting the contribution.
Categories
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution. Every test exercises the full stack from API client through model, tool calls and execution to result. Five principles apply: test on an unfamiliar project (never OpenHarness itself, since the agent would modify its own code), use real API calls with no mocks, hold multi-turn conversations of at least two turns that depend on earlier context, combine features such as hooks, skills and the agent loop, and verify tool execution by inspecting tool call lists and output files rather than only model text.
OpenHarness End-to-End Evals fits situations like: validating OpenHarness features with real model calls; running end-to-end agent loop tests against a cloned project; checking that sandbox behavior works with the real runtime.
Run `npx skills add HKUDS/OpenHarness --skill harness-eval -a claude-code`. Or copy the skill folder (.claude/skills/harness-eval in HKUDS/OpenHarness) into .claude/skills/harness-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add HKUDS/OpenHarness --skill harness-eval -a codex`. Or copy the skill folder (.claude/skills/harness-eval in HKUDS/OpenHarness) into .agents/skills/harness-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenHarness --skill harness-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-eval, .gemini/skills/harness-eval, .github/skills/harness-eval and .opencode/skills/harness-eval in your project.
Going by SKILL.md and its folder, OpenHarness End-to-End Evals needs the command-line tools its instructions call (python, apt-get, git and npm) and credentials named ANTHROPIC_API_KEY. Our summary lists: A real LLM endpoint and API key set through environment variables; Git, to clone an unfamiliar project as the workspace; For sandbox tests: the sandbox runtime package, bubblewrap and ripgrep.
SKILL.md names 2 domains. In commands or code: github.com and api.moonshot.cn; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
OpenHarness End-to-End Evals is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with OpenHarness End-to-End Evals: CodeGraph Agent Eval (colbymchenry/codegraph, 73k stars), Godot E2E (RandallLiuXin/GodotMaker, 549 stars), Interactive CLI Testing With tui-test (slopus/happy, 24k stars) and Test Pyramid (kubernetes-sigs/agent-sandbox, 4.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
HKUDS (a GitHub organization) maintains it in HKUDS/OpenHarness, which has 15,919 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on June 4, 2026.
Source: HKUDS/OpenHarness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.