OpenHarness End-to-End Evals
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
A skill your agent uses when writing or running performance benchmarks for Jazz packages.
$ npx skills add garden-co/classic-jazz --skill benchmarking -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install garden-co/classic-jazz benchmarking --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/garden-co/classic-jazz.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/benchmarking .claude/skills/benchmarking && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmarking" agent skill from https://github.com/garden-co/classic-jazz/tree/main/.cursor/skills/benchmarking into .claude/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/garden-co/classic-jazz/tree/main/.cursor/skills/benchmarkingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add garden-co/classic-jazz --skill benchmarking -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install garden-co/classic-jazz benchmarking --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/garden-co/classic-jazz.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.cursor/skills/benchmarking .agents/skills/benchmarking && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmarking" agent skill from https://github.com/garden-co/classic-jazz/tree/main/.cursor/skills/benchmarking into .agents/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add garden-co/classic-jazz --skill benchmarking -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install garden-co/classic-jazz benchmarking --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/garden-co/classic-jazz.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.cursor/skills/benchmarking .cursor/skills/benchmarking && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmarking" agent skill from https://github.com/garden-co/classic-jazz/tree/main/.cursor/skills/benchmarking into .cursor/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/garden-co/classic-jazz.git --path .cursor/skills/benchmarking--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add garden-co/classic-jazz --skill benchmarking -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install garden-co/classic-jazz benchmarking --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/garden-co/classic-jazz.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.cursor/skills/benchmarking .gemini/skills/benchmarking && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmarking" agent skill from https://github.com/garden-co/classic-jazz/tree/main/.cursor/skills/benchmarking into .gemini/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install garden-co/classic-jazz benchmarkingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add garden-co/classic-jazz --skill benchmarking -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/garden-co/classic-jazz.git skills-src && mkdir -p .github/skills && cp -r skills-src/.cursor/skills/benchmarking .github/skills/benchmarking && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmarking" agent skill from https://github.com/garden-co/classic-jazz/tree/main/.cursor/skills/benchmarking into .github/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add garden-co/classic-jazz --skill benchmarking -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install garden-co/classic-jazz benchmarking --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/garden-co/classic-jazz.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.cursor/skills/benchmarking .opencode/skills/benchmarking && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmarking" agent skill from https://github.com/garden-co/classic-jazz/tree/main/.cursor/skills/benchmarking into .opencode/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmarkingA skill your agent uses when writing or running performance benchmarks for Jazz packages.
Benchmarking is an agent skill from garden-co/classic-jazz. Use this skill when writing or running performance benchmarks for Jazz packages. Covers cronometro setup, file conventions, gotchas with worker threads, and how to compare implementations.
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA. It works with npm. The repository describes itself as: A new kind of database that's distributed across your frontend, containers, serverless functions and its own storage cloud. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4f90501. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pnpmnodeFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmarking loads about 3.2k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 694 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from garden-co/classic-jazz at commit 4f90501, republished under its MIT licence (© garden-co). 694 words, ~3,184 tokens.
.claude/skills/benchmarking/SKILL.md (or your agent's skills folder).jazz-performance)All benchmarks live in the bench/ directory at the repository root:
bench/
├── package.json # Dependencies: cronometro, cojson, jazz-tools, vitest
├── jazz-tools/ # jazz-tools benchmarks
│ └── *.bench.tsBenchmark files follow the pattern: <subject>.<operation>.bench.ts
Each file should focus on a single benchmark comparing multiple implementations (e.g., @latest vs @workspace).
Examples:
comap.create.jazz-tools.bench.ts — benchmarks CoMap creationfilestream.getChunks.bench.ts — benchmarks FileStream.getChunks()filestream.asBase64.bench.ts — benchmarks FileStream.asBase64()binaryCoStream.write.bench.ts — benchmarks binary stream writesBenchmarks use cronometro, which runs each test in an isolated worker thread for accurate measurement.
import cronometro from "cronometro";
const TOTAL_BYTES = 5 * 1024 * 1024;
let data: SomeType;
await cronometro(
{
"operation - @latest": {
async before() {
// Setup — runs once before the test iterations
data = prepareTestData(TOTAL_BYTES);
},
test() {
// The code being benchmarked — runs many times
latestImplementation(data);
},
async after() {
// Cleanup — runs once after all iterations
cleanup();
},
},
"operation - @workspace": {
async before() {
data = prepareTestData(TOTAL_BYTES);
},
test() {
workspaceImplementation(data);
},
async after() {
cleanup();
},
},
},
{
iterations: 50,
warmup: true,
print: {
colors: true,
compare: true,
},
onTestError: (testName: string, error: unknown) => {
console.error(`\nError in test "${testName}":`);
console.error(error);
},
},
);Each benchmark file should have a single cronometro() call that compares multiple implementations of the same operation. This makes results easier to read and compare:
import cronometro from "cronometro";
const TOTAL_BYTES = 5 * 1024 * 1024;
let data: InputType;
await cronometro(
{
"operationName - @latest": {
async before() {
data = generateInput(TOTAL_BYTES);
},
test() {
latestImplementation(data);
},
async after() {
cleanup();
},
},
"operationName - @workspace": {
async before() {
data = generateInput(TOTAL_BYTES);
},
test() {
workspaceImplementation(data);
},
async after() {
cleanup();
},
},
},
{
iterations: 50,
warmup: true,
print: { colors: true, compare: true },
onTestError: (testName: string, error: unknown) => {
console.error(`\nError in test "${testName}":`);
console.error(error);
},
},
);Key principles:
getChunks, asBase64, write)@latest vs @workspace (or old vs new)const TOTAL_BYTES = 5 * 1024 * 1024)"operation - @implementation"To compare current workspace code against the latest published version:
1. Add npm aliases to bench/package.json:
{
"dependencies": {
"cojson": "workspace:*",
"cojson-latest": "npm:cojson@0.20.7",
"jazz-tools": "workspace:*",
"jazz-tools-latest": "npm:jazz-tools@0.20.7"
}
}Then run pnpm install in bench/.
2. Import both versions:
import * as localTools from "jazz-tools";
import * as latestPublishedTools from "jazz-tools-latest";
import { WasmCrypto as LocalWasmCrypto } from "cojson/crypto/WasmCrypto";
import { WasmCrypto as LatestPublishedWasmCrypto } from "cojson-latest/crypto/WasmCrypto";3. Use @ts-expect-error when passing the published package since the types won't match the workspace version:
ctx = await createContext(
// @ts-expect-error version mismatch
latestPublishedTools,
LatestPublishedWasmCrypto,
);When benchmarking CoValues (not standalone functions), create a full Jazz context. Use this helper pattern:
async function createContext(tools: typeof localTools, wasmCrypto: typeof LocalWasmCrypto) {
const ctx = await tools.createJazzContextForNewAccount({
creationProps: { name: "Bench Account" },
peers: [],
crypto: await wasmCrypto.create(),
sessionProvider: new tools.MockSessionProvider(),
});
return { account: ctx.account, node: ctx.node };
}Key points:
peers: [] — benchmarks don't need network syncMockSessionProvider — avoids real session persistence(ctx.node as any).gracefulShutdown() in after() to clean upDefine a fixed data size constant at the top of the file, then generate test data inside the before hook:
const TOTAL_BYTES = 5 * 1024 * 1024; // 5MB
let chunks: Uint8Array[];
await cronometro({
"operationName - @workspace": {
async before() {
chunks = makeChunks(TOTAL_BYTES, CHUNK_SIZE);
},
test() {
doWork(chunks);
},
},
}, options);Choose a size large enough to measure meaningfully. Small data (e.g., 100KB) may complete so fast that measurement noise dominates. 5MB is typically a good default for file/stream operations.
All fixture generation must be done inside the before hook, not at module level. This ensures data is created in the same worker thread that runs the test.
Add a script entry to bench/package.json:
{
"scripts": {
"bench:mytest": "node --experimental-strip-types --no-warnings ./jazz-tools/mytest.jazz-tools.bench.ts"
}
}Then run from the bench/ directory:
cd bench
pnpm run bench:mytestnode --experimental-strip-types, NOT tsxCronometro spawns worker threads that re-import the benchmark file. Workers don't inherit tsx's custom ESM loader, so the TypeScript import fails silently and the benchmark hangs forever.
Use node --experimental-strip-types --no-warnings instead:
"bench:foo": "node --experimental-strip-types --no-warnings ./jazz-tools/foo.bench.ts"before/after hooks MUST be async or accept a callbackCronometro's lifecycle hooks expect either:
A plain synchronous function that does neither will silently prevent the test from ever starting, causing the benchmark to hang indefinitely:
// BAD — test never starts, benchmark hangs
{
before() {
data = generateInput(); // sync, no callback, no promise
},
test() { ... },
}
// GOOD — async function returns a Promise
{
async before() {
data = generateInput();
},
test() { ... },
}
// ALSO GOOD — callback style
{
before(cb: () => void) {
data = generateInput();
cb();
},
test() { ... },
}test() can be sync or asyncUnlike before/after, the test function works correctly as a plain synchronous function. Make it async only if the code under test is genuinely asynchronous.
--experimental-strip-typesNode's type stripping handles annotations, as casts, and ! assertions. But it does not support:
enum declarations (use const objects instead)namespace declarationsconstructor(private x: number))import = / export = syntaxKeep benchmark files to simple TypeScript that only uses type annotations, interfaces, type aliases, and casts.
This example shows a benchmark comparing getChunks() between the published package and workspace code:
import cronometro from "cronometro";
import * as localTools from "jazz-tools";
import * as latestPublishedTools from "jazz-tools-latest";
import { WasmCrypto as LocalWasmCrypto } from "cojson/crypto/WasmCrypto";
import { cojsonInternals } from "cojson";
import { WasmCrypto as LatestPublishedWasmCrypto } from "cojson-latest/crypto/WasmCrypto";
const CHUNK_SIZE = cojsonInternals.TRANSACTION_CONFIG.MAX_RECOMMENDED_TX_SIZE;
const TOTAL_BYTES = 5 * 1024 * 1024;
function makeChunks(totalBytes: number, chunkSize: number): Uint8Array[] {
const chunks: Uint8Array[] = [];
let remaining = totalBytes;
while (remaining > 0) {
const size = Math.min(chunkSize, remaining);
const chunk = new Uint8Array(size);
for (let i = 0; i < size; i++) {
chunk[i] = Math.floor(Math.random() * 256);
}
chunks.push(chunk);
remaining -= size;
}
return chunks;
}
type Tools = typeof localTools;
async function createContext(tools: Tools, wasmCrypto: typeof LocalWasmCrypto) {
const ctx = await tools.createJazzContextForNewAccount({
creationProps: { name: "Bench Account" },
peers: [],
crypto: await wasmCrypto.create(),
sessionProvider: new tools.MockSessionProvider(),
});
return { account: ctx.account, node: ctx.node, FileStream: tools.FileStream };
}
function populateStream(ctx: Awaited<ReturnType<typeof createContext>>, chunks: Uint8Array[]) {
let totalBytes = 0;
for (const c of chunks) totalBytes += c.length;
const stream = ctx.FileStream.create({ owner: ctx.account });
stream.start({ mimeType: "application/octet-stream", totalSizeBytes: totalBytes });
for (const chunk of chunks) stream.push(chunk);
stream.end();
return stream;
}
const benchOptions = {
iterations: 50,
warmup: true,
print: { colors: true, compare: true },
onTestError: (testName: string, error: unknown) => {
console.error(`\nError in test "${testName}":`);
console.error(error);
},
};
let readCtx: Awaited<ReturnType<typeof createContext>>;
let readStream: ReturnType<typeof populateStream>;
await cronometro(
{
"getChunks - @latest": {
async before() {
readCtx = await createContext(
// @ts-expect-error version mismatch
latestPublishedTools,
LatestPublishedWasmCrypto,
);
readStream = populateStream(readCtx, makeChunks(TOTAL_BYTES, CHUNK_SIZE));
},
test() {
readStream.getChunks();
},
async after() {
(readCtx.node as any).gracefulShutdown();
},
},
"getChunks - @workspace": {
async before() {
readCtx = await createContext(localTools, LocalWasmCrypto);
readStream = populateStream(readCtx, makeChunks(TOTAL_BYTES, CHUNK_SIZE));
},
test() {
readStream.getChunks();
},
async after() {
(readCtx.node as any).gracefulShutdown();
},
},
},
benchOptions,
);filestream.getChunks.bench.ts)cronometro() call comparing @latest vs @workspaceconst TOTAL_BYTES = 5 * 1024 * 1024)bench/jazz-tools/ with *.bench.ts namingbench/package.json using node --experimental-strip-types --no-warningsbefore/after hooks are async (not plain sync)iterations set to at least 50 for stable resultswarmup: true enabledonTestError handler included to surface worker failures"operation - @implementation" (e.g., "getChunks - @workspace")bench/package.json and pnpm install rungracefulShutdown() called in after() hookbefore() hooks (not at module level or inside test())© garden-co, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .cursor/skills/benchmarking of garden-co/classic-jazz.
Open the folder on GitHubat commit 4f90501
Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmarking this skillgarden-co/classic-jazz | 2.5k | — | ~3.2k | Automated safety check: Pass | MIT | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| Qwen Code E2E TestingQwenLM/qwen-code | 28k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Quality Scandiegosouzapw/OmniRoute | 74k | — | ~748 | Automated safety check: Pass | MIT | |
| E2E TestingInsForge/InsForge | 13k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Analogjsanalogjs/analog | 3.2k | — | ~891 | Automated safety check: Pass | MIT |
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
QwenLM/qwen-code
Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.
diegosouzapw/OmniRoute
Runs a scoped, read-only quality scan on a repository candidate and reports exact evidence, failures and frozen debt, without treating a static scan as release acceptance.
InsForge/InsForge
A skill your agent uses when an InsForge maintainer has finished an OSS repo change and is ready to open, update, or submit the InsForge PR.
analogjs/analog
Conventions for working in an AnalogJS app (an Angular meta-framework powered by Vite) — file-based routing with .page.ts files, RouteMeta, server and API routes on Nitro/h3, page load functions and…
rstudio/rstudio
Runs RStudio's Playwright end-to-end tests against the desktop app or a server build with the project's npm scripts, asking you which mode to use first.
garden-co/classic-jazz
A skill your agent uses when optimizing Jazz applications for speed, responsiveness, and scalability.
garden-co/classic-jazz
A skill your agent uses when you need to write, review, or debug automated tests for applications built on the Jazz framework.
garden-co/classic-jazz
A skill your agent uses when building, debugging, or optimizing Jazz applications.
garden-co/classic-jazz
A skill your agent uses when designing data schemas, implementing sharing workflows, or auditing access control in Jazz applications.
garden-co/classic-jazz
Design and implement collaborative data schemas using the Jazz framework.
garden-co/classic-jazz
Generate changeset files for versioning and changelog management in this monorepo.
Works with
Categories
A skill your agent uses when writing or running performance benchmarks for Jazz packages. Benchmarking is an agent skill from garden-co/classic-jazz. Use this skill when writing or running performance benchmarks for Jazz packages.
Benchmarking fits situations like: running performance benchmarks for Jazz packages.
Run `npx skills add garden-co/classic-jazz --skill benchmarking -a claude-code`. Or copy the skill folder (.cursor/skills/benchmarking in garden-co/classic-jazz) into .claude/skills/benchmarking in your project. Claude Code loads it when a task matches its description.
Run `npx skills add garden-co/classic-jazz --skill benchmarking -a codex`. Or copy the skill folder (.cursor/skills/benchmarking in garden-co/classic-jazz) into .agents/skills/benchmarking in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add garden-co/classic-jazz --skill benchmarking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmarking, .gemini/skills/benchmarking, .github/skills/benchmarking and .opencode/skills/benchmarking in your project.
Going by SKILL.md and its folder, Benchmarking needs the command-line tools its instructions call (pnpm and node).
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Benchmarking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Benchmarking: OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars), Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars), Quality Scan (diegosouzapw/OmniRoute, 74k stars) and E2E Testing (InsForge/InsForge, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
garden-co (a GitHub organization) maintains it in garden-co/classic-jazz, which has 2,534 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 7, 2026.
Source: garden-co/classic-jazz on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.