Agent Eval Engineering
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
$ npx skills add benchflow-ai/benchflow --skill benchflow -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install benchflow-ai/benchflow benchflow --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchflow .claude/skills/benchflow && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchflow" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow into .claude/skills/benchflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflowType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add benchflow-ai/benchflow --skill benchflow -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install benchflow-ai/benchflow benchflow --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/benchflow .agents/skills/benchflow && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchflow" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow into .agents/skills/benchflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/benchflow --skill benchflow -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install benchflow-ai/benchflow benchflow --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/benchflow .cursor/skills/benchflow && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchflow" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow into .cursor/skills/benchflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/benchflow-ai/benchflow.git --path .agents/skills/benchflow--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add benchflow-ai/benchflow --skill benchflow -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install benchflow-ai/benchflow benchflow --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/benchflow .gemini/skills/benchflow && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchflow" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow into .gemini/skills/benchflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install benchflow-ai/benchflow benchflowInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add benchflow-ai/benchflow --skill benchflow -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/benchflow .github/skills/benchflow && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchflow" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow into .github/skills/benchflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/benchflow --skill benchflow -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install benchflow-ai/benchflow benchflow --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/benchflow .opencode/skills/benchflow && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchflow" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow into .opencode/skills/benchflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchflowRun agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
Benchflow is an agent skill from benchflow-ai/benchflow. Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. Use when asked to benchmark an AI coding agent, run a benchmark suite, create tasks, view trajectories, or compare agent performance.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including reference files (for example `references/create-task.md`, `references/dogfood.md` and `references/review-and-test.md`).
It sits in AI & LLM Engineering, covering Agent evaluation and testing and LLM evaluation. It works with Google Gemini. The repository describes itself as: Research infra for creating RL environments, post-training, and evals. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e965eee. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBashFrom allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
acpx.shFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GEMINI_API_KEYANTHROPIC_API_KEYOPENAI_API_KEYDAYTONA_API_KEYLLM_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchflow loads about 1.9k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 58 tokens; SKILL.md has 451 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Write, Edit, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from benchflow-ai/benchflow at commit e965eee, republished under its Apache-2.0 licence (© benchflow-ai). 451 words, ~1,891 tokens.
.claude/skills/benchflow/SKILL.md (or your agent's skills folder). This skill also uses 18 other files; get the full folder from GitHub.BenchFlow runs AI coding agents against tasks in sandboxed environments and scores their output via ACP (Agent Communication Protocol).
Arguments passed: $ARGUMENTS
status — show current stateuv tool list | grep benchflowbench agent listjobs/ (the default --jobs-dir)run <task-path> — run a single taskbench eval run \
--tasks-dir <task-path> \
--agent gemini \
--model gemini-3.1-flash-lite-preview \
--sandbox daytonaOr via Python SDK:
import asyncio
import benchflow as bf
from benchflow import RolloutConfig, Scene
from benchflow._utils.benchmark_repos import resolve_source
async def main():
config = RolloutConfig(
task_path=resolve_source("benchflow-ai/skillsbench", path="tasks/edit-pdf"),
scenes=[Scene.single(agent="gemini", model="gemini-3.1-flash-lite-preview")],
environment="daytona",
)
result = await bf.run(config)
print(f"Reward: {result.rewards}, Tools: {result.n_tool_calls}")
asyncio.run(main())Note: resolve_source() is required for remote repos in the SDK. The CLI
handles this transparently via --source-repo / --source-path.
API keys are auto-inherited from os.environ into the sandbox.
eval <tasks-dir> — run a benchmark suitebench eval run \
--source-repo benchflow-ai/skillsbench \
--source-path tasks \
--agent gemini \
--model gemini-3.1-flash-lite-preview \
--sandbox daytona \
--concurrency 64Or via YAML config:
bench eval run --config benchmarks/harvey-lab/harvey-lab-gemini-flash-lite.yamlYAML format:
source:
repo: benchflow-ai/skillsbench
path: tasks
agent: gemini
model: gemini-3.1-flash-lite-preview
environment: daytona
concurrency: 64
max_retries: 1metrics <jobs-dir> — analyze resultsbench eval metrics jobs/ # aggregate pass-rate / tokens / cost (add --json to pipe)
bench eval list jobs/ # per-rollout tableview <rollout-dir> — view a trajectoryResults land under jobs/<job-name>/<rollout-name>/ (the default --jobs-dir is jobs/):
rollout-dir/
├── result.json # rewards, agent, timing
├── prompts.json # prompts sent
├── trajectory/
│ └── acp_trajectory.jsonl # tool calls + agent thoughts
└── verifier/
├── reward.txt # reward value
└── ctrf.json # test resultscreate-task — create a new benchmark taskbench tasks init my-task # native task.md format (default)
bench tasks init my-task --no-pytest --no-oracle
bench tasks check tasks/my-task # structural validationQuick structure (native task.md format, the default):
my-task/
├── task.md # YAML frontmatter (config) + prompt body
├── environment/
│ └── Dockerfile # sandbox setup
├── verifier/
│ ├── test.sh # verifier entrypoint -> writes /logs/verifier/reward.txt
│ └── test_outputs.py
└── oracle/ # optional reference solution (solve.sh)--format legacy is retired in v0.6.2: bench tasks init always scaffolds a
native task.md package. To bring an existing split-layout task forward, run
bench tasks migrate <dir> --remove-legacy.
skills — discover and evaluate agent skillsbench skills list # discover skills on disk
bench skills eval skills/citation-management \
--agent claude-agent-acp # score a skill against its evals/evals.jsonhub — check external-environment-hub compatibilitybench hub check # inventory/structurally-check representative Harbor-registry tasksagents — list available agentsbench agent list| Agent | Protocol | Auth |
|---|---|---|
gemini | ACP | GEMINI_API_KEY or host login |
antigravity (alias: agy) | ACP (bundled shim over agy --output-format stream-json) | GEMINI_API_KEY |
claude-agent-acp (alias: claude) | ACP | ANTHROPIC_API_KEY or host login |
codex-acp (alias: codex) | ACP | OPENAI_API_KEY or host login |
opencode | ACP | inferred from model |
openhands (alias: oh) | ACP | LLM_API_KEY |
harvey-lab-harness (alias: harvey-lab) | ACP | Provider key matching model |
Any agent can be prefixed with acpx/ to run via ACPX (https://acpx.sh/):
bench eval run --tasks-dir tasks/edit-pdf --agent acpx/gemini --model gemini-3.1-flash-lite-preview --sandbox daytonaACPX is a headless ACP client with persistent sessions and crash recovery. The underlying agent's install, env vars, credentials, and skill paths are preserved.
compare — multi-agent comparisonCompare by running one config per agent (the agent: key lives in each YAML)
and printing the aggregate scores:
import asyncio
from benchflow.evaluation import Evaluation
async def main():
for config_path in [
"benchmarks/harvey-lab/harvey-lab-gemini-flash-lite.yaml",
"benchmarks/harvey-lab/harvey-lab-harness-parity.yaml",
]:
result = await Evaluation.from_yaml(config_path).run()
print(f"{config_path}: {result.passed}/{result.total} ({result.score:.1%})")
asyncio.run(main())# Install benchflow from PyPI. BenchFlow CLI releases require Python 3.12+.
uv tool install --python 3.12 --upgrade benchflow
# (or from source: uv sync --extra dev --locked)
export GEMINI_API_KEY=... # or ANTHROPIC_API_KEY, OPENAI_API_KEY, etc.
export DAYTONA_API_KEY=... # for cloud sandboxes| Sandbox | Flag | Best for |
|---|---|---|
docker | --sandbox docker | Local dev, small runs (<=10 tasks) |
daytona | --sandbox daytona | Cloud runs with concurrency (needs DAYTONA_API_KEY) |
modal | --sandbox modal | Serverless, high concurrency (needs Modal auth) |
Use daytona for benchmarks. Docker is limited by network exhaustion.
Two approaches for deploying skills:
COPY skills /root/.claude/skills--skills-dirbench eval run \
--tasks-dir task-dir \
--agent claude-agent-acp \
--sandbox daytona \
--skills-dir skills/ \
--skill-mode with-skill--skill-mode with-skill is required whenever you pass --skills-dir (omitting
it errors). Skills are uploaded to /skills/ in the sandbox and symlinked to
agent-specific paths.
gemini-3.1-flash-lite-preview for testing. Use Pro/Sonnet for real benchmarks.jobs_dir skips completed tasks.None in prompts list gets replaced with instruction.md content.0.5 to reward.txt).--agent-env GEMINI_API_KEY=... in CLI; SDK auto-inherits from os.environ.© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 18 other files (references) in .agents/skills/benchflow of benchflow-ai/benchflow.
Open the folder on GitHubat commit e965eee
Benchflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchflow this skillbenchflow-ai/benchflow | 353 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 791 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Autocontextgreyhaven-ai/autocontext | 1.3k | — | ~892 | Automated safety check: Pass | Apache-2.0 | |
| Cross-Model Benchmarkgarrytan/gstack | 136k | — | ~4k | Automated safety check: Notes | MIT | |
| GAIA Agent Benchmarkingamd/gaia | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT |
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
greyhaven-ai/autocontext
Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.
garrytan/gstack
Sends one prompt to Claude, GPT through the Codex CLI and Gemini, then tabulates response time, token use and cost, with an optional judged quality score.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
tokencanopy/e2a
Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.
benchflow-ai/benchflow
SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.
benchflow-ai/benchflow
SkillsBench task PR review — classifies the task track (standard / research / multimodal), runs static policy checks against the track-specific rubric, benchmarks the task across oracle plus Claude…
benchflow-ai/benchflow
Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.
benchflow-ai/benchflow
Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation…
benchflow-ai/benchflow
Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it.
benchflow-ai/benchflow
Delegate complex coding tasks to a specialist model. An agent skill from benchflow-ai/benchflow.
Works with
Categories
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. Benchflow is an agent skill from benchflow-ai/benchflow. Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
Benchflow fits situations like: asked to benchmark an AI coding agent; run a benchmark suite; view trajectories; compare agent performance.
Run `npx skills add benchflow-ai/benchflow --skill benchflow -a claude-code`. Or copy the skill folder (.agents/skills/benchflow in benchflow-ai/benchflow) into .claude/skills/benchflow in your project. Claude Code loads it when a task matches its description.
Run `npx skills add benchflow-ai/benchflow --skill benchflow -a codex`. Or copy the skill folder (.agents/skills/benchflow in benchflow-ai/benchflow) into .agents/skills/benchflow in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/benchflow --skill benchflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchflow, .gemini/skills/benchflow, .github/skills/benchflow and .opencode/skills/benchflow in your project.
Going by SKILL.md and its folder, Benchflow needs a shell for the scripts in its folder, the command-line tools its instructions call (uv) and credentials named GEMINI_API_KEY, ANTHROPIC_API_KEY, OPENAI_API_KEY and DAYTONA_API_KEY. Our summary lists: Python 3; A Bash shell; Docker; A credential in GEMINI_API_KEY; A credential in ANTHROPIC_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash.
SKILL.md names 1 domain. As links in the text: acpx.sh. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Benchflow is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Benchflow: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 791 stars), Autocontext (greyhaven-ai/autocontext, 1.3k stars) and Cross-Model Benchmark (garrytan/gstack, 136k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
benchflow-ai (a GitHub organization) maintains it in benchflow-ai/benchflow, which has 353 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 6, 2026.
Source: benchflow-ai/benchflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.