Agent skill

Benchflow

by benchflow-ai in benchflow-ai/benchflow

Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Benchflow

skills CLI
$ npx skills add benchflow-ai/benchflow --skill benchflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/benchflow benchflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchflow .claude/skills/benchflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchflow
GitHub stars
353
Token cost
~1.9k tokens
SKILL.md length
451 words
Files
19 (incl. references)
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

  • Works in 5 steps: Check if benchflow is installed: uv tool… → Check if API keys are set… → Check available agents: bench agent list → …
  • Asked to benchmark an AI coding agent
  • SKILL.md covers Dispatch on arguments, Setup, Sandboxes and Skills in tasks, plus 1 more section
  • Runs Shell scripts from its folder; calls uv; needs GEMINI_API_KEY and ANTHROPIC_API_KEY

What it does

Benchflow is an agent skill from benchflow-ai/benchflow. Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. Use when asked to benchmark an AI coding agent, run a benchmark suite, create tasks, view trajectories, or compare agent performance.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including reference files (for example `references/create-task.md`, `references/dogfood.md` and `references/review-and-test.md`).

It sits in AI & LLM Engineering, covering Agent evaluation and testing and LLM evaluation. It works with Google Gemini. The repository describes itself as: Research infra for creating RL environments, post-training, and evals. The licence is Apache-2.0.

When your agent uses it

  • Asked to benchmark an AI coding agent
  • Run a benchmark suite
  • View trajectories
  • Compare agent performance

Example prompts

  • “/benchflow”

Requirements

  • Python 3
  • A Bash shell
  • Docker
  • A credential in GEMINI_API_KEY
  • A credential in ANTHROPIC_API_KEY
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Check if benchflow is installed: uv tool list | grep benchflow
  2. Check if API keys are set (GEMINI_API_KEY, ANTHROPIC_API_KEY, etc.)
  3. Check available agents: bench agent list
  4. Show recent eval results if any exist under jobs/ (the default --jobs-dir)
  5. Point to next action based on state

What it can do on your machine

Read from SKILL.md and the folder at commit e965eee. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • acpx.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • ANTHROPIC_API_KEY
    • OPENAI_API_KEY
    • DAYTONA_API_KEY
    • LLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchflow loads about 1.9k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 58 tokens; SKILL.md has 451 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/benchflow at commit e965eee, republished under its Apache-2.0 licence (© benchflow-ai). 451 words, ~1,891 tokens.

Download SKILL.mdSave it as .claude/skills/benchflow/SKILL.md (or your agent's skills folder). This skill also uses 18 other files; get the full folder from GitHub.
name
benchflow
description
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. Use when asked to benchmark an AI coding agent, run a benchmark suite, create tasks, view trajectories, or compare agent performance.
allowed-tools
Read, Write, Edit, Bash
user-invocable
true

BenchFlow — Agent Benchmarking

BenchFlow runs AI coding agents against tasks in sandboxed environments and scores their output via ACP (Agent Communication Protocol).

Arguments passed: $ARGUMENTS


Dispatch on arguments

No args or status — show current state
  1. Check if benchflow is installed: uv tool list | grep benchflow
  2. Check if API keys are set (GEMINI_API_KEY, ANTHROPIC_API_KEY, etc.)
  3. Check available agents: bench agent list
  4. Show recent eval results if any exist under jobs/ (the default --jobs-dir)
  5. Point to next action based on state
run <task-path> — run a single task
bash
bench eval run \
  --tasks-dir <task-path> \
  --agent gemini \
  --model gemini-3.1-flash-lite-preview \
  --sandbox daytona

Or via Python SDK:

python
import asyncio
import benchflow as bf
from benchflow import RolloutConfig, Scene
from benchflow._utils.benchmark_repos import resolve_source

async def main():
    config = RolloutConfig(
        task_path=resolve_source("benchflow-ai/skillsbench", path="tasks/edit-pdf"),
        scenes=[Scene.single(agent="gemini", model="gemini-3.1-flash-lite-preview")],
        environment="daytona",
    )
    result = await bf.run(config)
    print(f"Reward: {result.rewards}, Tools: {result.n_tool_calls}")

asyncio.run(main())

Note: resolve_source() is required for remote repos in the SDK. The CLI handles this transparently via --source-repo / --source-path.

API keys are auto-inherited from os.environ into the sandbox.

eval <tasks-dir> — run a benchmark suite
bash
bench eval run \
  --source-repo benchflow-ai/skillsbench \
  --source-path tasks \
  --agent gemini \
  --model gemini-3.1-flash-lite-preview \
  --sandbox daytona \
  --concurrency 64

Or via YAML config:

bash
bench eval run --config benchmarks/harvey-lab/harvey-lab-gemini-flash-lite.yaml

YAML format:

yaml
source:
  repo: benchflow-ai/skillsbench
  path: tasks
agent: gemini
model: gemini-3.1-flash-lite-preview
environment: daytona
concurrency: 64
max_retries: 1
metrics <jobs-dir> — analyze results
bash
bench eval metrics jobs/      # aggregate pass-rate / tokens / cost (add --json to pipe)
bench eval list jobs/         # per-rollout table
view <rollout-dir> — view a trajectory

Results land under jobs/<job-name>/<rollout-name>/ (the default --jobs-dir is jobs/):

rollout-dir/
├── result.json              # rewards, agent, timing
├── prompts.json             # prompts sent
├── trajectory/
│   └── acp_trajectory.jsonl # tool calls + agent thoughts
└── verifier/
    ├── reward.txt           # reward value
    └── ctrf.json            # test results
create-task — create a new benchmark task
bash
bench tasks init my-task                       # native task.md format (default)
bench tasks init my-task --no-pytest --no-oracle
bench tasks check tasks/my-task                # structural validation

Quick structure (native task.md format, the default):

my-task/
├── task.md            # YAML frontmatter (config) + prompt body
├── environment/
│   └── Dockerfile     # sandbox setup
├── verifier/
│   ├── test.sh        # verifier entrypoint -> writes /logs/verifier/reward.txt
│   └── test_outputs.py
└── oracle/            # optional reference solution (solve.sh)

--format legacy is retired in v0.6.2: bench tasks init always scaffolds a native task.md package. To bring an existing split-layout task forward, run bench tasks migrate <dir> --remove-legacy.

skills — discover and evaluate agent skills
bash
bench skills list                                   # discover skills on disk
bench skills eval skills/citation-management \
  --agent claude-agent-acp                          # score a skill against its evals/evals.json
hub — check external-environment-hub compatibility
bash
bench hub check          # inventory/structurally-check representative Harbor-registry tasks
agents — list available agents
bash
bench agent list
AgentProtocolAuth
geminiACPGEMINI_API_KEY or host login
antigravity (alias: agy)ACP (bundled shim over agy --output-format stream-json)GEMINI_API_KEY
claude-agent-acp (alias: claude)ACPANTHROPIC_API_KEY or host login
codex-acp (alias: codex)ACPOPENAI_API_KEY or host login
opencodeACPinferred from model
openhands (alias: oh)ACPLLM_API_KEY
harvey-lab-harness (alias: harvey-lab)ACPProvider key matching model

Any agent can be prefixed with acpx/ to run via ACPX (https://acpx.sh/):

bash
bench eval run --tasks-dir tasks/edit-pdf --agent acpx/gemini --model gemini-3.1-flash-lite-preview --sandbox daytona

ACPX is a headless ACP client with persistent sessions and crash recovery. The underlying agent's install, env vars, credentials, and skill paths are preserved.

Show full SKILL.md (155 more words)Show less
compare — multi-agent comparison

Compare by running one config per agent (the agent: key lives in each YAML) and printing the aggregate scores:

python
import asyncio
from benchflow.evaluation import Evaluation

async def main():
    for config_path in [
        "benchmarks/harvey-lab/harvey-lab-gemini-flash-lite.yaml",
        "benchmarks/harvey-lab/harvey-lab-harness-parity.yaml",
    ]:
        result = await Evaluation.from_yaml(config_path).run()
        print(f"{config_path}: {result.passed}/{result.total} ({result.score:.1%})")

asyncio.run(main())

Setup

bash
# Install benchflow from PyPI. BenchFlow CLI releases require Python 3.12+.
uv tool install --python 3.12 --upgrade benchflow
# (or from source: uv sync --extra dev --locked)
export GEMINI_API_KEY=...     # or ANTHROPIC_API_KEY, OPENAI_API_KEY, etc.
export DAYTONA_API_KEY=...    # for cloud sandboxes

Sandboxes

SandboxFlagBest for
docker--sandbox dockerLocal dev, small runs (<=10 tasks)
daytona--sandbox daytonaCloud runs with concurrency (needs DAYTONA_API_KEY)
modal--sandbox modalServerless, high concurrency (needs Modal auth)

Use daytona for benchmarks. Docker is limited by network exhaustion.

Skills in tasks

Two approaches for deploying skills:

Baked into Docker image (existing tasks)
dockerfile
COPY skills /root/.claude/skills
Runtime deployment via --skills-dir
bash
bench eval run \
  --tasks-dir task-dir \
  --agent claude-agent-acp \
  --sandbox daytona \
  --skills-dir skills/ \
  --skill-mode with-skill

--skill-mode with-skill is required whenever you pass --skills-dir (omitting it errors). Skills are uploaded to /skills/ in the sandbox and symlinked to agent-specific paths.

Tips

  • Use gemini-3.1-flash-lite-preview for testing. Use Pro/Sonnet for real benchmarks.
  • Evaluations resume — re-running the same jobs_dir skips completed tasks.
  • None in prompts list gets replaced with instruction.md content.
  • Partial rewards work (verifier can write 0.5 to reward.txt).
  • GEMINI_API_KEY requires explicit --agent-env GEMINI_API_KEY=... in CLI; SDK auto-inherits from os.environ.

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 18 other files (references) in .agents/skills/benchflow of benchflow-ai/benchflow.

  • SKILL.md
  • references/create-task.md
  • references/dogfood.md
  • references/review-and-test.md
  • references/smoke-test.md
  • tasks/benchflow-knowledge/environment/Dockerfile
  • tasks/benchflow-knowledge/instruction.md
  • tasks/benchflow-knowledge/task.toml
  • tasks/benchflow-knowledge/tests/test.sh
  • tasks/create-simple-task/environment/Dockerfile
  • tasks/create-simple-task/instruction.md
  • tasks/create-simple-task/task.toml
  • tasks/create-simple-task/tests/test.sh
  • … and 6 more

Open the folder on GitHubat commit e965eee

Compare with similar skills

Benchflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchflow this skillbenchflow-ai/benchflow353—~1.9kAutomated safety check: NotesApache-2.0
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples791—~2kAutomated safety check: PassApache-2.0
Autocontextgreyhaven-ai/autocontext1.3k—~892Automated safety check: PassApache-2.0
Cross-Model Benchmarkgarrytan/gstack136k—~4kAutomated safety check: NotesMIT
GAIA Agent Benchmarkingamd/gaia1.6k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    791 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Autocontext

    greyhaven-ai/autocontext

    Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

    1.3k GitHub stars~892 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cross-Model Benchmark

    garrytan/gstack

    Sends one prompt to Claude, GPT through the Codex CLI and Gemini, then tabulates response time, token use and cost, with an optional judged quality score.

    136k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Email Evals

    tokencanopy/e2a

    Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.

    192 GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from benchflow-ai/benchflow

All 8 skills in this repo
  • Task Creator

    benchflow-ai/benchflow

    SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

    353 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Task Review

    benchflow-ai/benchflow

    SkillsBench task PR review — classifies the task track (standard / research / multimodal), runs static policy checks against the track-specific rubric, benchmarks the task across oracle plus Claude…

    353 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check: notes
  • Benchflow Experiment Review

    benchflow-ai/benchflow

    Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.

    353 GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Benchflow Traj Upload Ops

    benchflow-ai/benchflow

    Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation…

    353 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Benchflow Traj Upload

    benchflow-ai/benchflow

    Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it.

    353 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes
  • Code Specialist

    benchflow-ai/benchflow

    Delegate complex coding tasks to a specialist model. An agent skill from benchflow-ai/benchflow.

    353 GitHub stars~225 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Benchflow

What does Benchflow do?

Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. Benchflow is an agent skill from benchflow-ai/benchflow. Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

When should I use Benchflow?

Benchflow fits situations like: asked to benchmark an AI coding agent; run a benchmark suite; view trajectories; compare agent performance.

How do I install Benchflow in Claude Code?

Run `npx skills add benchflow-ai/benchflow --skill benchflow -a claude-code`. Or copy the skill folder (.agents/skills/benchflow in benchflow-ai/benchflow) into .claude/skills/benchflow in your project. Claude Code loads it when a task matches its description.

How do I install Benchflow in Codex?

Run `npx skills add benchflow-ai/benchflow --skill benchflow -a codex`. Or copy the skill folder (.agents/skills/benchflow in benchflow-ai/benchflow) into .agents/skills/benchflow in your project. Codex loads it when a task matches its description.

Can I use Benchflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/benchflow --skill benchflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchflow, .gemini/skills/benchflow, .github/skills/benchflow and .opencode/skills/benchflow in your project.

What does Benchflow need to run?

Going by SKILL.md and its folder, Benchflow needs a shell for the scripts in its folder, the command-line tools its instructions call (uv) and credentials named GEMINI_API_KEY, ANTHROPIC_API_KEY, OPENAI_API_KEY and DAYTONA_API_KEY. Our summary lists: Python 3; A Bash shell; Docker; A credential in GEMINI_API_KEY; A credential in ANTHROPIC_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash.

Does Benchflow access the network?

SKILL.md names 1 domain. As links in the text: acpx.sh. This is read from the text; nothing was executed.

Is Benchflow safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Benchflow use?

Benchflow is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchflow use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to Benchflow?

Skills that share tags, products or a category with Benchflow: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 791 stars), Autocontext (greyhaven-ai/autocontext, 1.3k stars) and Cross-Model Benchmark (garrytan/gstack, 136k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchflow?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/benchflow, which has 353 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 6, 2026.

Source: benchflow-ai/benchflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.