Agent Eval Engineering
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
$ npx skills add amd/gaia --skill benchmarking-the-agent -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/gaia benchmarking-the-agent --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmarking-the-agent .claude/skills/benchmarking-the-agent && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmarking-the-agent" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/benchmarking-the-agent into .claude/skills/benchmarking-the-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking-the-agent", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/gaia/tree/main/.claude/skills/benchmarking-the-agentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/gaia --skill benchmarking-the-agent -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/gaia benchmarking-the-agent --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/benchmarking-the-agent .agents/skills/benchmarking-the-agent && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmarking-the-agent" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/benchmarking-the-agent into .agents/skills/benchmarking-the-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking-the-agent", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill benchmarking-the-agent -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/gaia benchmarking-the-agent --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/benchmarking-the-agent .cursor/skills/benchmarking-the-agent && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmarking-the-agent" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/benchmarking-the-agent into .cursor/skills/benchmarking-the-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking-the-agent", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/gaia.git --path .claude/skills/benchmarking-the-agent--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/gaia --skill benchmarking-the-agent -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/gaia benchmarking-the-agent --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/benchmarking-the-agent .gemini/skills/benchmarking-the-agent && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmarking-the-agent" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/benchmarking-the-agent into .gemini/skills/benchmarking-the-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking-the-agent", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/gaia benchmarking-the-agentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/gaia --skill benchmarking-the-agent -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/benchmarking-the-agent .github/skills/benchmarking-the-agent && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmarking-the-agent" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/benchmarking-the-agent into .github/skills/benchmarking-the-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking-the-agent", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill benchmarking-the-agent -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/gaia benchmarking-the-agent --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/benchmarking-the-agent .opencode/skills/benchmarking-the-agent && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmarking-the-agent" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/benchmarking-the-agent into .opencode/skills/benchmarking-the-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking-the-agent", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmarking-the-agentBenchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
The skill runs the same task suites through GAIA and through Claude Code with gaia eval tasks, scores them mechanically, then has an LLM judge grade the transcripts, answering whether the work gets done, whether the answer is honest about it and what it cost. The suites are everyday, with 14 tasks and the comparison table's suite, therock, with four real merged fixes, and core, full and adversarial, which CI gates on.
The main rule is to run serially: every run needs the local Lemonade model slot, so a pgrep check must show no other run and a run should never be killed mid-task. The budget is roughly 4 to 9 minutes per model for the 14-task suite plus about 2 minutes to judge. Recipes cover repeated runs of one model with a cost meter, the same model through Claude Code to isolate the harness, a Claude Code reference row, the TheRock tier, building the comparison report from run directories, and checking the judge with gaia eval tasks controls.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 05fb50b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
GAIA Agent Benchmarking loads about 1.8k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 994 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/gaia at commit 05fb50b, republished under its MIT licence (© amd). 994 words, ~1,834 tokens.
.claude/skills/benchmarking-the-agent/SKILL.md (or your agent's skills folder).gaia eval tasks runs the same tasks through GAIA and through Claude Code,
scores them mechanically, then has an LLM judge grade the transcripts. It
answers three questions: does the work get done, is the answer honest
about it, and what did it cost.
Everything here was learned by running it. The traps section is the part that
saves a day. The full reference is
docs/reference/eval.mdx.
| Suite | What it is |
|---|---|
everyday | The 14 tasks a developer actually asks for. The comparison table's suite. |
therock | Four real merged fixes in TheRock, from the commit before the fix. Two forbid the internet in the prompt. |
core / full / adversarial | What CI gates on: the traps and the committed expectations. |
Every run wants the local Lemonade model slot, and two at once race-evict each
other's models — the same reason CLAUDE.md gives for gaia eval agent. Before
starting anything:
pgrep -fl "[g]aia eval tasks" | wc -l # must be 0Budget 4–9 minutes per model for the 14-task suite depending on the model's
speed, plus about 2 minutes to judge it. Never pkill a run mid-task, and
never stop Lemonade or the daemon while one is in flight.
One model, three runs, real cost. The main loop of any harness change:
gaia eval tasks run --suite everyday --model fireworks.glm-5p3-flash \
--repeats 3 --meter fireworks --out runs/gaia-glmThe same model through Claude Code, so the difference is the harness:
gaia eval tasks run --suite everyday --harness claude-code \
--model fireworks.glm-5p3-flash --repeats 3 --out runs/cc-glmClaude Code on Anthropic's own model, as the reference row. Run Opus 5, not Sonnet: it is the strongest leg available, so a cost share reads as a fraction of the best agent and it sets an honest quality ceiling.
gaia eval tasks run --suite everyday --harness claude-code --model claude-opus-5 \
--out runs/cc-opusThe TheRock tier, a real 144K-line codebase:
gaia eval tasks run --suite therock --model fireworks.glm-5p3-flash --out runs/gaia-therockThe table. The first directory is the 100% row:
gaia eval tasks report runs/cc-opus runs/gaia-glm runs/cc-glm --out runs/reportThe judge's own controls, before you trust any quality number:
gaia eval tasks controls--repeats 3, and read
the min–max the table prints under each mean. A gap smaller than the range it
sits in is noise.metered is real money;
api_equivalent is Claude Code's list-price figure on a subscription, which
is a price of compute and not money spent; harness_counts is a rate-carded
estimate. The reports label them. Never sum or rank them together without
saying which is which.one_line for the worst.gaia.eval.bench.transcripts.tool_calls(transcript) yields (name, args, result) for either harness. Count what the agent actually did before blaming
a change: twice a confident hypothesis (trimmed tool descriptions; tool
ordering) died on inspection.gh
invocations halved while the same work got done in chained calls. A raw
call-count drop is not a behaviour regression.A wall-clock cap silently penalises slow models. At a 420 s cap, GLM-5.3
full averaged 182 s per task, got cut off three times, scored 40% and looked
incapable. The default is 1800 s; keep --run-timeout the same for every model
in a comparison and high enough that none of them binds.
Rate cards drift, and a wrong one rewrites every cost claim. A table once
had DeepSeek V4.1 Flash at $0.22/$0.007/$0.66 when the published card was
$0.30/$0.006/$1.20 — every Flash cost was ~30% low. Verify against
docs.fireworks.ai/serverless/pricing before quoting cost. Cards live in
src/gaia/eval/bench/metering.py::RATES; a model with no published card prints
n/a, never a guess. --meter fireworks bypasses the question entirely.
A metered run is account-wide. Two metered runs on one account at the same time pollute each other's deltas, and the meter trails by ~90 s.
Keyword rubrics are both lenient and brittle. They pass an answer
containing the right word without it establishing the point, and fail a correct
answer phrased differently. must_establish judged by the LLM replaced them.
Do not add new keyword checks.
The judge often costs more than the run it grades. A 14-task judge pass is
$0.4–0.9; GLM-5.3 Flash's entire run is $0.083. Do not re-judge a run you have
already judged — judge is a separate command for that reason.
Fresh input dominates GAIA's cost, and the first call of every task is uncached because each task is a new process. Within a process caching reaches ~95%; across processes Fireworks gives a new one nothing, while Claude Code's account-wide cache hits 91% across its runs. That asymmetry flatters Claude Code's token efficiency and is worth stating when reporting.
Graders must not write into the agent's workdir. A probe that writes
data.txt makes the judge blame the agent for a stray file. Grading runs after
the agent exits, and the diff is taken before the probes run — keep it that way
when adding a probe.
gaia eval tasks run --suite everyday --tasks 21-qa --model <id> --no-judge.metering.py::RATES, keyed by the bare model
name, or meter it with --meter fireworks.Numbers below are a snapshot from September 2026 — treat them as the shape to expect, and re-measure rather than quoting them.
On the 14-task suite, Claude Code at Opus 5 scored 14/14 at quality 4.97, truthfulness 5.00, 83 steps, 456 s. Against that ceiling GLM-5.3 Flash on GAIA reached 14/14 at quality 4.96 for $0.09 of real Fireworks charges, and DeepSeek V4.1 Flash 4.88 for $0.105. A model that costs more than the Sonnet reference and scores below it is not a candidate, however good its public benchmarks look.
Truthfulness is where the daylight is: the best legs score a clean 5.00, most others sit at 4.86. That axis — never claiming what the tool record does not support — is the one worth optimising next.
© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/benchmarking-the-agent of amd/gaia.
Open the folder on GitHubat commit 05fb50b
GAIA Agent Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GAIA Agent Benchmarking this skillamd/gaia | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Windmill AI Evalswindmill-labs/windmill | 18k | — | ~969 | Automated safety check: Notes | Custom licence | |
| Octocode Benchmark Runnerbgauryy/octocode | 949 | — | ~2.1k | Automated safety check: Pass | MIT | |
| SWE Benchmark Task Adderory/lumen | 305 | — | ~497 | Automated safety check: Pass | Custom licence | |
| Benchflowbenchflow-ai/benchflow | 353 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 |
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
windmill-labs/windmill
Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.
bgauryy/octocode
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
ory/lumen
Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.
benchflow-ai/benchflow
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
Dynatrace/dynatrace-for-ai
Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.
amd/gaia
Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.
amd/gaia
Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.
amd/gaia
Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.
amd/gaia
Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.
amd/gaia
Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.
amd/gaia
Turns a source document such as a README or spec into an executive slide deck as one self-contained HTML file that prints to PDF, one slide per page.
Categories
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks. The skill runs the same task suites through GAIA and through Claude Code with gaia eval tasks, scores them mechanically, then has an LLM judge grade the transcripts, answering whether the work gets done, whether the answer is honest about it and what it cost. The suites are everyday, with 14 tasks and the comparison table's suite, therock, with four real merged fixes, and core, full and adversarial, which CI gates on.
GAIA Agent Benchmarking fits situations like: comparing the GAIA agent with Claude Code on the same tasks; checking whether a change cost the agent quality; working out what a benchmark run costs; reproducing the harness by model comparison table.
Run `npx skills add amd/gaia --skill benchmarking-the-agent -a claude-code`. Or copy the skill folder (.claude/skills/benchmarking-the-agent in amd/gaia) into .claude/skills/benchmarking-the-agent in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/gaia --skill benchmarking-the-agent -a codex`. Or copy the skill folder (.claude/skills/benchmarking-the-agent in amd/gaia) into .agents/skills/benchmarking-the-agent in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill benchmarking-the-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmarking-the-agent, .gemini/skills/benchmarking-the-agent, .github/skills/benchmarking-the-agent and .opencode/skills/benchmarking-the-agent in your project.
SKILL.md names no scripts, command-line tools or credentials: GAIA Agent Benchmarking is instructions for the agent only. Our summary lists: The gaia CLI with a local Lemonade model slot; Claude Code, for the harness comparison runs.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
GAIA Agent Benchmarking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with GAIA Agent Benchmarking: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Windmill AI Evals (windmill-labs/windmill, 18k stars), Octocode Benchmark Runner (bgauryy/octocode, 949 stars) and SWE Benchmark Task Adder (ory/lumen, 305 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 8, 2026.
Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.