Agent skill

GAIA Agent Benchmarking

by amd in amd/gaia

Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

MITAuto-check passedAI & LLM Engineering

Install GAIA Agent Benchmarking

skills CLI
$ npx skills add amd/gaia --skill benchmarking-the-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/gaia benchmarking-the-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmarking-the-agent .claude/skills/benchmarking-the-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmarking-the-agent
GitHub stars
1.6k
Token cost
~1.8k tokens
SKILL.md length
994 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

  • Works in 3 steps: Confirm the gateway can reach it: `gaia… → Add its published card to… → Run it alongside at least one model you…
  • Comparing the GAIA agent with Claude Code on the same tasks
  • SKILL.md covers The suites, The one rule: serial, Recipes and Reading a result honestly, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill runs the same task suites through GAIA and through Claude Code with gaia eval tasks, scores them mechanically, then has an LLM judge grade the transcripts, answering whether the work gets done, whether the answer is honest about it and what it cost. The suites are everyday, with 14 tasks and the comparison table's suite, therock, with four real merged fixes, and core, full and adversarial, which CI gates on.

The main rule is to run serially: every run needs the local Lemonade model slot, so a pgrep check must show no other run and a run should never be killed mid-task. The budget is roughly 4 to 9 minutes per model for the 14-task suite plus about 2 minutes to judge. Recipes cover repeated runs of one model with a cost meter, the same model through Claude Code to isolate the harness, a Claude Code reference row, the TheRock tier, building the comparison report from run directories, and checking the judge with gaia eval tasks controls.

When your agent uses it

  • Comparing the GAIA agent with Claude Code on the same tasks
  • Checking whether a change cost the agent quality
  • Working out what a benchmark run costs
  • Reproducing the harness by model comparison table

Example prompts

  • “Run the everyday suite three times on one model and report the real cost.”
  • “Run the same model through Claude Code so we can see what the harness changes.”
  • “Build the comparison report from these three run directories.”

Requirements

  • The gaia CLI with a local Lemonade model slot
  • Claude Code, for the harness comparison runs

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the gateway can reach it: `gaia eval tasks run --suite everyday
  2. Add its published card to metering.py::RATES, keyed by the bare model
  3. Run it alongside at least one model you already have numbers for, so a

What it can do on your machine

Read from SKILL.md and the folder at commit 05fb50b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

GAIA Agent Benchmarking loads about 1.8k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 994 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/gaia at commit 05fb50b, republished under its MIT licence (© amd). 994 words, ~1,834 tokens.

Download SKILL.mdSave it as .claude/skills/benchmarking-the-agent/SKILL.md (or your agent's skills folder).
name
benchmarking-the-agent
description
Measure the flagship GAIA agent against Claude Code and across models — quality, truthfulness, steps, tokens, time and real cost — with `gaia eval tasks`. Use when asked to benchmark the agent, compare models or harnesses, check whether a change cost quality, work out what a run costs, or reproduce the harness × model table.

Benchmarking the flagship agent

gaia eval tasks runs the same tasks through GAIA and through Claude Code, scores them mechanically, then has an LLM judge grade the transcripts. It answers three questions: does the work get done, is the answer honest about it, and what did it cost.

Everything here was learned by running it. The traps section is the part that saves a day. The full reference is docs/reference/eval.mdx.

The suites

SuiteWhat it is
everydayThe 14 tasks a developer actually asks for. The comparison table's suite.
therockFour real merged fixes in TheRock, from the commit before the fix. Two forbid the internet in the prompt.
core / full / adversarialWhat CI gates on: the traps and the committed expectations.

The one rule: serial

Every run wants the local Lemonade model slot, and two at once race-evict each other's models — the same reason CLAUDE.md gives for gaia eval agent. Before starting anything:

bash
pgrep -fl "[g]aia eval tasks" | wc -l   # must be 0

Budget 4–9 minutes per model for the 14-task suite depending on the model's speed, plus about 2 minutes to judge it. Never pkill a run mid-task, and never stop Lemonade or the daemon while one is in flight.

Recipes

One model, three runs, real cost. The main loop of any harness change:

bash
gaia eval tasks run --suite everyday --model fireworks.glm-5p3-flash \
  --repeats 3 --meter fireworks --out runs/gaia-glm

The same model through Claude Code, so the difference is the harness:

bash
gaia eval tasks run --suite everyday --harness claude-code \
  --model fireworks.glm-5p3-flash --repeats 3 --out runs/cc-glm

Claude Code on Anthropic's own model, as the reference row. Run Opus 5, not Sonnet: it is the strongest leg available, so a cost share reads as a fraction of the best agent and it sets an honest quality ceiling.

bash
gaia eval tasks run --suite everyday --harness claude-code --model claude-opus-5 \
  --out runs/cc-opus

The TheRock tier, a real 144K-line codebase:

bash
gaia eval tasks run --suite therock --model fireworks.glm-5p3-flash --out runs/gaia-therock

The table. The first directory is the 100% row:

bash
gaia eval tasks report runs/cc-opus runs/gaia-glm runs/cc-glm --out runs/report

The judge's own controls, before you trust any quality number:

bash
gaia eval tasks controls

Reading a result honestly

  • One run is noise. A single model's quality spans about 0.3 across repeat runs. Never report a delta below that from n=1 — use --repeats 3, and read the min–max the table prints under each mean. A gap smaller than the range it sits in is noise.
  • Cost is three different kinds of dollar. metered is real money; api_equivalent is Claude Code's list-price figure on a subscription, which is a price of compute and not money spent; harness_counts is a rate-carded estimate. The reports label them. Never sum or rank them together without saying which is which.
  • Per-task deltas locate a regression; the aggregate only tells you one exists. Diff the per-task rows between two runs and read the judge's one_line for the worst.
  • Then check the transcript, not your theory. gaia.eval.bench.transcripts.tool_calls(transcript) yields (name, args, result) for either harness. Count what the agent actually did before blaming a change: twice a confident hypothesis (trimmed tool descriptions; tool ordering) died on inspection.
  • Watch for confounds. After compound shell commands landed, gh invocations halved while the same work got done in chained calls. A raw call-count drop is not a behaviour regression.
Show full SKILL.md (517 more words)Show less

Traps that cost real time

A wall-clock cap silently penalises slow models. At a 420 s cap, GLM-5.3 full averaged 182 s per task, got cut off three times, scored 40% and looked incapable. The default is 1800 s; keep --run-timeout the same for every model in a comparison and high enough that none of them binds.

Rate cards drift, and a wrong one rewrites every cost claim. A table once had DeepSeek V4.1 Flash at $0.22/$0.007/$0.66 when the published card was $0.30/$0.006/$1.20 — every Flash cost was ~30% low. Verify against docs.fireworks.ai/serverless/pricing before quoting cost. Cards live in src/gaia/eval/bench/metering.py::RATES; a model with no published card prints n/a, never a guess. --meter fireworks bypasses the question entirely.

A metered run is account-wide. Two metered runs on one account at the same time pollute each other's deltas, and the meter trails by ~90 s.

Keyword rubrics are both lenient and brittle. They pass an answer containing the right word without it establishing the point, and fail a correct answer phrased differently. must_establish judged by the LLM replaced them. Do not add new keyword checks.

The judge often costs more than the run it grades. A 14-task judge pass is $0.4–0.9; GLM-5.3 Flash's entire run is $0.083. Do not re-judge a run you have already judged — judge is a separate command for that reason.

Fresh input dominates GAIA's cost, and the first call of every task is uncached because each task is a new process. Within a process caching reaches ~95%; across processes Fireworks gives a new one nothing, while Claude Code's account-wide cache hits 91% across its runs. That asymmetry flatters Claude Code's token efficiency and is worth stating when reporting.

Graders must not write into the agent's workdir. A probe that writes data.txt makes the judge blame the agent for a stray file. Grading runs after the agent exits, and the diff is taken before the probes run — keep it that way when adding a probe.

Adding a model

  1. Confirm the gateway can reach it: gaia eval tasks run --suite everyday --tasks 21-qa --model <id> --no-judge.
  2. Add its published card to metering.py::RATES, keyed by the bare model name, or meter it with --meter fireworks.
  3. Run it alongside at least one model you already have numbers for, so a harness change cannot be mistaken for a model difference.

What good looks like

Numbers below are a snapshot from September 2026 — treat them as the shape to expect, and re-measure rather than quoting them.

On the 14-task suite, Claude Code at Opus 5 scored 14/14 at quality 4.97, truthfulness 5.00, 83 steps, 456 s. Against that ceiling GLM-5.3 Flash on GAIA reached 14/14 at quality 4.96 for $0.09 of real Fireworks charges, and DeepSeek V4.1 Flash 4.88 for $0.105. A model that costs more than the Sonnet reference and scores below it is not a candidate, however good its public benchmarks look.

Truthfulness is where the daylight is: the best legs score a clean 5.00, most others sit at 4.86. That axis — never claiming what the tool record does not support — is the one worth optimising next.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/benchmarking-the-agent of amd/gaia.

Open the folder on GitHubat commit 05fb50b

Compare with similar skills

GAIA Agent Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GAIA Agent Benchmarking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GAIA Agent Benchmarking this skillamd/gaia1.6k—~1.8kAutomated safety check: PassMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Windmill AI Evalswindmill-labs/windmill18k—~969Automated safety check: NotesCustom licence
Octocode Benchmark Runnerbgauryy/octocode949—~2.1kAutomated safety check: PassMIT
SWE Benchmark Task Adderory/lumen305—~497Automated safety check: PassCustom licence
Benchflowbenchflow-ai/benchflow353—~1.9kAutomated safety check: NotesApache-2.0

Similar skills

  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Windmill AI Evals

    windmill-labs/windmill

    Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

    18k GitHub stars~969 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

    305 GitHub stars~497 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Benchflow

    benchflow-ai/benchflow

    Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

    353 GitHub stars~1.9k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Dt Obs Genai

    Dynatrace/dynatrace-for-ai

    Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.

    161 GitHub stars~4.5k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed

More from amd/gaia

All 44 skills in this repo
  • Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

    1.6k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.

    1.6k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

    1.6k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Turns a source document such as a README or spec into an executive slide deck as one self-contained HTML file that prints to PDF, one slide per page.

    1.6k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Questions about GAIA Agent Benchmarking

What does GAIA Agent Benchmarking do?

Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks. The skill runs the same task suites through GAIA and through Claude Code with gaia eval tasks, scores them mechanically, then has an LLM judge grade the transcripts, answering whether the work gets done, whether the answer is honest about it and what it cost. The suites are everyday, with 14 tasks and the comparison table's suite, therock, with four real merged fixes, and core, full and adversarial, which CI gates on.

When should I use GAIA Agent Benchmarking?

GAIA Agent Benchmarking fits situations like: comparing the GAIA agent with Claude Code on the same tasks; checking whether a change cost the agent quality; working out what a benchmark run costs; reproducing the harness by model comparison table.

How do I install GAIA Agent Benchmarking in Claude Code?

Run `npx skills add amd/gaia --skill benchmarking-the-agent -a claude-code`. Or copy the skill folder (.claude/skills/benchmarking-the-agent in amd/gaia) into .claude/skills/benchmarking-the-agent in your project. Claude Code loads it when a task matches its description.

How do I install GAIA Agent Benchmarking in Codex?

Run `npx skills add amd/gaia --skill benchmarking-the-agent -a codex`. Or copy the skill folder (.claude/skills/benchmarking-the-agent in amd/gaia) into .agents/skills/benchmarking-the-agent in your project. Codex loads it when a task matches its description.

Can I use GAIA Agent Benchmarking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill benchmarking-the-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmarking-the-agent, .gemini/skills/benchmarking-the-agent, .github/skills/benchmarking-the-agent and .opencode/skills/benchmarking-the-agent in your project.

What does GAIA Agent Benchmarking need to run?

SKILL.md names no scripts, command-line tools or credentials: GAIA Agent Benchmarking is instructions for the agent only. Our summary lists: The gaia CLI with a local Lemonade model slot; Claude Code, for the harness comparison runs.

Does GAIA Agent Benchmarking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is GAIA Agent Benchmarking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does GAIA Agent Benchmarking use?

GAIA Agent Benchmarking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GAIA Agent Benchmarking use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to GAIA Agent Benchmarking?

Skills that share tags, products or a category with GAIA Agent Benchmarking: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Windmill AI Evals (windmill-labs/windmill, 18k stars), Octocode Benchmark Runner (bgauryy/octocode, 949 stars) and SWE Benchmark Task Adder (ory/lumen, 305 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GAIA Agent Benchmarking?

amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 8, 2026.

Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.