Agent skill

GAIA Agent Eval Scorecard

by amd in amd/gaia

Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

MITAuto-check passedDevelopment

Install GAIA Agent Eval Scorecard

skills CLI
$ npx skills add amd/gaia --skill adding-eval-scorecard -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/gaia adding-eval-scorecard --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/adding-eval-scorecard .claude/skills/adding-eval-scorecard && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
adding-eval-scorecard
GitHub stars
1.6k
Token cost
~2.6k tokens
SKILL.md length
1,075 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

  • Works in 5 steps: Locate the agent's surfaces → Write the adapter (harness → payload) → Run the REAL eval (hard gate — no… → …
  • Adding a release eval scorecard to a new GAIA hub agent
  • SKILL.md covers Phase 1 — Locate the agent's…, Phase 2 — Write the adapter…, Phase 3 — Run the REAL eval… and Phase 4 — Surface, link, and…, plus 3 more sections
  • Calls python and git

What it does

This is a phased checklist for adopting GAIA's eval scorecard system, which flows from a harness through a result payload to a generator and a scorecard, backed by a standalone presence-and-regression release gate. The email agent is named as the reference implementation to mirror, and core modules such as release_scorecard.py and scorecard_gate.py are marked reusable and not to be modified.

The first phase locates the agent's surfaces: the version comes only from the version field in the agent's gaia-agent.yaml, the canonical README is whichever file the agent's release workflow actually publishes, and the scorecard lives beside it as a single SCORECARD.md updated in place rather than in a separate folder. The checklist has a hard gate at the real-eval step: the scorecard must come from an actual eval run over the agent's existing harness, such as a benchmark command over its test fixtures, and the skill says to stop and propose a minimal harness rather than invent numbers if no eval vehicle exists yet.

When your agent uses it

  • Adding a release eval scorecard to a new GAIA hub agent
  • Wiring an existing agent's eval results into its README and release gate
  • Generalizing the scorecard format while adopting it for a new agent

Example prompts

  • “Add a scorecard for the calendar agent, following the email agent's pattern.”
  • “Generate the scorecard for the browsing agent from its existing benchmark harness.”
  • “Wire the release gate for the scorecard we just added to the email agent.”

Requirements

  • A GAIA hub agent with gaia-agent.yaml
  • An existing eval harness for that agent
  • Python with PyYAML

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Locate the agent's surfaces
  2. Write the adapter (harness → payload)
  3. Run the REAL eval (hard gate — no hand-authored numbers)
  4. Surface, link, and gate
  5. Verify (evidence before "done")

What it can do on your machine

Read from SKILL.md and the folder at commit 05fb50b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

GAIA Agent Eval Scorecard loads about 2.6k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,075 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/gaia at commit 05fb50b, republished under its MIT licence (© amd). 1,075 words, ~2,574 tokens.

Download SKILL.mdSave it as .claude/skills/adding-eval-scorecard/SKILL.md (or your agent's skills folder).
name
adding-eval-scorecard
description
Adopt the per-agent eval scorecard for a GAIA hub agent: write the harness→payload adapter, run the eval to produce a REAL scorecard, link + surface it from the agent's README, wire the release gate, and (for a new agent) generalize the format. Use when asked to 'add a scorecard', 'adopt the eval scorecard', 'generate the scorecard for <agent>', or wire scorecard CI for an agent. Builds on docs/reference/eval-scorecard.mdx and the email agent reference adapter.

Adding an Eval Scorecard to a GAIA Agent

Adopt the release eval scorecard (docs/reference/eval-scorecard.mdx) for one hub agent. The system is harness → result payload → generator → scorecard, with a standalone presence+regression release gate. The email agent is the reference implementation — mirror it.

Core modules (do not modify; reuse):

  • src/gaia/eval/release_scorecard.py — ResultPayload, compute_aggregate, render_scorecard, write_scorecard, validate_scorecard, carry_forward. Harness-agnostic (stdlib + PyYAML only).
  • src/gaia/eval/scorecard_gate.py — the standalone gate (python -m gaia.eval.scorecard_gate).
  • Reference adapter: hub/agents/email/python/packaging/gen_scorecard.py.

This is a phased checklist with a hard gate at the real-eval step — the scorecard MUST come from an actual eval run, never hand-authored numbers.

Phase 1 — Locate the agent's surfaces

  1. Version source of truth = the version: field in <agent>/gaia-agent.yaml. Never invent a parallel scheme.

  2. Canonical README (where the scorecard is linked + surfaced): for an npm-published agent it is the npm client README (e.g. hub/agents/<id>/npm/README.md), NOT a packaging/README.md. For a Python-only agent it is hub/agents/<id>/python/README.md. Confirm which by checking what release_agent_<id>.yml publishes (README: env) — the published README is the one to link.

  3. doc-root = the directory holding that canonical README. The scorecard lives at <doc-root>/SCORECARD.md — a single file updated in place, versioned via the publish snapshot (same as README.md). There is no scorecards/ directory.

  4. Eval vehicle: what existing harness produces this agent's accuracy metric? (email → gaia eval benchmark over tests/fixtures/email/.) If none exists, STOP and surface that — propose the minimal harness before building; do not invent numbers.

    Expect none. Email is the only agent with a corpus + harness; for every other agent this step is the long pole, not a lookup. Two traps: an agent_type in the scenario corpus may name a ChatAgent prompt profile rather than your package (so a green scenario is measuring something else), and the ground truth itself is human judgment the factory does not automate. Build the corpus + deterministic fixture harness first — see porting-agent-to-hub Phase 4 — then return here for the adapter and gate.

Phase 2 — Write the adapter (harness → payload)

Copy hub/agents/email/python/packaging/gen_scorecard.py as the template. The adapter:

  • imports ONLY gaia.eval.release_scorecard (never the harness or agent package — preserve loose coupling);
  • reads the harness output, builds a ResultPayload;
  • populates reproduction_command with the exact shell commands to reproduce this scorecard, including all required env vars (PYTHON_KEYRING_BACKEND, GAIA_AGENT_TOOL_TIMEOUT, PYTHONPATH);
  • defines "judged" explicitly and raises loudly if zero results are judged (no silent 0.0);
  • records dataset size (total labeled examples) and test_cases_run (subset executed) as DISTINCT fields;
  • stores repo-relative paths only (never a local absolute path — it ships in a published artifact);
  • records the eval limit/config so future regression checks are comparable;
  • optionally populates environment (gaia_commit, lemonade_version, model, hardware — a class descriptor, never a hostname) and breakdown (per_category accuracy + top_confusions) — additive blocks that pin the run and explain the misses without ever affecting aggregate.value. Capture live environment (git SHA, server health version) in main(), not build_payload, so the payload builder stays pure and unit-testable;
  • writes to <doc-root>/SCORECARD.md (the single file; --output-dir overrides to a directory, but the filename is always SCORECARD.md).

Add an offline unit test against a committed sample harness-output fixture (see tests/fixtures/eval/email_benchmark_scorecard.json + tests/unit/eval/test_release_scorecard.py::TestEmailAdapter) so the adapter is testable without a live model.

Phase 3 — Run the REAL eval (hard gate — no hand-authored numbers)

The accuracy number must come from an actual run. For the email agent:

bash
# Real eval needs Lemonade + the model. Prefer AMD hardware (Strix Halo / Ryzen AI);
# the [self-hosted, lemonade-eval] runner is the canonical environment.
GAIA_AGENT_TOOL_TIMEOUT=1800 \
PYTHON_KEYRING_BACKEND=keyring.backends.null.Keyring \
PYTHONPATH="$(pwd)" \
  <venv>/bin/gaia eval benchmark \
    --model Gemma-4-E4B-it-GGUF \
    --mbox-path tests/fixtures/email/_stub_inbox.mbox \
    --ground-truth tests/fixtures/email/<task>_ground_truth.json \
    --limit 220 --output-dir <persistent-dir>

PYTHONPATH="$(pwd)" \
<venv>/bin/python hub/agents/email/python/packaging/gen_scorecard.py \
    --benchmark-dir <persistent-dir> --limit 220
# → writes hub/agents/email/npm/SCORECARD.md in place

⚠️ Run evals SERIALLY — never two at once (CLAUDE.md). A scorecard refresh loop is the most likely thing to violate this: chain runs with &&, never background them with &. Concurrent runs race-evict each other's models on the single-tenant Lemonade backend and produce garbage (ctx_size errors, spurious INFRA_ERROR) that looks like a real regression. Before each run: ps aux | grep "gaia eval" | grep -v grep | wc -l must be 0.

Pass both fixture flags explicitly. ls tests/fixtures/email/ and pick the real files — ground truth is task-specific (action_items_, briefing_, drafting_, followups_, longthread_), and the CLI's --ground-truth default points at a ground_truth.json that does not exist. Omit the flag and the run prints one [WARN] and scores nothing, which reads as a successful eval right up until the scorecard is empty.

Headless gotchas (see memory project-email-benchmark-headless-gotchas):

  • PYTHON_KEYRING_BACKEND=keyring.backends.null.Keyring — the email agent's calendar-connector resolution blocks forever on the macOS Keychain (and can stall on Linux SecretService) in non-interactive contexts. Without this it hangs at 0% CPU during agent construction.
  • PYTHONPATH="$(pwd)" — the benchmark imports tests.fixtures.email.*; the console script doesn't add the repo root.
  • GAIA_AGENT_TOOL_TIMEOUT=1800 — triage of the full corpus is one tool call (~17 min for ~100 emails on a 4B model); a lower timeout (the 180s default, or even 900s) abandons it mid-run, yielding a degenerate 0-email FAIL run.
  • Write --output-dir to a persistent dir, not /tmp (cleared on session resume).
  • Record honestly: if the metric is low for a known reason (e.g. a taxonomy/label mismatch), put the explanation in the adapter's methodology string and link the tracking issue — never inflate the number.
Show full SKILL.md (334 more words)Show less
  1. Link + surface from the canonical README: a one-line Eval scorecard (vX.Y.Z): aggregate N/100 … ([./SCORECARD.md](./SCORECARD.md)). The relative link must resolve in-repo.
  2. npm files: if the agent publishes on npm, add SCORECARD.md to package.json files. Do not add a scorecards/ directory — only the single current file ships.
  3. Hub display: a published scorecard surfaces on the agent's hub page / Agent UI detail view (see workers/agent-hub + AgentDetailModal.tsx); ensure the publish step passes --eval-scorecard <doc-root>/SCORECARD.md to publish_to_r2.py.
  4. Release gate: add a scorecard-gate job to release_agent_<id>.yml and list it in publish.needs. The job runs on a GitHub-hosted runner (it only parses committed files — no eval):
    bash
    # Presence-only (no previous tag yet):
    python -m gaia.eval.scorecard_gate \
      --scorecard <doc-root>/SCORECARD.md
    
    # With best-effort previous-release baseline (recommended for CI):
    PREV="$(git describe --tags --abbrev=0 --match 'agent-pkg-<id>-*' "${GITHUB_REF_NAME}^" 2>/dev/null || true)"
    if [ -n "$PREV" ]; then
      python -m gaia.eval.scorecard_gate \
        --scorecard <doc-root>/SCORECARD.md --baseline-ref "$PREV"
    else
      python -m gaia.eval.scorecard_gate \
        --scorecard <doc-root>/SCORECARD.md
    fi
    The job must NOT have continue-on-error, an environment:, or a permissions: override (inherits contents: read; needs no secrets). Fetch full history (fetch-depth: 0) so git describe resolves.
  5. Auto-update/reject loop: for re-running on agent changes and refreshing the committed scorecard, see eval-scorecard.mdx "Keeping the scorecard current" and the self-hosted refresh workflow — reject-on-worse is the gate; better-or-equal refreshes the committed SCORECARD.md.

Phase 5 — Verify (evidence before "done")

Run and capture: the generated SCORECARD.md; the gate passing on it (exit 0); the gate blocking a manufactured regression (exit 1, via --baseline-file with a higher-scoring card) and a missing card (exit 1); a by-hand recompute of the aggregate from aggregate.components matching the recorded value. Run python util/lint.py --all and the eval unit tests. These are the PR's real-world proof.

Versioning

  • Patch release → carry_forward(prev_scorecard_path, new_version) reads the version from the front matter of the current SCORECARD.md (not from the filename) and copies results verbatim, sets inherited_from; do NOT re-run the eval.
  • Minor/major release → re-run the eval (Phase 3); carry_forward refuses a non-patch bump with a "re-run" error.

Output style

Follow CLAUDE.md → "How You Communicate".

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/adding-eval-scorecard of amd/gaia.

Open the folder on GitHubat commit 05fb50b

Compare with similar skills

GAIA Agent Eval Scorecard next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GAIA Agent Eval Scorecard compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GAIA Agent Eval Scorecard this skillamd/gaia1.6k—~2.6kAutomated safety check: PassMIT
Bumpav1155/houndarr292—~1.2kAutomated safety check: PassAGPL-3.0
Sonarcloud Reviewlucasvieirasilva/nx-plugins153—~2.7kAutomated safety check: NotesMIT
Blocking IO Guardbytedance/deer-flow83k—~1.7kAutomated safety check: PassMIT
Lc Javayennanliu/CS_basics142—~5.2kAutomated safety check: NotesNone
Lc Pythonyennanliu/CS_basics142—~4.4kAutomated safety check: NotesNone

Similar skills

  • Bump

    av1155/houndarr

    Bump Houndarr version and prepare a release PR. An agent skill from av1155/houndarr.

    292 GitHub stars~1.2k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Sonarcloud Review

    lucasvieirasilva/nx-plugins

    Fetches and triages SonarCloud findings (issues, security hotspots, quality gate) for the current pull request or branch of this repository via the SonarCloud Web API, summarizes them in a markdown…

    153 GitHub stars~2.7k tokensUpdated 9 days ago
    DevelopmentAuto-check: notes
  • Blocking IO Guard

    bytedance/deer-flow

    Adds a runtime test anchor for backend async code that could block the asyncio event loop, and proves the anchor fails when the blocking call returns.

    83k GitHub stars~1.7k tokensUpdated today
    DevelopmentAuto-check passed
  • Lc Java

    yennanliu/CS_basics

    File a LeetCode Java solution into this repo the way the existing ones are filed — put it in the package its pattern owns, write the file-level javadoc header (<number.

    142 GitHub stars~5.2k tokensUpdated today
    DevelopmentAuto-check: notes
  • Lc Python

    yennanliu/CS_basics

    File a LeetCode Python solution into this repo the way the existing ones are filed — find the problem's real slug, write the Python file in the house layout (problem docstring, V0 with an IDEA block…

    142 GitHub stars~4.4k tokensUpdated today
    DevelopmentAuto-check: notes
  • Adk Sample Creator

    google/adk-python

    Official

    Creates a new sample agent in the ADK Python repository — the sample directory, its agent.py, and its README.md — following the conventions the existing samples already use.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed

More from amd/gaia

All 44 skills in this repo
  • Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.

    1.6k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

    1.6k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Turns a source document such as a README or spec into an executive slide deck as one self-contained HTML file that prints to PDF, one slide per page.

    1.6k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Works with

Questions about GAIA Agent Eval Scorecard

What does GAIA Agent Eval Scorecard do?

Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate. This is a phased checklist for adopting GAIA's eval scorecard system, which flows from a harness through a result payload to a generator and a scorecard, backed by a standalone presence-and-regression release gate.py are marked reusable and not to be modified.

When should I use GAIA Agent Eval Scorecard?

GAIA Agent Eval Scorecard fits situations like: adding a release eval scorecard to a new GAIA hub agent; wiring an existing agent's eval results into its README and release gate; generalizing the scorecard format while adopting it for a new agent.

How do I install GAIA Agent Eval Scorecard in Claude Code?

Run `npx skills add amd/gaia --skill adding-eval-scorecard -a claude-code`. Or copy the skill folder (.claude/skills/adding-eval-scorecard in amd/gaia) into .claude/skills/adding-eval-scorecard in your project. Claude Code loads it when a task matches its description.

How do I install GAIA Agent Eval Scorecard in Codex?

Run `npx skills add amd/gaia --skill adding-eval-scorecard -a codex`. Or copy the skill folder (.claude/skills/adding-eval-scorecard in amd/gaia) into .agents/skills/adding-eval-scorecard in your project. Codex loads it when a task matches its description.

Can I use GAIA Agent Eval Scorecard in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill adding-eval-scorecard -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adding-eval-scorecard, .gemini/skills/adding-eval-scorecard, .github/skills/adding-eval-scorecard and .opencode/skills/adding-eval-scorecard in your project.

What does GAIA Agent Eval Scorecard need to run?

Going by SKILL.md and its folder, GAIA Agent Eval Scorecard needs the command-line tools its instructions call (python and git). Our summary lists: A GAIA hub agent with gaia-agent.yaml; An existing eval harness for that agent; Python with PyYAML.

Does GAIA Agent Eval Scorecard access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is GAIA Agent Eval Scorecard safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does GAIA Agent Eval Scorecard use?

GAIA Agent Eval Scorecard is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GAIA Agent Eval Scorecard use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to GAIA Agent Eval Scorecard?

Skills that share tags, products or a category with GAIA Agent Eval Scorecard: Bump (av1155/houndarr, 292 stars), Sonarcloud Review (lucasvieirasilva/nx-plugins, 153 stars), Blocking IO Guard (bytedance/deer-flow, 83k stars) and Lc Java (yennanliu/CS_basics, 142 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GAIA Agent Eval Scorecard?

amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 8, 2026.

Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.