Bump
av1155/houndarr
Bump Houndarr version and prepare a release PR. An agent skill from av1155/houndarr.
Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.
$ npx skills add amd/gaia --skill adding-eval-scorecard -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/gaia adding-eval-scorecard --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/adding-eval-scorecard .claude/skills/adding-eval-scorecard && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "adding-eval-scorecard" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/adding-eval-scorecard into .claude/skills/adding-eval-scorecard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adding-eval-scorecard", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/gaia/tree/main/.claude/skills/adding-eval-scorecardType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/gaia --skill adding-eval-scorecard -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/gaia adding-eval-scorecard --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/adding-eval-scorecard .agents/skills/adding-eval-scorecard && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "adding-eval-scorecard" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/adding-eval-scorecard into .agents/skills/adding-eval-scorecard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adding-eval-scorecard", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill adding-eval-scorecard -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/gaia adding-eval-scorecard --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/adding-eval-scorecard .cursor/skills/adding-eval-scorecard && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "adding-eval-scorecard" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/adding-eval-scorecard into .cursor/skills/adding-eval-scorecard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adding-eval-scorecard", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/gaia.git --path .claude/skills/adding-eval-scorecard--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/gaia --skill adding-eval-scorecard -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/gaia adding-eval-scorecard --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/adding-eval-scorecard .gemini/skills/adding-eval-scorecard && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "adding-eval-scorecard" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/adding-eval-scorecard into .gemini/skills/adding-eval-scorecard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adding-eval-scorecard", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/gaia adding-eval-scorecardInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/gaia --skill adding-eval-scorecard -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/adding-eval-scorecard .github/skills/adding-eval-scorecard && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "adding-eval-scorecard" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/adding-eval-scorecard into .github/skills/adding-eval-scorecard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adding-eval-scorecard", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill adding-eval-scorecard -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/gaia adding-eval-scorecard --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/adding-eval-scorecard .opencode/skills/adding-eval-scorecard && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "adding-eval-scorecard" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/adding-eval-scorecard into .opencode/skills/adding-eval-scorecard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adding-eval-scorecard", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
adding-eval-scorecardAdds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.
This is a phased checklist for adopting GAIA's eval scorecard system, which flows from a harness through a result payload to a generator and a scorecard, backed by a standalone presence-and-regression release gate. The email agent is named as the reference implementation to mirror, and core modules such as release_scorecard.py and scorecard_gate.py are marked reusable and not to be modified.
The first phase locates the agent's surfaces: the version comes only from the version field in the agent's gaia-agent.yaml, the canonical README is whichever file the agent's release workflow actually publishes, and the scorecard lives beside it as a single SCORECARD.md updated in place rather than in a separate folder. The checklist has a hard gate at the real-eval step: the scorecard must come from an actual eval run over the agent's existing harness, such as a benchmark command over its test fixtures, and the skill says to stop and propose a minimal harness rather than invent numbers if no eval vehicle exists yet.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 05fb50b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythongitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
GAIA Agent Eval Scorecard loads about 2.6k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,075 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/gaia at commit 05fb50b, republished under its MIT licence (© amd). 1,075 words, ~2,574 tokens.
.claude/skills/adding-eval-scorecard/SKILL.md (or your agent's skills folder).Adopt the release eval scorecard (docs/reference/eval-scorecard.mdx) for one hub agent. The system is harness → result payload → generator → scorecard, with a standalone presence+regression release gate. The email agent is the reference implementation — mirror it.
Core modules (do not modify; reuse):
src/gaia/eval/release_scorecard.py — ResultPayload, compute_aggregate, render_scorecard, write_scorecard, validate_scorecard, carry_forward. Harness-agnostic (stdlib + PyYAML only).src/gaia/eval/scorecard_gate.py — the standalone gate (python -m gaia.eval.scorecard_gate).hub/agents/email/python/packaging/gen_scorecard.py.This is a phased checklist with a hard gate at the real-eval step — the scorecard MUST come from an actual eval run, never hand-authored numbers.
Version source of truth = the version: field in <agent>/gaia-agent.yaml. Never invent a parallel scheme.
Canonical README (where the scorecard is linked + surfaced): for an npm-published agent it is the npm client README (e.g. hub/agents/<id>/npm/README.md), NOT a packaging/README.md. For a Python-only agent it is hub/agents/<id>/python/README.md. Confirm which by checking what release_agent_<id>.yml publishes (README: env) — the published README is the one to link.
doc-root = the directory holding that canonical README. The scorecard lives at <doc-root>/SCORECARD.md — a single file updated in place, versioned via the publish snapshot (same as README.md). There is no scorecards/ directory.
Eval vehicle: what existing harness produces this agent's accuracy metric? (email → gaia eval benchmark over tests/fixtures/email/.) If none exists, STOP and surface that — propose the minimal harness before building; do not invent numbers.
Expect none. Email is the only agent with a corpus + harness; for every other agent this step is the long pole, not a lookup. Two traps: an agent_type in the scenario corpus may name a ChatAgent prompt profile rather than your package (so a green scenario is measuring something else), and the ground truth itself is human judgment the factory does not automate. Build the corpus + deterministic fixture harness first — see porting-agent-to-hub Phase 4 — then return here for the adapter and gate.
Copy hub/agents/email/python/packaging/gen_scorecard.py as the template. The adapter:
gaia.eval.release_scorecard (never the harness or agent package — preserve loose coupling);ResultPayload;reproduction_command with the exact shell commands to reproduce this scorecard, including all required env vars (PYTHON_KEYRING_BACKEND, GAIA_AGENT_TOOL_TIMEOUT, PYTHONPATH);limit/config so future regression checks are comparable;environment (gaia_commit, lemonade_version, model, hardware — a class descriptor, never a hostname) and breakdown (per_category accuracy + top_confusions) — additive blocks that pin the run and explain the misses without ever affecting aggregate.value. Capture live environment (git SHA, server health version) in main(), not build_payload, so the payload builder stays pure and unit-testable;<doc-root>/SCORECARD.md (the single file; --output-dir overrides to a directory, but the filename is always SCORECARD.md).Add an offline unit test against a committed sample harness-output fixture (see tests/fixtures/eval/email_benchmark_scorecard.json + tests/unit/eval/test_release_scorecard.py::TestEmailAdapter) so the adapter is testable without a live model.
The accuracy number must come from an actual run. For the email agent:
# Real eval needs Lemonade + the model. Prefer AMD hardware (Strix Halo / Ryzen AI);
# the [self-hosted, lemonade-eval] runner is the canonical environment.
GAIA_AGENT_TOOL_TIMEOUT=1800 \
PYTHON_KEYRING_BACKEND=keyring.backends.null.Keyring \
PYTHONPATH="$(pwd)" \
<venv>/bin/gaia eval benchmark \
--model Gemma-4-E4B-it-GGUF \
--mbox-path tests/fixtures/email/_stub_inbox.mbox \
--ground-truth tests/fixtures/email/<task>_ground_truth.json \
--limit 220 --output-dir <persistent-dir>
PYTHONPATH="$(pwd)" \
<venv>/bin/python hub/agents/email/python/packaging/gen_scorecard.py \
--benchmark-dir <persistent-dir> --limit 220
# → writes hub/agents/email/npm/SCORECARD.md in place⚠️ Run evals SERIALLY — never two at once (CLAUDE.md). A scorecard refresh loop is the
most likely thing to violate this: chain runs with &&, never background them with &.
Concurrent runs race-evict each other's models on the single-tenant Lemonade backend and
produce garbage (ctx_size errors, spurious INFRA_ERROR) that looks like a real
regression. Before each run: ps aux | grep "gaia eval" | grep -v grep | wc -l must be 0.
Pass both fixture flags explicitly. ls tests/fixtures/email/ and pick the real files
— ground truth is task-specific (action_items_, briefing_, drafting_,
followups_, longthread_), and the CLI's --ground-truth default points at a
ground_truth.json that does not exist. Omit the flag and the run prints one [WARN] and
scores nothing, which reads as a successful eval right up until the scorecard is empty.
Headless gotchas (see memory project-email-benchmark-headless-gotchas):
PYTHON_KEYRING_BACKEND=keyring.backends.null.Keyring — the email agent's calendar-connector resolution blocks forever on the macOS Keychain (and can stall on Linux SecretService) in non-interactive contexts. Without this it hangs at 0% CPU during agent construction.PYTHONPATH="$(pwd)" — the benchmark imports tests.fixtures.email.*; the console script doesn't add the repo root.GAIA_AGENT_TOOL_TIMEOUT=1800 — triage of the full corpus is one tool call (~17 min for ~100 emails on a 4B model); a lower timeout (the 180s default, or even 900s) abandons it mid-run, yielding a degenerate 0-email FAIL run.--output-dir to a persistent dir, not /tmp (cleared on session resume).methodology string and link the tracking issue — never inflate the number.Eval scorecard (vX.Y.Z): aggregate N/100 … ([./SCORECARD.md](./SCORECARD.md)). The relative link must resolve in-repo.files: if the agent publishes on npm, add SCORECARD.md to package.json files. Do not add a scorecards/ directory — only the single current file ships.workers/agent-hub + AgentDetailModal.tsx); ensure the publish step passes --eval-scorecard <doc-root>/SCORECARD.md to publish_to_r2.py.scorecard-gate job to release_agent_<id>.yml and list it in publish.needs. The job runs on a GitHub-hosted runner (it only parses committed files — no eval):# Presence-only (no previous tag yet):
python -m gaia.eval.scorecard_gate \
--scorecard <doc-root>/SCORECARD.md
# With best-effort previous-release baseline (recommended for CI):
PREV="$(git describe --tags --abbrev=0 --match 'agent-pkg-<id>-*' "${GITHUB_REF_NAME}^" 2>/dev/null || true)"
if [ -n "$PREV" ]; then
python -m gaia.eval.scorecard_gate \
--scorecard <doc-root>/SCORECARD.md --baseline-ref "$PREV"
else
python -m gaia.eval.scorecard_gate \
--scorecard <doc-root>/SCORECARD.md
ficontinue-on-error, an environment:, or a permissions: override (inherits contents: read; needs no secrets). Fetch full history (fetch-depth: 0) so git describe resolves.eval-scorecard.mdx "Keeping the scorecard current" and the self-hosted refresh workflow — reject-on-worse is the gate; better-or-equal refreshes the committed SCORECARD.md.Run and capture: the generated SCORECARD.md; the gate passing on it (exit 0); the gate blocking a manufactured regression (exit 1, via --baseline-file with a higher-scoring card) and a missing card (exit 1); a by-hand recompute of the aggregate from aggregate.components matching the recorded value. Run python util/lint.py --all and the eval unit tests. These are the PR's real-world proof.
carry_forward(prev_scorecard_path, new_version) reads the version from the front matter of the current SCORECARD.md (not from the filename) and copies results verbatim, sets inherited_from; do NOT re-run the eval.carry_forward refuses a non-patch bump with a "re-run" error.Follow CLAUDE.md → "How You Communicate".
© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/adding-eval-scorecard of amd/gaia.
Open the folder on GitHubat commit 05fb50b
GAIA Agent Eval Scorecard next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GAIA Agent Eval Scorecard this skillamd/gaia | 1.6k | — | ~2.6k | Automated safety check: Pass | MIT | |
| Bumpav1155/houndarr | 292 | — | ~1.2k | Automated safety check: Pass | AGPL-3.0 | |
| Sonarcloud Reviewlucasvieirasilva/nx-plugins | 153 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Blocking IO Guardbytedance/deer-flow | 83k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Lc Javayennanliu/CS_basics | 142 | — | ~5.2k | Automated safety check: Notes | None | |
| Lc Pythonyennanliu/CS_basics | 142 | — | ~4.4k | Automated safety check: Notes | None |
av1155/houndarr
Bump Houndarr version and prepare a release PR. An agent skill from av1155/houndarr.
lucasvieirasilva/nx-plugins
Fetches and triages SonarCloud findings (issues, security hotspots, quality gate) for the current pull request or branch of this repository via the SonarCloud Web API, summarizes them in a markdown…
bytedance/deer-flow
Adds a runtime test anchor for backend async code that could block the asyncio event loop, and proves the anchor fails when the blocking call returns.
yennanliu/CS_basics
File a LeetCode Java solution into this repo the way the existing ones are filed — put it in the package its pattern owns, write the file-level javadoc header (<number.
yennanliu/CS_basics
File a LeetCode Python solution into this repo the way the existing ones are filed — find the problem's real slug, write the Python file in the house layout (problem docstring, V0 with an IDEA block…
google/adk-python
Creates a new sample agent in the ADK Python repository — the sample directory, its agent.py, and its README.md — following the conventions the existing samples already use.
amd/gaia
Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.
amd/gaia
Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
amd/gaia
Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.
amd/gaia
Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.
amd/gaia
Turns a source document such as a README or spec into an executive slide deck as one self-contained HTML file that prints to PDF, one slide per page.
Works with
Categories
Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate. This is a phased checklist for adopting GAIA's eval scorecard system, which flows from a harness through a result payload to a generator and a scorecard, backed by a standalone presence-and-regression release gate.py are marked reusable and not to be modified.
GAIA Agent Eval Scorecard fits situations like: adding a release eval scorecard to a new GAIA hub agent; wiring an existing agent's eval results into its README and release gate; generalizing the scorecard format while adopting it for a new agent.
Run `npx skills add amd/gaia --skill adding-eval-scorecard -a claude-code`. Or copy the skill folder (.claude/skills/adding-eval-scorecard in amd/gaia) into .claude/skills/adding-eval-scorecard in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/gaia --skill adding-eval-scorecard -a codex`. Or copy the skill folder (.claude/skills/adding-eval-scorecard in amd/gaia) into .agents/skills/adding-eval-scorecard in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill adding-eval-scorecard -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adding-eval-scorecard, .gemini/skills/adding-eval-scorecard, .github/skills/adding-eval-scorecard and .opencode/skills/adding-eval-scorecard in your project.
Going by SKILL.md and its folder, GAIA Agent Eval Scorecard needs the command-line tools its instructions call (python and git). Our summary lists: A GAIA hub agent with gaia-agent.yaml; An existing eval harness for that agent; Python with PyYAML.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
GAIA Agent Eval Scorecard is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with GAIA Agent Eval Scorecard: Bump (av1155/houndarr, 292 stars), Sonarcloud Review (lucasvieirasilva/nx-plugins, 153 stars), Blocking IO Guard (bytedance/deer-flow, 83k stars) and Lc Java (yennanliu/CS_basics, 142 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 8, 2026.
Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.