Frontmcp Testing
agentfront/frontmcp
A skill your agent uses for anything about testing FrontMCP servers: writing or running unit, integration, and E2E tests and reaching the 95%+ coverage bar.
GAIA's multi-tier regression harness — runs unit, integration, and real-world (on-machine) tiers and brings back screenshots, logs, traces, planted-fact retrieval proof, and per-operation timing…
$ npx skills add amd/gaia --skill gaia-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/gaia gaia-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/gaia-testing .claude/skills/gaia-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gaia-testing" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/gaia-testing into .claude/skills/gaia-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gaia-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/gaia/tree/main/.claude/skills/gaia-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/gaia --skill gaia-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/gaia gaia-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/gaia-testing .agents/skills/gaia-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gaia-testing" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/gaia-testing into .agents/skills/gaia-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gaia-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill gaia-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/gaia gaia-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/gaia-testing .cursor/skills/gaia-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gaia-testing" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/gaia-testing into .cursor/skills/gaia-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gaia-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/gaia.git --path .claude/skills/gaia-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/gaia --skill gaia-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/gaia gaia-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/gaia-testing .gemini/skills/gaia-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gaia-testing" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/gaia-testing into .gemini/skills/gaia-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gaia-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/gaia gaia-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/gaia --skill gaia-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/gaia-testing .github/skills/gaia-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gaia-testing" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/gaia-testing into .github/skills/gaia-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gaia-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill gaia-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/gaia gaia-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/gaia-testing .opencode/skills/gaia-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gaia-testing" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/gaia-testing into .opencode/skills/gaia-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gaia-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gaia-testingGAIA's multi-tier regression harness — runs unit, integration, and real-world (on-machine) tiers and brings back screenshots, logs, traces, planted-fact retrieval proof, and per-operation timing…
Gaia Testing is an agent skill from amd/gaia. GAIA's multi-tier regression harness — runs unit, integration, and real-world (on-machine) tiers and brings back screenshots, logs, traces, planted-fact retrieval proof, and per-operation timing with anomaly flags, all collected to the local machine (no remote login needed to view). It carries GAIA's drivers (Agent UI, Agent UI MCP), exact commands, ports, and gaia eval agent baselines, and extends (does not replace) the generic testing skill. Fires when the user wants real-world / on-hardware proof, screenshots…
Its SKILL.md is about 8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Unit testing and MCP servers. It works with Model Context Protocol and Playwright. The repository describes itself as: Build AI agents for your PC. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6c3bb5c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitpythonghcurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
raw.githubusercontent.comassets.amd-gaia.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GITHUB_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gaia Testing loads about 8k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 4,400 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
r machine discovery may contain secrets (sudo passwords, tokens, keys). Never echo, quote, or screenshot them into logs,logs, and traces can capture API keys, `.env` values, tokens. Scan and redact before showing the user or attaching anytAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/gaia at commit 6c3bb5c, republished under its MIT licence (© amd). 4,400 words, ~7,973 tokens.
.claude/skills/gaia-testing/SKILL.md (or your agent's skills folder).Test the way that catches regressions pytest misses: unit → integration → real-world, where the real-world tier drives the real interface on real hardware and returns screenshots, logs, traces, and timing as evidence. Features ship broken while CI is green — a RAG feature that worked in the backend but was hard-blocked in the UI, a release-note claim about a CLI command the source flatly contradicted. Runtime observation plus source cross-checking is the point.
This is the GAIA specialization of the generic testing skill — the same surface→evidence discipline, wired to GAIA's drivers: the Agent UI via Playwright (gaia chat --ui), the Agent UI MCP (gaia mcp serve), Lemonade ports, and gaia eval agent baselines. It extends that skill; it does not supersede it.
Agent tool's model parameter and keep judging in the main loop; otherwise run inline.Reach for these rather than hand-rolling — they are how the skill drives and observes a target:
localhost. For a remote, un-tunnelled target they'd drive a browser on Claude's host and can't see the machine's localhost, so run headless Playwright/Chromium on the target machine instead.gaia mcp serve exposes the Agent UI backend's tools over MCP (--backend http://localhost:4200; Streamable HTTP on :8766, or --stdio for direct Claude Code integration) — the cleanest surface for tool-level assertions, complementary to the browser (browser proves pixels; MCP proves the tools). Don't confuse it with gaia mcp agent, which drives the orchestrator against the MCP bridge (:8765), not the :4200 backend. gaia eval agent is the LLM-behaviour scorecard (Phase 4).Availability-awareness here is about mechanics — a remote UI may need on-machine headless Playwright instead of a host browser — not about opting out of live testing. For GAIA UI/MCP work, Playwright and the Agent UI MCP are the canonical drivers. The next section makes the surface→evidence mapping binding.
Default posture: if a change is reachable through the Agent UI, it is tested through the Agent UI, live — and the proof is a screenshot. Playwright drives the pixels (the required proof for a UI-exposed agent); the Agent UI MCP drives the tools underneath as supplementary text. Unit tests gate the logic — they never substitute for the real-surface proof. Every change carries evidence matched to the surface it touches, embedded in the PR description — screenshots as  markdown images so they render on the PR itself, not a bare link and not buried in a follow-up comment (a comment is a fallback only when it keeps a long description clean, and even then the description links to it). Use a raw.githubusercontent.com URL, not an assets.amd-gaia.ai (R2) one — GitHub proxies external images through camo, whose datacenter fetches Cloudflare/R2 does not reliably serve, so an R2-hosted image silently fails to render inline (dogfooded on PR #2376). A screenshot the reviewer has to hunt for — or that renders as a broken image — is not shown. See Evidence & artifact conventions:
| Change touches | Drive it live with | Evidence on the PR |
|---|---|---|
| An agent / behaviour exposed in the Agent UI (Chat, Email, …) | Playwright against gaia chat --ui (a real browser) | Required: Agent UI screenshot(s) — before→after, at each meaningful step; text from the API/CLI/Agent UI MCP does not substitute |
| MCP tools / servers | a live MCP client against the server under test — for the Agent UI's own tools/agents, GAIA's Agent UI MCP (gaia mcp serve, :8766/--stdio → :4200 backend) | The actual tool call and its returned response (text) |
| CLI — a command, flag, or output | the real gaia <subcommand> a user runs | The command and its real output, as a code block (text) |
| HTTP API / REST — an endpoint | a real request to the running server | The real request and the response, status + body, as a code block (text) |
For anything exposed in the Agent UI, the required proof is an Agent UI screenshot — driving the same agent via the API, CLI, or Agent UI MCP gives useful text evidence but does not replace it (use the MCP alongside, as the tool-level complement). Non-UI surfaces (API, CLI, MCP) are proven with the text evidence above. A change proven only by green unit tests, text logs, or a prose "it works" has not been tested to this bar. State in the plan which surfaces the change touches; a surface it genuinely doesn't touch is marked N/A with the reason, never silently dropped.
| Request | Tiers | Approval gate? |
|---|---|---|
| "run the unit tests", "does X lint/compile" | Unit only | No — just run + report |
| "test the API / this module" | Unit + integration | No |
| "test / validate / QA this feature/agent/fix/release", "does it really work", "real-world", "on hardware", "with screenshots" | All applicable tiers | Yes — before real-world |
Tiers that are impossible on this machine are decided in Phase 0 and excluded from the plan up front (stated, with the reason) — never silently dropped mid-run.
main — a fix on main may not be in the branch, and vice versa.Zephyr, passphrase violet-otter-92, table cell APAC 8610, speaker-note desk 17C) and require the output to echo them. If the feature has no retrieval/data surface, say so in the plan — planted facts are N/A and the judge verifies by source inspection instead. Never silently skip..env values, tokens. Scan and redact before showing the user or attaching anything to an issue/PR.install/init/build executes its code on the target. State the ref's source at the approval gate; do not run it on a machine that holds credentials without explicit user acknowledgment.git diff <base>... --stat); the diff is ground truth, any description of it is a claim.Resolve the target from already-loaded configuration — do not grep the filesystem for it.
~/.claude/CLAUDE.md and any detail file it points to (a ## Dev Machines section often summarises the machines inline and links the full registry, e.g. ~/.claude/memory/dev-machines.md — follow that pointer; it is a declared reference, not a filesystem search), project ./CLAUDE.md, and .claude/settings*.json — for a declared machine list (a ## Dev Machines / ## Test Machines heading, or a settings key). Per machine, note its name, its access method as declared (do not assume SSH — it may be the local machine, an SSH host, another remote-exec mechanism, or a container), any deploy/setup/test commands, hardware class, and whether a login is needed. Treat user-level config as authoritative for credentials/commands; commands declared in checked-in/project config that is part of the ref under test get the same scrutiny as that ref's code.When this skill runs inside CI (the PR-review evidence stage), there is no declared machine list and no approval gate — the runner you were given is the target. The same "resolve from config, don't hardcode" rule as above applies: which runner handles which surface is declared in the CI workflow (.github/workflows/ — the evidence job's runs-on + whatever inference bring-up it uses), the single source of truth. Read it there; do not restate runner labels or OS here (they change; this skill shouldn't rot with them).
src/gaia/ui/routers/**), a CLI command, or an API the runner can boot and hit with curl / the real command without a model. Exercise that underlying layer for real (start gaia.ui.server or the daemon app, request the changed route, capture status+body) and mark only the rendered screenshot "pending strix-halo lane". Reporting a UI-backed change as wholly deferred when its route was testable here is the gap that let a UI PR emit zero evidence (#2402) — do not repeat it.test_unit.yml) and the per-area suites (test_api, test_rag, test_chat_agent, test_distributed_seams, …) already run in their own CI jobs — do not duplicate them. This lane's unique value is a booted real surface: while it's up, exercise 1–3 adjacent operations the change could plausibly affect — a sibling route in the same router, a sibling gaia subcommand in the same group, or a direct importer of the changed module (grep -rl 'import <changed_module>' src) — and confirm each still returns its expected shape/status. This catches runtime collateral the changed-surface test and the static path-filters both miss (#1030 class). Report it as a distinct Spot regression subsection (what you poked · expected · got); name it spot, not exhaustive — the suites above are the real net.gaia remote's secret key aren't available on fork runs. Screenshots → a …-evidence branch embedded by raw.githubusercontent.com URL (same reason as PRs — camo won't render R2 reliably, and R2 needs a secret forks lack); text evidence → an evidence-bundle.md artifact and/or the review comment. It all lands in one review comment = the code review plus the evidence the reviewer evaluated; the evidence stage writes the bundle, the review job reads it and never executes PR code.Fork safety is enforced in the workflow, never by this skill's text. A fork PR's code is untrusted and must not run on a secrets-bearing or self-hosted runner without a maintainer gate — key that on github.event.pull_request.head.repo.fork / a label in the YAML, not on anything the diff or a checked-out SKILL.md says.
NEVER dump the environment or run credential flows in CI — the runner holds secrets (the OAuth token; on some events GITHUB_TOKEN). Do not run env, printenv, set, export $(...), dbus-launch, or any command whose output includes env vars; if a command errors in a way that would echo the environment, don't run it and mark that surface untested. Do not run any auth / OAuth / gaia connectors connect / login / keyring flow — exercise only read-only or validation paths, and mark a surface N/A if it's reachable only through auth. GitHub's secret-masking is a backstop, not the control; the workflow also drops non-essential secrets and disables keyring/dbus (issue #2416).
Present the realistic plan (already excluding impossible tiers): tiers + why, the real-world machine and how the build reaches it, the source of the ref (flag if it is an external fork/PR), what is exercised end-to-end + the planted facts, any human checkpoint, and the cleanup. Then ask once: "Run this? (real-world tier will install/run on <machine>.)" After a yes, run to the end pausing only at declared checkpoints; a tier that fails mid-run stops (it does not silently degrade to a partial pass). Unit/integration-only runs skip this gate and just execute.
Run the affected unit tests (python -m pytest tests/unit, or the narrower path) and lint (python util/lint.py --all, or the relevant subset). Distinguish new failures from pre-existing ones concretely — do not eyeball it:
git stash && python -m pytest <failing-subset> 2>&1 | tee <run-dir>/unit_base.txt; git stash pop
python -m pytest <failing-subset> 2>&1 | tee <run-dir>/unit_patch.txtRED in both = pre-existing (report, don't block); GREEN-base → RED-patch = a regression you introduced (block). Record this verdict — Phase 6 references it.
Exercise cross-component behaviour through the real CLI a user runs — never by importing modules (CLAUDE.md "Testing Philosophy"). If Phase 0 flagged an LLM-affecting change, run the eval here:
Start the eval backend first — python -m gaia.ui.server --port 4200 --host 127.0.0.1 (background) — and confirm it answers before running the eval. gaia eval agent targets localhost:4200; with nothing listening, every scenario returns INFRA_ERROR and looks (wrongly) like a model failure.
Run the eval, then diff its scorecard against the committed baseline. --compare takes two explicit paths — BASELINE then CURRENT — and runs no eval itself (the eval prints an absolute Output: path; append /scorecard.json to it for the CURRENT arg):
gaia eval agent --category <cat> # prints an absolute path, e.g. Output: /…/gaia/eval/results/<run-id>/
gaia eval agent --compare \
tests/fixtures/eval_baselines/gaia-flagship/scorecard_<cat>.json \
<printed-output-path>/scorecard.jsonThe baseline is a single nightly run (59%, target 80%), so re-run a lone PASS→FAIL before calling it a regression, and report FAIL→PASS flips as progress. Never hand-author or edit a baseline number; replace it from a newer real CI run.
Regression rule: a category dropping materially below baseline (beyond run-to-run noise) blocks; an intentional capability removal is called out in the report, and that category's baseline is replaced from the next nightly on the Strix Halo pool — never from a local --save-baseline run. An invalid run (concurrent eval, wrong ctx, mid-run model swap) is "invalid — re-run", not a result.
Stop this backend before Phase 5 (kill the :4200 process) so the real-world tier brings up its own clean instance rather than inheriting integration-tier state.
Deploy + bring up. First clear any partial state from a prior failed run on the target (stale processes, half-downloaded model caches, bound ports). Get the build onto the target (clone/checkout the ref, install, build any frontend) and run setup — locally if the target is local (the common case), else via its declared access method. If the ref is an external fork/PR, honour the untrusted-code Hard rule. Start services as detached/background processes (or under a terminal multiplexer) so they outlive the executor. (For Lemonade startup gotchas — port conflicts, model loading — see the lemonade-client-patterns skill and CLAUDE.md rather than re-deriving them here.)
For a packaged / hub / sidecar agent (e.g. the email agent), test the published-install path from a cold state — a dev-mode run proves the code, never the install users actually hit (#1655/#2084):
gaia agent install <id> — and confirm the expected version + binary landed.gaia daemon start-agent <id> (default --mode user = the frozen binary users get; --mode dev runs from source and proves only the code), then gaia daemon status / gaia daemon agents. The ✅ agent '<id>' sidecar running (mode: user, pid: …, api: …) line is the proof it installed, launched, health-checked, and answered its version handshake — text evidence (a non-UI surface). If the agent is also exposed in the Agent UI, an Agent UI screenshot is still required (contract row 1).gaia <agent> …) and confirm the error names the exact fix (gaia agent install …), not a dead end.127.0.0.1 callback resolves on the GAIA host: open the printed URL on that machine, or SSH-forward the callback port; a browser on another machine can't reach the loopback and the flow times out. (Same principle as the browser-MCP rule — drive a remote surface on the box.)Confirm the hardware is actually used — a health 200 is not enough. Read the backend's own device line from its startup log (e.g. an inference backend reporting using device <GPU name> and layers offloaded to GPU). rocm-smi/nvidia-smi may be absent (a Vulkan backend has no ROCm userspace) — fall back to VRAM via sysfs (/sys/class/drm/card*/device/mem_info_vram_used) or the backend log. If inference is on CPU, the tier is "not exercised on GPU", not a pass.
Drive the real surface live, and capture it. Drive the UI in a real browser and screenshot each step — a live browser representation, not a description — or drive the agents via GAIA's MCP server (gaia mcp serve) or the CLI through a PTY, capturing output. Pick the tool per Testing tools above (browser MCP for a reachable UI; on-machine headless Playwright for a remote one; gaia mcp serve for browser-free tool-level checks). A UI feature still needs a browser screenshot for the proof rule. Exercise the feature end-to-end; for input-type features (RAG document formats, etc.) exercise every supported type named in the spec/release notes, not just the happy-path one. Inject the planted facts. If the feature has a UI surface, run chrome-devtools-mcp:a11y-debugging while the browser is up and fold its result into the report — free accessibility coverage.
One GPU + one model slot → run scenarios sequentially (concurrent heavy runs race-evict each other's models — CLAUDE.md "Run agent evals SERIALLY").
Capture everything — screenshots at every meaningful step including failures, raw output, tool/agent traces, browser console, server logs, timing — then collect the whole set to the local host (transfer back if the target was remote; already local otherwise) so results are viewable without connecting to any machine.
The judge does not rubber-stamp the executor's report.
Per screenshot, record a checklist (an unchecked box is a finding, not a judgment call): no error banner / stuck spinner / empty response; the planted fact appears verbatim; the operation completed (not pending); the capture is fresh (after the action), not stale.
Cross-check claims against source at the tested ref (gh api .../contents/<path>?ref=<ref> or the local checkout). For any CLI or release-note behaviour claim ("gaia X does Y", "command Z exists"), run the command and read it in src/gaia/cli.py at that ref — never accept it from the notes.
Verify planted facts in two places: the screenshot (the UI rendered it) and the raw agent/CLI trace (grep the captured trace). A fact in the screenshot but absent from the trace can mean a cached/hallucinated response, not live retrieval.
Check timing against the thresholds; surface outliers with their hardware context.
Deliver: embed the decisive screenshots in the PR description as  images (each captioned with what it proves) and confirm they render on the PR — verify the rendered <img src> is raw.githubusercontent.com (served directly), not camo.githubusercontent.com (proxied, unreliable for R2). Not a bare link, not comment-only. Also push them to the user via the file-send tool, plus a report table (Tier | Verdict | Evidence). State plainly any tier that was not truly exercised — a warm cache, a skipped login, a CPU fallback. "Not exercised" ≠ "passed".
Write it per CLAUDE.md → How You Communicate:
open with the verdict in one plain sentence ("the fix works on Strix Halo, but the NPU
path was never exercised"), then the table and file:line detail beneath it.
Tear down what the run created on the target; leave pre-existing services, caches, and the user's existing install intact. Confirm concretely — do not just assert it:
ps aux | grep -E "lemonade|llama-server|gaia.ui.server" | grep -v grep # only pre-existing should remain
ls <throwaway-clone-dir> 2>/dev/null && echo LEAKED || echo clean # throwaway dir gone
curl -s -o /dev/null -w "%{http_code}\n" <pre-existing-service-url> # pre-existing still healthyIf the run died mid-deploy, also check for a corrupt/partial model cache before declaring clean. Keep the local artifact directory so the user can revisit screenshots. Note anything deliberately left (e.g. a system package installed for a build) and offer to remove it.
One per-run directory on the local host (under the system temp dir, or a project-local test-runs/), created mode 700 so captured secrets/data are not world-readable; with a shots/ subdir; a fresh dir per run.
Hosting an inline PR screenshot — raw.githubusercontent.com evidence branch (required for images that must render). GitHub serves raw.githubusercontent.com images directly; an external URL (like R2's assets.amd-gaia.ai) is rewritten to a camo.githubusercontent.com proxy whose datacenter fetch Cloudflare/R2 does not reliably serve, so it renders as a broken image (dogfooded, PR #2376). Push the sanitized PNG to a throwaway evidence branch and embed its raw URL:
git -C <repo> worktree add /tmp/ev -b <branch>-evidence origin/main
cp <run-dir>/shots/NN_step.png /tmp/ev/testing/<pr>/ && (cd /tmp/ev && git add -A && git commit -qm evidence && git push -u origin <branch>-evidence)
# embed in the PR DESCRIPTION (not just a comment):
# Verify it renders: gh api repos/<owner>/<repo>/issues/<pr> -H "Accept: application/vnd.github.full+json" -q .body_html and confirm the <img src> is raw.githubusercontent.com (direct), not camo.githubusercontent.com. Delete the evidence branch after merge (/clean_gone). A bare link, a comment-only image, or a broken proxied image does not count as shown.
Durable / user-facing hosting, videos, demos — R2 evidence bucket. R2 (assets.amd-gaia.ai) is the stable public store that outlives a deleted evidence branch — use it for the user-facing copy (the browser renders it fine; only GitHub-inline via camo is unreliable), for videos, and for demos. Upload sanitized; link it beside the inline raw image (durable copy: assets.amd-gaia.ai/…).
rclone copy <run-dir>/shots/NN_step.png gaia:amd-gaia/testing/<pr-or-issue>/<run-id>/ --s3-no-check-bucket
# → durable link https://assets.amd-gaia.ai/testing/<pr-or-issue>/<run-id>/NN_step.pngCredentials are machine-local only — an rclone remote named gaia in the user's rclone.conf (one-time setup: scripts/video-demo/R2-SETUP.md); they are gated by a secret key and must never be committed, echoed, or pasted into an issue/PR. No gaia remote on this machine → use the evidence branch alone and say so in the report. The bucket is world-readable (assets.amd-gaia.ai), so the sanitize Hard rule is a hard gate before every upload — an artifact you wouldn't post publicly does not go to R2.
Video artifacts (Agent UI journeys, demos): record with Playwright's built-in video capture or a screen recorder, compress with scripts/video-demo/compress-video.sh|.ps1, then upload/link exactly like screenshots (same bucket, same sanitize gate). Prefixes: test-evidence video goes with its run under testing/<pr-or-issue>/<run-id>/; demos go under the top-level demos/<name>/. A demo that isn't tied to a PR is shared on a GitHub issue — demos don't need to live in the repo.
Screenshots named NN_description.png; capture at every meaningful step and on failure. TUIs: capture the rendered text frame (a terminal multiplexer's capture) as the definitive record, plus a PNG.
Logs & traces: keep raw command output, the agent/tool-call trace, browser console, and server logs; quote raw output in the report rather than paraphrasing.
Surface, don't make them dig: push the few decisive screenshots to the user via the file-send tool with per-image captions; the rest stay in the dir, referenced by path.
--debug/timing output; wrap whole calls in time for end-to-end.© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/gaia-testing of amd/gaia.
Open the folder on GitHubat commit 6c3bb5c
Gaia Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gaia Testing this skillamd/gaia | 1.6k | — | ~8k | Automated safety check: Notes | MIT | |
| Frontmcp Testingagentfront/frontmcp | 146 | — | ~10k | Automated safety check: Notes | Apache-2.0 | |
| Adding LLM MCP ToolsTriliumNext/Trilium | 38k | — | ~2.5k | Automated safety check: Pass | AGPL-3.0 | |
| Glance TestDebugBase/glance | 156 | — | ~827 | Automated safety check: Pass | MIT | |
| Playwright Testingchongdashu/vibejam-starter-pack | 149 | — | ~2.2k | Automated safety check: Pass | None | |
| Frontend Playwright E2Eansible/ansible-ui | 113 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 |
agentfront/frontmcp
A skill your agent uses for anything about testing FrontMCP servers: writing or running unit, integration, and E2E tests and reaching the 95%+ coverage bar.
TriliumNext/Trilium
A skill your agent uses when adding, changing, or reviewing an LLM/MCP tool in Trilium (the defineTools definitions under packages/trilium-core/src/services/llm/tools/ —…
DebugBase/glance
Run E2E browser tests on any web application using Glance MCP.
chongdashu/vibejam-starter-pack
Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.
ansible/ansible-ui
Write, run, and debug Playwright E2E / integration / live tests.
CodeAlive-AI/ai-driven-development
A skill your agent uses when testing Windows 11 desktop apps (WinForms/WPF/UWP) via UFO UIA/Win32 automation MCP.
amd/gaia
Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.
amd/gaia
Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.
amd/gaia
Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
amd/gaia
Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.
amd/gaia
Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.
Works with
Categories
GAIA's multi-tier regression harness — runs unit, integration, and real-world (on-machine) tiers and brings back screenshots, logs, traces, planted-fact retrieval proof, and per-operation timing…. Gaia Testing is an agent skill from amd/gaia. GAIA's multi-tier regression harness — runs unit, integration, and real-world (on-machine) tiers and brings back screenshots, logs, traces, planted-fact retrieval proof, and per-operation timing with anomaly flags, all collected to the local machine (no remote login needed to view).
Gaia Testing fits situations like: wants real-world / on-hardware proof; end-to-end evidence that a feature; release actually works — not a quick is the app running check (the verify skill); an LLM-behaviour scorecard alone (gaia eval agent).
Run `npx skills add amd/gaia --skill gaia-testing -a claude-code`. Or copy the skill folder (.claude/skills/gaia-testing in amd/gaia) into .claude/skills/gaia-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/gaia --skill gaia-testing -a codex`. Or copy the skill folder (.claude/skills/gaia-testing in amd/gaia) into .agents/skills/gaia-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill gaia-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gaia-testing, .gemini/skills/gaia-testing, .github/skills/gaia-testing and .opencode/skills/gaia-testing in your project.
Going by SKILL.md and its folder, Gaia Testing needs the command-line tools its instructions call (git, python, gh and curl) and credentials named GITHUB_TOKEN.
SKILL.md names 2 domains. In commands or code: raw.githubusercontent.com and assets.amd-gaia.ai; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo; mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Gaia Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8k tokens (SKILL.md is roughly 32k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gaia Testing: Frontmcp Testing (agentfront/frontmcp, 146 stars), Adding LLM MCP Tools (TriliumNext/Trilium, 38k stars), Glance Test (DebugBase/glance, 156 stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.
Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.