Advanced Evaluation
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.
$ npx skills add anthropics/commerce-agents --skill commerce-evals -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install anthropics/commerce-agents commerce-evals --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/anthropics/commerce-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/commerce-builder/skills/commerce-evals .claude/skills/commerce-evals && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "commerce-evals" agent skill from https://github.com/anthropics/commerce-agents/tree/main/plugins/commerce-builder/skills/commerce-evals into .claude/skills/commerce-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "commerce-evals", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/anthropics/commerce-agents/tree/main/plugins/commerce-builder/skills/commerce-evalsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add anthropics/commerce-agents --skill commerce-evals -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install anthropics/commerce-agents commerce-evals --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anthropics/commerce-agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/commerce-builder/skills/commerce-evals .agents/skills/commerce-evals && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "commerce-evals" agent skill from https://github.com/anthropics/commerce-agents/tree/main/plugins/commerce-builder/skills/commerce-evals into .agents/skills/commerce-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "commerce-evals", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add anthropics/commerce-agents --skill commerce-evals -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install anthropics/commerce-agents commerce-evals --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anthropics/commerce-agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/commerce-builder/skills/commerce-evals .cursor/skills/commerce-evals && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "commerce-evals" agent skill from https://github.com/anthropics/commerce-agents/tree/main/plugins/commerce-builder/skills/commerce-evals into .cursor/skills/commerce-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "commerce-evals", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/anthropics/commerce-agents.git --path plugins/commerce-builder/skills/commerce-evals--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add anthropics/commerce-agents --skill commerce-evals -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install anthropics/commerce-agents commerce-evals --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anthropics/commerce-agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/commerce-builder/skills/commerce-evals .gemini/skills/commerce-evals && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "commerce-evals" agent skill from https://github.com/anthropics/commerce-agents/tree/main/plugins/commerce-builder/skills/commerce-evals into .gemini/skills/commerce-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "commerce-evals", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install anthropics/commerce-agents commerce-evalsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add anthropics/commerce-agents --skill commerce-evals -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/anthropics/commerce-agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/commerce-builder/skills/commerce-evals .github/skills/commerce-evals && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "commerce-evals" agent skill from https://github.com/anthropics/commerce-agents/tree/main/plugins/commerce-builder/skills/commerce-evals into .github/skills/commerce-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "commerce-evals", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add anthropics/commerce-agents --skill commerce-evals -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install anthropics/commerce-agents commerce-evals --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anthropics/commerce-agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/commerce-builder/skills/commerce-evals .opencode/skills/commerce-evals && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "commerce-evals" agent skill from https://github.com/anthropics/commerce-agents/tree/main/plugins/commerce-builder/skills/commerce-evals into .opencode/skills/commerce-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "commerce-evals", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
commerce-evalsAuthoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.
Commerce Evals is an agent skill from anthropics/commerce-agents, published by the product's own GitHub organization. Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures. Load when writing eval cases or rubrics or deciding how a suite runs.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation and Quizzes and assessments. The repository describes itself as: Reference blueprint for building shopping and merchant agents with Claude. Examples in retail, commerce, telecom, and entertainment included. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit fd4d592. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are json).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Commerce Evals loads about 1.8k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 923 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from anthropics/commerce-agents at commit fd4d592, republished under its Apache-2.0 licence (© anthropics). 923 words, ~1,813 tokens.
.claude/skills/commerce-evals/SKILL.md (or your agent's skills folder).The repo ships no eval harness; the suite is yours, because a case only means something against your catalog,
orders, and fixtures. Paths below are in the reference repo: commerce_common/ is commerce-common/commerce_common/.
Gate behavior that needs no model (provenance, caps, guardrails, approval) is unit-tested with FakeClient in
commerce_common/testing.py; evals cover what the model decides.
This block is the one home of the case shape and the scorer names; /author-commerce-evals refers to it.
{
"id": "<flow>-<nnn>-<behavior>",
"priority": "critical | high | medium | low", "difficulty": "easy | medium | hard", "tags": ["..."],
"skip": "<reason, when a case cannot run yet>",
"state": {"seen_products": ["..."], "cart": [...], "memory": [...], "staged_changes": [...]},
"turns": ["<the customer's or operator's message>", "..."],
"expected": {
"calls_tool": ["..."], "calls_one_of": ["..."], "never_calls": ["..."],
"first_tool": "...", "first_tool_not": "...",
"ui_components": ["..."], "no_ui": true,
"cart_contains": ["..."], "cart_item_count": 0, "cart_not_contains": ["..."],
"staged_change_kinds": ["..."], "no_applied_changes": true,
"memory_contains": ["..."], "memory_not_contains": ["..."],
"skill_loaded": "...", "skill_not_loaded": "...", "no_skill_load": true,
"reply_includes": ["..."], "reply_omits": ["..."], "max_tool_calls": 0,
"rubric": "PASS if <condition>. FAIL if <condition>."
},
"notes": "<what the case pins and the fixture fact that decides it>"
}state is the precondition; turns is one message unless the behavior under test is carrying state across turns;
expected holds only the keys the case is about. Ids in state and expected are real ids from your fixtures.
priority and difficulty let a report say which failures matter; a case that cannot run yet carries skip
with its reason rather than being deleted.
skill_loaded, one named component, never_calls on a
presentation tool) is asserted only where the route is the behavior (a grounding read first, a write that must
never happen); elsewhere calls_one_of names the acceptable set. When a live run takes a route the case did not
expect and the answer was right, widen the case to the acceptable set; do not re-pin it to the route observed.memory_not_contains, or never_calls on save_memory). A case where a stored fact should change the pick uses a
query whose results contain both the item the fact favors and the one it rules out; run the search before writing
the case.max_tool_calls is set from what a well-behaved agent needs; a multi-item request fans out several searches in one round, and the present_suggestions call that ends the turn is not counted.commerce_common/streaming.py): tool_call names and arguments,
tool_result with status blocked and the gate in reason, ui component names, the last cart_update or
change_update, and the reply text. Every key above except rubric is a code grader.rubric goes to a judge. One judge call per dimension (budget respected, no invented availability, trade-off stated),
returning structured output with a verdict and a reason; the transcript is passed to it as quoted material, tool
results and component payloads included; when the transcript exceeds the judge's window, truncate from the start so
the graded turn survives, and record the truncation on the outcome. Pin the judge model at temperature zero; a change to the judge model or a rubric
invalidates every stored verdict scored with it, so the recording carries a fingerprint of both.| When | What runs | What decides |
|---|---|---|
| Every merge to the agent, a skill, a tool description, or a fixture | The regression set | Each case over several trials; a pass threshold per set |
| While changing one flow | That flow's targeted set | The failure set, read beside the previous run's as the baseline |
| Choosing or upgrading a model (commerce-architecture's model fields) | Everything | The two failure sets side by side |
| In production | A judged sample of live traffic against the same rubrics | Trend per dimension |
Diff failure sets; a topline moving a point between live runs is noise. A case that fails after a change means the change broke the behavior or the case encoded a stale one; fix whichever it is and say which in the commit.
Listings, reviews, and messages carrying instructions live in eval-only fixtures the runner merges into the backend
for the run, under a third-party brand or seller; none of their ids appears in demo data, seeds, or captures. Each
such case asserts the negative in code (never_calls, cart_not_contains, no_applied_changes,
memory_not_contains, reply_omits), and every vector it asserts is one the driven turn actually puts in front of
the model (a review is only read on a details call). Cover at least an instruction to write to the cart or stage a
change, one to remember something, and a false claim (a code, a guarantee). The should-serve counterpart is a
separate benign eval-only listing in the same niche, so an agent that refuses everything fails it.
© anthropics, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/commerce-builder/skills/commerce-evals of anthropics/commerce-agents.
Open the folder on GitHubat commit fd4d592
Commerce Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Commerce Evals this skillanthropics/commerce-agents | 3.2k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Advanced Evaluationguanyang/open-agent-hub | 975 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | |
| Agentic Evalgithub/awesome-copilot | 40k | 4 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Metric Designagentscope-ai/OpenJudge | 868 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | |
| Agent Evalssickn33/agentic-awesome-skills | 47k | 2 repos | ~3.1k | Automated safety check: Warn | MIT | |
| Clawpathy AutoresearchClawBio/ClawBio | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT |
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
github/awesome-copilot
Patterns and techniques for evaluating and improving AI agent outputs.
agentscope-ai/OpenJudge
A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining…
sickn33/agentic-awesome-skills
Build automated evaluation suites for AI agents using golden datasets, rubrics, and regression gates.
ClawBio/ClawBio
Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.
agentsope/SkillAlchemy
Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.
anthropics/commerce-agents
Creating and improving listing content, covering titles, descriptions, attribute completeness, categorization fixes, image callouts written as text, edits written from material the operator…
anthropics/commerce-agents
How the reference commerce agents of either role are structured, covering the loop, where each rule lives, skills, the backend interface, delegates, and model fields.
anthropics/commerce-agents
The reference merchant agent, covering its flows, staged changes and host approval, metrics grounding, the analysis delegate, store memory, components, and marketplace rules.
anthropics/commerce-agents
The reference agents' cache-stable request assembly, covering the static system and per-request context split, the fixed tool list, the rolling conversation breakpoint, which config fields are…
anthropics/commerce-agents
The rules the reference agents enforce in code for third-party content, writes, grounding, identity, and memory, each with its module, plus two adversarial-eval rules.
anthropics/commerce-agents
The reference presentation-tool contract, covering server-side enrichment, suggestion chips, the event stream, progressive rendering, both roles' built-in components, and adding a vertical component.
Categories
Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures. Commerce Evals is an agent skill from anthropics/commerce-agents, published by the product's own GitHub organization. Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.
Commerce Evals fits situations like: tasks that involve LLM evaluation; tasks that involve Quizzes and assessments.
Run `npx skills add anthropics/commerce-agents --skill commerce-evals -a claude-code`. Or copy the skill folder (plugins/commerce-builder/skills/commerce-evals in anthropics/commerce-agents) into .claude/skills/commerce-evals in your project. Claude Code loads it when a task matches its description.
Run `npx skills add anthropics/commerce-agents --skill commerce-evals -a codex`. Or copy the skill folder (plugins/commerce-builder/skills/commerce-evals in anthropics/commerce-agents) into .agents/skills/commerce-evals in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anthropics/commerce-agents --skill commerce-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/commerce-evals, .gemini/skills/commerce-evals, .github/skills/commerce-evals and .opencode/skills/commerce-evals in your project.
SKILL.md names no scripts, command-line tools or credentials: Commerce Evals is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Commerce Evals is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Commerce Evals: Advanced Evaluation (guanyang/open-agent-hub, 975 stars), Agentic Eval (github/awesome-copilot, 40k stars), Metric Design (agentscope-ai/OpenJudge, 868 stars) and Agent Evals (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
anthropics (a GitHub organization, an official publisher) maintains it in anthropics/commerce-agents, which has 3,179 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 2, 2026.
Source: anthropics/commerce-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.