Official agent skill

Commerce Evals

by anthropics in anthropics/commerce-agents

Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Commerce Evals

skills CLI
$ npx skills add anthropics/commerce-agents --skill commerce-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install anthropics/commerce-agents commerce-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/anthropics/commerce-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/commerce-builder/skills/commerce-evals .claude/skills/commerce-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
commerce-evals
GitHub stars
3.2k
Token cost
~1.8k tokens
SKILL.md length
923 words
Files
1
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.

  • Tasks that involve LLM evaluation
  • SKILL.md covers The case shape, Authoring rules, Scorers and Run pattern, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Quizzes and assessments

What it does

Commerce Evals is an agent skill from anthropics/commerce-agents, published by the product's own GitHub organization. Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures. Load when writing eval cases or rubrics or deciding how a suite runs.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and Quizzes and assessments. The repository describes itself as: Reference blueprint for building shopping and merchant agents with Claude. Examples in retail, commerce, telecom, and entertainment included. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve Quizzes and assessments

Example prompts

  • “/commerce-evals”

What it can do on your machine

Read from SKILL.md and the folder at commit fd4d592. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Commerce Evals loads about 1.8k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 923 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from anthropics/commerce-agents at commit fd4d592, republished under its Apache-2.0 licence (© anthropics). 923 words, ~1,813 tokens.

Download SKILL.mdSave it as .claude/skills/commerce-evals/SKILL.md (or your agent's skills folder).
name
commerce-evals
description
Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures. Load when writing eval cases or rubrics or deciding how a suite runs.

Commerce agent evals

The repo ships no eval harness; the suite is yours, because a case only means something against your catalog, orders, and fixtures. Paths below are in the reference repo: commerce_common/ is commerce-common/commerce_common/. Gate behavior that needs no model (provenance, caps, guardrails, approval) is unit-tested with FakeClient in commerce_common/testing.py; evals cover what the model decides.

The case shape

This block is the one home of the case shape and the scorer names; /author-commerce-evals refers to it.

json
{
  "id": "<flow>-<nnn>-<behavior>",
  "priority": "critical | high | medium | low",   "difficulty": "easy | medium | hard",   "tags": ["..."],
  "skip": "<reason, when a case cannot run yet>",
  "state": {"seen_products": ["..."], "cart": [...], "memory": [...], "staged_changes": [...]},
  "turns": ["<the customer's or operator's message>", "..."],
  "expected": {
    "calls_tool": ["..."],            "calls_one_of": ["..."],          "never_calls": ["..."],
    "first_tool": "...",              "first_tool_not": "...",
    "ui_components": ["..."],         "no_ui": true,
    "cart_contains": ["..."],         "cart_item_count": 0,             "cart_not_contains": ["..."],
    "staged_change_kinds": ["..."],   "no_applied_changes": true,
    "memory_contains": ["..."],       "memory_not_contains": ["..."],
    "skill_loaded": "...",            "skill_not_loaded": "...",        "no_skill_load": true,
    "reply_includes": ["..."],        "reply_omits": ["..."],           "max_tool_calls": 0,
    "rubric": "PASS if <condition>. FAIL if <condition>."
  },
  "notes": "<what the case pins and the fixture fact that decides it>"
}

state is the precondition; turns is one message unless the behavior under test is carrying state across turns; expected holds only the keys the case is about. Ids in state and expected are real ids from your fixtures. priority and difficulty let a report say which failures matter; a case that cannot run yet carries skip with its reason rather than being deleted.

Authoring rules

  • Preconditions go in injected state (the products already seen, the cart, the memory facts, the staged queue), which the runner loads into the session state and the memory store before the turn. Earlier turns are for state the behavior itself carries, and for nothing else.
  • Every positive has a negative: for each case that asserts a component, a skill load, a gate, a disclosure, or a memory write, a case in the same niche asserts its absence. A refusal case has a should-serve counterpart.
  • Memory is three cases: a remark worth keeping is written; an identifier or excluded content is refused (nothing stored, no error to the person); a stored fact changes the next session's pick.
  • Grade the final tool arguments and the state they produced: the ids on the last presentation call, the fields on the staged change, the cart and the memory store after the turn. The reply's wording is graded only for strings that must or must not appear. Which route the agent took (skill_loaded, one named component, never_calls on a presentation tool) is asserted only where the route is the behavior (a grounding read first, a write that must never happen); elsewhere calls_one_of names the acceptable set. When a live run takes a route the case did not expect and the answer was right, widen the case to the acceptable set; do not re-pin it to the route observed.
  • A rubric is one PASS condition and one FAIL condition that no response satisfies both of; it names the fixture fact that decides it (the updated delivery date, the price today); variants you accept are written into it; it says nothing about tone, length, or the order components appear in.
  • A turn that mentions health, a one-off errand, or hostile content asserts the memory end-state (memory_not_contains, or never_calls on save_memory). A case where a stored fact should change the pick uses a query whose results contain both the item the fact favors and the one it rules out; run the search before writing the case.
  • max_tool_calls is set from what a well-behaved agent needs; a multi-item request fans out several searches in one round, and the present_suggestions call that ends the turn is not counted.
Show full SKILL.md (412 more words)Show less

Scorers

  • Code graders read the events a turn yields (commerce_common/streaming.py): tool_call names and arguments, tool_result with status blocked and the gate in reason, ui component names, the last cart_update or change_update, and the reply text. Every key above except rubric is a code grader.
  • rubric goes to a judge. One judge call per dimension (budget respected, no invented availability, trade-off stated), returning structured output with a verdict and a reason; the transcript is passed to it as quoted material, tool results and component payloads included; when the transcript exceeds the judge's window, truncate from the start so the graded turn survives, and record the truncation on the outcome. Pin the judge model at temperature zero; a change to the judge model or a rubric invalidates every stored verdict scored with it, so the recording carries a fingerprint of both.
  • A judge reply that does not parse into a verdict is a judge failure on the case, kept apart from an agent failure.

Run pattern

WhenWhat runsWhat decides
Every merge to the agent, a skill, a tool description, or a fixtureThe regression setEach case over several trials; a pass threshold per set
While changing one flowThat flow's targeted setThe failure set, read beside the previous run's as the baseline
Choosing or upgrading a model (commerce-architecture's model fields)EverythingThe two failure sets side by side
In productionA judged sample of live traffic against the same rubricsTrend per dimension

Diff failure sets; a topline moving a point between live runs is noise. A case that fails after a change means the change broke the behavior or the case encoded a stale one; fix whichever it is and say which in the commit.

Poisoned fixtures

Listings, reviews, and messages carrying instructions live in eval-only fixtures the runner merges into the backend for the run, under a third-party brand or seller; none of their ids appears in demo data, seeds, or captures. Each such case asserts the negative in code (never_calls, cart_not_contains, no_applied_changes, memory_not_contains, reply_omits), and every vector it asserts is one the driven turn actually puts in front of the model (a review is only read on a details call). Cover at least an instruction to write to the cart or stage a change, one to remember something, and a false claim (a code, a guarantee). The should-serve counterpart is a separate benign eval-only listing in the same niche, so an agent that refuses everything fails it.

© anthropics, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/commerce-builder/skills/commerce-evals of anthropics/commerce-agents.

Open the folder on GitHubat commit fd4d592

Compare with similar skills

Commerce Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Commerce Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Commerce Evals this skillanthropics/commerce-agents3.2k—~1.8kAutomated safety check: PassApache-2.0
Advanced Evaluationguanyang/open-agent-hub9752 repos~4.2kAutomated safety check: PassMIT
Agentic Evalgithub/awesome-copilot40k4 repos~1.5kAutomated safety check: PassMIT
Metric Designagentscope-ai/OpenJudge868—~5.1kAutomated safety check: PassApache-2.0
Agent Evalssickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: WarnMIT
Clawpathy AutoresearchClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Advanced Evaluation

    guanyang/open-agent-hub

    This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Agentic Eval

    github/awesome-copilot

    Official

    Patterns and techniques for evaluating and improving AI agent outputs.

    40k GitHub starsUsed in 4 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Metric Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining…

    868 GitHub stars~5.1k tokensUpdated 27 days ago
    AI & LLM EngineeringAuto-check passed
  • Agent Evals

    sickn33/agentic-awesome-skills

    Build automated evaluation suites for AI agents using golden datasets, rubrics, and regression gates.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check: warnings
  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agentsop Metric Design

    agentsope/SkillAlchemy

    Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.

    459 GitHub stars~6.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from anthropics/commerce-agents

All 16 skills in this repo
  • Catalog Listings

    anthropics/commerce-agents

    Official

    Creating and improving listing content, covering titles, descriptions, attribute completeness, categorization fixes, image callouts written as text, edits written from material the operator…

    3.2k GitHub stars~1k tokensUpdated 6 days ago
    Auto-check passed
  • Commerce Architecture

    anthropics/commerce-agents

    Official

    How the reference commerce agents of either role are structured, covering the loop, where each rule lives, skills, the backend interface, delegates, and model fields.

    3.2k GitHub stars~1.7k tokensUpdated 6 days ago
    Auto-check passed
  • Commerce Merchant Operations

    anthropics/commerce-agents

    Official

    The reference merchant agent, covering its flows, staged changes and host approval, metrics grounding, the analysis delegate, store memory, components, and marketplace rules.

    3.2k GitHub stars~2.1k tokensUpdated 6 days ago
    Auto-check passed
  • Commerce Prompt Caching

    anthropics/commerce-agents

    Official

    The reference agents' cache-stable request assembly, covering the static system and per-request context split, the fixed tool list, the rolling conversation breakpoint, which config fields are…

    3.2k GitHub stars~2k tokensUpdated 6 days ago
    Auto-check passed
  • Commerce Trust Safety

    anthropics/commerce-agents

    Official

    The rules the reference agents enforce in code for third-party content, writes, grounding, identity, and memory, each with its module, plus two adversarial-eval rules.

    3.2k GitHub stars~1.8k tokensUpdated 6 days ago
    Auto-check passed
  • Commerce UI Tools

    anthropics/commerce-agents

    Official

    The reference presentation-tool contract, covering server-side enrichment, suggestion chips, the event stream, progressive rendering, both roles' built-in components, and adding a vertical component.

    3.2k GitHub stars~1.7k tokensUpdated 6 days ago
    Auto-check passed

Questions about Commerce Evals

What does Commerce Evals do?

Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures. Commerce Evals is an agent skill from anthropics/commerce-agents, published by the product's own GitHub organization. Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.

When should I use Commerce Evals?

Commerce Evals fits situations like: tasks that involve LLM evaluation; tasks that involve Quizzes and assessments.

How do I install Commerce Evals in Claude Code?

Run `npx skills add anthropics/commerce-agents --skill commerce-evals -a claude-code`. Or copy the skill folder (plugins/commerce-builder/skills/commerce-evals in anthropics/commerce-agents) into .claude/skills/commerce-evals in your project. Claude Code loads it when a task matches its description.

How do I install Commerce Evals in Codex?

Run `npx skills add anthropics/commerce-agents --skill commerce-evals -a codex`. Or copy the skill folder (plugins/commerce-builder/skills/commerce-evals in anthropics/commerce-agents) into .agents/skills/commerce-evals in your project. Codex loads it when a task matches its description.

Can I use Commerce Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anthropics/commerce-agents --skill commerce-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/commerce-evals, .gemini/skills/commerce-evals, .github/skills/commerce-evals and .opencode/skills/commerce-evals in your project.

What does Commerce Evals need to run?

SKILL.md names no scripts, command-line tools or credentials: Commerce Evals is instructions for the agent only.

Does Commerce Evals access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Commerce Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Commerce Evals use?

Commerce Evals is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Commerce Evals use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Commerce Evals?

Skills that share tags, products or a category with Commerce Evals: Advanced Evaluation (guanyang/open-agent-hub, 975 stars), Agentic Eval (github/awesome-copilot, 40k stars), Metric Design (agentscope-ai/OpenJudge, 868 stars) and Agent Evals (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Commerce Evals?

anthropics (a GitHub organization, an official publisher) maintains it in anthropics/commerce-agents, which has 3,179 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 2, 2026.

Source: anthropics/commerce-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.