LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/).
$ npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Arize-ai/phoenix pxi-eval-dataset --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/pxi-eval-dataset .claude/skills/pxi-eval-dataset && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pxi-eval-dataset" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/pxi-eval-dataset into .claude/skills/pxi-eval-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pxi-eval-dataset", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/pxi-eval-datasetType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Arize-ai/phoenix pxi-eval-dataset --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/pxi-eval-dataset .agents/skills/pxi-eval-dataset && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pxi-eval-dataset" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/pxi-eval-dataset into .agents/skills/pxi-eval-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pxi-eval-dataset", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Arize-ai/phoenix pxi-eval-dataset --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/pxi-eval-dataset .cursor/skills/pxi-eval-dataset && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pxi-eval-dataset" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/pxi-eval-dataset into .cursor/skills/pxi-eval-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pxi-eval-dataset", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Arize-ai/phoenix.git --path .agents/skills/pxi-eval-dataset--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Arize-ai/phoenix pxi-eval-dataset --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/pxi-eval-dataset .gemini/skills/pxi-eval-dataset && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pxi-eval-dataset" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/pxi-eval-dataset into .gemini/skills/pxi-eval-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pxi-eval-dataset", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Arize-ai/phoenix pxi-eval-datasetInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/pxi-eval-dataset .github/skills/pxi-eval-dataset && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pxi-eval-dataset" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/pxi-eval-dataset into .github/skills/pxi-eval-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pxi-eval-dataset", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Arize-ai/phoenix pxi-eval-dataset --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/pxi-eval-dataset .opencode/skills/pxi-eval-dataset && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pxi-eval-dataset" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/pxi-eval-dataset into .opencode/skills/pxi-eval-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pxi-eval-dataset", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pxi-eval-datasetGenerate synthetic evaluation datasets for the PXI eval harness (evals/pxi/).
Pxi Eval Dataset is an agent skill from Arize-ai/phoenix. Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/). Use whenever the user asks to create, author, draft, expand, or audit an eval dataset for a PXI tool, skill, or behavior — including phrases like "write evals for <tool", "test PXI behavior", "synthetic dataset for PXI", "cover this tool with eval examples", or "find gaps in our PXI eval coverage". Inspects whichever evaluators currently live under evals/pxi/evaluators/ at use time and pauses to recommend a new evaluator if the behavior…
Its SKILL.md is about 7.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/annotate_via_codex.sh`).
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: AI Observability & Evaluation. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 52f76fc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
uvcodexFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pxi Eval Dataset loads about 7.1k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 3,580 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Arize-ai/phoenix at commit 52f76fc, republished under its Apache-2.0 licence (© Arize-ai). 3,580 words, ~7,063 tokens.
.claude/skills/pxi-eval-dataset/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Produce a small, well-targeted YAML dataset that drops into
evals/pxi/datasets/<name>.yaml, runs through
evals/pxi/harness/run_experiment.py, and is scored by deterministic
code evaluators under evals/pxi/evaluators/.
The aim is a minimal but representative set of synthetic examples — think unit tests, not a benchmark. 10–50 examples, each covering a distinct dimension. Add more only when a new example tests something no existing example does.
Every example must be scorable by deterministic / heuristic / code logic.
Every example must include a non-empty splits: list. Use
splits: [regression] for small committed regression suites unless the
user explicitly asks for a dev or val dataset; keep regression,
dev, and val disjoint.
Confirm with the user what's under test: a specific PXI tool name (e.g.
set_time_range), a skill, or a higher-level behavior. For tools:
src/phoenix/server/agents/toolsets/external/tools/<name>.py — the
ToolDefinition's description and parameters_json_schema are what
the LLM actually sees.src/phoenix/server/agents/toolsets/__init__.py::build_toolset
and src/phoenix/server/agents/toolsets/external/__init__.py::build_external_toolset
to learn the availability conditions — which ChatContext must be
present for the tool to be exposed (e.g. set_spans_filter only
exists when deps.contexts.project is set).List evals/pxi/evaluators/ and read each module. For every
evaluator, note:
@create_evaluator(...) decorator,expected (e.g. expected.tools.required,
expected.tool_call_args[<tool>]),Also peek at evals/pxi/tests/test_evaluators.py for canonical input
shapes.
Do NOT hard-code knowledge of which evaluators exist — re-read the directory every time. The set will grow.
Decide whether the behavior under test can be scored by what's there today.
Yes → continue. The chosen evaluator names go in the dataset's
evaluators: field (required, top-level). The runner uses ONLY the
listed evaluators; unrelated ones are not invoked, so the dashboard
stays free of vacuous passes from evaluators that don't apply to
this dataset. Available names live in
evals/pxi/evaluators/__init__.py (EVALUATORS_BY_NAME).
No → stop and summarize the gap to the user. Propose the shape of a new evaluator:
expected.<field> it would read,Then ask: "Should I add this evaluator before we generate examples?" If yes:
evals/pxi/evaluators/<file>.py,evals/pxi/tests/test_evaluators.py,evals/pxi/evaluators/__init__.py,If no, scope the dataset down to assertions the existing evaluators can score, and tell the user what coverage that costs.
Walk this checklist for the target. For each row, write down which queries you'll add to cover it. Skip rows that don't apply (e.g. booleans on a tool with no boolean field).
Polarity and difficulty targets are stated once in step 5 below; aim for those across the whole dataset, not within each row.
If you cannot fill at least 10 rows, the target is probably too narrow — confirm scope with the user.
For each coverage dimension, write one or more queries that feel like they came from a real Phoenix user. Mix across:
Anti-patterns to avoid:
Ground truth must be generated in a fresh subprocess, not inline.
When annotation happens in the same context that drafted the queries, the
agent builds a prior from its own examples: it anchors on condition
patterns it established early, collapses ambiguous cases into overconfident
annotations, and fills in tool_call_args based on what earlier examples
"look like" rather than what the tool spec actually says. This is context
rot — annotation signal weakens as the dataset grows.
The remedy: annotate each example in a separate subprocess whose context contains the full PXI toolset, the author's design intent for the query, and the schema contract. Each annotation is an independent reasoning trace from the toolset spec.
Prepare once, reuse across examples. Before launching annotation subprocesses, extract the full toolset spec once (see section 1 below) and cache it. Every annotation subprocess across the dataset reuses the same toolset block; only the query and design-intent block change per-example.
Full toolset spec — the complete toolset the PXI agent sees at
runtime, NOT just the tool under test. For each tool, include the
name, description, and parameters_json_schema. The annotator has
to decide whether the right answer is the focal tool, a different
tool, or no tool at all — and cannot do that without seeing every
tool's spec. This is the single biggest hedge against annotator
hallucination: with only one tool visible, the annotator will
force-fit it to ambiguous or negative queries.
Source the toolset from build_toolset and build_external_toolset
in src/phoenix/server/agents/toolsets/ (and
src/phoenix/server/agents/toolsets/external/). The eval harness
configures ChatContext such that all external tools are available
— see evals/pxi/harness/agent_task.py. List every tool that survives
that filtering, not just the focal one.
Mark the focal tool explicitly at the top of this section
(★ FOCAL TOOL: <name> — this dataset evaluates this tool's behavior). The annotator's job is still to reason from the spec for
whichever tool best fits the query, but the highlight tells them
which tool's evaluator family will judge the answer.
Evaluator contract — provide the dataset-specific
expected-block schema derived from the selected evaluators (what
each expected field means, matching semantics, and which fields
may be omitted or expressed as variants). Add the explicit
instruction: if you are unsure about a value, omit the key rather
than guessing.
Query and design intent — the input.query string, plus the
dataset author's framing if assigned:
difficulty: obvious | moderate | trickypolarity: positive | negativecategory: <coverage dimension> (the taxonomy from step 4 —
parameter, value, combination, negative, ambiguity, plus any
finer-grained category like span_kind, status, latency,
negative_chitchat, negative_wrong_tool, etc.)notes: the author wroteFrame this for the annotator as the dataset author's intent, NOT ground truth. The annotator should reason from the toolset spec and may disagree with the author — that disagreement is one of the most valuable signals this protocol produces, because it surfaces mis-classified queries.
Task instruction (goal-oriented, adapt only the schema reference):
"Your goal: determine the correct
expected:block for this dataset example — the expected output needed to validate whether the PXI agent took the right actions in response to the user's query, given the full toolset spec, evaluation goal, evaluator set, and expected-block schema.Constraints:
- Follow the expected-block schema provided for this dataset. The schema is determined by the evaluation goal and selected evaluators; do not assume every dataset validates only tool selection or tool arguments.
- When the schema includes tool expectations, consider every tool in the toolset spec, not just the focal one. The right answer may be the focal tool, a different tool, or no tool at all.
- Use only values directly implied by the relevant spec and the query. Omit any key whose value the spec doesn't determine — don't invent a plausible default.
- The author's design intent (
difficulty,polarity,category,notes) is a hint, not ground truth. If your reading of the spec contradicts the author's labeled polarity, output what the spec implies and record the disagreement inmetadata.notes.- Return exactly one best
expected:block. If multiple outputs seem valid, choose the single strongest answer and mention the alternatives in your brief reasoning; the orchestrator decides later whether variants belong in the final example.Output format (exactly two parts, in this order):
- Brief reasoning — one or two sentences explaining which expected output you concluded should be used and why. Just enough that an orchestrator comparing three independent annotations can see where you agree or diverge with the others.
- The
expected:block as YAML matching the schema reference provided in this prompt. No other prose."
This frames the task as a goal with constraints, not a procedural recipe. Cross-model fan-out is only useful if each annotator reasons independently — prescribing the exact reasoning steps collapses that diversity. The brief-reasoning requirement gives the orchestrator enough signal to adjudicate disagreements without dictating how each annotator gets there.
Before launching annotation subprocesses, define the expected-block
schema for this dataset from the evaluation goal and selected
evaluators. Include that schema verbatim in every annotation prompt.
For example, a dataset that validates tool selection and tool
arguments may use expected.tools and expected.tool_call_args, while
another dataset may validate response text, refusal behavior, or some
other evaluator-specific field.
When the schema includes tool expectations, use these conventions:
tools.required for the must-be-called tool(s).tools.exact_match: true for strict-no-extras cases.tools.required: [] and tools.exact_match: true. Do not enumerate
every tool in tools.forbidden; the tool list can grow and make
that brittle. Do not supply tool_call_args.tool_call_args[<tool>] only when argument correctness matters;
leave it off for "any args are fine" cases — especially ambiguous
cases where you want to assert the tool fires but not pin its args.Collect three independent opinions for every dataset example: Sonnet, Opus, and Codex. Even obvious and pure-negative examples get the full fan-out so the dataset has a consistent annotation provenance and easy-to-audit agreement metadata.
Cross-model diversity — different families, different training, and (for Codex) different agent runtime — catches single-model biases that seed-level diversity misses.
How each opinion is invoked. The Claude Code Agent tool is
Anthropic-only, so the OpenAI-side opinion comes from Codex via a
Bash wrapper. Using Codex (rather than a one-shot Chat Completions
call) gives the OpenAI annotator the same kind of environment access
a Claude subagent has — sandbox-constrained file reads, repo
navigation, the ability to double-check its reading of the toolset
against the actual source. This keeps the three opinions
genuinely symmetric.
| Opinion | Invocation |
|---|---|
| Sonnet | Agent tool, model: "sonnet" |
| Opus | Agent tool, model: "opus" |
| Codex | Bash tool, piping the prompt into .agents/skills/pxi-eval-dataset/scripts/annotate_via_codex.sh |
The Codex wrapper at .agents/skills/pxi-eval-dataset/scripts/annotate_via_codex.sh reads a
prompt on stdin and writes Codex's final message to stdout. It runs
codex exec with --sandbox read-only (the annotator can read the
repo but cannot modify anything), --ephemeral (no session
persistence between invocations), and --skip-git-repo-check.
Authentication is whatever Codex is already configured with
(codex login status); no OPENAI_API_KEY plumbing required. Example
invocation:
cat /tmp/annotation_prompt.txt | .agents/skills/pxi-eval-dataset/scripts/annotate_via_codex.sh
# Or pin a specific model:
cat /tmp/annotation_prompt.txt | .agents/skills/pxi-eval-dataset/scripts/annotate_via_codex.sh --model o3Run all three opinions in parallel. Send a single message
containing two Agent tool calls (Sonnet, Opus) and one Bash
tool call (the Codex wrapper) — they will run concurrently. Do not
share context between them.
After the three annotation subprocesses return, run a fourth fresh
subprocess as the orchestrator. Its job is to merge the candidates
into a single expected: block, with authority to drop spec-invalid
proposals. Use a separate subprocess so the orchestrator sees
anonymized candidate blocks and cannot be biased by knowing which
model produced which annotation.
Orchestrator subprocess context:
difficulty, polarity, category, notes). The orchestrator
should also be willing to override the author's polarity if the
toolset spec and the majority of annotation opinions disagree with
it; flag any such override with metadata.notes: 'orchestrator overrode author polarity'.expected: blocks (anonymized — do not tell the
orchestrator which model produced which proposal; this prevents
reputation bias)Orchestrator decision rules:
expected: block where the models agree.tool_call_args[<tool>] so the evaluator checks tool selection
without pinning arbitrary arguments.The orchestrator must justify any "spec-invalid" rejection against a specific clause of the tool spec. It does NOT vote: it adjudicates against the spec.
After orchestration, write the annotation provenance into the example's metadata. This makes noisy / low-confidence ground truth easy to audit later.
metadata:
category: preset
annotation:
agreement: high # high | medium | low — see rule below
n_opinions: 3 # number of annotation subprocesses runAgreement categorization rule:
high — all three opinions produced equivalent expected: blocks
(equal tools.{required,forbidden} lists after sorting, and per-tool
tool_call_args dicts equal after applying _normalize_arg_value
from evals/pxi/evaluators/tools.py to each value).medium — majority agreed; minority differed but the
orchestrator was able to merge (as variants) or drop (as spec-invalid).low — no majority; the orchestrator either fell back to
omit-args or made a judgment call with weak signal.Why this matters: an agreement: low example may still have a
correct ground truth, but it's a candidate for human review — the
models hedged, which can indicate either a genuinely ambiguous query
or a spec the LLMs interpreted differently. A follow-up audit script
can flag these for inspection.
Equality-detection caveat: the heuristic catches DSL clause
reordering (because it reuses _normalize_arg_value) but does NOT
catch semantically-equivalent rewrites (e.g. latency_ms >= 5000 vs
latency_ms > 4999). Such cases will surface as agreement: medium
or low, which is arguably the right signal — the model hedged.
When more than one set of arguments is reasonable per the spec, write
tool_call_args[<tool>] as a list of dicts instead of a single
dict. The observed call passes if any variant matches under the same
subset rules as the single-dict form. This lets each variant pin a
different combination of keys (e.g. a preset variant vs. a custom
variant with startTime):
tool_call_args:
set_time_range:
- { timeRangeKey: 1d } # preset variant
- { timeRangeKey: 7d } # preset variant
- { timeRangeKey: custom } # custom range, key omittedUse variants for genuinely-ambiguous queries (e.g. "show me recent traces" when only a small number of choices are defensible). Keep the list to at most three variants. If there are many valid choices, omit the argument assertion instead. Do not use variants to paper over agent inconsistency on queries that have one clear correct answer.
When the orchestrator returns adjudicated expected: blocks (or new
metadata.annotation blocks during a backfill), prefer targeted
Edit operations on the specific examples that change. Do NOT
re-serialize the whole YAML file with pyyaml or ruamel.yaml:
pyyaml.safe_dump strips comments and rewrites indentation (e.g.
sequence-indent=2 and unquoted list-item style) — a 38-example
file comes back with a 1000+ line diff that obscures the real
semantic change.ruamel.yaml can round-trip, but only when you pin both
preserve_quotes=True AND the exact indent(mapping=M, sequence=S, offset=O) settings of the source file. Mismatched indent settings
silently corrupt list-item structure (e.g. child fields land at the
list-item column instead of inside the item). Recover requires a
full revert.The safe pattern for bulk merges:
{example_id: (new_expected, new_metadata)} map.Edit call
targeting the example's existing block as old_string and the
merged block as new_string. Each Edit's diff stays scoped to the
one example.load_dataset after the batch.This preserves section-header comments by construction and produces a diff that reviews cleanly.
Section-header comments are convenience navigation only. The
load-bearing per-example commentary lives in
metadata.annotation.notes and metadata.triage.notes, which
round-trip safely through any merge approach. If a bulk rewrite loses
section comments, the dataset is still fully self-documenting — no
need to reconstruct them.
Save to evals/pxi/datasets/<name>.yaml. Then:
# Parse + validate schema:
uv run python -c "from evals.pxi.harness.datasets import load_dataset; load_dataset('<name>')"
# Run the experiment end-to-end against the real PXI agent:
uv run python -m evals.pxi.harness.run_experiment --dataset <name>Run the experiment end-to-end and triage every failure before calling the dataset done. Don't assume a failed example means a broken agent — eval suites have three plausible failure sources, and mixing them up is the easiest way to ship bad ground truth or chase phantom agent regressions.
For each failed example, classify the failure into one bucket:
| Category | What it looks like | What to do |
|---|---|---|
| Dataset / annotation issue | Agent's actual output is valid and reasonable per the toolset spec, but doesn't match the dataset's expected: block. The dataset author (or annotation protocol) was wrong, too strict, or missed a valid variant. | Fix the dataset: relax the expectation, add a variant via the variant-list syntax, omit tool_call_args if the case is ambiguous, or correct an outright wrong annotation. Note in metadata.notes if a polarity flip is involved. |
| Genuine agent issue | The harness caught a real problem — the agent called the wrong tool, missed a tool, hallucinated arg values, or violated the spec. | Leave the example as-is; the failure is doing its job. Surface the failure to whoever owns the agent's behavior (file an issue, add to the regression log). |
| Harness / evaluator issue | The evaluator's matching logic, the runner's plumbing, or the Phoenix-side experiment integration is broken. Symptom: the agent's output and the expected block agree by any reasonable reading, but the evaluator labels it failing — or vice versa. | Fix the evaluator or harness, add a unit test in test_evaluators.py covering the case, then re-run. Do NOT paper over a harness bug by editing the dataset. |
id, observed tool calls, and evaluator
label from the experiment output (the runner prints these per
evaluator).parameters_json_schema and the
relevant evaluator code to decide which category the failure
belongs in. The categorization is not always obvious — when
genuinely uncertain, prefer "harness/evaluator issue" and read the
evaluator first, since a broken evaluator silently corrupts both
of the other categories.{example_id: category} — is enough.
Without this list, the fix loop tends to oscillate (fix dataset →
evaluator flips on a different example → "fix" evaluator → first
example breaks again).annotation.agreement: low in metadata — these
are exactly the candidates for human review even if they passed.For local-only experimentation use the harness env vars (see
evals/pxi/harness/README.md). There is no --limit flag — keep the
dataset small while iterating, or commit a temporary copy.
YAML gotchas that cost an iteration if you miss them:
splits: [...]. For the normal small
regression datasets this skill creates, set splits: [regression].
The harness rejects missing or empty split lists.ConfigDict(extra="forbid") on
EvalDataset, so a typo in a top-level key (example instead of
examples) raises rather than being silently ignored.", [, {, :, *, &, !,
or %, and any value containing a leading/trailing colon or # —
e.g. condition: "status_code == 'ERROR'". Plain DSL strings without
punctuation can stay unquoted, but quoting consistently is cheaper
than diagnosing a parser error.ids must be unique across the dataset; duplicate ids fail
validation with the offending ids listed.Reference the live datasets in evals/pxi/datasets/ for current
dataset shapes, examples, and naming conventions. Do not copy a schema
from this skill into prompts. Instead, derive the expected-block schema
from the selected evaluators and include that schema in each
annotation and orchestration prompt.
Validator is in evals/pxi/harness/datasets.py. Matching semantics in
evals/pxi/evaluators/tools.py:
tool_call_args: an observed call passes if it
has every expected key with a matching value; extra keys are ignored.tool_call_args[<tool>] is a list of dicts:
the observed call passes if ANY variant matches. Use for genuinely
ambiguous queries where multiple arg shapes are equally correct.and
are normalized to a frozenset, so
"span_kind == 'LLM' and latency_ms >= 5000" matches
"latency_ms >= 5000 and span_kind == 'LLM'". Pure-or expressions
are normalized the same way; mixed and/or fall back to exact
string comparison.© Arize-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in .agents/skills/pxi-eval-dataset of Arize-ai/phoenix.
Open the folder on GitHubat commit 52f76fc
Pxi Eval Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pxi Eval Dataset this skillArize-ai/phoenix | 12k | — | ~7.1k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
Arize-ai/phoenix
A skill your agent uses when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only…
Arize-ai/phoenix
Build and maintain documentation sites with Mintlify. An agent skill from Arize-ai/phoenix.
Arize-ai/phoenix
Frontend development guidelines for the Phoenix AI observability platform.
Arize-ai/phoenix
Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.
Arize-ai/phoenix
Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).
Arize-ai/phoenix
Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).
Categories
Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/). Pxi Eval Dataset is an agent skill from Arize-ai/phoenix. Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/).
Pxi Eval Dataset fits situations like: the user asks to create; audit an eval dataset for a PXI tool; behavior — including phrases like write evals for <tool; test PXI behavior.
Run `npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a claude-code`. Or copy the skill folder (.agents/skills/pxi-eval-dataset in Arize-ai/phoenix) into .claude/skills/pxi-eval-dataset in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a codex`. Or copy the skill folder (.agents/skills/pxi-eval-dataset in Arize-ai/phoenix) into .agents/skills/pxi-eval-dataset in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arize-ai/phoenix --skill pxi-eval-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pxi-eval-dataset, .gemini/skills/pxi-eval-dataset, .github/skills/pxi-eval-dataset and .opencode/skills/pxi-eval-dataset in your project.
Going by SKILL.md and its folder, Pxi Eval Dataset needs a shell for the scripts in its folder, the command-line tools its instructions call (uv and codex) and credentials named OPENAI_API_KEY. Our summary lists: A Bash shell.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Pxi Eval Dataset is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.1k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Pxi Eval Dataset: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Arize-ai (a GitHub organization) maintains it in Arize-ai/phoenix, which has 11,764 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 9, 2026.
Source: Arize-ai/phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.