Agent skill

Phoenix Evals New Metric

by Arize-ai in Arize-ai/phoenix

Create a new built-in classification evaluator for Phoenix evals.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Phoenix Evals New Metric

skills CLI
$ npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Arize-ai/phoenix phoenix-evals-new-metric --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/phoenix-evals-new-metric .claude/skills/phoenix-evals-new-metric && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phoenix-evals-new-metric
GitHub stars
12k
Token cost
~2.3k tokens
SKILL.md length
1,047 words
Files
1
Skills in repo
39
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create a new built-in classification evaluator for Phoenix evals.

  • Works in 9 steps: Gather Requirements → Create the YAML Config → Compile Prompts → …
  • The user asks to create a new eval
  • SKILL.md covers Step 0: Gather Requirements, Step 1: Create the YAML Config, Step 2: Compile Prompts and Step 3: Create the Python…, plus 7 more sections
  • Calls pnpm and make

What it does

Phoenix Evals New Metric is an agent skill from Arize-ai/phoenix. Create a new built-in classification evaluator for Phoenix evals. Use this skill whenever the user asks to create a new eval, build a new metric, add a new builtin evaluator, create an LLM-as-a-judge metric, or add a new classification evaluator to Phoenix.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM observability and LLM evaluation. It works with Python. The repository describes itself as: AI Observability & Evaluation. The licence is Apache-2.0.

When your agent uses it

  • The user asks to create a new eval
  • Build a new metric
  • Add a new builtin evaluator
  • Create an LLM-as-a-judge metric

Example prompts

  • “/phoenix-evals-new-metric”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Gather Requirements
  2. Create the YAML Config
  3. Compile Prompts
  4. Create the Python Evaluator
  5. Create the TypeScript Evaluator
  6. Build JS
  7. Write the Benchmark
  8. Run the Benchmark
  9. Create Documentation

What it can do on your machine

Read from SKILL.md and the folder at commit e471315. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phoenix Evals New Metric loads about 2.3k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 1,047 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Arize-ai/phoenix at commit e471315, republished under its Apache-2.0 licence (© Arize-ai). 1,047 words, ~2,333 tokens.

Download SKILL.mdSave it as .claude/skills/phoenix-evals-new-metric/SKILL.md (or your agent's skills folder).
name
phoenix-evals-new-metric
description
Create a new built-in classification evaluator for Phoenix evals. Use this skill whenever the user asks to create a new eval, build a new metric, add a new builtin evaluator, create an LLM-as-a-judge metric, or add a new classification evaluator to Phoenix.
license
Apache-2.0
metadata.author
oss@arize.com
metadata.version
1.0.0
metadata.internal
true

Creating a New Built-in Classification Evaluator

A built-in evaluator is a YAML config (source of truth) that gets compiled into Python and TypeScript code, wrapped in evaluator classes, benchmarked, and documented. The whole pipeline is linear — follow these steps in order.

Step 0: Gather Requirements

Before writing anything, clarify with the user:

  1. What does this evaluator measure? Get a one-sentence description of the quality dimension.
  2. What input data is available? This determines the template placeholders (e.g., {{input}}, {{output}}, {{reference}}, {{tool_definitions}}). If the user is vague, ask follow-up questions — the placeholders are the contract between the evaluator and the caller.
  3. What labels make sense? Binary is most common (e.g., correct/incorrect, faithful/unfaithful), but some metrics use more. Labels map to scores.
  4. Should this appear in the dataset experiments UI? If yes, it needs the promoted_dataset_evaluator label. Currently only correctness, tool_selection, and tool_invocation have this — some may new evaluators don't need it.

Step 1: Create the YAML Config

Create prompts/classification_evaluator_configs/{NAME}_CLASSIFICATION_EVALUATOR_CONFIG.yaml.

Read an existing config to match the current schema. Start with CORRECTNESS_CLASSIFICATION_EVALUATOR_CONFIG.yaml for a simple example, or TOOL_SELECTION_CLASSIFICATION_EVALUATOR_CONFIG.yaml if your evaluator needs structured span data.

Key Decision Points

choices — Maps label strings to numeric scores. For binary evaluators, use positive/negative labels (e.g., correct: 1.0 / incorrect: 0.0). The labels you pick here flow through to the Python class, TS factory, and benchmarks.

optimization_direction — Use maximize when the positive label is the desired outcome (most evaluators). Use minimize only if the metric measures something undesirable (e.g., hallucination). This affects how Phoenix displays the metric in the UI.

labels — Optional list. Add promoted_dataset_evaluator only if this evaluator should appear in the dataset experiments UI sidebar.

substitutions — Only needed if the evaluator is a promoted_dataset_evaluator and works with structured span data (tool definitions, tool calls, message arrays). These reference formatter snippets defined in prompts/formatters/server.yaml. Read that file if you need substitutions — it defines what structured data formats are available. Most evaluators that only use simple text fields (input, output, reference) don't need substitutions.

Prompt Writing Tips
  • Be explicit about what makes each label correct — the LLM judge needs a clear rubric
  • Separate concerns: if evaluating X, explicitly state you're NOT evaluating Y
  • Wrap inputs in XML-style tags (e.g., <context>, <output>) for clear data formatting
  • Tell the judge to reason before deciding — this improves accuracy
  • Use {{placeholder}} (Mustache syntax) for template variables

Step 2: Compile Prompts

bash
make codegen-prompts

This generates code in three places:

  • packages/phoenix-evals/src/phoenix/evals/__generated__/classification_evaluator_configs/ (Python)
  • src/phoenix/__generated__/classification_evaluator_configs/ (Python, server copy)
  • js/packages/phoenix-evals/src/__generated__/default_templates/ (TypeScript)

Verify the generated files look correct before moving on.

Step 3: Create the Python Evaluator

Create packages/phoenix-evals/src/phoenix/evals/metrics/{name}.py.

Read correctness.py in that directory — it's the canonical example. Your evaluator follows the same pattern: subclass ClassificationEvaluator, pull constants from the generated config, define a Pydantic input schema with fields matching your template placeholders.

After creating the file, add it to the exports in metrics/__init__.py — both the import and the __all__ list. Read the current __init__.py to see the existing pattern.

Step 4: Create the TypeScript Evaluator

Create js/packages/phoenix-evals/src/llm/create{Name}Evaluator.ts.

Read createCorrectnessEvaluator.ts — it's the canonical example. The pattern is a factory function that wraps createClassificationEvaluator with defaults from the generated config.

Then:

  1. Add the export to js/packages/phoenix-evals/src/llm/index.ts
  2. Add a vitest test under js/packages/phoenix-evals/test/llm/ — read createFaithfulnessEvaluator.test.ts there for the test pattern

Step 5: Build JS

bash
cd js && pnpm build

Fix any TypeScript errors before proceeding.

Step 6: Write the Benchmark

Create js/benchmarks/evals-benchmarks/src/{name}.eval.ts.

Read existing benchmarks in that directory to match the current patterns:

  • tool_invocation.eval.ts — multi-category analysis and aggregate metrics
  • aggregateMetrics.ts — shared macro precision/recall/F1 accumulation
Show full SKILL.md (474 more words)Show less
Benchmark Requirements
  • 30-50 synthetic examples organized by category
  • 2-4 examples per category covering: success cases, failure modes, and edge cases
  • Accuracy evaluator that compares predicted vs expected labels
  • Failed examples printer — this is critical for debugging. For each misclassified example, print: category, input, output (truncated), expected vs actual label, and the LLM judge's explanation
  • Per-category accuracy breakdown in the output
  • For binary evaluators, a confusion matrix is helpful

The task function must return input and output text in its result so the failed examples printer has access to them.

Consider using a separate agent session for synthetic dataset generation if the examples need realistic domain-specific content — this keeps the dataset creation focused and avoids context-switching.

Step 7: Run the Benchmark

bash
# Terminal 1: Start Phoenix using the normal configured database. Do not set
# PHOENIX_WORKING_DIR unless the user explicitly requests an isolated instance.
phoenix serve

# Terminal 2: Run the benchmark
cd js
pnpm --filter evals-benchmarks... build
pnpm --filter evals-benchmarks exec vitest run \
  src/{name}.eval.ts --config phoenix.vitest.config.ts

Target >80% accuracy. If accuracy is low, look at the failed examples output to decide whether to adjust the prompt (Step 1) or the benchmark examples (Step 6). Iterate until accuracy is acceptable.

Step 8: Create Documentation

Create docs/phoenix/evaluation/pre-built-metrics/{name}.mdx.

Read faithfulness.mdx in that directory — it's the template. Follow the same section structure:

  1. Overview — when to use, what it measures
  2. Supported Levels — span/trace/session, relevant span kinds
  3. Input Requirements — required fields table
  4. Output Interpretation — labels, scores, direction
  5. Usage Examples — Python and TypeScript in tabs
  6. Using Input Mapping — lambda example if applicable
  7. Viewing/Modifying the Prompt — link to GitHub config, custom prompt usage
  8. Configuration — link to LLM config docs
  9. Using with Phoenix — links to traces and experiments docs
  10. Benchmarks — "Coming soon" placeholder (until benchmark results are published)
  11. API Reference — links to Python and TypeScript API docs
  12. Related — links to related evaluators
Navigation Updates

After creating the docs page, update these three files:

  1. docs.json — add the page to the Evaluation > Pre-built Metrics nav group
  2. docs/phoenix/evaluation/pre-built-metrics.mdx — add a card to the landing page grid
  3. docs/phoenix/sitemap.xml — add the new URL

Read each file to see the existing pattern before editing.

Checklist

Before calling it done, verify:

  • YAML config created with clear rubric and appropriate labels/choices
  • make codegen-prompts ran successfully
  • Python evaluator class with input schema matching template placeholders
  • Python exports updated in metrics/__init__.py
  • TypeScript evaluator factory with types
  • TypeScript export added to llm/index.ts
  • Vitest test for TypeScript evaluator
  • JS packages rebuilt (cd js && pnpm build)
  • Benchmark with 30-50 examples, category breakdown, failed examples printer
  • Benchmark accuracy >80%
  • Documentation page following the template structure
  • docs.json nav updated
  • Landing page card added
  • Sitemap updated

Retrospection

After completing the workflow, verify these instructions matched reality:

  • Did any file paths, export patterns, or command names change from what's described here?
  • Did the YAML config schema gain or lose fields since this was written?
  • Did the benchmark or docs patterns evolve from the referenced examples?
  • Did make codegen-prompts generate to different locations?

If anything drifted, update this SKILL.md before finishing so the next person (or agent) doesn't hit the same surprises.

© Arize-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/phoenix-evals-new-metric of Arize-ai/phoenix.

Open the folder on GitHubat commit e471315

Compare with similar skills

Phoenix Evals New Metric next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phoenix Evals New Metric compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phoenix Evals New Metric this skillArize-ai/phoenix12k—~2.3kAutomated safety check: PassApache-2.0
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence
Phoenix Evalsgithub/awesome-copilot40k2 repos~1.1kAutomated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
LLM Trace Review Interfaceai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0
Analyzing Claude Code Sessionsamd/gaia1.6k—~2.3kAutomated safety check: PassMIT

Similar skills

  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Phoenix Evals

    github/awesome-copilot

    Official

    Build and run evaluators for AI/LLM applications using Phoenix.

    40k GitHub starsUsed in 2 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Trace Review Interface

    ai-evals-course/evals-skills

    Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

    1.5k GitHub stars~1.4k tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Trackio Experiment Tracking

    huggingface/skills

    Official

    Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space.

    11k GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from Arize-ai/phoenix

All 39 skills in this repo
  • Harbor Exec

    Arize-ai/phoenix

    A skill your agent uses when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only…

    12k GitHub stars~909 tokensUpdated yesterday
    Auto-check passed
  • Mintlify

    Arize-ai/phoenix

    Build and maintain documentation sites with Mintlify. An agent skill from Arize-ai/phoenix.

    12k GitHub starsUsed in 8 repos~3.4k tokens
    Auto-check passed
  • Phoenix Frontend

    Arize-ai/phoenix

    Frontend development guidelines for the Phoenix AI observability platform.

    12k GitHub stars~709 tokensUpdated yesterday
    Auto-check passed
  • Phoenix Graphql

    Arize-ai/phoenix

    Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.

    12k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Phoenix Server

    Arize-ai/phoenix

    Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).

    12k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Phoenix Storybook

    Arize-ai/phoenix

    Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).

    12k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Phoenix Evals New Metric

What does Phoenix Evals New Metric do?

Create a new built-in classification evaluator for Phoenix evals. Phoenix Evals New Metric is an agent skill from Arize-ai/phoenix. Create a new built-in classification evaluator for Phoenix evals.

When should I use Phoenix Evals New Metric?

Phoenix Evals New Metric fits situations like: the user asks to create a new eval; build a new metric; add a new builtin evaluator; create an LLM-as-a-judge metric.

How do I install Phoenix Evals New Metric in Claude Code?

Run `npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a claude-code`. Or copy the skill folder (.agents/skills/phoenix-evals-new-metric in Arize-ai/phoenix) into .claude/skills/phoenix-evals-new-metric in your project. Claude Code loads it when a task matches its description.

How do I install Phoenix Evals New Metric in Codex?

Run `npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a codex`. Or copy the skill folder (.agents/skills/phoenix-evals-new-metric in Arize-ai/phoenix) into .agents/skills/phoenix-evals-new-metric in your project. Codex loads it when a task matches its description.

Can I use Phoenix Evals New Metric in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phoenix-evals-new-metric, .gemini/skills/phoenix-evals-new-metric, .github/skills/phoenix-evals-new-metric and .opencode/skills/phoenix-evals-new-metric in your project.

What does Phoenix Evals New Metric need to run?

Going by SKILL.md and its folder, Phoenix Evals New Metric needs the command-line tools its instructions call (pnpm and make). Our summary lists: Python 3.

Does Phoenix Evals New Metric access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Phoenix Evals New Metric safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phoenix Evals New Metric use?

Phoenix Evals New Metric is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phoenix Evals New Metric use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phoenix Evals New Metric?

Skills that share tags, products or a category with Phoenix Evals New Metric: Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), Phoenix Evals (github/awesome-copilot, 40k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and LLM Trace Review Interface (ai-evals-course/evals-skills, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phoenix Evals New Metric?

Arize-ai (a GitHub organization) maintains it in Arize-ai/phoenix, which has 11,770 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 10, 2026.

Source: Arize-ai/phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.