Agent skill

Evaluators

by Arize-ai in Arize-ai/phoenix

Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

Custom licenceAuto-check passedEducation

Install Evaluators

skills CLI
$ npx skills add Arize-ai/phoenix --skill evaluators -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Arize-ai/phoenix evaluators --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/phoenix/server/agents/prompts/skills/evaluators .claude/skills/evaluators && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluators
GitHub stars
12k
Token cost
~1.7k tokens
SKILL.md length
882 words
Files
1
Skills in repo
39
Repo updated
First seen
Licence
Custom licence

At a glance

Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

  • Works in 7 steps: Derive the grading task from the stated… → Inventory before creating. Read the… → Decide the labels. Choose a small,… → …
  • The user wants to create a new evaluator
  • SKILL.md covers The Authoring Loop, Reference Provenance, Choosing The Judgment Structure and Matching The Field Topology, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Evaluators is an agent skill from Arize-ai/phoenix. Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output. Trigger when the user wants to create a new evaluator, improve an existing one's logic or rubric, choose labels, or decide what to measure on a dataset or experiment. Do NOT trigger on: (1) manual prompt drafting (use playground), (2) running or comparing experiments themselves (use experiments), (3) cross-trace failure diagnosis with no evaluator in scope (use phoenix-error-analysis).

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering LLM observability and Quizzes and assessments. The repository describes itself as: AI Observability & Evaluation.

When your agent uses it

  • The user wants to create a new evaluator
  • Improve an existing ones logic
  • Decide what to measure on a dataset
  • Manual prompt drafting (use playground)

Example prompts

  • “/evaluators”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Derive the grading task from the stated purpose — a hypothesis and its evaluator are one design
  2. Inventory before creating. Read the dataset's existing evaluators and check input-shape
  3. Decide the labels. Choose a small, mutually exclusive, collectively exhaustive set — often
  4. Locate the signal in the run's fields — a top-level key, a chat-style messages array,
  5. Write the judgment: a function that reads the field and returns the label or score, or a judge
  6. Calibrate against several representative cases covering the named failure modes — one preview is
  7. Iterate until the representative cases label correctly and the tradeoff is acceptable.

What it can do on your machine

Read from SKILL.md and the folder at commit 3383f07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluators loads about 1.7k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 882 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 882 words (~1,720 tokens).

“A Phoenix evaluator scores a run: it reads some subset of the run's input, output, reference, and metadata and returns named annotations — a label, a score, or both. The two artifact kinds — a code evaluator (a Python or…”

— opening of SKILL.md by Arize-ai, Custom licence
name
evaluators
summary
Design or refine a code or LLM evaluator — labels, logic or rubric, the field it reads, and representative preview cases.

Read the full SKILL.md on GitHub

Files

Just SKILL.md in src/phoenix/server/agents/prompts/skills/evaluators of Arize-ai/phoenix.

Open the folder on GitHubat commit 3383f07

Compare with similar skills

Evaluators next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluators compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluators this skillArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch65k—~2kAutomated safety check: PassMIT
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch65k—~2.1kAutomated safety check: PassMIT
Claude Certification Tutorrohitg00/ai-engineering-from-scratch65k—~3kAutomated safety check: PassMIT
Generate Verifiers Envadithya-s-k/FineEnvs4211 repos~2.3kAutomated safety check: PassApache-2.0
Laya Integrationwdobry/laya-playground182—~3.2kAutomated safety check: PassMIT

Similar skills

  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    65k GitHub stars~2k tokensUpdated yesterday
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    65k GitHub stars~2.1k tokensUpdated yesterday
    EducationAuto-check passed
  • Claude Certification Tutor

    rohitg00/ai-engineering-from-scratch

    Guides a learner through one of four independent Claude certification tracks with onboarding, lessons, practice labs, mock exams and remediation.

    65k GitHub stars~3k tokensUpdated yesterday
    EducationAuto-check passed
  • Generate Verifiers Env

    adithya-s-k/FineEnvs

    Builds a Verifiers (PrimeIntellect) variant of an RL environment.

    421 GitHub starsUsed in 1 repo~2.3k tokens
    EducationAuto-check passed
  • Laya Integration

    wdobry/laya-playground

    Add fast, local, typed decisions to any project with Laya, an open-source non-generative decision model (pip install laya).

    182 GitHub stars~3.2k tokensUpdated 17 days ago
    EducationAuto-check passed
  • 01 Auto Arena

    agentscope-ai/OpenJudge

    Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

    867 GitHub starsUsed in 1 repo~2.5k tokens
    EducationAuto-check passed

More from Arize-ai/phoenix

All 39 skills in this repo
  • Harbor Exec

    Arize-ai/phoenix

    A skill your agent uses when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only…

    12k GitHub stars~909 tokensUpdated today
    Auto-check passed
  • Mintlify

    Arize-ai/phoenix

    Build and maintain documentation sites with Mintlify. An agent skill from Arize-ai/phoenix.

    12k GitHub starsUsed in 7 repos~3.4k tokens
    Auto-check passed
  • Phoenix Frontend

    Arize-ai/phoenix

    Frontend development guidelines for the Phoenix AI observability platform.

    12k GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Phoenix Graphql

    Arize-ai/phoenix

    Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.

    12k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Phoenix Server

    Arize-ai/phoenix

    Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).

    12k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Phoenix Storybook

    Arize-ai/phoenix

    Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).

    12k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Evaluators

What does Evaluators do?

Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output. Evaluators is an agent skill from Arize-ai/phoenix. Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

When should I use Evaluators?

Evaluators fits situations like: the user wants to create a new evaluator; improve an existing ones logic; decide what to measure on a dataset; manual prompt drafting (use playground).

How do I install Evaluators in Claude Code?

Run `npx skills add Arize-ai/phoenix --skill evaluators -a claude-code`. Or copy the skill folder (src/phoenix/server/agents/prompts/skills/evaluators in Arize-ai/phoenix) into .claude/skills/evaluators in your project. Claude Code loads it when a task matches its description.

How do I install Evaluators in Codex?

Run `npx skills add Arize-ai/phoenix --skill evaluators -a codex`. Or copy the skill folder (src/phoenix/server/agents/prompts/skills/evaluators in Arize-ai/phoenix) into .agents/skills/evaluators in your project. Codex loads it when a task matches its description.

Can I use Evaluators in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arize-ai/phoenix --skill evaluators -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluators, .gemini/skills/evaluators, .github/skills/evaluators and .opencode/skills/evaluators in your project.

What does Evaluators need to run?

SKILL.md names no scripts, command-line tools or credentials: Evaluators is instructions for the agent only. Our summary lists: Python 3.

Does Evaluators access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluators safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluators use?

Evaluators has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Evaluators use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluators?

Skills that share tags, products or a category with Evaluators: AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 65k stars), AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 65k stars), Claude Certification Tutor (rohitg00/ai-engineering-from-scratch, 65k stars) and Generate Verifiers Env (adithya-s-k/FineEnvs, 421 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluators?

Arize-ai (a GitHub organization) maintains it in Arize-ai/phoenix, which has 11,738 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 7, 2026.

Source: Arize-ai/phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.