Official agent skill

Phoenix Evals

by github in github/awesome-copilot

Build and run evaluators for AI/LLM applications using Phoenix.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Phoenix Evals

skills CLI
$ npx skills add github/awesome-copilot --skill phoenix-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot phoenix-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/phoenix-evals .claude/skills/phoenix-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phoenix-evals
GitHub stars
40k
Used in
3 other repos
Token cost
~1.1k tokens
SKILL.md length
189 words
Files
35 (incl. references)
Skills in repo
417
Repo updated
First seen
Licence
Apache-2.0

At a glance

Build and run evaluators for AI/LLM applications using Phoenix.

  • Tasks that involve LLM observability
  • SKILL.md covers Quick Reference, Workflows, Reference Categories and Key Principles
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve LLM evaluation

What it does

Phoenix Evals is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Build and run evaluators for AI/LLM applications using Phoenix.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 35 other files, including reference files (for example `references/axial-coding.md`, `references/common-mistakes-python.md` and `references/error-analysis-multi-turn.md`). Compatibility notes: Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.

It sits in AI & LLM Engineering, covering LLM observability and LLM evaluation. It works with Python and TypeScript. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM observability
  • Tasks that involve LLM evaluation

Example prompts

  • “/phoenix-evals”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.

    From compatibility in the SKILL.md frontmatter.

Context cost

Phoenix Evals loads about 1.1k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 19 tokens; SKILL.md has 189 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its Apache-2.0 licence (© github). 189 words, ~1,148 tokens.

Download SKILL.mdSave it as .claude/skills/phoenix-evals/SKILL.md (or your agent's skills folder). This skill also uses 34 other files; get the full folder from GitHub.
name
phoenix-evals
description
Build and run evaluators for AI/LLM applications using Phoenix.
compatibility
Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.
license
Apache-2.0
metadata.author
oss@arize.com
metadata.version
1.0.0
metadata.languages
Python, TypeScript

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

Workflows

Starting Fresh: observe-tracing-setup → error-analysis → axial-coding → evaluators-overview

Building Evaluator: fundamentals → common-mistakes-python → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

RAG Systems: evaluators-rag → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)

Production: production-overview → production-guardrails → production-continuous

Reference Categories

PrefixDescription
fundamentals-*Types, scores, anti-patterns
observe-*Tracing, sampling
error-analysis-*Finding failures
axial-coding-*Categorizing failures
evaluators-*Code, LLM, RAG evaluators
experiments-*Datasets, running experiments
validation-*Validating evaluator accuracy against human labels
production-*CI/CD, monitoring

Key Principles

PrincipleAction
Error analysis firstCan't automate what you haven't observed
Custom > genericBuild from your failures
Code firstDeterministic before LLM
Validate judges>80% TPR/TNR
Binary > LikertPass/fail, not 1-5

© github, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 34 other files (references) in skills/phoenix-evals of github/awesome-copilot.

  • SKILL.md
  • references/axial-coding.md
  • references/common-mistakes-python.md
  • references/error-analysis-multi-turn.md
  • references/error-analysis.md
  • references/evaluate-dataframe-python.md
  • references/evaluators-code-python.md
  • references/evaluators-code-typescript.md
  • references/evaluators-custom-templates.md
  • references/evaluators-llm-python.md
  • references/evaluators-llm-typescript.md
  • references/evaluators-overview.md
  • references/evaluators-pre-built.md
  • references/evaluators-rag.md
  • references/experiments-datasets-python.md
  • references/experiments-datasets-typescript.md
  • references/experiments-overview.md
  • references/experiments-running-python.md
  • references/experiments-running-typescript.md
  • references/experiments-synthetic-python.md
  • … and 15 more

Open the folder on GitHubat commit 727ff2e

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in github/awesome-copilot, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Phoenix Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phoenix Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phoenix Evals this skillgithub/awesome-copilot40k3 repos~1.1kAutomated safety check: PassApache-2.0
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence
Phoenix Skills AuditArize-ai/phoenix12k—~5.1kAutomated safety check: PassCustom licence
Langchain Dependencieslangchain-ai/langchain-skills1.3k1 repos~3.6kAutomated safety check: PassMIT
Opik External Integrationscomet-ml/opik22k—~1.3kAutomated safety check: PassApache-2.0
Opik SDK Integrationscomet-ml/opik22k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Phoenix Skills Audit

    Arize-ai/phoenix

    Audit recent changes to Phoenix's user-facing surfaces (Python clients, TypeScript clients, CLI, REST/GraphQL APIs) and patch the three external-facing agent skills — phoenix-tracing, phoenix-cli…

    12k GitHub stars~5.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Langchain Dependencies

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.

    1.3k GitHub starsUsed in 1 repo~3.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Builds Opik tracing integrations that live outside the Opik repository, either as standalone opik-* packages or as contributions into projects such as LiteLLM or Dify.

    22k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Builds, updates and verifies integrations inside the Opik Python and TypeScript SDKs so users can trace a framework or provider with one call.

    22k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Generates onboarding code snippets for Phoenix tracing integrations and wires them into the project onboarding UI.

    12k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Phoenix Evals

What does Phoenix Evals do?

Build and run evaluators for AI/LLM applications using Phoenix. Phoenix Evals is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Build and run evaluators for AI/LLM applications using Phoenix.

When should I use Phoenix Evals?

Phoenix Evals fits situations like: tasks that involve LLM observability; tasks that involve LLM evaluation.

How do I install Phoenix Evals in Claude Code?

Run `npx skills add github/awesome-copilot --skill phoenix-evals -a claude-code`. Or copy the skill folder (skills/phoenix-evals in github/awesome-copilot) into .claude/skills/phoenix-evals in your project. Claude Code loads it when a task matches its description.

How do I install Phoenix Evals in Codex?

Run `npx skills add github/awesome-copilot --skill phoenix-evals -a codex`. Or copy the skill folder (skills/phoenix-evals in github/awesome-copilot) into .agents/skills/phoenix-evals in your project. Codex loads it when a task matches its description.

Can I use Phoenix Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill phoenix-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phoenix-evals, .gemini/skills/phoenix-evals, .github/skills/phoenix-evals and .opencode/skills/phoenix-evals in your project.

What does Phoenix Evals need to run?

SKILL.md names no scripts, command-line tools or credentials: Phoenix Evals is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client..

Does Phoenix Evals access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Phoenix Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phoenix Evals use?

Phoenix Evals is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phoenix Evals use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.

What are the alternatives to Phoenix Evals?

Skills that share tags, products or a category with Phoenix Evals: Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), Phoenix Skills Audit (Arize-ai/phoenix, 12k stars), Langchain Dependencies (langchain-ai/langchain-skills, 1.3k stars) and Opik External Integrations (comet-ml/opik, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phoenix Evals?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.