Failproof AI SDK Integration
FailproofAI/failproofai
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
Create a new built-in classification evaluator for Phoenix evals.
$ npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Arize-ai/phoenix phoenix-evals-new-metric --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/phoenix-evals-new-metric .claude/skills/phoenix-evals-new-metric && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "phoenix-evals-new-metric" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-evals-new-metric into .claude/skills/phoenix-evals-new-metric/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-evals-new-metric", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-evals-new-metricType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Arize-ai/phoenix phoenix-evals-new-metric --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/phoenix-evals-new-metric .agents/skills/phoenix-evals-new-metric && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "phoenix-evals-new-metric" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-evals-new-metric into .agents/skills/phoenix-evals-new-metric/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-evals-new-metric", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Arize-ai/phoenix phoenix-evals-new-metric --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/phoenix-evals-new-metric .cursor/skills/phoenix-evals-new-metric && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "phoenix-evals-new-metric" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-evals-new-metric into .cursor/skills/phoenix-evals-new-metric/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-evals-new-metric", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Arize-ai/phoenix.git --path .agents/skills/phoenix-evals-new-metric--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Arize-ai/phoenix phoenix-evals-new-metric --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/phoenix-evals-new-metric .gemini/skills/phoenix-evals-new-metric && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "phoenix-evals-new-metric" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-evals-new-metric into .gemini/skills/phoenix-evals-new-metric/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-evals-new-metric", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Arize-ai/phoenix phoenix-evals-new-metricInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/phoenix-evals-new-metric .github/skills/phoenix-evals-new-metric && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "phoenix-evals-new-metric" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-evals-new-metric into .github/skills/phoenix-evals-new-metric/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-evals-new-metric", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Arize-ai/phoenix phoenix-evals-new-metric --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/phoenix-evals-new-metric .opencode/skills/phoenix-evals-new-metric && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "phoenix-evals-new-metric" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-evals-new-metric into .opencode/skills/phoenix-evals-new-metric/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-evals-new-metric", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
phoenix-evals-new-metricCreate a new built-in classification evaluator for Phoenix evals.
Phoenix Evals New Metric is an agent skill from Arize-ai/phoenix. Create a new built-in classification evaluator for Phoenix evals. Use this skill whenever the user asks to create a new eval, build a new metric, add a new builtin evaluator, create an LLM-as-a-judge metric, or add a new classification evaluator to Phoenix.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM observability and LLM evaluation. It works with Python. The repository describes itself as: AI Observability & Evaluation. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e471315. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pnpmmakeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Phoenix Evals New Metric loads about 2.3k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 1,047 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Arize-ai/phoenix at commit e471315, republished under its Apache-2.0 licence (© Arize-ai). 1,047 words, ~2,333 tokens.
.claude/skills/phoenix-evals-new-metric/SKILL.md (or your agent's skills folder).A built-in evaluator is a YAML config (source of truth) that gets compiled into Python and TypeScript code, wrapped in evaluator classes, benchmarked, and documented. The whole pipeline is linear — follow these steps in order.
Before writing anything, clarify with the user:
{{input}}, {{output}}, {{reference}}, {{tool_definitions}}). If the user is vague, ask follow-up questions — the placeholders are the contract between the evaluator and the caller.promoted_dataset_evaluator label. Currently only correctness, tool_selection, and tool_invocation have this — some may new evaluators don't need it.Create prompts/classification_evaluator_configs/{NAME}_CLASSIFICATION_EVALUATOR_CONFIG.yaml.
Read an existing config to match the current schema. Start with CORRECTNESS_CLASSIFICATION_EVALUATOR_CONFIG.yaml for a simple example, or TOOL_SELECTION_CLASSIFICATION_EVALUATOR_CONFIG.yaml if your evaluator needs structured span data.
choices — Maps label strings to numeric scores. For binary evaluators, use positive/negative labels (e.g., correct: 1.0 / incorrect: 0.0). The labels you pick here flow through to the Python class, TS factory, and benchmarks.
optimization_direction — Use maximize when the positive label is the desired outcome (most evaluators). Use minimize only if the metric measures something undesirable (e.g., hallucination). This affects how Phoenix displays the metric in the UI.
labels — Optional list. Add promoted_dataset_evaluator only if this evaluator should appear in the dataset experiments UI sidebar.
substitutions — Only needed if the evaluator is a promoted_dataset_evaluator and works with structured span data (tool definitions, tool calls, message arrays). These reference formatter snippets defined in prompts/formatters/server.yaml. Read that file if you need substitutions — it defines what structured data formats are available. Most evaluators that only use simple text fields (input, output, reference) don't need substitutions.
<context>, <output>) for clear data formatting{{placeholder}} (Mustache syntax) for template variablesmake codegen-promptsThis generates code in three places:
packages/phoenix-evals/src/phoenix/evals/__generated__/classification_evaluator_configs/ (Python)src/phoenix/__generated__/classification_evaluator_configs/ (Python, server copy)js/packages/phoenix-evals/src/__generated__/default_templates/ (TypeScript)Verify the generated files look correct before moving on.
Create packages/phoenix-evals/src/phoenix/evals/metrics/{name}.py.
Read correctness.py in that directory — it's the canonical example. Your evaluator follows the same pattern: subclass ClassificationEvaluator, pull constants from the generated config, define a Pydantic input schema with fields matching your template placeholders.
After creating the file, add it to the exports in metrics/__init__.py — both the import and the __all__ list. Read the current __init__.py to see the existing pattern.
Create js/packages/phoenix-evals/src/llm/create{Name}Evaluator.ts.
Read createCorrectnessEvaluator.ts — it's the canonical example. The pattern is a factory function that wraps createClassificationEvaluator with defaults from the generated config.
Then:
js/packages/phoenix-evals/src/llm/index.tsjs/packages/phoenix-evals/test/llm/ — read
createFaithfulnessEvaluator.test.ts there for the test patterncd js && pnpm buildFix any TypeScript errors before proceeding.
Create js/benchmarks/evals-benchmarks/src/{name}.eval.ts.
Read existing benchmarks in that directory to match the current patterns:
tool_invocation.eval.ts — multi-category analysis and aggregate metricsaggregateMetrics.ts — shared macro precision/recall/F1 accumulationThe task function must return input and output text in its result so the failed examples printer has access to them.
Consider using a separate agent session for synthetic dataset generation if the examples need realistic domain-specific content — this keeps the dataset creation focused and avoids context-switching.
# Terminal 1: Start Phoenix using the normal configured database. Do not set
# PHOENIX_WORKING_DIR unless the user explicitly requests an isolated instance.
phoenix serve
# Terminal 2: Run the benchmark
cd js
pnpm --filter evals-benchmarks... build
pnpm --filter evals-benchmarks exec vitest run \
src/{name}.eval.ts --config phoenix.vitest.config.tsTarget >80% accuracy. If accuracy is low, look at the failed examples output to decide whether to adjust the prompt (Step 1) or the benchmark examples (Step 6). Iterate until accuracy is acceptable.
Create docs/phoenix/evaluation/pre-built-metrics/{name}.mdx.
Read faithfulness.mdx in that directory — it's the template. Follow the same section structure:
After creating the docs page, update these three files:
docs.json — add the page to the Evaluation > Pre-built Metrics nav groupdocs/phoenix/evaluation/pre-built-metrics.mdx — add a card to the landing page griddocs/phoenix/sitemap.xml — add the new URLRead each file to see the existing pattern before editing.
Before calling it done, verify:
make codegen-prompts ran successfullymetrics/__init__.pyllm/index.tscd js && pnpm build)docs.json nav updatedAfter completing the workflow, verify these instructions matched reality:
make codegen-prompts generate to different locations?If anything drifted, update this SKILL.md before finishing so the next person (or agent) doesn't hit the same surprises.
© Arize-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/phoenix-evals-new-metric of Arize-ai/phoenix.
Open the folder on GitHubat commit e471315
Phoenix Evals New Metric next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Phoenix Evals New Metric this skillArize-ai/phoenix | 12k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Failproof AI SDK IntegrationFailproofAI/failproofai | 5.3k | — | ~6k | Automated safety check: Pass | Custom licence | |
| Phoenix Evalsgithub/awesome-copilot | 40k | 2 repos | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| LLM Trace Review Interfaceai-evals-course/evals-skills | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Analyzing Claude Code Sessionsamd/gaia | 1.6k | — | ~2.3k | Automated safety check: Pass | MIT |
FailproofAI/failproofai
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
github/awesome-copilot
Build and run evaluators for AI/LLM applications using Phoenix.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
ai-evals-course/evals-skills
Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.
amd/gaia
Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.
huggingface/skills
Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space.
Arize-ai/phoenix
A skill your agent uses when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only…
Arize-ai/phoenix
Build and maintain documentation sites with Mintlify. An agent skill from Arize-ai/phoenix.
Arize-ai/phoenix
Frontend development guidelines for the Phoenix AI observability platform.
Arize-ai/phoenix
Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.
Arize-ai/phoenix
Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).
Arize-ai/phoenix
Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).
Works with
Categories
Create a new built-in classification evaluator for Phoenix evals. Phoenix Evals New Metric is an agent skill from Arize-ai/phoenix. Create a new built-in classification evaluator for Phoenix evals.
Phoenix Evals New Metric fits situations like: the user asks to create a new eval; build a new metric; add a new builtin evaluator; create an LLM-as-a-judge metric.
Run `npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a claude-code`. Or copy the skill folder (.agents/skills/phoenix-evals-new-metric in Arize-ai/phoenix) into .claude/skills/phoenix-evals-new-metric in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a codex`. Or copy the skill folder (.agents/skills/phoenix-evals-new-metric in Arize-ai/phoenix) into .agents/skills/phoenix-evals-new-metric in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arize-ai/phoenix --skill phoenix-evals-new-metric -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phoenix-evals-new-metric, .gemini/skills/phoenix-evals-new-metric, .github/skills/phoenix-evals-new-metric and .opencode/skills/phoenix-evals-new-metric in your project.
Going by SKILL.md and its folder, Phoenix Evals New Metric needs the command-line tools its instructions call (pnpm and make). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Phoenix Evals New Metric is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Phoenix Evals New Metric: Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), Phoenix Evals (github/awesome-copilot, 40k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and LLM Trace Review Interface (ai-evals-course/evals-skills, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Arize-ai (a GitHub organization) maintains it in Arize-ai/phoenix, which has 11,770 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 10, 2026.
Source: Arize-ai/phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.