Agent skill

Arize Phoenix

by Arize-ai in Arize-ai/phoenix

Open-source AI observability platform for tracing, evaluating, and improving LLM applications with OpenTelemetry integration

MITAuto-check passedDevOps & Cloud

Install Arize Phoenix

skills CLI
$ npx skills add Arize-ai/phoenix --skill arize-phoenix -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Arize-ai/phoenix arize-phoenix --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/phoenix .claude/skills/arize-phoenix && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arize-phoenix
GitHub stars
12k
Used in
1 other repo
Token cost
~3.8k tokens
SKILL.md length
1,852 words
Files
948
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Open-source AI observability platform for tracing, evaluating, and improving LLM applications with OpenTelemetry integration

  • Works in 6 steps: Choose integration - Select appropriate… → Install package - Install Phoenix client… → Configure endpoint - Set Phoenix… → …
  • Tasks that involve Observability
  • SKILL.md covers When to Use This Skill, Capabilities, Skills and Workflows, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Arize Phoenix is an agent skill from Arize-ai/phoenix. Open-source AI observability platform for tracing, evaluating, and improving LLM applications with OpenTelemetry integration

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 951 other files.

It sits in DevOps & Cloud, covering Observability and LLM observability. It works with Arize Phoenix and OpenTelemetry. The repository describes itself as: AI Observability & Evaluation. The licence is MIT.

When your agent uses it

  • Tasks that involve Observability
  • Tasks that involve LLM observability

Example prompts

  • “/arize-phoenix”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Choose integration - Select appropriate Phoenix integration for your framework (LangChain, LlamaIndex, OpenAI, etc.)
  2. Install package - Install Phoenix client and OpenTelemetry packages for your language (Python or TypeScript)
  3. Configure endpoint - Set Phoenix endpoint URL and optionally configure project name and session tracking
  4. Instrument application - Add auto-instrumentation or manual instrumentation to capture LLM calls, tool executions, and retrievals
  5. View traces - Open Phoenix UI to see execution flow, latency, token usage, and detailed span information
  6. Add annotations - Add scores, labels, or human feedback to traces for quality measurement

What it can do on your machine

Read from SKILL.md and the folder at commit 856100b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arize.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Arize Phoenix loads about 3.8k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 1,852 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Arize-ai/phoenix at commit 856100b, republished under its MIT licence (© Arize-ai). 1,852 words, ~3,774 tokens.

Download SKILL.mdSave it as .claude/skills/arize-phoenix/SKILL.md (or your agent's skills folder). This skill also uses 947 other files; get the full folder from GitHub.
name
arize-phoenix
description
Open-source AI observability platform for tracing, evaluating, and improving LLM applications with OpenTelemetry integration
license
MIT
metadata.author
Arize AI
metadata.category
ai-observability

Arize Phoenix

Phoenix is an open-source AI observability platform built on OpenTelemetry that helps developers understand, debug, and improve AI applications. It provides comprehensive tracing, evaluation, prompt engineering, and experimentation capabilities for LLM-based systems. Phoenix captures detailed execution information from AI applications, measures output quality with evaluators, enables systematic prompt iteration, and supports data-driven experimentation to optimize AI performance.

When to Use This Skill

  • Debugging AI application failures by inspecting LLM calls, tool executions, and retrieval operations
  • Measuring and improving AI output quality using LLM-based or code-based evaluators
  • Iterating on prompts using real production examples and testing variations systematically
  • Comparing different versions of AI applications (prompts, models, architectures) using experiments
  • Monitoring LLM costs, token usage, latency, and error rates in production
  • Building datasets from production traces for evaluation and fine-tuning
  • Tracking multi-turn conversations and maintaining context across interactions
  • Optimizing RAG systems by analyzing retrieval quality and document relevance
  • Evaluating agent performance including tool call accuracy and actionability
  • Managing prompt versions and deploying them across different environments

Capabilities

Agents can leverage Phoenix to:

  • Trace AI application execution with detailed visibility into LLM calls, tool executions, retrieval operations, embeddings, and prompt templates
  • Evaluate output quality using pre-built or custom evaluators with LLM-as-a-judge or code-based evaluation logic
  • Annotate traces with human feedback, scores, labels, and quality signals for continuous improvement
  • Experiment systematically by comparing different versions of applications using datasets and evaluators
  • Monitor performance metrics including latency, token usage, costs, and error rates across projects
  • Iterate on prompts using the playground, span replay, and dataset-based testing
  • Organize traces into projects and sessions for better management and analysis
  • Integrate with 20+ AI frameworks and LLM providers via OpenTelemetry instrumentation

Skills

Tracing
  • Capture traces via OpenTelemetry (OTLP) protocol with automatic instrumentation for major frameworks
  • View execution flow showing every LLM call, tool execution, retrieval operation, embedding generation, and response generation
  • Inspect LLM parameters including temperature, system prompts, function calls, and invocation parameters
  • Analyze retrieval operations with document scores, order, and embedding text for RAG systems
  • Track token usage with detailed breakdowns by token type (input/output) and model
  • Monitor latency at trace, span, and component levels with quantile analysis
  • Organize with projects to separate traces by environment, application, team, or use case
  • Group with sessions to track multi-turn conversations and maintain context across interactions
  • Add metadata to traces with custom attributes, tags, and structured data for filtering and analysis
  • Annotate traces with scores, labels, human feedback, and LLM evaluations for quality measurement
  • Export and import traces for backup, migration, or analysis in external tools
  • Track costs with automatic calculation based on token usage and model pricing
Evaluation
  • Run LLM-as-a-judge evaluations using any LLM provider (OpenAI, Anthropic, Gemini, custom endpoints) to assess output quality
  • Build custom evaluators with Python or TypeScript using custom prompts, scoring logic, and evaluation criteria
  • Use pre-built evaluators for common tasks including faithfulness, relevance, toxicity, summarization, agent evaluation, and RAG quality
  • Write code-based evaluators for deterministic checks like exact match, regex patterns, or custom Python/TypeScript logic
  • Execute evaluations at scale with automatic concurrency, rate limit handling, error management, and batching via executors
  • Map complex inputs using input schemas and mappings to transform nested data structures for evaluators
  • View evaluator traces with complete transparency into prompts, model reasoning, scores, and execution metadata
  • Run batch evaluations on traces, datasets, or custom data sources with automatic retry and error handling
  • Integrate evaluations into workflows by running evals on production traces or test datasets
Datasets & Experiments
  • Create datasets from traces, code, CSV files, or manually curated examples with inputs and optional reference outputs
  • Build golden datasets with reference outputs (ground truth) for objective evaluation using code-based evaluators
  • Version datasets with automatic tracking of inserts, updates, and deletes for reproducibility
  • Run experiments by executing task functions against datasets with evaluators to compare different versions
  • Compare experiments side-by-side in the UI to see performance differences, score distributions, and individual example results
  • Use repetitions to run experiments multiple times for statistical confidence and account for LLM variability
  • Organize with splits to separate datasets into train/test/validation splits for proper evaluation workflows
  • Export datasets in JSONL or CSV formats for fine-tuning, analysis, or sharing
  • View experiment results in the Phoenix UI with task function traces, scores per example, and aggregate performance metrics
Prompt Engineering
  • Manage prompts with versioning, storage, and deployment across different environments
  • Test prompts interactively in the Prompt Playground with various models, parameters, and tools
  • Replay LLM spans from production traces in the playground to debug failures and test improvements
  • Test at scale by running prompts against datasets to evaluate performance systematically
  • Compare prompt versions side-by-side to see which performs better on your data
  • Optimize automatically using automated prompt optimization features
  • Sync prompts via SDK to keep prompts in sync across applications and environments programmatically
  • Tag prompts for deployment control across development, staging, and production environments
  • Track prompt changes with version history, author information, and timestamps
Projects & Organization
  • Create projects to organize traces by environment (development, staging, production), application, or team
  • Set up sessions to track multi-turn conversations with chatbot-like UI showing conversation history
  • View metrics dashboards with pre-defined metrics including latency, errors, token usage, costs, and model performance
  • Filter and search traces by metadata, attributes, annotations, or custom tags
  • Configure data retention policies to control how long trace and evaluation data is stored
API & Programmatic Access
  • Use Python SDK (arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) for programmatic access
  • Use TypeScript SDK (arizeai-phoenix-client, arizeai-phoenix-evals, arizeai-phoenix-otel) for JavaScript/TypeScript applications
  • Access REST API for annotations, datasets, experiments, traces, spans, prompts, projects, and users
  • Instrument manually using OpenTelemetry decorators, wrappers, or direct OpenInference SDKs
  • Generate API keys for programmatic access with role-based permissions
Authentication & Security
  • Configure RBAC with role-based access control for user permissions and project access
  • Set up authentication including SSO and user management for self-hosted instances
  • Manage API keys for secure programmatic access to Phoenix APIs and SDKs
  • Control data privacy with self-hosting options for VPC deployment or local execution

Workflows

Workflow 1: Instrument and Trace an AI Application
  1. Choose integration - Select appropriate Phoenix integration for your framework (LangChain, LlamaIndex, OpenAI, etc.)
  2. Install package - Install Phoenix client and OpenTelemetry packages for your language (Python or TypeScript)
  3. Configure endpoint - Set Phoenix endpoint URL and optionally configure project name and session tracking
  4. Instrument application - Add auto-instrumentation or manual instrumentation to capture LLM calls, tool executions, and retrievals
  5. View traces - Open Phoenix UI to see execution flow, latency, token usage, and detailed span information
  6. Add annotations - Add scores, labels, or human feedback to traces for quality measurement
Show full SKILL.md (776 more words)Show less
Workflow 2: Evaluate AI Output Quality
  1. Choose evaluator type - Select LLM-as-a-judge for subjective quality or code-based for objective checks
  2. Configure LLM provider - Set up evaluator LLM (OpenAI, Anthropic, Gemini, or custom endpoint)
  3. Define evaluation logic - Use pre-built evaluator or create custom evaluator with prompts/scoring logic
  4. Run evaluation - Execute evaluator on traces, datasets, or custom data with automatic batching and concurrency
  5. Review results - View evaluator traces, scores, explanations, and labels in Phoenix UI
  6. Iterate - Adjust evaluator prompts or logic based on results and human feedback
Workflow 3: Run Experiments to Compare Versions
  1. Create dataset - Build dataset with inputs and optional reference outputs from traces, code, or CSV
  2. Define task function - Create Python function that wraps your AI application logic and returns outputs
  3. Select evaluators - Choose code-based evaluators for ground truth comparison or LLM-as-a-judge for subjective quality
  4. Run experiment - Execute task function against dataset with evaluators to generate scores
  5. Compare results - View experiment results in UI with aggregate metrics, score distributions, and per-example analysis
  6. Iterate - Make changes to prompts, models, or architecture and run new experiment to compare performance
Workflow 4: Optimize Prompts with Playground
  1. Identify prompt - Find prompt in traces or load existing prompt from prompt management
  2. Open playground - Load prompt into Prompt Playground with current parameters and tools
  3. Test variations - Modify prompt text, model parameters, tools, or response format and test with real inputs
  4. View traces - All playground runs are automatically recorded as traces for analysis
  5. Test at scale - Run prompt variations against dataset examples to evaluate performance systematically
  6. Save and deploy - Save best-performing prompt version, tag for environment, and deploy via SDK
Workflow 5: Debug Production Issues
  1. Identify problematic trace - Search or filter traces to find failed or low-quality executions
  2. Inspect execution flow - View detailed span information including LLM calls, tool executions, and retrievals
  3. Replay span - Load problematic LLM span into Prompt Playground to test fixes
  4. Test improvements - Modify prompts, parameters, or tools in playground and compare outputs
  5. Add to dataset - Add problematic examples to dataset for future testing
  6. Run experiment - Test improved version against dataset to verify fix before deployment

Integrations

LLM Providers

OpenAI, Anthropic, Amazon Bedrock, Google (Gemini), Groq, MistralAI, VertexAI, LiteLLM, OpenRouter, Together, Vercel AI

Python Frameworks

AG2, Agno, AutoGen, BeeAI, CrewAI, DSPy, Google ADK, Graphite, Guardrails AI, Haystack, Hugging Face smolagents, Instructor, LlamaIndex, LangChain, LangGraph, MCP, NVIDIA, Portkey, Pydantic AI

TypeScript Frameworks

BeeAI, LangChain.js, Mastra, MCP, Vercel AI SDK

Java Frameworks

LangChain4j, Spring AI, Arconia

Platforms

Dify, Flowise, LangFlow, Prompt Flow

Vector Databases

MongoDB, OpenSearch, Pinecone, Qdrant, Weaviate, Zilliz/Milvus, Couchbase

Evaluation Integrations

Cleanlab, Ragas, UQLM

Observability Protocols

OpenTelemetry (OTLP), OpenInference

Developer Tools

Claude Code, Cursor, Phoenix MCP Server

Cloud Platforms

AWS (CloudFormation), Kubernetes (Helm), Docker, Railway

Context

OpenTelemetry: Phoenix tracing is built on OpenTelemetry (OTLP), an industry-standard observability protocol. This means instrumentation code written for Phoenix can be reused with other observability platforms, avoiding vendor lock-in.

OpenInference: Phoenix uses OpenInference instrumentation, an extension of OpenTelemetry specifically designed for AI/LLM applications. OpenInference adds semantic conventions for LLM spans, retrieval operations, and embeddings.

Traces and Spans: A trace represents the complete execution path of a request through an AI application. Spans are individual units of work within a trace (e.g., a single LLM call, tool execution, or retrieval operation). Spans can be nested to show hierarchical execution flow.

Projects: Projects provide organizational structure for traces, allowing separation by environment, application, or team. Each project has its own metrics dashboard and data isolation.

Sessions: Sessions group related traces into conversational threads, enabling tracking of multi-turn conversations with context maintained across interactions.

Evaluators: Evaluators measure the quality of AI outputs. LLM-based evaluators use LLMs as judges to assess subjective quality. Code-based evaluators use deterministic logic for objective checks. All evaluators return scores with optional labels, explanations, and metadata.

Datasets: Datasets are collections of examples with inputs and optional reference outputs. Golden datasets contain reference outputs (ground truth) for objective evaluation. Datasets are versioned automatically.

Experiments: Experiments run task functions (wrapped AI application logic) against datasets with evaluators to systematically compare different versions. Experiments track scores per example and aggregate metrics.

Prompts: In Phoenix, a prompt includes the prompt template, invocation parameters (temperature, etc.), tools, and response format. Prompts are versioned and can be tagged for deployment across environments.

Executors: Executors handle evaluation execution with automatic concurrency, rate limit management, error handling, and batching. They can achieve up to 20x speedup compared to direct API calls.

Self-Hosting: Phoenix can be self-hosted on Docker, Kubernetes, AWS, Railway, or locally. Self-hosted instances support authentication, email configuration, and data retention policies.

For additional documentation: https://arize.com/docs/phoenix/llms.txt

© Arize-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 947 other files in docs/phoenix of Arize-ai/phoenix.

  • SKILL.md
  • agent-assisted-setup.mdx
  • cookbook.mdx
  • cookbook/agent-workflow-patterns.mdx
  • cookbook/agent-workflow-patterns/autogen.mdx
  • cookbook/agent-workflow-patterns/crewai.mdx
  • cookbook/agent-workflow-patterns/google-genai-sdk-manual-orchestration.mdx
  • cookbook/agent-workflow-patterns/langgraph.mdx
  • cookbook/agent-workflow-patterns/openai-agents.mdx
  • cookbook/agent-workflow-patterns/smolagents.mdx
  • cookbook/ai-engineering-workflows/aligning-evals-with-human-feedback.mdx
  • cookbook/ai-engineering-workflows/analyzing-customer-review-evals-with-repetition-experiments.mdx
  • cookbook/ai-engineering-workflows/iterative-evaluation-and-experimentation-workflow-python.mdx
  • cookbook/ai-engineering-workflows/iterative-evaluation-and-experimentation-workflow-typescript.mdx
  • cookbook/datasets-and-experiments/analyzing-customer-review-evals-with-repetition-experiments.mdx
  • cookbook/datasets-and-experiments/building-your-own-eval-harness.mdx
  • cookbook/datasets-and-experiments/comparing-llamaindex-query-engines-with-a-pairwise-evaluator.mdx
  • … and 931 more

Open the folder on GitHubat commit 856100b

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Arize-ai/phoenix, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Arize Phoenix next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Arize Phoenix compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Arize Phoenix this skillArize-ai/phoenix12k1 repos~3.8kAutomated safety check: PassMIT
Agent Platform Alert Configurationgoogle/skills21k—~4.2kAutomated safety check: PassApache-2.0
Arize Instrumentationgithub/awesome-copilot40k—~6.2kAutomated safety check: NotesMIT
Ag2 Telemetryag2ai/build-with-ag2252—~1.9kAutomated safety check: PassApache-2.0
Sentry Elixir SDKgetsentry/sentry-for-ai268—~3.5kAutomated safety check: PassApache-2.0
Phoenix LLM ObservabilityOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT

Similar skills

  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Arize Instrumentation

    github/awesome-copilot

    Official

    Adds Arize AX tracing to an LLM application for the first time.

    40k GitHub stars~6.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Ag2 Telemetry

    ag2ai/build-with-ag2

    Add OpenTelemetry traces to an AG2 beta Agent via TelemetryMiddleware (autogen.beta.middleware.builtin).

    252 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Sentry Elixir SDK

    getsentry/sentry-for-ai

    Official

    Full Sentry SDK setup for Elixir. An agent skill from getsentry/sentry-for-ai.

    268 GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Phoenix LLM Observability

    Orchestra-Research/AI-Research-SKILLs

    Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Agentsop Observability Setup

    agentsope/SkillAlchemy

    Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.

    459 GitHub stars~4.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from Arize-ai/phoenix

All 39 skills in this repo
  • Harbor Exec

    Arize-ai/phoenix

    A skill your agent uses when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only…

    12k GitHub stars~909 tokensUpdated today
    Auto-check passed
  • Mintlify

    Arize-ai/phoenix

    Build and maintain documentation sites with Mintlify. An agent skill from Arize-ai/phoenix.

    12k GitHub starsUsed in 8 repos~3.4k tokens
    Auto-check passed
  • Phoenix Frontend

    Arize-ai/phoenix

    Frontend development guidelines for the Phoenix AI observability platform.

    12k GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Phoenix Graphql

    Arize-ai/phoenix

    Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.

    12k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Phoenix Server

    Arize-ai/phoenix

    Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).

    12k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Phoenix Storybook

    Arize-ai/phoenix

    Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).

    12k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Arize Phoenix

What does Arize Phoenix do?

Open-source AI observability platform for tracing, evaluating, and improving LLM applications with OpenTelemetry integration. Arize Phoenix is an agent skill from Arize-ai/phoenix.

When should I use Arize Phoenix?

Arize Phoenix fits situations like: tasks that involve Observability; tasks that involve LLM observability.

How do I install Arize Phoenix in Claude Code?

Run `npx skills add Arize-ai/phoenix --skill arize-phoenix -a claude-code`. Or copy the skill folder (docs/phoenix in Arize-ai/phoenix) into .claude/skills/arize-phoenix in your project. Claude Code loads it when a task matches its description.

How do I install Arize Phoenix in Codex?

Run `npx skills add Arize-ai/phoenix --skill arize-phoenix -a codex`. Or copy the skill folder (docs/phoenix in Arize-ai/phoenix) into .agents/skills/arize-phoenix in your project. Codex loads it when a task matches its description.

Can I use Arize Phoenix in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arize-ai/phoenix --skill arize-phoenix -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arize-phoenix, .gemini/skills/arize-phoenix, .github/skills/arize-phoenix and .opencode/skills/arize-phoenix in your project.

What does Arize Phoenix need to run?

SKILL.md names no scripts, command-line tools or credentials: Arize Phoenix is instructions for the agent only. Our summary lists: Python 3.

Does Arize Phoenix access the network?

SKILL.md names 1 domain. As links in the text: arize.com. This is read from the text; nothing was executed.

Is Arize Phoenix safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Arize Phoenix use?

Arize Phoenix is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Arize Phoenix use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Arize Phoenix?

Skills that share tags, products or a category with Arize Phoenix: Agent Platform Alert Configuration (google/skills, 21k stars), Arize Instrumentation (github/awesome-copilot, 40k stars), Ag2 Telemetry (ag2ai/build-with-ag2, 252 stars) and Sentry Elixir SDK (getsentry/sentry-for-ai, 268 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Arize Phoenix?

Arize-ai (a GitHub organization) maintains it in Arize-ai/phoenix, which has 11,744 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 8, 2026.

Source: Arize-ai/phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.