Agent skill

Phoenix LLM Observability

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

MITAuto-check passedAI & LLM Engineering

Install Phoenix LLM Observability

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill phoenix-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs phoenix-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/17-observability/phoenix .claude/skills/phoenix-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phoenix-observability
GitHub stars
13k
Used in
2 other repos
Token cost
~2.9k tokens
SKILL.md length
338 words
Files
3 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

  • Works in 6 steps: Use projects: Separate traces by… → Add metadata: Include user IDs, session… → Evaluate regularly: Run automated… → …
  • Debugging an LLM application by inspecting detailed traces of each call
  • SKILL.md covers When to use Phoenix, Quick start, Core concepts and Framework instrumentation, plus 4 more sections
  • Calls pip, docker and psql; needs PHOENIX_SECRET and PHOENIX_ADMIN_SECRET

What it does

Phoenix is an open-source tool for seeing what an LLM application actually does at runtime. The skill shows how to install it with pip, start the server from a notebook or with the phoenix serve command, and send traces to it through OpenTelemetry-based instrumentors for OpenAI, LangChain and LlamaIndex. It also explains traces, spans and projects, which group related traces by name.

On the evaluation side it covers LLM-as-judge evaluators, versioned datasets that act as regression test sets, experiments that compare prompts, models or settings, and a playground for trying prompts against several models. Phoenix can be self-hosted on PostgreSQL or SQLite. The skill names LangSmith, Weights & Biases, Arize Cloud and MLflow as better fits for other needs, and keeps extra detail in advanced-usage and troubleshooting reference files.

When your agent uses it

  • Debugging an LLM application by inspecting detailed traces of each call
  • Running repeatable evaluations of model output against a stored dataset
  • Watching a production LLM system for problems as they happen
  • Comparing prompts or models in an experiment before changing production

Example prompts

  • “Add Phoenix tracing to our OpenAI-based support bot and show me how to view the traces.”
  • “Set up a Phoenix dataset and an LLM-as-judge evaluator for our summarization prompts.”
  • “Start a self-hosted Phoenix server backed by PostgreSQL for the staging environment.”
  • “Instrument our LangChain retrieval chain so each step shows up as a span in Phoenix.”

Requirements

  • Python with the `arize-phoenix` package
  • PostgreSQL or SQLite for a self-hosted server

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Use projects: Separate traces by environment (dev/staging/prod)
  2. Add metadata: Include user IDs, session IDs for debugging
  3. Evaluate regularly: Run automated evaluations in CI/CD
  4. Version datasets: Track test set changes over time
  5. Monitor costs: Track token usage via Phoenix dashboards
  6. Self-host: Use PostgreSQL for production deployments

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • docker
    • psql

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.arize.com
    • github.com
    • hub.docker.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • PHOENIX_SECRET
    • PHOENIX_ADMIN_SECRET

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phoenix LLM Observability loads about 2.9k tokens when it runs, and up to ~9.4k if it reads all its reference files. Until then it costs about 64 tokens; SKILL.md has 338 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 338 words, ~2,867 tokens.

Download SKILL.mdSave it as .claude/skills/phoenix-observability/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
phoenix-observability
description
Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Observability, Phoenix, Arize, Tracing, Evaluation, Monitoring, LLM Ops, OpenTelemetry
dependencies
arize-phoenix>=12.0.0

Phoenix - AI Observability Platform

Open-source AI observability and evaluation platform for LLM applications with tracing, evaluation, datasets, experiments, and real-time monitoring.

When to use Phoenix

Use Phoenix when:

  • Debugging LLM application issues with detailed traces
  • Running systematic evaluations on datasets
  • Monitoring production LLM systems in real-time
  • Building experiment pipelines for prompt/model comparison
  • Self-hosted observability without vendor lock-in

Key features:

  • Tracing: OpenTelemetry-based trace collection for any LLM framework
  • Evaluation: LLM-as-judge evaluators for quality assessment
  • Datasets: Versioned test sets for regression testing
  • Experiments: Compare prompts, models, and configurations
  • Playground: Interactive prompt testing with multiple models
  • Open-source: Self-hosted with PostgreSQL or SQLite

Use alternatives instead:

  • LangSmith: Managed platform with LangChain-first integration
  • Weights & Biases: Deep learning experiment tracking focus
  • Arize Cloud: Managed Phoenix with enterprise features
  • MLflow: General ML lifecycle, model registry focus

Quick start

Installation
bash
pip install arize-phoenix

# With specific backends
pip install arize-phoenix[embeddings]  # Embedding analysis
pip install arize-phoenix-otel         # OpenTelemetry config
pip install arize-phoenix-evals        # Evaluation framework
pip install arize-phoenix-client       # Lightweight REST client
Launch Phoenix server
python
import phoenix as px

# Launch in notebook (ThreadServer mode)
session = px.launch_app()

# View UI
session.view()  # Embedded iframe
print(session.url)  # http://localhost:6006
Command-line server (production)
bash
# Start Phoenix server
phoenix serve

# With PostgreSQL
export PHOENIX_SQL_DATABASE_URL="postgresql://user:pass@host/db"
phoenix serve --port 6006
Basic tracing
python
from phoenix.otel import register
from openinference.instrumentation.openai import OpenAIInstrumentor

# Configure OpenTelemetry with Phoenix
tracer_provider = register(
    project_name="my-llm-app",
    endpoint="http://localhost:6006/v1/traces"
)

# Instrument OpenAI SDK
OpenAIInstrumentor().instrument(tracer_provider=tracer_provider)

# All OpenAI calls are now traced
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}]
)

Core concepts

Traces and spans

A trace represents a complete execution flow, while spans are individual operations within that trace.

python
from phoenix.otel import register
from opentelemetry import trace

# Setup tracing
tracer_provider = register(project_name="my-app")
tracer = trace.get_tracer(__name__)

# Create custom spans
with tracer.start_as_current_span("process_query") as span:
    span.set_attribute("input.value", query)

    # Child spans are automatically nested
    with tracer.start_as_current_span("retrieve_context"):
        context = retriever.search(query)

    with tracer.start_as_current_span("generate_response"):
        response = llm.generate(query, context)

    span.set_attribute("output.value", response)
Projects

Projects organize related traces:

python
import os
os.environ["PHOENIX_PROJECT_NAME"] = "production-chatbot"

# Or per-trace
from phoenix.otel import register
tracer_provider = register(project_name="experiment-v2")

Framework instrumentation

OpenAI
python
from phoenix.otel import register
from openinference.instrumentation.openai import OpenAIInstrumentor

tracer_provider = register()
OpenAIInstrumentor().instrument(tracer_provider=tracer_provider)
LangChain
python
from phoenix.otel import register
from openinference.instrumentation.langchain import LangChainInstrumentor

tracer_provider = register()
LangChainInstrumentor().instrument(tracer_provider=tracer_provider)

# All LangChain operations traced
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="gpt-4o")
response = llm.invoke("Hello!")
LlamaIndex
python
from phoenix.otel import register
from openinference.instrumentation.llama_index import LlamaIndexInstrumentor

tracer_provider = register()
LlamaIndexInstrumentor().instrument(tracer_provider=tracer_provider)
Anthropic
python
from phoenix.otel import register
from openinference.instrumentation.anthropic import AnthropicInstrumentor

tracer_provider = register()
AnthropicInstrumentor().instrument(tracer_provider=tracer_provider)

Evaluation framework

Built-in evaluators
python
from phoenix.evals import (
    OpenAIModel,
    HallucinationEvaluator,
    RelevanceEvaluator,
    ToxicityEvaluator,
    llm_classify
)

# Setup model for evaluation
eval_model = OpenAIModel(model="gpt-4o")

# Evaluate hallucination
hallucination_eval = HallucinationEvaluator(eval_model)
results = hallucination_eval.evaluate(
    input="What is the capital of France?",
    output="The capital of France is Paris.",
    reference="Paris is the capital of France."
)
Custom evaluators
python
from phoenix.evals import llm_classify

# Define custom evaluation
def evaluate_helpfulness(input_text, output_text):
    template = """
    Evaluate if the response is helpful for the given question.

    Question: {input}
    Response: {output}

    Is this response helpful? Answer 'helpful' or 'not_helpful'.
    """

    result = llm_classify(
        model=eval_model,
        template=template,
        input=input_text,
        output=output_text,
        rails=["helpful", "not_helpful"]
    )
    return result
Run evaluations on dataset
python
from phoenix import Client
from phoenix.evals import run_evals

client = Client()

# Get spans to evaluate
spans_df = client.get_spans_dataframe(
    project_name="my-app",
    filter_condition="span_kind == 'LLM'"
)

# Run evaluations
eval_results = run_evals(
    dataframe=spans_df,
    evaluators=[
        HallucinationEvaluator(eval_model),
        RelevanceEvaluator(eval_model)
    ],
    provide_explanation=True
)

# Log results back to Phoenix
client.log_evaluations(eval_results)

Datasets and experiments

Create dataset
python
from phoenix import Client

client = Client()

# Create dataset
dataset = client.create_dataset(
    name="qa-test-set",
    description="QA evaluation dataset"
)

# Add examples
client.add_examples_to_dataset(
    dataset_name="qa-test-set",
    examples=[
        {
            "input": {"question": "What is Python?"},
            "output": {"answer": "A programming language"}
        },
        {
            "input": {"question": "What is ML?"},
            "output": {"answer": "Machine learning"}
        }
    ]
)
Run experiment
python
from phoenix import Client
from phoenix.experiments import run_experiment

client = Client()

def my_model(input_data):
    """Your model function."""
    question = input_data["question"]
    return {"answer": generate_answer(question)}

def accuracy_evaluator(input_data, output, expected):
    """Custom evaluator."""
    return {
        "score": 1.0 if expected["answer"].lower() in output["answer"].lower() else 0.0,
        "label": "correct" if expected["answer"].lower() in output["answer"].lower() else "incorrect"
    }

# Run experiment
results = run_experiment(
    dataset_name="qa-test-set",
    task=my_model,
    evaluators=[accuracy_evaluator],
    experiment_name="baseline-v1"
)

print(f"Average accuracy: {results.aggregate_metrics['accuracy']}")

Client API

Query traces and spans
python
from phoenix import Client

client = Client(endpoint="http://localhost:6006")

# Get spans as DataFrame
spans_df = client.get_spans_dataframe(
    project_name="my-app",
    filter_condition="span_kind == 'LLM'",
    limit=1000
)

# Get specific span
span = client.get_span(span_id="abc123")

# Get trace
trace = client.get_trace(trace_id="xyz789")
Log feedback
python
from phoenix import Client

client = Client()

# Log user feedback
client.log_annotation(
    span_id="abc123",
    name="user_rating",
    annotator_kind="HUMAN",
    score=0.8,
    label="helpful",
    metadata={"comment": "Good response"}
)
Export data
python
# Export to pandas
df = client.get_spans_dataframe(project_name="my-app")

# Export traces
traces = client.list_traces(project_name="my-app")

Production deployment

Docker
bash
docker run -p 6006:6006 arizephoenix/phoenix:latest
With PostgreSQL
bash
# Set database URL
export PHOENIX_SQL_DATABASE_URL="postgresql://user:pass@host:5432/phoenix"

# Start server
phoenix serve --host 0.0.0.0 --port 6006
Environment variables
VariableDescriptionDefault
PHOENIX_PORTHTTP server port6006
PHOENIX_HOSTServer bind address127.0.0.1
PHOENIX_GRPC_PORTgRPC/OTLP port4317
PHOENIX_SQL_DATABASE_URLDatabase connectionSQLite temp
PHOENIX_WORKING_DIRData storage directoryOS temp
PHOENIX_ENABLE_AUTHEnable authenticationfalse
PHOENIX_SECRETJWT signing secretRequired if auth enabled
With authentication
bash
export PHOENIX_ENABLE_AUTH=true
export PHOENIX_SECRET="your-secret-key-min-32-chars"
export PHOENIX_ADMIN_SECRET="admin-bootstrap-token"

phoenix serve

Best practices

  1. Use projects: Separate traces by environment (dev/staging/prod)
  2. Add metadata: Include user IDs, session IDs for debugging
  3. Evaluate regularly: Run automated evaluations in CI/CD
  4. Version datasets: Track test set changes over time
  5. Monitor costs: Track token usage via Phoenix dashboards
  6. Self-host: Use PostgreSQL for production deployments

Common issues

Traces not appearing:

python
from phoenix.otel import register

# Verify endpoint
tracer_provider = register(
    project_name="my-app",
    endpoint="http://localhost:6006/v1/traces"  # Correct endpoint
)

# Force flush
from opentelemetry import trace
trace.get_tracer_provider().force_flush()

High memory in notebook:

python
# Close session when done
session = px.launch_app()
# ... do work ...
session.close()
px.close_app()

Database connection issues:

bash
# Verify PostgreSQL connection
psql $PHOENIX_SQL_DATABASE_URL -c "SELECT 1"

# Check Phoenix logs
phoenix serve --log-level debug

References

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in 17-observability/phoenix of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/advanced-usage.md
  • references/troubleshooting.md

Open the folder on GitHubat commit 773a529

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Phoenix LLM Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phoenix LLM Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phoenix LLM Observability this skillOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Agent Eval Casesagentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT
Langfusedavila7/claude-code-templates32k5 repos~1.4kAutomated safety check: PassMIT
Phoenix Integration SnippetsArize-ai/phoenix12k—~1.4kAutomated safety check: PassApache-2.0
Phoenix Release PleaseArize-ai/phoenix12k—~708Automated safety check: PassApache-2.0
Langfusesickn33/agentic-awesome-skills47k2 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Agent Eval Cases

    agentailor/fullstack-langgraph-nextjs-agent

    Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

    132 GitHub stars~5.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Langfuse

    davila7/claude-code-templates

    Expert in Langfuse - the open-source LLM observability platform.

    32k GitHub starsUsed in 5 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Generates onboarding code snippets for Phoenix tracing integrations and wires them into the project onboarding UI.

    12k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Phoenix Release Please

    Arize-ai/phoenix

    Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer.

    12k GitHub stars~708 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Langfuse

    sickn33/agentic-awesome-skills

    Expert in Langfuse - the open-source LLM observability platform.

    47k GitHub starsUsed in 2 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Ag2 Telemetry

    ag2ai/build-with-ag2

    Add OpenTelemetry traces to an AG2 beta Agent via TelemetryMiddleware (autogen.beta.middleware.builtin).

    252 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about Phoenix LLM Observability

What does Phoenix LLM Observability do?

Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server. Phoenix is an open-source tool for seeing what an LLM application actually does at runtime. The skill shows how to install it with pip, start the server from a notebook or with the phoenix serve command, and send traces to it through OpenTelemetry-based instrumentors for OpenAI, LangChain and LlamaIndex.

When should I use Phoenix LLM Observability?

Phoenix LLM Observability fits situations like: debugging an LLM application by inspecting detailed traces of each call; running repeatable evaluations of model output against a stored dataset; watching a production LLM system for problems as they happen; comparing prompts or models in an experiment before changing production.

How do I install Phoenix LLM Observability in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill phoenix-observability -a claude-code`. Or copy the skill folder (17-observability/phoenix in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/phoenix-observability in your project. Claude Code loads it when a task matches its description.

How do I install Phoenix LLM Observability in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill phoenix-observability -a codex`. Or copy the skill folder (17-observability/phoenix in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/phoenix-observability in your project. Codex loads it when a task matches its description.

Can I use Phoenix LLM Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill phoenix-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phoenix-observability, .gemini/skills/phoenix-observability, .github/skills/phoenix-observability and .opencode/skills/phoenix-observability in your project.

What does Phoenix LLM Observability need to run?

Going by SKILL.md and its folder, Phoenix LLM Observability needs the command-line tools its instructions call (pip, docker and psql) and credentials named PHOENIX_SECRET and PHOENIX_ADMIN_SECRET. Our summary lists: Python with the `arize-phoenix` package; PostgreSQL or SQLite for a self-hosted server.

Does Phoenix LLM Observability access the network?

SKILL.md names 3 domains. As links in the text: docs.arize.com, github.com and hub.docker.com. This is read from the text; nothing was executed.

Is Phoenix LLM Observability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phoenix LLM Observability use?

Phoenix LLM Observability is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phoenix LLM Observability use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.6k tokens, read only when the agent opens those files.

What are the alternatives to Phoenix LLM Observability?

Skills that share tags, products or a category with Phoenix LLM Observability: Agent Eval Cases (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Langfuse (davila7/claude-code-templates, 32k stars), Phoenix Integration Snippets (Arize-ai/phoenix, 12k stars) and Phoenix Release Please (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phoenix LLM Observability?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.