Phoenix LLM Observability
Orchestra-Research/AI-Research-SKILLs
Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.
Configure and interpret the Phoenix plugin for Harbor agent evaluations.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add Arize-ai/phoenix --skill phoenix-harbor -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Arize-ai/phoenix phoenix-harbor --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/phoenix-harbor .claude/skills/phoenix-harbor && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "phoenix-harbor" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-harbor into .claude/skills/phoenix-harbor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-harbor", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-harborType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Arize-ai/phoenix --skill phoenix-harbor -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Arize-ai/phoenix phoenix-harbor --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/phoenix-harbor .agents/skills/phoenix-harbor && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "phoenix-harbor" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-harbor into .agents/skills/phoenix-harbor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-harbor", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Arize-ai/phoenix --skill phoenix-harbor -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Arize-ai/phoenix phoenix-harbor --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/phoenix-harbor .cursor/skills/phoenix-harbor && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "phoenix-harbor" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-harbor into .cursor/skills/phoenix-harbor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-harbor", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Arize-ai/phoenix.git --path .agents/skills/phoenix-harbor--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Arize-ai/phoenix --skill phoenix-harbor -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Arize-ai/phoenix phoenix-harbor --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/phoenix-harbor .gemini/skills/phoenix-harbor && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "phoenix-harbor" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-harbor into .gemini/skills/phoenix-harbor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-harbor", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Arize-ai/phoenix phoenix-harborInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Arize-ai/phoenix --skill phoenix-harbor -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/phoenix-harbor .github/skills/phoenix-harbor && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "phoenix-harbor" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-harbor into .github/skills/phoenix-harbor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-harbor", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Arize-ai/phoenix --skill phoenix-harbor -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Arize-ai/phoenix phoenix-harbor --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/phoenix-harbor .opencode/skills/phoenix-harbor && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "phoenix-harbor" agent skill from https://github.com/Arize-ai/phoenix/tree/main/.agents/skills/phoenix-harbor into .opencode/skills/phoenix-harbor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phoenix-harbor", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
phoenix-harborConfigure and interpret the Phoenix plugin for Harbor agent evaluations.
Phoenix Harbor is an agent skill from Arize-ai/phoenix. Configure and interpret the Phoenix plugin for Harbor agent evaluations. Use when adding arize-phoenix to Harbor jobs, choosing ATIF tracing, mapping Harbor tasks and rewards to Phoenix experiments, comparing agents or models, resuming jobs, or troubleshooting Harbor records in Phoenix.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM observability. It works with Arize Phoenix and OpenTelemetry. The repository describes itself as: AI Observability & Evaluation. The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 52f76fc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
arize.comharborframework.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
PHOENIX_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Phoenix Harbor loads about 3.4k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 1,713 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
If Phoenix reports a conflict, do not tell the user to ignore it. The plugin validates the stored run's trial output andAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Arize-ai/phoenix at commit 52f76fc, republished under its Apache-2.0 licence (© Arize-ai). 1,713 words, ~3,394 tokens.
.claude/skills/phoenix-harbor/SKILL.md (or your agent's skills folder).Use the Phoenix Harbor plugin to record Harbor agent evaluations as versioned Phoenix datasets, experiments, runs, scores, and ATIF traces.
Harbor runs agents and verifiers. Phoenix records and compares their results. Do not describe Phoenix as executing Harbor tasks or recalculating Harbor rewards.
arize-phoenix-client installed with the harbor extraInstall the client and Harbor in the same Python environment:
uv pip install "arize-phoenix-client[harbor]"Use atif unless the agent has no ATIF trajectory or the user does not want traces. ATIF is the default. It reads trajectory files after the final trial attempt, so the sandbox needs no Phoenix endpoint, credentials, instrumentation, or outbound network access.
Use null to record datasets, experiments, runs, and evaluations without traces:
--plugin-kwarg trace_mode=nullLive OpenTelemetry Protocol (OTLP) support is deferred to a follow-up. This release accepts atif or null, and does not link live OpenTelemetry agent traces to experiment runs.
Set the endpoint for the Phoenix plugin process. ATIF agents and their sandboxes do not connect to Phoenix. Prefer environment variables so credentials do not enter shell history or job configuration:
export PHOENIX_COLLECTOR_ENDPOINT=http://localhost:6006
export PHOENIX_API_KEY=your-api-keyOmit PHOENIX_API_KEY when the Phoenix instance does not require authentication. The endpoint and api_key plugin kwargs override these values when the user asks for per-job settings.
Add --plugin arize-phoenix to the user's existing harbor run command. Preserve their dataset, agent, model, environment, concurrency, retry, and task selections.
harbor run \
-d terminal-bench/terminal-bench-2 \
-a terminus-2 \
-m openai/gpt-5-mini \
--plugin arize-phoenix \
--yesDo not invent or replace Harbor settings that are unrelated to Phoenix.
Use this mapping when explaining a job or checking its results:
| Harbor | Phoenix |
|---|---|
| One task collection | One versioned dataset |
| One task | One dataset example |
| One distinct agent and model configuration | One experiment |
| One planned task attempt | One repetition |
| One final logical trial | One experiment run |
| Final textual agent turn | Experiment run output |
| Final verifier reward | Experiment evaluation with the original key and CODE annotator kind |
| Step verifier reward | Evaluation named <step_name>.<reward_key> |
| Trial or step exception | Run error and infra_ok=0 |
| Saved ATIF trajectories | One trace linked to the run, with one step span per attempted step in a multi-step task |
Each single-step or multi-step Harbor task becomes one Phoenix dataset example. A multi-step example input includes its ordered step names and instructions. Phoenix examples keep output empty unless the task declares a reference file.
The plugin records only the terminal physical attempt for a logical trial. An attempt that Harbor will retry does not create a Phoenix run. Completion order does not define repetition numbers.
When a saved terminal ATIF trajectory ends with a user-facing textual agent turn, the plugin records it in chat-message format so Phoenix experiment comparisons render it as Markdown. Structured messages contribute their text parts in order; media parts are omitted. The output stays empty for missing or invalid trajectories, terminal tool calls, media-only turns, and state-only tasks. Multi-step tasks use the last attempted step, and continued trajectories use the terminal continuation. Output extraction still runs with trace_mode=null; that setting disables trace creation, not result display.
Successful runs written by older plugin versions keep their legacy Harbor metadata output because Phoenix runs are immutable. Resume recognizes that exact legacy shape and reuses the run.
A task can declare a checked-in reference file in its root task.toml:
[metadata.arize-phoenix]
reference_output_path = "tests/expected.json"The plugin reads this UTF-8 JSON file during setup. A JSON string becomes
{"messages": [{"role": "assistant", "content": "the reference text"}]}.
A JSON object is stored unchanged as the example output. Other top-level types
must be wrapped in an object. No setting means an empty output, even if an
expected.json file exists.
Paths are relative to the downloaded task root on the Harbor host. Absolute
paths, .. components, and symlinks escaping that root are rejected. Missing,
unreadable, or invalid configured files fail setup before trials run. Keep
references beside verifier assets, out of the agent workspace.
This works for local datasets, direct tasks, and published tasks. Multi-step tasks have one reference for the whole task, usually its expected final result. Reference-content changes create a dataset version; existing experiments keep their original version. References do not change Harbor's grading.
The plugin infers a Phoenix dataset name for each supported single-source job. Provide dataset=<name> only when a job contains several direct tasks, which have no shared collection name, or when you want to customize the dataset's display name in Phoenix.
The inferred names are:
| Harbor source | Phoenix dataset name |
|---|---|
| Named registry dataset | The selected dataset name |
| Published package | The selected <organization>/<dataset> name |
| Local dataset path | The resolved directory name |
| Repository dataset | The resolved registry metadata name |
| One direct task | harbor-task/<task-name> |
To name several direct tasks or override an inferred name, add this setting to the Harbor command:
--plugin-kwarg dataset=release-candidate-tasksStop and explain the constraint if the job has any unsupported source shape:
dataset=<name>;The plugin synchronizes the complete resolved task set at job start. An unchanged set reuses the dataset version. A task addition, removal, or content change creates a version. Existing experiments stay pinned to their creation-time version.
The default template is:
{job.name} · {agent.name} · {agent.model}For one agent configuration, an exact name is valid:
--plugin-kwarg experiment_name=release-candidateFor several agent configurations, use experiment_name_template. Available fields are:
{job.name}{job.id}{dataset.name}{agent.name}{agent.model}{agent.short_digest}Agent names do not need to be unique. Two agents with the same name but different effective configurations each get an experiment. If rendered names collide, the plugin appends the short agent digest. Stable identity comes from the Harbor job ID and effective agent configuration, not the display name.
Keep behavioral outcomes separate from execution health. Phoenix does not run another evaluator. The plugin records Harbor's completed verifier rewards as named experiment evaluations with the CODE annotator kind.
| Evaluation | Interpretation |
|---|---|
reward | Present only when the final Harbor verifier emits a literal reward key. A value of 0 is behavioral failure, not an infrastructure error. |
infra_ok | Present on every run. 1 means Harbor recorded no trial or step exception. 0 means at least one exception occurred. |
<reward_key> | A task-specific final-verifier score in its original numeric scale. |
<step_name>.<reward_key> | A task-specific step score for multi-step diagnosis. |
Do not infer reward from another lone key. Check its coverage before computing cross-task summaries.
A run may contain rewards and still have infra_ok=0. Harbor can produce verifier output before or alongside a step exception. Preserve both facts when explaining the result.
For a multi-step task, trial-level reward evaluations include multi_step_reward_strategy metadata. Harbor's omitted default resolves to mean; preserve an explicit final. Step evaluations and infra_ok do not include this field.
For comparisons:
reward.reward among behaviorally completed runs.infra_ok to find environment, timeout, agent-process, or verifier reliability problems.One logical trial maps to one trace and one Phoenix session. The trace starts with a plugin-owned harbor.trial CHAIN span. Multi-step trials add one harbor.step span per attempted step:
harbor.trial <task> CHAIN
harbor.step 1 <step name> CHAIN, multi-step trials only
<agent> AGENT
turn 1 AGENT, multi-turn trajectories only
iteration 1 CHAIN
<model> LLM
<tool> TOOL
<subagent> AGENTSingle-step trajectories attach directly to the trial root. Each multi-step harbor.step span carries its instruction, timing, exception status, and any verifier rewards. A step remains visible even when its trajectory is missing. All step spans and trajectories share the trial root.
Agent, model, and tool spans use their ATIF names. Fresh agent operations use iteration N; context-management operations use compaction N; and other operational system steps use system event N. Multi-turn trajectories add turn N spans. Steps with llm_call_count: 0 keep their operation and tool spans but do not create an LLM span. Continuation roots use <agent> (continuation N).
The converter supports ATIF v1.0 through v1.7. It reconstructs LLM inputs from ATIF and marks them with metadata.atif.input_source = "reconstructed". Copied prompt history contributes to those inputs without creating spans. It pairs an observation with a tool call only when source_call_id matches. Keep multiple results in order. Unmatched step observations stay on the operation span; unassigned feedback remains structured in the reconstructed input without an invented role or tool association. Structured text and image parts remain serialized, but media bytes are not uploaded. ATIF v1.8 audio fields are unsupported.
Only LLM spans carry llm.* attributes. Trajectory final_metrics remain in agent-root metadata to avoid double-counting tokens. Producer-specific cache-write and reasoning token counts map to the corresponding OpenInference token-detail attributes when present.
ATIF timestamps are point events. Zero-duration LLM or TOOL spans can mean no unambiguous duration was available. Do not interpret them as proof that the operation took no time. Declared tool order does not prove serial execution.
ATIF discovery and conversion are best-effort. If the trajectory is missing or invalid, the plugin warns and records the run without a trace. A later replay cannot attach a trace to an immutable successful run.
Selecting the plugin makes successful Phoenix recording required.
Sequential resume and replay reuse matching datasets, experiments, successful runs, evaluations, and traces. Failed runs can be retried. If another job creates a newer version of the shared dataset, recover the original experiment and keep it pinned to its creation-time version. Do not run multiple ingesters for the same Harbor job because experiment recovery is not atomic across processes.
If Phoenix reports a conflict, do not tell the user to ignore it. The plugin validates the stored run's trial output and trace identity. A mismatch requires a new Harbor job or resolution of the conflicting Phoenix record.
The plugin does not support:
For the public guide, use Phoenix's Harbor documentation. For Harbor command and task configuration, use the Harbor documentation.
© Arize-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/phoenix-harbor of Arize-ai/phoenix.
Open the folder on GitHubat commit 52f76fc
Phoenix Harbor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Phoenix Harbor this skillArize-ai/phoenix | 12k | — | ~3.4k | Automated safety check: Warn | Apache-2.0 | |
| Phoenix LLM ObservabilityOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Sentry Elixir SDKgetsentry/sentry-for-ai | 268 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Agentsop Observability Setupagentsope/SkillAlchemy | 466 | — | ~4.4k | Automated safety check: Pass | MIT | |
| TMA1 Observability Querytma1-ai/tma1 | 119 | — | ~5.1k | Automated safety check: Notes | Apache-2.0 | |
| Olore Langfuse Latestolorehq/olore | 104 | — | ~1.2k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.
getsentry/sentry-for-ai
Full Sentry SDK setup for Elixir. An agent skill from getsentry/sentry-for-ai.
agentsope/SkillAlchemy
Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.
tma1-ai/tma1
Answers questions about agent spend, token use, traces, events, errors and tool usage by running read-only SQL against a local TMA1 observability store.
olorehq/olore
Local Langfuse documentation reference (latest). An agent skill from olorehq/olore.
jeremylongshore/tons-of-skills-marketplace
Langfuse SDK best practices, patterns, and idiomatic usage. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Arize-ai/phoenix
A skill your agent uses when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only…
Arize-ai/phoenix
Build and maintain documentation sites with Mintlify. An agent skill from Arize-ai/phoenix.
Arize-ai/phoenix
Frontend development guidelines for the Phoenix AI observability platform.
Arize-ai/phoenix
Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.
Arize-ai/phoenix
Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).
Arize-ai/phoenix
Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).
Works with
Categories
Configure and interpret the Phoenix plugin for Harbor agent evaluations. Phoenix Harbor is an agent skill from Arize-ai/phoenix. Configure and interpret the Phoenix plugin for Harbor agent evaluations.
Phoenix Harbor fits situations like: adding arize-phoenix to Harbor jobs; choosing ATIF tracing; mapping Harbor tasks and rewards to Phoenix experiments; comparing agents.
Run `npx skills add Arize-ai/phoenix --skill phoenix-harbor -a claude-code`. Or copy the skill folder (.agents/skills/phoenix-harbor in Arize-ai/phoenix) into .claude/skills/phoenix-harbor in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Arize-ai/phoenix --skill phoenix-harbor -a codex`. Or copy the skill folder (.agents/skills/phoenix-harbor in Arize-ai/phoenix) into .agents/skills/phoenix-harbor in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arize-ai/phoenix --skill phoenix-harbor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phoenix-harbor, .gemini/skills/phoenix-harbor, .github/skills/phoenix-harbor and .opencode/skills/phoenix-harbor in your project.
Going by SKILL.md and its folder, Phoenix Harbor needs the command-line tools its instructions call (uv) and credentials named PHOENIX_API_KEY. Our summary lists: Python 3; A credential in PHOENIX_API_KEY.
SKILL.md names 2 domains. As links in the text: arize.com and harborframework.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): contains instruction-override wording (e.g. “without asking the user”). Read the flagged lines before installing; the check is not a guarantee either way.
Phoenix Harbor is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Phoenix Harbor: Phoenix LLM Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars), Sentry Elixir SDK (getsentry/sentry-for-ai, 268 stars), Agentsop Observability Setup (agentsope/SkillAlchemy, 466 stars) and TMA1 Observability Query (tma1-ai/tma1, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Arize-ai (a GitHub organization) maintains it in Arize-ai/phoenix, which has 11,764 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 9, 2026.
Source: Arize-ai/phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.