Agent skill

Phoenix Harbor

by Arize-ai in Arize-ai/phoenix

Configure and interpret the Phoenix plugin for Harbor agent evaluations.

Apache-2.0Auto-check: warningsAI & LLM Engineering

Install Phoenix Harbor

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add Arize-ai/phoenix --skill phoenix-harbor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Arize-ai/phoenix phoenix-harbor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/phoenix-harbor .claude/skills/phoenix-harbor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phoenix-harbor
GitHub stars
12k
Token cost
~3.4k tokens
SKILL.md length
1,713 words
Files
1
Skills in repo
39
Repo updated
First seen
Licence
Apache-2.0

At a glance

Configure and interpret the Phoenix plugin for Harbor agent evaluations.

  • Works in 4 steps: Check the fraction of runs with reward. → Compare reward among behaviorally… → Compare infra_ok to find environment,… → …
  • Adding arize-phoenix to Harbor jobs
  • SKILL.md covers Requirements, Choose the trace mode, Set the Phoenix destination and Add the plugin, plus 8 more sections
  • Calls uv; needs PHOENIX_API_KEY

What it does

Phoenix Harbor is an agent skill from Arize-ai/phoenix. Configure and interpret the Phoenix plugin for Harbor agent evaluations. Use when adding arize-phoenix to Harbor jobs, choosing ATIF tracing, mapping Harbor tasks and rewards to Phoenix experiments, comparing agents or models, resuming jobs, or troubleshooting Harbor records in Phoenix.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM observability. It works with Arize Phoenix and OpenTelemetry. The repository describes itself as: AI Observability & Evaluation. The licence is Apache-2.0.

When your agent uses it

  • Adding arize-phoenix to Harbor jobs
  • Choosing ATIF tracing
  • Mapping Harbor tasks and rewards to Phoenix experiments
  • Comparing agents

Example prompts

  • “/phoenix-harbor”

Requirements

  • Python 3
  • A credential in PHOENIX_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Check the fraction of runs with reward.
  2. Compare reward among behaviorally completed runs.
  3. Compare infra_ok to find environment, timeout, agent-process, or verifier reliability problems.
  4. Use step scores and the linked trace to locate the failure within a task.

What it can do on your machine

Read from SKILL.md and the folder at commit 52f76fc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arize.com
    • harborframework.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • PHOENIX_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phoenix Harbor loads about 3.4k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 1,713 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningContains instruction-override wording (e.g. “without asking the user”)SKILL.md:234
    If Phoenix reports a conflict, do not tell the user to ignore it. The plugin validates the stored run's trial output and

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Arize-ai/phoenix at commit 52f76fc, republished under its Apache-2.0 licence (© Arize-ai). 1,713 words, ~3,394 tokens.

Download SKILL.mdSave it as .claude/skills/phoenix-harbor/SKILL.md (or your agent's skills folder).
name
phoenix-harbor
description
Configure and interpret the Phoenix plugin for Harbor agent evaluations. Use when adding `arize-phoenix` to Harbor jobs, choosing ATIF tracing, mapping Harbor tasks and rewards to Phoenix experiments, comparing agents or models, resuming jobs, or troubleshooting Harbor records in Phoenix.
license
Apache-2.0
metadata.author
oss@arize.com
metadata.version
1.0.0

Phoenix for Harbor

Use the Phoenix Harbor plugin to record Harbor agent evaluations as versioned Phoenix datasets, experiments, runs, scores, and ATIF traces.

Harbor runs agents and verifiers. Phoenix records and compares their results. Do not describe Phoenix as executing Harbor tasks or recalculating Harbor rewards.

Requirements

  • Python 3.12 or newer
  • Harbor 0.21.0 or newer
  • arize-phoenix-client installed with the harbor extra
  • Phoenix server 15.0 or newer

Install the client and Harbor in the same Python environment:

bash
uv pip install "arize-phoenix-client[harbor]"

Choose the trace mode

Use atif unless the agent has no ATIF trajectory or the user does not want traces. ATIF is the default. It reads trajectory files after the final trial attempt, so the sandbox needs no Phoenix endpoint, credentials, instrumentation, or outbound network access.

Use null to record datasets, experiments, runs, and evaluations without traces:

bash
--plugin-kwarg trace_mode=null

Live OpenTelemetry Protocol (OTLP) support is deferred to a follow-up. This release accepts atif or null, and does not link live OpenTelemetry agent traces to experiment runs.

Set the Phoenix destination

Set the endpoint for the Phoenix plugin process. ATIF agents and their sandboxes do not connect to Phoenix. Prefer environment variables so credentials do not enter shell history or job configuration:

bash
export PHOENIX_COLLECTOR_ENDPOINT=http://localhost:6006
export PHOENIX_API_KEY=your-api-key

Omit PHOENIX_API_KEY when the Phoenix instance does not require authentication. The endpoint and api_key plugin kwargs override these values when the user asks for per-job settings.

Add the plugin

Add --plugin arize-phoenix to the user's existing harbor run command. Preserve their dataset, agent, model, environment, concurrency, retry, and task selections.

bash
harbor run \
  -d terminal-bench/terminal-bench-2 \
  -a terminus-2 \
  -m openai/gpt-5-mini \
  --plugin arize-phoenix \
  --yes

Do not invent or replace Harbor settings that are unrelated to Phoenix.

Predict the Phoenix records

Use this mapping when explaining a job or checking its results:

HarborPhoenix
One task collectionOne versioned dataset
One taskOne dataset example
One distinct agent and model configurationOne experiment
One planned task attemptOne repetition
One final logical trialOne experiment run
Final textual agent turnExperiment run output
Final verifier rewardExperiment evaluation with the original key and CODE annotator kind
Step verifier rewardEvaluation named <step_name>.<reward_key>
Trial or step exceptionRun error and infra_ok=0
Saved ATIF trajectoriesOne trace linked to the run, with one step span per attempted step in a multi-step task

Each single-step or multi-step Harbor task becomes one Phoenix dataset example. A multi-step example input includes its ordered step names and instructions. Phoenix examples keep output empty unless the task declares a reference file.

The plugin records only the terminal physical attempt for a logical trial. An attempt that Harbor will retry does not create a Phoenix run. Completion order does not define repetition numbers.

When a saved terminal ATIF trajectory ends with a user-facing textual agent turn, the plugin records it in chat-message format so Phoenix experiment comparisons render it as Markdown. Structured messages contribute their text parts in order; media parts are omitted. The output stays empty for missing or invalid trajectories, terminal tool calls, media-only turns, and state-only tasks. Multi-step tasks use the last attempted step, and continued trajectories use the terminal continuation. Output extraction still runs with trace_mode=null; that setting disables trace creation, not result display.

Successful runs written by older plugin versions keep their legacy Harbor metadata output because Phoenix runs are immutable. Resume recognizes that exact legacy shape and reuses the run.

Add optional reference outputs

A task can declare a checked-in reference file in its root task.toml:

toml
[metadata.arize-phoenix]
reference_output_path = "tests/expected.json"

The plugin reads this UTF-8 JSON file during setup. A JSON string becomes {"messages": [{"role": "assistant", "content": "the reference text"}]}. A JSON object is stored unchanged as the example output. Other top-level types must be wrapped in an object. No setting means an empty output, even if an expected.json file exists.

Paths are relative to the downloaded task root on the Harbor host. Absolute paths, .. components, and symlinks escaping that root are rejected. Missing, unreadable, or invalid configured files fail setup before trials run. Keep references beside verifier assets, out of the agent workspace.

This works for local datasets, direct tasks, and published tasks. Multi-step tasks have one reference for the whole task, usually its expected final result. Reference-content changes create a dataset version; existing experiments keep their original version. References do not change Harbor's grading.

Name the dataset

The plugin infers a Phoenix dataset name for each supported single-source job. Provide dataset=<name> only when a job contains several direct tasks, which have no shared collection name, or when you want to customize the dataset's display name in Phoenix.

The inferred names are:

Harbor sourcePhoenix dataset name
Named registry datasetThe selected dataset name
Published packageThe selected <organization>/<dataset> name
Local dataset pathThe resolved directory name
Repository datasetThe resolved registry metadata name
One direct taskharbor-task/<task-name>

To name several direct tasks or override an inferred name, add this setting to the Harbor command:

bash
--plugin-kwarg dataset=release-candidate-tasks

Stop and explain the constraint if the job has any unsupported source shape:

  • more than one configured dataset;
  • both a configured dataset and direct tasks;
  • several direct tasks without dataset=<name>;
  • duplicate task IDs; or
  • a regrade job or another job derived from previous Harbor results.

The plugin synchronizes the complete resolved task set at job start. An unchanged set reuses the dataset version. A task addition, removal, or content change creates a version. Existing experiments stay pinned to their creation-time version.

Name experiments

The default template is:

text
{job.name} · {agent.name} · {agent.model}

For one agent configuration, an exact name is valid:

bash
--plugin-kwarg experiment_name=release-candidate

For several agent configurations, use experiment_name_template. Available fields are:

  • {job.name}
  • {job.id}
  • {dataset.name}
  • {agent.name}
  • {agent.model}
  • {agent.short_digest}

Agent names do not need to be unique. Two agents with the same name but different effective configurations each get an experiment. If rendered names collide, the plugin appends the short agent digest. Stable identity comes from the Harbor job ID and effective agent configuration, not the display name.

Show full SKILL.md (763 more words)Show less

Interpret scores

Keep behavioral outcomes separate from execution health. Phoenix does not run another evaluator. The plugin records Harbor's completed verifier rewards as named experiment evaluations with the CODE annotator kind.

EvaluationInterpretation
rewardPresent only when the final Harbor verifier emits a literal reward key. A value of 0 is behavioral failure, not an infrastructure error.
infra_okPresent on every run. 1 means Harbor recorded no trial or step exception. 0 means at least one exception occurred.
<reward_key>A task-specific final-verifier score in its original numeric scale.
<step_name>.<reward_key>A task-specific step score for multi-step diagnosis.

Do not infer reward from another lone key. Check its coverage before computing cross-task summaries.

A run may contain rewards and still have infra_ok=0. Harbor can produce verifier output before or alongside a step exception. Preserve both facts when explaining the result.

For a multi-step task, trial-level reward evaluations include multi_step_reward_strategy metadata. Harbor's omitted default resolves to mean; preserve an explicit final. Step evaluations and infra_ok do not include this field.

For comparisons:

  1. Check the fraction of runs with reward.
  2. Compare reward among behaviorally completed runs.
  3. Compare infra_ok to find environment, timeout, agent-process, or verifier reliability problems.
  4. Use step scores and the linked trace to locate the failure within a task.

Interpret ATIF traces

One logical trial maps to one trace and one Phoenix session. The trace starts with a plugin-owned harbor.trial CHAIN span. Multi-step trials add one harbor.step span per attempted step:

text
harbor.trial <task>                  CHAIN
  harbor.step 1 <step name>          CHAIN, multi-step trials only
    <agent>                          AGENT
      turn 1                         AGENT, multi-turn trajectories only
        iteration 1                  CHAIN
          <model>                    LLM
          <tool>                     TOOL
            <subagent>               AGENT

Single-step trajectories attach directly to the trial root. Each multi-step harbor.step span carries its instruction, timing, exception status, and any verifier rewards. A step remains visible even when its trajectory is missing. All step spans and trajectories share the trial root.

Agent, model, and tool spans use their ATIF names. Fresh agent operations use iteration N; context-management operations use compaction N; and other operational system steps use system event N. Multi-turn trajectories add turn N spans. Steps with llm_call_count: 0 keep their operation and tool spans but do not create an LLM span. Continuation roots use <agent> (continuation N).

The converter supports ATIF v1.0 through v1.7. It reconstructs LLM inputs from ATIF and marks them with metadata.atif.input_source = "reconstructed". Copied prompt history contributes to those inputs without creating spans. It pairs an observation with a tool call only when source_call_id matches. Keep multiple results in order. Unmatched step observations stay on the operation span; unassigned feedback remains structured in the reconstructed input without an invented role or tool association. Structured text and image parts remain serialized, but media bytes are not uploaded. ATIF v1.8 audio fields are unsupported.

Only LLM spans carry llm.* attributes. Trajectory final_metrics remain in agent-root metadata to avoid double-counting tokens. Producer-specific cache-write and reasoning token counts map to the corresponding OpenInference token-detail attributes when present.

ATIF timestamps are point events. Zero-duration LLM or TOOL spans can mean no unambiguous duration was available. Do not interpret them as proof that the operation took no time. Declared tool order does not prove serial execution.

ATIF discovery and conversion are best-effort. If the trajectory is missing or invalid, the plugin warns and records the run without a trace. A later replay cannot attach a trace to an immutable successful run.

Handle failures and resume

Selecting the plugin makes successful Phoenix recording required.

  • Setup failures stop the job before trials run.
  • Run or evaluation write failures stop the job. Already recorded trials remain in Phoenix.
  • Harbor keeps terminal results that the plugin can ingest during resume.
  • ATIF conversion or upload failure does not stop the job. The run is recorded without a trace.

Sequential resume and replay reuse matching datasets, experiments, successful runs, evaluations, and traces. Failed runs can be retried. If another job creates a newer version of the shared dataset, recover the original experiment and keep it pinned to its creation-time version. Do not run multiple ingesters for the same Harbor job because experiment recovery is not atomic across processes.

If Phoenix reports a conflict, do not tell the user to ignore it. The plugin validates the stored run's trial output and trace identity. A mismatch requires a new Harbor job or resolution of the conflicting Phoenix record.

Current boundaries

The plugin does not support:

  • Harbor regrade jobs and other jobs derived from previous Harbor results;
  • post-hoc ingestion of a completed job;
  • live OTLP trace linkage, which is deferred to a follow-up;
  • several configured datasets in one job;
  • mixed configured datasets and direct tasks; or
  • concurrent ingestion of one Harbor job.

For the public guide, use Phoenix's Harbor documentation. For Harbor command and task configuration, use the Harbor documentation.

© Arize-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/phoenix-harbor of Arize-ai/phoenix.

Open the folder on GitHubat commit 52f76fc

Compare with similar skills

Phoenix Harbor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phoenix Harbor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phoenix Harbor this skillArize-ai/phoenix12k—~3.4kAutomated safety check: WarnApache-2.0
Phoenix LLM ObservabilityOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Sentry Elixir SDKgetsentry/sentry-for-ai268—~3.5kAutomated safety check: PassApache-2.0
Agentsop Observability Setupagentsope/SkillAlchemy466—~4.4kAutomated safety check: PassMIT
TMA1 Observability Querytma1-ai/tma1119—~5.1kAutomated safety check: NotesApache-2.0
Olore Langfuse Latestolorehq/olore104—~1.2kAutomated safety check: PassMIT

Similar skills

  • Phoenix LLM Observability

    Orchestra-Research/AI-Research-SKILLs

    Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Sentry Elixir SDK

    getsentry/sentry-for-ai

    Official

    Full Sentry SDK setup for Elixir. An agent skill from getsentry/sentry-for-ai.

    268 GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agentsop Observability Setup

    agentsope/SkillAlchemy

    Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.

    466 GitHub stars~4.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Answers questions about agent spend, token use, traces, events, errors and tool usage by running read-only SQL against a local TMA1 observability store.

    119 GitHub stars~5.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Local Langfuse documentation reference (latest). An agent skill from olorehq/olore.

    104 GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Langfuse SDK Patterns

    jeremylongshore/tons-of-skills-marketplace

    Langfuse SDK best practices, patterns, and idiomatic usage. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from Arize-ai/phoenix

All 39 skills in this repo
  • Harbor Exec

    Arize-ai/phoenix

    A skill your agent uses when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only…

    12k GitHub stars~909 tokensUpdated today
    Auto-check passed
  • Mintlify

    Arize-ai/phoenix

    Build and maintain documentation sites with Mintlify. An agent skill from Arize-ai/phoenix.

    12k GitHub starsUsed in 8 repos~3.4k tokens
    Auto-check passed
  • Phoenix Frontend

    Arize-ai/phoenix

    Frontend development guidelines for the Phoenix AI observability platform.

    12k GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Phoenix Graphql

    Arize-ai/phoenix

    Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.

    12k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Phoenix Server

    Arize-ai/phoenix

    Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).

    12k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Phoenix Storybook

    Arize-ai/phoenix

    Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).

    12k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Phoenix Harbor

What does Phoenix Harbor do?

Configure and interpret the Phoenix plugin for Harbor agent evaluations. Phoenix Harbor is an agent skill from Arize-ai/phoenix. Configure and interpret the Phoenix plugin for Harbor agent evaluations.

When should I use Phoenix Harbor?

Phoenix Harbor fits situations like: adding arize-phoenix to Harbor jobs; choosing ATIF tracing; mapping Harbor tasks and rewards to Phoenix experiments; comparing agents.

How do I install Phoenix Harbor in Claude Code?

Run `npx skills add Arize-ai/phoenix --skill phoenix-harbor -a claude-code`. Or copy the skill folder (.agents/skills/phoenix-harbor in Arize-ai/phoenix) into .claude/skills/phoenix-harbor in your project. Claude Code loads it when a task matches its description.

How do I install Phoenix Harbor in Codex?

Run `npx skills add Arize-ai/phoenix --skill phoenix-harbor -a codex`. Or copy the skill folder (.agents/skills/phoenix-harbor in Arize-ai/phoenix) into .agents/skills/phoenix-harbor in your project. Codex loads it when a task matches its description.

Can I use Phoenix Harbor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arize-ai/phoenix --skill phoenix-harbor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phoenix-harbor, .gemini/skills/phoenix-harbor, .github/skills/phoenix-harbor and .opencode/skills/phoenix-harbor in your project.

What does Phoenix Harbor need to run?

Going by SKILL.md and its folder, Phoenix Harbor needs the command-line tools its instructions call (uv) and credentials named PHOENIX_API_KEY. Our summary lists: Python 3; A credential in PHOENIX_API_KEY.

Does Phoenix Harbor access the network?

SKILL.md names 2 domains. As links in the text: arize.com and harborframework.com. This is read from the text; nothing was executed.

Is Phoenix Harbor safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): contains instruction-override wording (e.g. “without asking the user”). Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Phoenix Harbor use?

Phoenix Harbor is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phoenix Harbor use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phoenix Harbor?

Skills that share tags, products or a category with Phoenix Harbor: Phoenix LLM Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars), Sentry Elixir SDK (getsentry/sentry-for-ai, 268 stars), Agentsop Observability Setup (agentsope/SkillAlchemy, 466 stars) and TMA1 Observability Query (tma1-ai/tma1, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phoenix Harbor?

Arize-ai (a GitHub organization) maintains it in Arize-ai/phoenix, which has 11,764 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 9, 2026.

Source: Arize-ai/phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.