Agent skill

Experiments

by Arize-ai in Arize-ai/phoenix

Run, read, and compare dataset-backed experiments to find evidence that a prompt or pipeline is improving.

Custom licenceAuto-check passedAI & LLM Engineering

Install Experiments

skills CLI
$ npx skills add Arize-ai/phoenix --skill experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Arize-ai/phoenix experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Arize-ai/phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/phoenix/server/agents/prompts/skills/experiments .claude/skills/experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiments
GitHub stars
12k
Token cost
~1.8k tokens
SKILL.md length
911 words
Files
1
Skills in repo
39
Repo updated
First seen
Licence
Custom licence

At a glance

Run, read, and compare dataset-backed experiments to find evidence that a prompt or pipeline is improving.

  • Works in 9 steps: Confirm the dataset represents the task:… → Make sure the starting prompt is well… → Run the prompt over the dataset as a… → …
  • The user wants to iterate over a dataset with experiments
  • SKILL.md covers Before You Start: Read What…, Workflow: Iterate Over A Dataset, Recording What You Learned and Boundaries, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiments is an agent skill from Arize-ai/phoenix. Run, read, and compare dataset-backed experiments to find evidence that a prompt or pipeline is improving. Trigger when the user wants to iterate over a dataset with experiments, compare experiment runs, read experiment quality/latency/cost, or decide whether a change actually helped. Running a prompt over a dataset is implicitly an experiment — load this skill when dataset-backed work begins, before authoring evaluators for the experiment and before starting the recorded run, not only when reading results. Do…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM observability and Quizzes and assessments. The repository describes itself as: AI Observability & Evaluation.

When your agent uses it

  • The user wants to iterate over a dataset with experiments
  • Compare experiment runs
  • Read experiment quality/latency/cost
  • Decide whether a change actually helped

Example prompts

  • “/experiments”

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the dataset represents the task: the input fields the run consumes, the expected outputs,
  2. Make sure the starting prompt is well formed before running it — task, variables, output format,
  3. Run the prompt over the dataset as a recorded experiment, staging the scaffold (hypothesis,
  4. Read the results across all three axes together — output quality (evaluator annotations, including
  5. Score what you observe. Anything example-level and scorable defaults to an evaluator at the moment
  6. Form one specific hypothesis for the next candidate — a named failure mode and the single change
  7. Compare the new experiment against its baseline per-example and aligned, not by aggregate means
  8. Report what the comparison showed: a verdict on the hypothesis, a summary across quality, latency,
  9. Continue hypothesis → run → compare → report until the evidence meets the stated goal, then save

What it can do on your machine

Read from SKILL.md and the folder at commit 3383f07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiments loads about 1.8k tokens when it runs. Until then it costs about 201 tokens; SKILL.md has 911 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~201
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 911 words (~1,760 tokens).

“An experiment is one run of a prompt or pipeline over every example in a dataset, captured with its outputs and any evaluator annotations so it can be reviewed and compared later. Experiments turn "this prompt feels better" into evidence…”

— opening of SKILL.md by Arize-ai, Custom licence
name
experiments
summary
Iterate over a dataset with experiments — run, read results across quality, latency, and cost, and compare candidates to drive improvement.

Read the full SKILL.md on GitHub

Files

Just SKILL.md in src/phoenix/server/agents/prompts/skills/experiments of Arize-ai/phoenix.

Open the folder on GitHubat commit 3383f07

Compare with similar skills

Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiments this skillArize-ai/phoenix12k—~1.8kAutomated safety check: PassCustom licence
Agent Prompt Quality Barmastra-ai/mastra29k—~2kAutomated safety check: PassCustom licence
Advanced Evaluationguanyang/open-agent-hub9732 repos~4.2kAutomated safety check: PassMIT
Agentic Evalgithub/awesome-copilot40k4 repos~1.5kAutomated safety check: PassMIT
Clawpathy AutoresearchClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT
Agentsop Metric Designagentsope/SkillAlchemy457—~6.4kAutomated safety check: PassMIT

Similar skills

  • Agent Prompt Quality Bar

    mastra-ai/mastra

    Universal quality bar and final audit rubric for any agent system prompt.

    29k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Advanced Evaluation

    guanyang/open-agent-hub

    This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

    973 GitHub starsUsed in 2 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Agentic Eval

    github/awesome-copilot

    Official

    Patterns and techniques for evaluating and improving AI agent outputs.

    40k GitHub starsUsed in 4 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agentsop Metric Design

    agentsope/SkillAlchemy

    Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.

    457 GitHub stars~6.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Suede AI Eval

    JasonColapietro/suede-creator-skills

    Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates.

    127 GitHub stars~3.3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from Arize-ai/phoenix

All 39 skills in this repo
  • Harbor Exec

    Arize-ai/phoenix

    A skill your agent uses when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only…

    12k GitHub stars~909 tokensUpdated today
    Auto-check passed
  • Mintlify

    Arize-ai/phoenix

    Build and maintain documentation sites with Mintlify. An agent skill from Arize-ai/phoenix.

    12k GitHub starsUsed in 7 repos~3.4k tokens
    Auto-check passed
  • Phoenix Frontend

    Arize-ai/phoenix

    Frontend development guidelines for the Phoenix AI observability platform.

    12k GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Phoenix Graphql

    Arize-ai/phoenix

    Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.

    12k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Phoenix Server

    Arize-ai/phoenix

    Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).

    12k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Phoenix Storybook

    Arize-ai/phoenix

    Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).

    12k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Experiments

What does Experiments do?

Run, read, and compare dataset-backed experiments to find evidence that a prompt or pipeline is improving. Experiments is an agent skill from Arize-ai/phoenix. Run, read, and compare dataset-backed experiments to find evidence that a prompt or pipeline is improving.

When should I use Experiments?

Experiments fits situations like: the user wants to iterate over a dataset with experiments; compare experiment runs; read experiment quality/latency/cost; decide whether a change actually helped.

How do I install Experiments in Claude Code?

Run `npx skills add Arize-ai/phoenix --skill experiments -a claude-code`. Or copy the skill folder (src/phoenix/server/agents/prompts/skills/experiments in Arize-ai/phoenix) into .claude/skills/experiments in your project. Claude Code loads it when a task matches its description.

How do I install Experiments in Codex?

Run `npx skills add Arize-ai/phoenix --skill experiments -a codex`. Or copy the skill folder (src/phoenix/server/agents/prompts/skills/experiments in Arize-ai/phoenix) into .agents/skills/experiments in your project. Codex loads it when a task matches its description.

Can I use Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arize-ai/phoenix --skill experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiments, .gemini/skills/experiments, .github/skills/experiments and .opencode/skills/experiments in your project.

What does Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiments is instructions for the agent only.

Does Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiments use?

Experiments has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Experiments use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiments?

Skills that share tags, products or a category with Experiments: Agent Prompt Quality Bar (mastra-ai/mastra, 29k stars), Advanced Evaluation (guanyang/open-agent-hub, 973 stars), Agentic Eval (github/awesome-copilot, 40k stars) and Clawpathy Autoresearch (ClawBio/ClawBio, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiments?

Arize-ai (a GitHub organization) maintains it in Arize-ai/phoenix, which has 11,738 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 7, 2026.

Source: Arize-ai/phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.