Official agent skill

Arize Evaluator

by github in github/awesome-copilot

Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

OfficialMITAuto-check: notesAI & LLM Engineering

Install Arize Evaluator

skills CLI
$ npx skills add github/awesome-copilot --skill arize-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot arize-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arize-evaluator .claude/skills/arize-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arize-evaluator
GitHub stars
40k
Used in
2 other repos
Token cost
~8.1k tokens
SKILL.md length
3,140 words
Files
3 (incl. references)
Skills in repo
417
Repo updated
First seen
Licence
MIT

At a glance

Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

  • Works in 12 steps: Confirm the project name → Understand what to evaluate → Confirm or create an AI integration → …
  • The user mentions create evaluator
  • SKILL.md covers Prerequisites, Concepts, Data Granularity and Basic CRUD, plus 2 more sections
  • Calls python3; needs OPENAI_API_KEY and ANTHROPIC_API_KEY

What it does

Arize Evaluator is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring. Use when the user mentions create evaluator, LLM judge, hallucination, faithfulness, correctness, relevance, run eval, score spans, score experiment, trigger-run, column mapping, continuous monitoring, or improve evaluator prompt.

Its SKILL.md is about 8.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/ax-profiles.md` and `references/ax-setup.md`). Compatibility notes: Requires the ax CLI and a configured Arize profile with an AI integration.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.

When your agent uses it

  • The user mentions create evaluator
  • Score experiment
  • Continuous monitoring
  • Improve evaluator prompt

Example prompts

  • “Use the arize-evaluator skill to handle LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on…”
  • “/arize-evaluator”

Requirements

  • A credential in OPENAI_API_KEY
  • A credential in ANTHROPIC_API_KEY
  • Compatibility (from SKILL.md): Requires the ax CLI and a configured Arize profile with an AI integration.

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Confirm the project name
  2. Understand what to evaluate
  3. Confirm or create an AI integration
  4. Create the evaluator
  5. Ask — backfill, continuous, or both?
  6. Determine column mappings from real span data
  7. Create the task
  8. Trigger a backfill run (if requested)
  9. Find the dataset and experiment names
  10. Understand what to evaluate
  11. Confirm or create an AI integration
  12. Create the evaluator

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arize.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the ax CLI and a configured Arize profile with an AI integration.

    From compatibility in the SKILL.md frontmatter.

Context cost

Arize Evaluator loads about 8.1k tokens when it runs, and up to ~9.6k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 3,140 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~8.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:27
    - **Security:** Never read `.env` files or search the filesystem for credentials. Use `ax profiles` for Arize credential

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 3,140 words, ~8,053 tokens.

Download SKILL.mdSave it as .claude/skills/arize-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
arize-evaluator
description
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring. Use when the user mentions create evaluator, LLM judge, hallucination, faithfulness, correctness, relevance, run eval, score spans, score experiment, trigger-run, column mapping, continuous monitoring, or improve evaluator prompt.
compatibility
Requires the ax CLI and a configured Arize profile with an AI integration.
metadata.author
arize
metadata.version
1.0

Arize Evaluator Skill

SPACE — All --space flags and the ARIZE_SPACE env var accept a space name (e.g., my-workspace) or a base64 space ID (e.g., U3BhY2U6...). Find yours with ax spaces list.

This skill covers designing, creating, and running LLM-as-judge evaluators on Arize. An evaluator defines the judge; a task is how you run it against real data.


Prerequisites

Proceed directly with the task — run the ax command you need. Do NOT check versions, env vars, or profiles upfront.

If an ax command fails, troubleshoot based on the error:

  • command not found or version error → see references/ax-setup.md
  • 401 Unauthorized / missing API key → run ax profiles show to inspect the current profile. If the profile is missing or the API key is wrong, follow references/ax-profiles.md to create/update it. If the user doesn't have their key, direct them to https://app.arize.com/admin > API Keys
  • Space unknown → run ax spaces list to pick by name, or ask the user
  • LLM provider call fails (missing OPENAI_API_KEY / ANTHROPIC_API_KEY) → run ax ai-integrations list --space SPACE to check for platform-managed credentials. If none exist, ask the user to provide the key or create an integration via the arize-ai-provider-integration skill
  • Security: Never read .env files or search the filesystem for credentials. Use ax profiles for Arize credentials and ax ai-integrations for LLM provider keys. If credentials are not available through these channels, ask the user.
  • CRITICAL — Never fabricate evaluation results: If an evaluation task fails, is cancelled, or produces no scores, report the failure clearly and explain what went wrong. Do NOT perform a "manual evaluation," invent quality scores, estimate percentages, or present any agent-generated analysis as if it came from the Arize evaluation system. Instead suggest: (1) fix the identified issue and retry, (2) try running from the Arize UI, (3) verify integration credentials with ax ai-integrations list, (4) contact support at https://arize.com/support

Concepts

What is an Evaluator?

An evaluator is an LLM-as-judge definition. It contains:

FieldDescription
TemplateThe judge prompt. Uses {variable} placeholders (e.g. {input}, {output}, {context}) that get filled in at run time via a task's column mappings.
Classification choicesThe set of allowed output labels (e.g. factual / hallucinated). Binary is the default and most common. Each choice can optionally carry a numeric score.
AI IntegrationStored LLM provider credentials (OpenAI, Anthropic, Bedrock, etc.) the evaluator uses to call the judge model.
ModelThe specific judge model (e.g. gpt-4o, claude-sonnet-4-5).
Invocation paramsOptional JSON of model settings like {"temperature": 0}. Low temperature is recommended for reproducibility.
Optimization directionWhether higher scores are better (maximize) or worse (minimize). Sets how the UI renders trends.
Data granularityWhether the evaluator runs at the span, trace, or session level. Most evaluators run at the span level.

Evaluators are versioned — every prompt or model change creates a new immutable version. The most recent version is active.

What is a Task?

A task is how you run one or more evaluators against real data. Tasks are attached to a project (live traces/spans) or a dataset (experiment runs). A task contains:

FieldDescription
EvaluatorsList of evaluators to run. You can run multiple in one task.
Column mappingsMaps each evaluator's template variables to actual field paths on spans or experiment runs (e.g. "input" → "attributes.input.value"). This is what makes evaluators portable across projects and experiments.
Query filterSQL-style expression to select which spans/runs to evaluate (e.g. "span_kind = 'LLM'"). Optional but important for precision.
ContinuousFor project tasks: whether to automatically score new spans as they arrive.
Sampling rateFor continuous project tasks: fraction of new spans to evaluate (0–1).

Data Granularity

The --data-granularity flag controls what unit of data the evaluator scores. It defaults to span and only applies to project tasks (not dataset/experiment tasks — those evaluate experiment runs directly).

LevelWhat it evaluatesUse forResult column prefix
span (default)Individual spansQ&A correctness, hallucination, relevanceeval.{name}.label / .score / .explanation
traceAll spans in a trace, grouped by context.trace_idAgent trajectory, task correctness — anything that needs the full call chaintrace_eval.{name}.label / .score / .explanation
sessionAll traces in a session, grouped by attributes.session.id and ordered by start timeMulti-turn coherence, overall tone, conversation qualitysession_eval.{name}.label / .score / .explanation
How trace and session aggregation works

For trace granularity, spans sharing the same context.trace_id are grouped together. Column values used by the evaluator template are comma-joined into a single string (each value truncated to 100K characters) before being passed to the judge model.

For session granularity, the same trace-level grouping happens first, then traces are ordered by start_time and grouped by attributes.session.id. Session-level values are capped at 100K characters total.

The {conversation} template variable

At session granularity, {conversation} is a special template variable that renders as a JSON array of {input, output} turns across all traces in the session, built from attributes.input.value / attributes.llm.input_messages (input side) and attributes.output.value / attributes.llm.output_messages (output side).

At span or trace granularity, {conversation} is treated as a regular template variable and resolved via column mappings like any other.

Multi-evaluator tasks

A task can contain evaluators at different granularities. At runtime the system uses the highest granularity (session > trace > span) for data fetching and automatically splits into one child run per evaluator. Per-evaluator query_filter in the task's evaluators JSON further narrows which spans are included (e.g., only tool-call spans within a session).


Basic CRUD

AI Integrations

AI integrations store the LLM provider credentials the evaluator uses. For full CRUD — listing, creating for all providers (OpenAI, Anthropic, Azure, Bedrock, Vertex, Gemini, NVIDIA NIM, custom), updating, and deleting — use the arize-ai-provider-integration skill.

Quick reference for the common case (OpenAI):

bash
# Check for an existing integration first
ax ai-integrations list --space SPACE

# Create if none exists
ax ai-integrations create \
  --name "My OpenAI Integration" \
  --provider openAI \
  --api-key $OPENAI_API_KEY

Copy the returned integration ID — it is required for ax evaluators create --ai-integration-id.

Evaluators
bash
# List / Get
ax evaluators list --space SPACE
ax evaluators get ID                    # accepts name or ID
ax evaluators get NAME --space SPACE   # required when using name instead of ID
ax evaluators list-versions NAME_OR_ID
ax evaluators get-version VERSION_ID

# Create (creates the evaluator and its first version)
ax evaluators create \
  --name "Answer Correctness" \
  --space SPACE \
  --description "Judges if the model answer is correct" \
  --template-name "correctness" \
  --commit-message "Initial version" \
  --ai-integration-id INT_ID \
  --model-name "gpt-4o" \
  --include-explanations \
  --use-function-calling \
  --classification-choices '{"correct": 1, "incorrect": 0}' \
  --template 'You are an evaluator. Given the user question and the model response, decide if the response correctly answers the question.

User question: {input}

Model response: {output}

Respond with exactly one of these labels: correct, incorrect'

# Create a new version (for prompt or model changes — versions are immutable)
ax evaluators create-version NAME_OR_ID \
  --commit-message "Added context grounding" \
  --template-name "correctness" \
  --ai-integration-id INT_ID \
  --model-name "gpt-4o" \
  --include-explanations \
  --classification-choices '{"correct": 1, "incorrect": 0}' \
  --template 'Updated prompt...

{input} / {output} / {context}'

# Update metadata only (name, description — not prompt)
ax evaluators update NAME_OR_ID \
  --name "New Name" \
  --description "Updated description"

# Delete (permanent — removes all versions)
ax evaluators delete NAME_OR_ID

Key flags for create:

FlagRequiredDescription
--nameyesEvaluator name (unique within space)
--spaceyesSpace name or ID to create in
--template-nameyesEval column name — alphanumeric, spaces, hyphens, underscores
--commit-messageyesDescription of this version
--ai-integration-idyesAI integration ID (from above)
--model-nameyesJudge model (e.g. gpt-4o)
--templateyesPrompt with {variable} placeholders (single-quoted in bash)
--classification-choicesyesJSON object mapping choice labels to numeric scores e.g. '{"correct": 1, "incorrect": 0}'
--descriptionnoHuman-readable description
--include-explanationsnoInclude reasoning alongside the label
--use-function-callingnoPrefer structured function-call output
--invocation-paramsnoJSON of model params e.g. '{"temperature": 0}'
--data-granularitynospan (default), trace, or session. Only relevant for project tasks, not dataset/experiment tasks. See Data Granularity section.
--directionnoOptimization direction: maximize or minimize. Sets how the UI renders trends.
--provider-paramsnoJSON object of provider-specific parameters
Tasks

PROJECT_NAME, DATASET_NAME, and evaluator_id all accept a name or base64 ID.

bash
# List / Get
ax tasks list --space SPACE
ax tasks list --project PROJECT_NAME
ax tasks list --dataset DATASET_NAME --space SPACE
ax tasks get TASK_ID

# Create (project — continuous)
ax tasks create \
  --name "Correctness Monitor" \
  --task-type template_evaluation \
  --project PROJECT_NAME \
  --evaluators '[{"evaluator_id": "EVAL_ID", "column_mappings": {"input": "attributes.input.value", "output": "attributes.output.value"}}]' \
  --is-continuous \
  --sampling-rate 0.1

# Create (project — one-time / backfill)
ax tasks create \
  --name "Correctness Backfill" \
  --task-type template_evaluation \
  --project PROJECT_NAME \
  --evaluators '[{"evaluator_id": "EVAL_ID", "column_mappings": {"input": "attributes.input.value", "output": "attributes.output.value"}}]' \
  --no-continuous

# Create (experiment / dataset)
ax tasks create \
  --name "Experiment Scoring" \
  --task-type template_evaluation \
  --dataset DATASET_NAME --space SPACE \
  --experiment-ids "EXP_ID_1,EXP_ID_2" \   # base64 IDs from `ax experiments list --space SPACE -o json`
  --evaluators '[{"evaluator_id": "EVAL_ID", "column_mappings": {"output": "output"}}]' \
  --no-continuous

# Trigger a run (project task — use data window)
ax tasks trigger-run TASK_ID \
  --data-start-time "2026-03-20T00:00:00" \
  --data-end-time "2026-03-21T23:59:59" \
  --wait

# Trigger a run (experiment task — use experiment IDs)
ax tasks trigger-run TASK_ID \
  --experiment-ids "EXP_ID_1" \   # base64 ID from `ax experiments list --space SPACE -o json`
  --wait

# Monitor
ax tasks list-runs TASK_ID
ax tasks get-run RUN_ID
ax tasks wait-for-run RUN_ID --timeout 300
ax tasks cancel-run RUN_ID --force

Time format for trigger-run: 2026-03-21T09:00:00 — no trailing Z.

Additional trigger-run flags:

FlagDescription
--max-spansCap processed spans (default 10,000)
--override-evaluationsRe-score spans that already have labels
--wait / -wBlock until the run finishes
--timeoutSeconds to wait with --wait (default 600)
--poll-intervalPoll interval in seconds when waiting (default 5)

Run status guide:

StatusMeaning
completed, 0 spansThe eval index lags 1–2 hours — spans ingested recently may not be indexed yet. Shift the window to data at least 2 hours old, or widen the time range to cover more historical data.
cancelled ~1sIntegration credentials invalid
cancelled ~3minFound spans but LLM call failed — check model name or key
completed, N > 0Success — check scores in UI

Workflow A: Create an evaluator for a project

Use this when the user says something like "create an evaluator for my Playground Traces project".

Step 1: Confirm the project name

ax spans export accepts a project name directly — no ID lookup needed. If you don't know the project name, list available projects:

bash
ax projects list --space SPACE -o json

Find the entry whose "name" matches (case-insensitive) and use that name as PROJECT in subsequent commands. If you later hit a validation error with a name, fall back to using the project's "id" (a base64 string) instead.

Step 2: Understand what to evaluate

If the user specified the evaluator type (hallucination, correctness, relevance, etc.) → skip to Step 3.

If not, sample recent spans to base the evaluator on actual data:

bash
ax spans export PROJECT --space SPACE -l 10 --days 30 --stdout

Inspect attributes.input, attributes.output, span kinds, and any existing annotations. Identify failure modes (e.g. hallucinated facts, off-topic answers, missing context) and propose 1–3 concrete evaluator ideas. Let the user pick.

Each suggestion must include: the evaluator name (bold), a one-sentence description of what it judges, and the binary label pair in parentheses. Format each like:

  1. Name — Description of what is being judged. (label_a / label_b)

Example:

  1. Response Correctness — Does the agent's response correctly address the user's financial query? (correct / incorrect)
  2. Hallucination — Does the response fabricate facts not grounded in retrieved context? (factual / hallucinated)
Step 3: Confirm or create an AI integration
bash
ax ai-integrations list --space SPACE -o json

If a suitable integration exists, note its ID. If not, create one using the arize-ai-provider-integration skill. Ask the user which provider/model they want for the judge.

Step 4: Create the evaluator

Use the template design best practices below. Keep the evaluator name and variables generic — the task (Step 6) handles project-specific wiring via column_mappings.

bash
ax evaluators create \
  --name "Hallucination" \
  --space SPACE \
  --template-name "hallucination" \
  --commit-message "Initial version" \
  --ai-integration-id INT_ID \
  --model-name "gpt-4o" \
  --include-explanations \
  --use-function-calling \
  --classification-choices '{"factual": 1, "hallucinated": 0}' \
  --template 'You are an evaluator. Given the user question and the model response, decide if the response is factual or contains unsupported claims.

User question: {input}

Model response: {output}

Respond with exactly one of these labels: hallucinated, factual'
Step 5: Ask — backfill, continuous, or both?

Recommended approach: Always start with a small backfill (~100 historical spans) to validate the evaluator before turning on continuous monitoring. This lets you catch column mapping errors, wrong span kinds, and template issues on known data before scoring all future production spans. Only enable continuous after a backfill confirms correct scoring.

Before creating the task, ask:

"Would you like to: (a) Run a backfill on historical spans (one-time)? (b) Set up continuous evaluation on new spans going forward? (c) Both — backfill first to validate, then keep scoring new spans automatically? (recommended)"

Step 6: Determine column mappings from real span data

Do not guess paths. Pull a sample and inspect what fields are actually present:

bash
ax spans export PROJECT --space SPACE -l 5 --days 7 --stdout

For each template variable ({input}, {output}, {context}), find the matching JSON path. Common starting points — always verify on your actual data before using:

Template varLLM spanCHAIN span
inputattributes.input.valueattributes.input.value
outputattributes.llm.output_messages.0.message.contentattributes.output.value
contextattributes.retrieval.documents.contents—
tool_outputattributes.input.value (fallback)attributes.output.value

Validate span kind alignment: If the evaluator prompt assumes LLM final text but the task targets CHAIN spans (or vice versa), runs can cancel or score the wrong text. Make sure the query_filter on the task matches the span kind you mapped.

query_filter only works on indexed attributes: The query_filter in the evaluators JSON is evaluated against the eval index, not the raw span store. Attributes under attributes.metadata.* or custom keys may not be indexed and will silently match nothing. Use well-known indexed attributes like span_kind or attributes.llm.model_name for filtering. If a filter returns 0 spans despite data existing, try removing the filter as a diagnostic step.

Full example --evaluators JSON:

json
[
  {
    "evaluator_id": "EVAL_ID",
    "query_filter": "span_kind = 'LLM'",
    "column_mappings": {
      "input": "attributes.input.value",
      "output": "attributes.llm.output_messages.0.message.content",
      "context": "attributes.retrieval.documents.contents"
    }
  }
]

Include a mapping for every variable the template references. Omitting one causes runs to produce no valid scores.

Step 7: Create the task

Backfill only (a):

bash
ax tasks create \
  --name "Hallucination Backfill" \
  --task-type template_evaluation \
  --project PROJECT \
  --evaluators '[{"evaluator_id": "EVAL_ID", "column_mappings": {"input": "attributes.input.value", "output": "attributes.output.value"}}]' \
  --no-continuous

Continuous only (b):

bash
ax tasks create \
  --name "Hallucination Monitor" \
  --task-type template_evaluation \
  --project PROJECT \
  --evaluators '[{"evaluator_id": "EVAL_ID", "column_mappings": {"input": "attributes.input.value", "output": "attributes.output.value"}}]' \
  --is-continuous \
  --sampling-rate 0.1

Both (c): Use --is-continuous on create, then also trigger a backfill run in Step 8.

Step 8: Trigger a backfill run (if requested)

Eval index lag: The eval index is built asynchronously from the primary trace store and can lag 1–2 hours. For your first test run, use a time window ending at least 2 hours in the past. If you set --data-end-time to "now" on spans ingested in the last hour, the run will complete successfully but score 0 spans.

First find what time range has data:

bash
ax spans export PROJECT --space SPACE -l 100 --days 1 --stdout   # try last 24h first
ax spans export PROJECT --space SPACE -l 100 --days 7 --stdout   # widen if empty

Use the start_time / end_time fields from real spans to set the window. For the first validation run, cap --max-spans at ~100 to get quick feedback:

bash
ax tasks trigger-run TASK_ID \
  --data-start-time "2026-03-20T00:00:00" \
  --data-end-time "2026-03-21T23:59:59" \
  --max-spans 100 \
  --wait

Review scores and explanations before widening to the full backfill or enabling continuous.


Show full SKILL.md (1,238 more words)Show less

Workflow B: Create an evaluator for an experiment

Use this when the user says something like "create an evaluator for my experiment" or "evaluate my dataset runs".

If the user says "dataset" but doesn't have an experiment: A task must target an experiment (not a bare dataset). Ask:

"Evaluation tasks run against experiment runs, not datasets directly. Would you like help creating an experiment on that dataset first?"

If yes, use the arize-experiment skill to create one, then return here.

Step 1: Find the dataset and experiment names
bash
ax datasets list --space SPACE
ax experiments list --dataset DATASET_NAME --space SPACE -o json

Note the dataset name and the experiment name(s) to score. These accept names or IDs in subsequent commands — names are preferred.

Step 2: Understand what to evaluate

If the user specified the evaluator type → skip to Step 3.

If not, inspect a recent experiment run to base the evaluator on actual data:

bash
ax experiments export EXPERIMENT_NAME --dataset DATASET_NAME --space SPACE --stdout | python3 -c "import sys,json; runs=json.load(sys.stdin); print(json.dumps(runs[0], indent=2))"

Look at the output, input, evaluations, and metadata fields. Identify gaps (metrics the user cares about but doesn't have yet) and propose 1–3 evaluator ideas. Each suggestion must include: the evaluator name (bold), a one-sentence description, and the binary label pair in parentheses — same format as Workflow A, Step 2.

Step 3: Confirm or create an AI integration

Same as Workflow A, Step 3.

Step 4: Create the evaluator

Same as Workflow A, Step 4. Keep variables generic.

Step 5: Determine column mappings from real run data

Run data shape differs from span data. Inspect:

bash
ax experiments export EXPERIMENT_NAME --dataset DATASET_NAME --space SPACE --stdout | python3 -c "import sys,json; runs=json.load(sys.stdin); print(json.dumps(runs[0], indent=2))"

Common mapping for experiment runs:

  • output → "output" (top-level field on each run)
  • input → check if it's on the run or embedded in the linked dataset examples

If input is not on the run JSON, export dataset examples to find the path:

bash
ax datasets export DATASET_NAME --space SPACE --stdout | python3 -c "import sys,json; ex=json.load(sys.stdin); print(json.dumps(ex[0], indent=2))"
Step 6: Create the task
bash
ax tasks create \
  --name "Experiment Correctness" \
  --task-type template_evaluation \
  --dataset DATASET_NAME --space SPACE \
  --experiment-ids "EXP_ID" \   # base64 ID from `ax experiments list --space SPACE -o json`
  --evaluators '[{"evaluator_id": "EVAL_ID", "column_mappings": {"output": "output"}}]' \
  --no-continuous
Step 7: Trigger and monitor
bash
ax tasks trigger-run TASK_ID \
  --experiment-ids "EXP_ID" \   # base64 ID from `ax experiments list --space SPACE -o json`
  --wait

ax tasks list-runs TASK_ID
ax tasks get-run RUN_ID

Best Practices for Template Design

1. Use generic, portable variable names

Use {input}, {output}, and {context} — not names tied to a specific project or span attribute (e.g. do not use {attributes_input_value}). The evaluator itself stays abstract; the task's column_mappings is where you wire it to the actual fields in a specific project or experiment. This lets the same evaluator run across multiple projects and experiments without modification.

2. Default to binary labels

Use exactly two clear string labels (e.g. hallucinated / factual, correct / incorrect, pass / fail). Binary labels are:

  • Easiest for the judge model to produce consistently
  • Most common in the industry
  • Simplest to interpret in dashboards

If the user insists on more than two choices, that's fine — but recommend binary first and explain the tradeoff (more labels → more ambiguity → lower inter-rater reliability).

3. Be explicit about what the model must return

The template must tell the judge model to respond with only the label string — nothing else. The label strings in the prompt must exactly match the labels in --classification-choices (same spelling, same casing).

Good:

Respond with exactly one of these labels: hallucinated, factual

Bad (too open-ended):

Is this hallucinated? Answer yes or no.
4. Keep temperature low

Pass --invocation-params '{"temperature": 0}' for reproducible scoring. Higher temperatures introduce noise into evaluation results.

5. Use --include-explanations for debugging

During initial setup, always include explanations so you can verify the judge is reasoning correctly before trusting the labels at scale.

6. Pass the template in single quotes in bash

Single quotes prevent the shell from interpolating {variable} placeholders. Double quotes will cause issues:

bash
# Correct
--template 'Judge this: {input} → {output}'

# Wrong — shell may interpret { } or fail
--template "Judge this: {input} → {output}"
7. Always set --classification-choices to match your template labels

The labels in --classification-choices must exactly match the labels referenced in --template (same spelling, same casing). Omitting --classification-choices causes task runs to fail with "missing rails and classification choices."


Troubleshooting

ProblemSolution
ax: command not foundSee references/ax-setup.md
401 UnauthorizedAPI key may not have access to this space. Verify at https://app.arize.com/admin > API Keys
Evaluator not foundax evaluators list --space SPACE
Integration not foundax ai-integrations list --space SPACE
Task not foundax tasks list --space SPACE
project and dataset-id are mutually exclusiveUse only one when creating a task
experiment-ids required for dataset tasksAdd --experiment-ids to create and trigger-run
sampling-rate only valid for project tasksRemove --sampling-rate from dataset tasks
Validation error on ax spans exportProject name usually works; if you still get a validation error, look up the base64 project ID via ax projects list --space SPACE -o json and use the id field instead
Template validation errorsUse single-quoted --template '...' in bash; single braces {var}, not double {{var}}
Run stuck in pendingax tasks get-run RUN_ID; then ax tasks cancel-run RUN_ID
Run cancelled ~1sIntegration credentials invalid — check AI integration
Run cancelled ~3minFound spans but LLM call failed — wrong model name or bad key
Run completed, 0 spansWiden time window; eval index may not cover older data
No scores in UIFix column_mappings to match real paths on your spans/runs
Scores look wrongAdd --include-explanations and inspect judge reasoning on a few samples
Evaluator cancels on wrong span kindMatch query_filter and column_mappings to LLM vs CHAIN spans
Time format error on trigger-runUse 2026-03-21T09:00:00 — no trailing Z
Run failed: "missing rails and classification choices"Add --classification-choices '{"label_a": 1, "label_b": 0}' to ax evaluators create — labels must match the template
Run completed, all spans skippedQuery filter matched spans but column mappings are wrong or template variables don't resolve — export a sample span and verify paths
query_filter set but 0 spans scoredThe filter attribute may not be indexed in the eval index. attributes.metadata.* and custom attributes are often not indexed. Use span_kind or attributes.llm.model_name instead, or remove the filter to confirm spans exist in the window.
Diagnosing cancelled runs

When a task run is cancelled (status cancelled), follow this checklist in order:

1. Check integration credentials

bash
ax ai-integrations list --space SPACE -o json

Verify the integration ID used by the evaluator exists and has valid credentials. If the integration was deleted or the API key expired, the run cancels within ~1 second.

2. Verify the model name

bash
ax evaluators get EVALUATOR_NAME --space SPACE -o json

Check the model_name field. A typo or deprecated model causes the LLM call to fail and the run to cancel after ~3 minutes.

3. Export a sample span/run and compare paths to column_mappings

For project tasks:

bash
ax spans export PROJECT --space SPACE -l 1 --days 7 --stdout | python3 -m json.tool

For experiment tasks:

bash
ax experiments export EXPERIMENT_NAME --dataset DATASET_NAME --space SPACE --stdout | python3 -c "import sys,json; runs=json.load(sys.stdin); print(json.dumps(runs[0], indent=2)) if runs else print('No runs')"

Compare the exported JSON paths against the task's column_mappings. For each template variable, confirm the mapped path actually exists. Common mismatches:

  • Mapping output to attributes.output.value on an experiment run (should be just output)
  • Mapping input to attributes.input.value on a CHAIN span when the actual path is attributes.llm.input_messages
  • Mapping context to a path that doesn't exist on the span kind being filtered

4. Check that data_start_time is not epoch

If trigger-run used a start time of 0, 1970-01-01, or an empty string, the time window is invalid. Always derive from real span timestamps:

bash
ax spans export PROJECT --space SPACE -l 5 --days 30 --stdout | python3 -c "
import sys, json
spans = json.load(sys.stdin)
for s in spans:
    print(s.get('start_time', 'N/A'), s.get('end_time', 'N/A'))
"

5. Verify span kind matches evaluator scope

If the evaluator was created with --data-granularity trace but the task's query_filter is span_kind = 'LLM', the run may find no qualifying data and cancel. Ensure the granularity and filter are consistent.

6. Check that all template variables resolve

Every {variable} in the evaluator template must have a corresponding column_mappings entry that resolves to a non-null value. Test resolution against a real span:

bash
ax spans export PROJECT --space SPACE -l 3 --days 7 --stdout | python3 -c "
import sys, json
spans = json.load(sys.stdin)
# Replace these paths with your actual column_mappings values
mappings = {'input': 'attributes.input.value', 'output': 'attributes.output.value'}
for i, span in enumerate(spans):
    print(f'--- Span {i} ---')
    for var, path in mappings.items():
        parts = path.split('.')
        val = span
        for p in parts:
            val = val.get(p) if isinstance(val, dict) else None
        status = 'FOUND' if val else 'MISSING'
        print(f'  {var} ({path}): {status} — {str(val)[:80] if val else \"null\"}')
"

If any variable shows MISSING on all spans, fix the column mapping or adjust query_filter to target a different span kind.


  • arize-ai-provider-integration: Full CRUD for LLM provider integrations (create, update, delete credentials)
  • arize-trace: Export spans to discover column paths and time ranges
  • arize-experiment: Create experiments and export runs for experiment column mappings
  • arize-dataset: Export dataset examples to find input fields when runs omit them
  • arize-link: Deep links to evaluators and tasks in the Arize UI

Save Credentials for Future Use

See references/ax-profiles.md § Save Credentials for Future Use.

© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/arize-evaluator of github/awesome-copilot.

  • SKILL.md
  • references/ax-profiles.md
  • references/ax-setup.md

Open the folder on GitHubat commit 727ff2e

Used in 2 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in github/awesome-copilot, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Arize Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Arize Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Arize Evaluator this skillgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Arize Evaluator

What does Arize Evaluator do?

Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…. Arize Evaluator is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring.

When should I use Arize Evaluator?

Arize Evaluator fits situations like: the user mentions create evaluator; score experiment; continuous monitoring; improve evaluator prompt.

How do I install Arize Evaluator in Claude Code?

Run `npx skills add github/awesome-copilot --skill arize-evaluator -a claude-code`. Or copy the skill folder (skills/arize-evaluator in github/awesome-copilot) into .claude/skills/arize-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Arize Evaluator in Codex?

Run `npx skills add github/awesome-copilot --skill arize-evaluator -a codex`. Or copy the skill folder (skills/arize-evaluator in github/awesome-copilot) into .agents/skills/arize-evaluator in your project. Codex loads it when a task matches its description.

Can I use Arize Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill arize-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arize-evaluator, .gemini/skills/arize-evaluator, .github/skills/arize-evaluator and .opencode/skills/arize-evaluator in your project.

What does Arize Evaluator need to run?

Going by SKILL.md and its folder, Arize Evaluator needs the command-line tools its instructions call (python3) and credentials named OPENAI_API_KEY and ANTHROPIC_API_KEY. Our summary lists: A credential in OPENAI_API_KEY; A credential in ANTHROPIC_API_KEY. Compatibility (from SKILL.md): Requires the ax CLI and a configured Arize profile with an AI integration..

Does Arize Evaluator access the network?

SKILL.md names 1 domain. As links in the text: arize.com. This is read from the text; nothing was executed.

Is Arize Evaluator safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Arize Evaluator use?

Arize Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Arize Evaluator use?

About 8.1k tokens (SKILL.md is roughly 32k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Arize Evaluator?

Skills that share tags, products or a category with Arize Evaluator: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Arize Evaluator?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.