Agent skill

LLM Trace Review Interface

by ai-evals-course in ai-evals-course/evals-skills

Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

Apache-2.0Auto-check passedAI & LLM Engineering

Install LLM Trace Review Interface

skills CLI
$ npx skills add ai-evals-course/evals-skills --skill build-review-interface -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-evals-course/evals-skills build-review-interface --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/build-review-interface .claude/skills/build-review-interface && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
build-review-interface
GitHub stars
1.5k
Token cost
~1.4k tokens
SKILL.md length
718 words
Files
2
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

  • Works in 9 steps: Load the app and verify traces are… → Click Pass on a trace, verify the label… → Click Fail on a trace, add a note,… → …
  • Building a custom annotation tool to label LLM traces
  • SKILL.md covers Overview, Data Display, Feedback Collection and Navigation and Status, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill has the agent write an HTML page that loads traces from a JSON or CSV file and shows one at a time, with Pass and Fail buttons, a free-text notes field, a Defer button for unsure cases and Next and Previous navigation. Labels are saved to a local CSV, SQLite or JSON file and auto-saved on every action.

Most of the guidance is about display. Render emails like emails, code with highlighting and markdown as markdown, collapse repetitive parts such as a shared system prompt, surface key metadata as badges, color-code messages by role and keep the full trace reachable with intermediate steps collapsed. Model output is sanitized by stripping raw HTML and disabling images. Annotation happens at the trace level, and failure-category tags come later, after error analysis.

When your agent uses it

  • Building a custom annotation tool to label LLM traces
  • Collecting human pass/fail judgments for an eval dataset
  • Replacing raw trace dumps with a readable review page

Example prompts

  • “Build a review page for the traces in data/support_traces.json with pass/fail buttons and notes.”
  • “Make an annotation interface for our chatbot logs where tool calls are collapsed and labels save to CSV.”
  • “Create a trace reviewer for our email-drafting agent that renders emails like real emails.”

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Load the app and verify traces are displayed
  2. Click Pass on a trace, verify the label is saved
  3. Click Fail on a trace, add a note, verify both are saved
  4. Click Defer, verify it is recorded
  5. Navigate forward and backward with buttons and keyboard shortcuts
  6. Verify the trace counter updates correctly
  7. Verify auto-save by reloading the page and checking labels persist
  8. Expand collapsed sections (system prompts, tool calls) and verify content is accessible
  9. Test that all keyboard shortcuts trigger the correct actions

What it can do on your machine

Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Trace Review Interface loads about 1.4k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 718 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 718 words, ~1,352 tokens.

Download SKILL.mdSave it as .claude/skills/build-review-interface/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
build-review-interface
description
Build a custom browser-based annotation interface tailored to your data for reviewing LLM traces and collecting structured feedback. Use when you need to build an annotation tool, review traces, or collect human labels.

Build a Custom Annotation Interface

Overview

Build an HTML page that loads traces from a data source (JSON/CSV file), displays one trace at a time with Pass/Fail buttons, a free-text notes field, and Next/Previous navigation. Save labels to a local file (CSV/SQLite/JSON). Then customize to the domain using the guidelines below.

Data Display

Format all data in the most human-readable representation for the domain. Emails should look like emails. Code should have syntax highlighting. Markdown should be rendered. Tables should be tables. JSON should be pretty-printed and collapsible.

  • Collapse repetitive elements. If every trace shares the same system prompt, put it in a <details> toggle.
  • Extract and surface key metadata. If traces contain a property name, client type, or session ID buried in the data, extract it and display it prominently as a header or badge.
  • Color-code by role or status. Use left-border colors to distinguish user messages, assistant messages, tool calls, and system prompts at a glance.
  • Group related elements visually. Tool calls and their responses should be visually linked (indentation, shared border).
  • Collapse what doesn't help judgment. Verbose tool response JSON, intermediate reasoning steps, and debugging context go behind toggles.
  • Highlight what matters most. Make the primary content reviewers judge visually dominant. Bold key entities (prices, dates, names). Use font size and spacing to create hierarchy.
  • Show the full trace. Include all intermediate steps (tool calls, retrieved context, reasoning), not just the final output. Collapse them by default but keep them accessible.
  • Sanitize rendered content. Strip raw HTML from LLM outputs before rendering. Disable images in rendered markdown if they could be tracking pixels.

Feedback Collection

Annotate at the trace level. The reviewer judges the whole trace, not individual spans.

  • Binary Pass/Fail buttons as the primary action.
  • Free-text notes field for the reviewer to describe what went wrong (or right).
  • Defer button for uncertain cases.
  • Auto-save on every action.

Once you have established failure categories from error analysis, you can later add predefined failure mode tags as clickable checkboxes, dropdowns or picklists so reviewers can select from known categories in addition to writing notes. But don't add these in the initial build.

Navigation and Status

  • Next/Previous buttons and keyboard arrow keys.
  • Trace counter showing position and progress ("12 of 87 remaining").
  • Jump to specific trace by ID.
  • Counts of labeled vs unlabeled traces.

Keyboard Shortcuts

Arrow keys = Navigate traces
1 = Pass              2 = Fail
D = Defer             U = Undo last action
Cmd+S = Save          Cmd+Enter = Save and next

Selecting Traces to Load

Build the app to accept traces from any source (JSON/CSV file). Keep sampling logic outside the app in a separate script. Start with random sampling.

Show full SKILL.md (302 more words)Show less

Additional Features

Reference panel: Toggle-able panel showing ground truth, expected answers, or rubric definitions alongside the trace.

Filtering: Filter traces by metadata dimensions relevant to the product (channel, user type, pipeline version).

Clustering: Group traces by metadata or semantic similarity. Show representative traces per cluster with drill-down.

Design Checklist

  • Same layout, controls, and terminology on every trace
  • Pass and Fail buttons are visually distinct (color, size)
  • Keyboard shortcuts work for all primary actions
  • Full trace accessible even when sections are collapsed
  • Labels persist automatically without explicit save
  • Trace-level annotation (not span-level) as the default
  • All data rendered in its native format (markdown as HTML, code with highlighting, JSON pretty-printed, tables as HTML tables, URLs as clickable links)

Testing

After building the interface, verify it with Playwright.

Visual review: Take screenshots of the interface with representative trace data loaded. Review each screenshot for:

  • Layout and spacing: is the visual hierarchy clear? Can you immediately see what matters?
  • Readability: is all data rendered in its native format? Are there any raw JSON blobs, unrendered markdown, or unstyled content?
  • Aesthetics: does the interface look professional and clean? Would a domain expert use this?
  • Responsiveness: does the layout hold at different window sizes?

Functional test: Write a Playwright script that performs a full annotation workflow:

  1. Load the app and verify traces are displayed
  2. Click Pass on a trace, verify the label is saved
  3. Click Fail on a trace, add a note, verify both are saved
  4. Click Defer, verify it is recorded
  5. Navigate forward and backward with buttons and keyboard shortcuts
  6. Verify the trace counter updates correctly
  7. Verify auto-save by reloading the page and checking labels persist
  8. Expand collapsed sections (system prompts, tool calls) and verify content is accessible
  9. Test that all keyboard shortcuts trigger the correct actions

© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/build-review-interface of ai-evals-course/evals-skills.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 80d5f7b

Compare with similar skills

LLM Trace Review Interface next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Trace Review Interface compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Trace Review Interface this skillai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0
Phoenix LLM ObservabilityOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Phoenix Evals New MetricArize-ai/phoenix12k—~2.3kAutomated safety check: PassApache-2.0
Phoenix Error AnalysisArize-ai/phoenix12k—~6.4kAutomated safety check: PassApache-2.0
Phoenix CLIgithub/awesome-copilot40k1 repos~4kAutomated safety check: PassApache-2.0
DatasetsArize-ai/phoenix12k—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Phoenix LLM Observability

    Orchestra-Research/AI-Research-SKILLs

    Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Phoenix Evals New Metric

    Arize-ai/phoenix

    Create a new built-in classification evaluator for Phoenix evals.

    12k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Phoenix Error Analysis

    Arize-ai/phoenix

    Find out what is going wrong in LLM or agent traffic by reading sampled Phoenix traces, spans, or sessions, writing free-form notes (open coding), then grouping the notes into a few narrow…

    12k GitHub stars~6.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Phoenix CLI

    github/awesome-copilot

    Official

    Debug LLM applications using the Phoenix CLI. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~4k tokens
    AI & LLM EngineeringAuto-check passed
  • Datasets

    Arize-ai/phoenix

    Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments.

    12k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Error Analysis

    yonatangross/orchestkit

    Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes.

    290 GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from ai-evals-course/evals-skills

All 9 skills in this repo
  • LLM Eval Pipeline Audit

    ai-evals-course/evals-skills

    Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

    1.5k GitHub stars~2.5k tokensUpdated 15 days ago
    Auto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 15 days ago
    Auto-check passed
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 15 days ago
    Auto-check passed
  • LLM Judge Validation

    ai-evals-course/evals-skills

    Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

    1.5k GitHub stars~2.2k tokensUpdated 15 days ago
    Auto-check passed
  • LLM-as-Judge Prompt Writer

    ai-evals-course/evals-skills

    Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

    1.5k GitHub stars~1.9k tokensUpdated 15 days ago
    Auto-check passed
  • Write Code Eval

    ai-evals-course/evals-skills

    Write code evaluators for known failure modes with objective rules.

    1.5k GitHub stars~385 tokensUpdated 15 days ago
    Auto-check passed

Works with

Questions about LLM Trace Review Interface

What does LLM Trace Review Interface do?

Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data. The skill has the agent write an HTML page that loads traces from a JSON or CSV file and shows one at a time, with Pass and Fail buttons, a free-text notes field, a Defer button for unsure cases and Next and Previous navigation. Labels are saved to a local CSV, SQLite or JSON file and auto-saved on every action.

When should I use LLM Trace Review Interface?

LLM Trace Review Interface fits situations like: building a custom annotation tool to label LLM traces; collecting human pass/fail judgments for an eval dataset; replacing raw trace dumps with a readable review page.

How do I install LLM Trace Review Interface in Claude Code?

Run `npx skills add ai-evals-course/evals-skills --skill build-review-interface -a claude-code`. Or copy the skill folder (skills/build-review-interface in ai-evals-course/evals-skills) into .claude/skills/build-review-interface in your project. Claude Code loads it when a task matches its description.

How do I install LLM Trace Review Interface in Codex?

Run `npx skills add ai-evals-course/evals-skills --skill build-review-interface -a codex`. Or copy the skill folder (skills/build-review-interface in ai-evals-course/evals-skills) into .agents/skills/build-review-interface in your project. Codex loads it when a task matches its description.

Can I use LLM Trace Review Interface in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill build-review-interface -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-review-interface, .gemini/skills/build-review-interface, .github/skills/build-review-interface and .opencode/skills/build-review-interface in your project.

What does LLM Trace Review Interface need to run?

SKILL.md names no scripts, command-line tools or credentials: LLM Trace Review Interface is instructions for the agent only.

Does LLM Trace Review Interface access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is LLM Trace Review Interface safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Trace Review Interface use?

LLM Trace Review Interface is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Trace Review Interface use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Trace Review Interface?

Skills that share tags, products or a category with LLM Trace Review Interface: Phoenix LLM Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars), Phoenix Evals New Metric (Arize-ai/phoenix, 12k stars), Phoenix Error Analysis (Arize-ai/phoenix, 12k stars) and Phoenix CLI (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Trace Review Interface?

ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,472 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.

Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.