Agent skill

LLM Error Discovery Review App

by ai-evals-course in ai-evals-course/evals-skills

Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate.

Apache-2.0Auto-check passedAI & LLM Engineering

Install LLM Error Discovery Review App

skills CLI
$ npx skills add ai-evals-course/evals-skills --skill error-discovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-evals-course/evals-skills error-discovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/error-discovery .claude/skills/error-discovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
error-discovery
GitHub stars
1.5k
Token cost
~3.7k tokens
SKILL.md length
2,210 words
Files
3
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate.

  • Works in 5 steps: Understand the domain and data → Design the visual encoding → Build the review interface → …
  • Reviewing a sample of LLM outputs or agent traces to find recurring failures
  • SKILL.md covers Progress updates, Phase 1: Understand the domain…, Phase 2: Design the visual… and Phase 3: Build the review…, plus 2 more sections
  • Calls codex and claude

What it does

Phases one to four live in the main file. The agent loads your JSONL, CSV or JSON data, inspects a handful of records and decides which shape it has: a single text field, input and output pairs, multi-turn traces with roles and tool calls, code, structured output or a mix. It then designs and builds a review app suited to that shape and selects diverse samples, for example by clustering on a few features.

The interactive review itself is phase five, described in a separate `review-loop.md` that the agent follows once the app is running and you start annotating. The skill is meant for interactive sessions only and tells the agent to narrate what it is doing at every step, so there are no long silent stretches while the interface is built.

When your agent uses it

  • Reviewing a sample of LLM outputs or agent traces to find recurring failures
  • Setting up a labeling interface before writing evals
  • Choosing which records to read when a dataset is too large to review in full

Example prompts

  • “Run error analysis on traces.jsonl and build me a review UI.”
  • “Pick a diverse set of records from this chatbot log so I can annotate failure modes.”
  • “Help me sort the failures in these summarization outputs into categories.”

Requirements

  • A dataset of LLM outputs or traces in JSONL, CSV or JSON form

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Understand the domain and data
  2. Design the visual encoding
  3. Build the review interface
  4. Cluster and select initial samples
  5. Run the interactive review loop

What it can do on your machine

Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • codex
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Error Discovery Review App loads about 3.7k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 2,210 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 2,210 words, ~3,732 tokens.

Download SKILL.mdSave it as .claude/skills/error-discovery/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
error-discovery
description
Run error analysis on a dataset. Build a review UI, select diverse samples, monitor annotations, and organize failure modes.

Error Discovery Skill

You are running an interactive error analysis session. The user has a dataset (JSONL, CSV, JSON, etc.) of LLM outputs or traces and wants to discover failure modes by reviewing samples.

This skill is meant for interactive sessions only.

This skill has two parts:

  • This file covers phases 1 through 4. You read the data, design the UI, build it, and select samples.
  • review-loop.md covers phase 5. You run the interactive review session.

Read this file first to build everything. Once the app is running and the human starts reviewing, follow review-loop.md.

Progress updates

Each phase takes time, especially building the interface. Tell the user what you are doing at each step. Before starting a phase, say what you are about to do and why. When a step finishes, say what you did and what comes next. For example: "Reading 10 sample records to understand the data shape", "Building the HTML app with three views: article, map, and progress", "Clustering on 5 features to select diverse samples." Do not go silent for long stretches.

Phase 1: Understand the domain and data

Before building anything, study the dataset thoroughly.

1a: Read and inventory the data

Load the file. Examine 5 to 10 records across the distribution. For each record, identify:

  • All fields, their types, and what they represent
  • Which fields are the primary content the human needs to judge
  • Which fields are metadata (context for understanding, but not the thing being judged)
  • Which fields vary across records and which are constant
1b: Identify the content structure

The data could take many forms. Determine which pattern fits:

  • Single text field, e.g., article, summary, email, essay, translation
  • Input/output pair, e.g., prompt + completion, question + answer, instruction + result
  • Multi-turn trace, e.g., a sequence of messages with different roles (system, human, assistant, tool calls, tool results, thinking/reasoning blocks). This is common for agent traces, chatbot logs, and agentic workflows.
  • Code, e.g., source files, patches, diffs, with or without surrounding context
  • Structured output, e.g., JSON, function calls, extracted entities, classifications
  • Composite, e.g., a task description + an agent trace + a final output

For multi-turn traces specifically, identify:

  • What roles or authors exist (system, user, assistant, tool, thinking, etc.)
  • Whether there are tool call / tool result pairs
  • Whether there are thinking/reasoning blocks
  • What the logical grouping of turns is, e.g., one "step" = thinking + tool call + tool result
1c: Identify dimensions of variation

Think about what differs between data points and within each data point. You will use these dimensions to decide the visual design.

Between data points (vary across records):

  • Metadata, e.g., topic, model, difficulty, source, label, task type
  • Structural, e.g., length, number of turns, number of tool calls
  • Outcome, e.g., success/failure, score

Within each data point (vary across parts of one record):

  • Author/role of each segment (human vs agent vs system vs tool)
  • Content type of each segment (natural language vs code vs JSON vs thinking)
  • Importance/relevance of each segment (boilerplate system prompt vs the actual response)
1d: Think about what "bad" means

Think of a few plausible failure categories in this domain. Expect the human to discover most of them during review.

Phase 2: Design the visual encoding

Before writing any code, design how every dimension of variation maps to a visual property. Use Gestalt principles and information visualization fundamentals.

Core Gestalt principles to apply
  • Similarity (color, shape). Viewers perceive things that share a visual property as related. Use this for categorical dimensions. Give different message roles different colors, and different labels different badge colors.
  • Proximity (spacing). Viewers perceive things that are close together as grouped. Use this for logical units. Place turns in a conversation step close together, and separate steps by more space.
  • Common region (containers, backgrounds). Viewers perceive things inside a shared boundary as grouped. Use this for multi-part content, e.g., a tool call and its result share a container.
  • Figure/ground (opacity, contrast). Important content should be high-contrast (figure). Less important content should recede (ground). Use this for emphasis. Show boilerplate at low opacity.
Visual encoding rules

Map each dimension of variation to exactly one visual channel. Do not use the same channel for two different things.

Color hue. Use for the categorical dimension with the most important distinction.

  • For multi-turn traces, use color for message role. Each role (human, assistant, system, tool, thinking) gets its own distinct hue at full saturation. Pick 3 to 5 easily distinguishable hues.
  • For single-content, use color for label or source category.
  • Do not use color hue for quantitative data.

Opacity / saturation. Use to show whether the human needs to read this content carefully.

  • Reduce opacity only for content the human genuinely does not need to read, e.g., repeated boilerplate that is identical across records, verbose tool schemas, auto-generated headers. Ask yourself: "if I removed this, would the human miss anything?"
  • Do not mute content by role. System messages, tool results, and thinking blocks can all contain the actual bug. Mute only specific content that is redundant or mechanical, regardless of which role produced it.
  • Normal content gets full opacity.

Spacing. Use for hierarchical structure.

  • Tight spacing (4 to 8px) between items within a logical group, e.g., consecutive messages in one "step" of an agent trace.
  • Medium spacing (16 to 24px) between logical groups, e.g., between steps.
  • Large spacing (32 to 48px) between major sections.

Typography. Use for content type.

  • Natural language prose: proportional font, normal size.
  • Code: monospace font, slightly smaller.
  • Metadata: small, muted, compact.
  • Thinking/reasoning: italic or a distinct font treatment to signal "internal."

Border / container. Use for grouping related parts.

  • Put a tool call and its tool result in a shared container with a subtle border.
  • Separate an input/output pair with a clear divider, or place them side by side.

Structural outlier flags. Show in the header only, not inline.

  • Pre-compute dataset-level averages for key structural features (length, heading count, paragraph count, etc.).
  • For each record, flag dimensions where it is a clear statistical outlier (e.g., top/bottom 10%).
  • Show these as small, compact badges in the header, e.g., "6 headings (more than 89%)" or "shorter than 95%."
  • Keep it minimal. Most records should have zero or one flag. Do not annotate inline content.

Phase 3: Build the review interface

Architecture
  • Python HTTP server (stdlib http.server, no dependencies):
    • GET / serves the HTML app
    • GET /api/samples returns the current sample set
    • POST /api/samples lets the agent push new samples
    • GET /api/annotations returns the current annotations
    • POST /api/annotations lets the app save annotations on every change
    • GET /api/graph returns the 2D projection of all records for the cluster map
    • GET /api/patterns returns the agent's current failure mode taxonomy
    • POST /api/patterns lets the agent push the updated taxonomy
    • GET /api/suggestions returns agent-suggested annotations
    • POST /api/suggestions lets the agent push suggestions
  • On-disk files in an error_discovery_data/ directory:
    • samples.json, annotations.json, graph.json, patterns.json, suggestions.json
  • The HTML app auto-saves to the server on every annotation. It polls for new samples and suggestions.
Show full SKILL.md (1,068 more words)Show less
HTML app structure

Three views, toggled from the top bar:

  1. Article/content view. The main review interface where the human reads and annotates.
  2. Map view. A 2D scatter plot (PCA or UMAP projection) of all records. It shows clusters, which items are in the sample, and which have been annotated. The reviewer can click a sample node to go to its content view.
  3. Progress view. Two sections:
    • Failure modes: a treemap of modes the agent has categorized so far. Each block is a failure mode, sized by annotation count. Inside each block, list the notes. The reviewer can click a note to go to that annotation in the content view.
    • Agent suggestions to review: a list of pending suggestions with checkboxes. Each row shows the failure mode, quoted text, and source record. The reviewer can click the text to go to that spot in the article. At the top, a "Select all" checkbox and "Accept selected" / "Dismiss selected" buttons. This lets the reviewer select all, uncheck the few they disagree with, and accept the rest in one click.

Content view design. Apply the visual encoding from Phase 2:

  • Header: title + label + topic. Keep it minimal. Add structural outlier flags (from Phase 2) only when the record is a genuine outlier.

  • Body: render the primary content using the appropriate treatment per content type.

    For multi-turn traces (agent logs, conversations, chat):

    • Each message/turn is a block. Left-align all blocks but use a colored left-border or background tint per role.
    • Role label (small, bold) at the top of each block, e.g., "System", "User", "Assistant", "Tool Call", "Tool Result", "Thinking".
    • Assign a distinct hue to each role. Be consistent across all records. All roles at full opacity by default.
    • Render thinking/reasoning blocks in a visually distinct way (e.g., lighter background, italic, or slightly indented) to show that they are internal monologue, while keeping full readability.
    • Show tool call function names prominently. Put parameters in collapsible formatted JSON.
    • Make tool results collapsible by default if they are long, with a summary line visible.
    • Only reduce opacity for content that is literally identical across records, e.g., the same system prompt repeated verbatim in every trace. If system messages vary, keep them fully visible.
    • Group related turns: a thinking block + the tool call it produces + the tool result, visually grouped with tight spacing and a shared container.

    For single text content (articles, summaries, etc.):

    • Render markdown as formatted HTML (use marked.js or similar).
    • Render plain text with paragraph breaks.

    For code / diffs:

    • Use syntax highlighting (highlight.js or Prism via CDN).
    • Show additions with green background and deletions with red background.
    • Show line numbers.

    For input/output pairs:

    • Stack them with a clear divider, or place them side by side if both are short.
    • Label each section.
  • Inline annotation: the reviewer selects text, a floating popover appears with a text input, they press Enter to save, and the span is highlighted. When the popover appears and the input is focused, the browser clears the native text selection. To prevent this, wrap the selected range in a temporary highlight span (e.g., class "pending-highlight" with a visible background) BEFORE focusing the input. Remove the temporary highlight when the annotation is saved, cancelled, or the popover is dismissed by clicking outside. This way the reviewer always sees what text they are annotating.

  • Margin notes: annotations and suggestions must appear as side notes in a right margin column, aligned vertically with their corresponding highlighted text. Use a two-column layout: the article body on the left (flex: 1, max-width ~720px) and a margin-notes column on the right (width ~240px). Each margin note is position: absolute inside the margin column, with its top offset calculated from the highlight element's position relative to the margin container (use getBoundingClientRect on both the highlight and the margin container, take the difference — do NOT add scrollTop, as that double-counts the scroll offset). Stack notes with a minimum gap so they do not overlap. Include hover linking: hovering a margin note outlines its highlight, and hovering a highlight outlines its margin note. Each margin note shows the quoted text, the reviewer's note, and edit/delete buttons (visible on hover). Do NOT use hover tooltips as the primary way to show annotation content — margin notes replace tooltips.

  • Agent suggestions: visually distinct from human annotations in both the inline highlight (dashed border, muted tint) and the margin note (different left-border color, an "agent suggestion" tag). The margin note shows accept/dismiss buttons (always visible, not hover-gated). The reviewer can accept (which promotes it to an annotation) or dismiss.

  • No quality labels, no dropdowns, no structured forms. Free-text notes only.

  • Auto-save to server on every change. Keep a localStorage backup.

  • Poll for new samples and suggestions periodically. Show a banner or toast when new ones arrive.

Map view:

  • 2D scatter of all records from PCA/UMAP projection.
  • Color by cluster (match hull colors), not by label.
  • Use shape to distinguish categories, e.g., circles for AI and squares for human.
  • Show sample items as larger nodes with a dark border.
  • Show annotated items in a distinct color, e.g., orange.
  • Draw cluster hulls or convex boundaries as subtle background shapes.
  • On hover, show a tooltip with title, metadata, and annotation count.
  • On click (sample nodes only), go to the content view for that item.

Phase 4: Cluster and select initial samples

  1. Extract features appropriate to the content type (see Phase 1c for the dimensions).
  2. Cluster using KMeans or similar on normalized features. Target 6 to 10 clusters.
  3. Build the initial sample (15 to 25 items for datasets over 50):
    • Cluster representatives (about 60 to 70%): 1 to 2 items closest to each centroid, mixing categories/labels.
    • Random samples (about 30 to 40%): from the full dataset regardless of cluster. The clustering may not capture every important dimension, so random picks help cover what it misses.
    • Remove duplicates.
  4. Prioritize diversity. The goal is discovering failure modes, not estimating how common they are.

Phase 5: Run the interactive review loop

Once the app is running and the human starts reviewing, follow review-loop.md for the ongoing interactive session, including how to monitor annotations as they arrive.

If the session ends before a human can review (a non-interactive run, for example codex exec or claude -p), build and smoke-test the app, then stop the server and give the command to launch it later; never claim the review loop ran or that anything is still running.

© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/error-discovery of ai-evals-course/evals-skills.

  • SKILL.md
  • agents/openai.yaml
  • review-loop.md

Open the folder on GitHubat commit 80d5f7b

Compare with similar skills

LLM Error Discovery Review App next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Error Discovery Review App compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Error Discovery Review App this skillai-evals-course/evals-skills1.5k—~3.7kAutomated safety check: PassApache-2.0
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0
Hugging Face Dataset Viewerhuggingface/skills11k3 repos~1.1kAutomated safety check: PassApache-2.0
Audit Sft Data Qualitytokenbender/agent-guides367—~2.7kAutomated safety check: PassApache-2.0
VLM BCQ Gap AnalysisNVIDIA/skills3.5k—~1.3kAutomated safety check: NotesApache-2.0
Mcaf ML AI Deliverymanagedcode/Storage138—~1kAutomated safety check: PassMIT

Similar skills

  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Dataset Viewer

    huggingface/skills

    Official

    Explores Hugging Face datasets through the read-only Dataset Viewer API: list splits, preview and page through rows, search, filter, and fetch parquet links and statistics.

    11k GitHub starsUsed in 3 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Audit Sft Data Quality

    tokenbender/agent-guides

    Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.

    367 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report.

    3.5k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Mcaf ML AI Delivery

    managedcode/Storage

    Apply ML/AI project delivery guidance for data exploration, feasibility, experimentation, testing, responsible AI, and operating ML systems.

    138 GitHub stars~1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed

More from ai-evals-course/evals-skills

All 9 skills in this repo
  • LLM Trace Review Interface

    ai-evals-course/evals-skills

    Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

    1.5k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • LLM Eval Pipeline Audit

    ai-evals-course/evals-skills

    Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

    1.5k GitHub stars~2.5k tokensUpdated 14 days ago
    Auto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 14 days ago
    Auto-check passed
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • LLM Judge Validation

    ai-evals-course/evals-skills

    Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

    1.5k GitHub stars~2.2k tokensUpdated 14 days ago
    Auto-check passed
  • LLM-as-Judge Prompt Writer

    ai-evals-course/evals-skills

    Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

    1.5k GitHub stars~1.9k tokensUpdated 14 days ago
    Auto-check passed

Questions about LLM Error Discovery Review App

What does LLM Error Discovery Review App do?

Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate. Phases one to four live in the main file. The agent loads your JSONL, CSV or JSON data, inspects a handful of records and decides which shape it has: a single text field, input and output pairs, multi-turn traces with roles and tool calls, code, structured output or a mix.

When should I use LLM Error Discovery Review App?

LLM Error Discovery Review App fits situations like: reviewing a sample of LLM outputs or agent traces to find recurring failures; setting up a labeling interface before writing evals; choosing which records to read when a dataset is too large to review in full.

How do I install LLM Error Discovery Review App in Claude Code?

Run `npx skills add ai-evals-course/evals-skills --skill error-discovery -a claude-code`. Or copy the skill folder (skills/error-discovery in ai-evals-course/evals-skills) into .claude/skills/error-discovery in your project. Claude Code loads it when a task matches its description.

How do I install LLM Error Discovery Review App in Codex?

Run `npx skills add ai-evals-course/evals-skills --skill error-discovery -a codex`. Or copy the skill folder (skills/error-discovery in ai-evals-course/evals-skills) into .agents/skills/error-discovery in your project. Codex loads it when a task matches its description.

Can I use LLM Error Discovery Review App in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill error-discovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/error-discovery, .gemini/skills/error-discovery, .github/skills/error-discovery and .opencode/skills/error-discovery in your project.

What does LLM Error Discovery Review App need to run?

Going by SKILL.md and its folder, LLM Error Discovery Review App needs the command-line tools its instructions call (codex and claude). Our summary lists: A dataset of LLM outputs or traces in JSONL, CSV or JSON form.

Does LLM Error Discovery Review App access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is LLM Error Discovery Review App safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Error Discovery Review App use?

LLM Error Discovery Review App is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Error Discovery Review App use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Error Discovery Review App?

Skills that share tags, products or a category with LLM Error Discovery Review App: Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 792 stars), Hugging Face Dataset Viewer (huggingface/skills, 11k stars), Audit Sft Data Quality (tokenbender/agent-guides, 367 stars) and VLM BCQ Gap Analysis (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Error Discovery Review App?

ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,472 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.

Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.