Quality Flywheel
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate.
$ npx skills add ai-evals-course/evals-skills --skill error-discovery -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ai-evals-course/evals-skills error-discovery --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/error-discovery .claude/skills/error-discovery && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "error-discovery" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/error-discovery into .claude/skills/error-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-discovery", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ai-evals-course/evals-skills/tree/main/skills/error-discoveryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ai-evals-course/evals-skills --skill error-discovery -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ai-evals-course/evals-skills error-discovery --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/error-discovery .agents/skills/error-discovery && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "error-discovery" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/error-discovery into .agents/skills/error-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-discovery", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-evals-course/evals-skills --skill error-discovery -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ai-evals-course/evals-skills error-discovery --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/error-discovery .cursor/skills/error-discovery && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "error-discovery" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/error-discovery into .cursor/skills/error-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-discovery", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ai-evals-course/evals-skills.git --path skills/error-discovery--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ai-evals-course/evals-skills --skill error-discovery -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ai-evals-course/evals-skills error-discovery --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/error-discovery .gemini/skills/error-discovery && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "error-discovery" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/error-discovery into .gemini/skills/error-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-discovery", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ai-evals-course/evals-skills error-discoveryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ai-evals-course/evals-skills --skill error-discovery -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/error-discovery .github/skills/error-discovery && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "error-discovery" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/error-discovery into .github/skills/error-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-discovery", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-evals-course/evals-skills --skill error-discovery -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ai-evals-course/evals-skills error-discovery --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/error-discovery .opencode/skills/error-discovery && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "error-discovery" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/error-discovery into .opencode/skills/error-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-discovery", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
error-discoveryGuides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate.
Phases one to four live in the main file. The agent loads your JSONL, CSV or JSON data, inspects a handful of records and decides which shape it has: a single text field, input and output pairs, multi-turn traces with roles and tool calls, code, structured output or a mix. It then designs and builds a review app suited to that shape and selects diverse samples, for example by clustering on a few features.
The interactive review itself is phase five, described in a separate `review-loop.md` that the agent follows once the app is running and you start annotating. The skill is meant for interactive sessions only and tells the agent to narrate what it is doing at every step, so there are no long silent stretches while the interface is built.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
codexclaudeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
LLM Error Discovery Review App loads about 3.7k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 2,210 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 2,210 words, ~3,732 tokens.
.claude/skills/error-discovery/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.You are running an interactive error analysis session. The user has a dataset (JSONL, CSV, JSON, etc.) of LLM outputs or traces and wants to discover failure modes by reviewing samples.
This skill is meant for interactive sessions only.
This skill has two parts:
Read this file first to build everything. Once the app is running and the human starts reviewing, follow review-loop.md.
Each phase takes time, especially building the interface. Tell the user what you are doing at each step. Before starting a phase, say what you are about to do and why. When a step finishes, say what you did and what comes next. For example: "Reading 10 sample records to understand the data shape", "Building the HTML app with three views: article, map, and progress", "Clustering on 5 features to select diverse samples." Do not go silent for long stretches.
Before building anything, study the dataset thoroughly.
Load the file. Examine 5 to 10 records across the distribution. For each record, identify:
The data could take many forms. Determine which pattern fits:
For multi-turn traces specifically, identify:
Think about what differs between data points and within each data point. You will use these dimensions to decide the visual design.
Between data points (vary across records):
Within each data point (vary across parts of one record):
Think of a few plausible failure categories in this domain. Expect the human to discover most of them during review.
Before writing any code, design how every dimension of variation maps to a visual property. Use Gestalt principles and information visualization fundamentals.
Map each dimension of variation to exactly one visual channel. Do not use the same channel for two different things.
Color hue. Use for the categorical dimension with the most important distinction.
Opacity / saturation. Use to show whether the human needs to read this content carefully.
Spacing. Use for hierarchical structure.
Typography. Use for content type.
Border / container. Use for grouping related parts.
Structural outlier flags. Show in the header only, not inline.
http.server, no dependencies):GET / serves the HTML appGET /api/samples returns the current sample setPOST /api/samples lets the agent push new samplesGET /api/annotations returns the current annotationsPOST /api/annotations lets the app save annotations on every changeGET /api/graph returns the 2D projection of all records for the cluster mapGET /api/patterns returns the agent's current failure mode taxonomyPOST /api/patterns lets the agent push the updated taxonomyGET /api/suggestions returns agent-suggested annotationsPOST /api/suggestions lets the agent push suggestionserror_discovery_data/ directory:samples.json, annotations.json, graph.json, patterns.json, suggestions.jsonThree views, toggled from the top bar:
Content view design. Apply the visual encoding from Phase 2:
Header: title + label + topic. Keep it minimal. Add structural outlier flags (from Phase 2) only when the record is a genuine outlier.
Body: render the primary content using the appropriate treatment per content type.
For multi-turn traces (agent logs, conversations, chat):
For single text content (articles, summaries, etc.):
For code / diffs:
For input/output pairs:
Inline annotation: the reviewer selects text, a floating popover appears with a text input, they press Enter to save, and the span is highlighted. When the popover appears and the input is focused, the browser clears the native text selection. To prevent this, wrap the selected range in a temporary highlight span (e.g., class "pending-highlight" with a visible background) BEFORE focusing the input. Remove the temporary highlight when the annotation is saved, cancelled, or the popover is dismissed by clicking outside. This way the reviewer always sees what text they are annotating.
Margin notes: annotations and suggestions must appear as side notes in a right margin column, aligned vertically with their corresponding highlighted text. Use a two-column layout: the article body on the left (flex: 1, max-width ~720px) and a margin-notes column on the right (width ~240px). Each margin note is position: absolute inside the margin column, with its top offset calculated from the highlight element's position relative to the margin container (use getBoundingClientRect on both the highlight and the margin container, take the difference — do NOT add scrollTop, as that double-counts the scroll offset). Stack notes with a minimum gap so they do not overlap. Include hover linking: hovering a margin note outlines its highlight, and hovering a highlight outlines its margin note. Each margin note shows the quoted text, the reviewer's note, and edit/delete buttons (visible on hover). Do NOT use hover tooltips as the primary way to show annotation content — margin notes replace tooltips.
Agent suggestions: visually distinct from human annotations in both the inline highlight (dashed border, muted tint) and the margin note (different left-border color, an "agent suggestion" tag). The margin note shows accept/dismiss buttons (always visible, not hover-gated). The reviewer can accept (which promotes it to an annotation) or dismiss.
No quality labels, no dropdowns, no structured forms. Free-text notes only.
Auto-save to server on every change. Keep a localStorage backup.
Poll for new samples and suggestions periodically. Show a banner or toast when new ones arrive.
Map view:
Once the app is running and the human starts reviewing, follow review-loop.md for the ongoing interactive session, including how to monitor annotations as they arrive.
If the session ends before a human can review (a non-interactive run, for example codex exec or claude -p), build and smoke-test the app, then stop the server and give the command to launch it later; never claim the review loop ran or that anything is still running.
© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in skills/error-discovery of ai-evals-course/evals-skills.
Open the folder on GitHubat commit 80d5f7b
LLM Error Discovery Review App next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| LLM Error Discovery Review App this skillai-evals-course/evals-skills | 1.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Dataset Viewerhuggingface/skills | 11k | 3 repos | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Audit Sft Data Qualitytokenbender/agent-guides | 367 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| VLM BCQ Gap AnalysisNVIDIA/skills | 3.5k | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | |
| Mcaf ML AI Deliverymanagedcode/Storage | 138 | — | ~1k | Automated safety check: Pass | MIT |
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
huggingface/skills
Explores Hugging Face datasets through the read-only Dataset Viewer API: list splits, preview and page through rows, search, filter, and fetch parquet links and statistics.
tokenbender/agent-guides
Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.
NVIDIA/skills
Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report.
managedcode/Storage
Apply ML/AI project delivery guidance for data exploration, feasibility, experimentation, testing, responsible AI, and operating ML systems.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
ai-evals-course/evals-skills
Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.
ai-evals-course/evals-skills
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
ai-evals-course/evals-skills
Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.
ai-evals-course/evals-skills
Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.
ai-evals-course/evals-skills
Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.
Categories
Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate. Phases one to four live in the main file. The agent loads your JSONL, CSV or JSON data, inspects a handful of records and decides which shape it has: a single text field, input and output pairs, multi-turn traces with roles and tool calls, code, structured output or a mix.
LLM Error Discovery Review App fits situations like: reviewing a sample of LLM outputs or agent traces to find recurring failures; setting up a labeling interface before writing evals; choosing which records to read when a dataset is too large to review in full.
Run `npx skills add ai-evals-course/evals-skills --skill error-discovery -a claude-code`. Or copy the skill folder (skills/error-discovery in ai-evals-course/evals-skills) into .claude/skills/error-discovery in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ai-evals-course/evals-skills --skill error-discovery -a codex`. Or copy the skill folder (skills/error-discovery in ai-evals-course/evals-skills) into .agents/skills/error-discovery in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill error-discovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/error-discovery, .gemini/skills/error-discovery, .github/skills/error-discovery and .opencode/skills/error-discovery in your project.
Going by SKILL.md and its folder, LLM Error Discovery Review App needs the command-line tools its instructions call (codex and claude). Our summary lists: A dataset of LLM outputs or traces in JSONL, CSV or JSON form.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
LLM Error Discovery Review App is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with LLM Error Discovery Review App: Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 792 stars), Hugging Face Dataset Viewer (huggingface/skills, 11k stars), Audit Sft Data Quality (tokenbender/agent-guides, 367 stars) and VLM BCQ Gap Analysis (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,472 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.
Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.