Agent skill

Comparative Evaluation

by Owl-Listener in Owl-Listener/ai-design-skills

A/B testing, side-by-side comparison, and preference ranking for AI outputs.

MITAuto-check passedMarketing & SEO

Install Comparative Evaluation

skills CLI
$ npx skills add Owl-Listener/ai-design-skills --skill comparative-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Owl-Listener/ai-design-skills comparative-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Owl-Listener/ai-design-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluation/comparative-evaluation .claude/skills/comparative-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
comparative-evaluation
GitHub stars
180
Token cost
~614 tokens
SKILL.md length
318 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

A/B testing, side-by-side comparison, and preference ranking for AI outputs.

  • Tasks that involve A/B testing
  • SKILL.md covers Comparison Methods, Designing A/B Tests for AI, Side-by-Side Evaluation Design and When to Use Comparative vs.…, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Comparative Evaluation is an agent skill from Owl-Listener/ai-design-skills. A/B testing, side-by-side comparison, and preference ranking for AI outputs.

Its SKILL.md is about 610 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: AI Design Skills Collection: agentic skills, commands, and plugins for designing AI products — from interaction patterns to alignment, evaluation, agent orchestration, and prompt… The licence is MIT.

When your agent uses it

  • Tasks that involve A/B testing

Example prompts

  • “/comparative-evaluation”

What it can do on your machine

Read from SKILL.md and the folder at commit f41b650. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Comparative Evaluation loads about 614 tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 318 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~614

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Owl-Listener/ai-design-skills at commit f41b650, republished under its MIT licence (© Owl-Listener). 318 words, ~614 tokens.

Download SKILL.mdSave it as .claude/skills/comparative-evaluation/SKILL.md (or your agent's skills folder).
name
comparative-evaluation
description
A/B testing, side-by-side comparison, and preference ranking for AI outputs.

Comparative Evaluation

Absolute quality scores are useful but limited. Comparative evaluation — putting outputs side by side and asking which is better — often reveals quality differences that rubrics miss.

Comparison Methods

  • A/B testing: Show different users different versions and compare outcomes
  • Side-by-side evaluation: Show evaluators two outputs for the same input and ask which is better
  • Preference ranking: Show evaluators multiple outputs and rank them from best to worst
  • Paired comparison: Compare every pair of options to build a complete ranking
  • Elo rating: Use tournament-style comparisons to develop continuous quality scores

Designing A/B Tests for AI

A/B testing AI is different from A/B testing UI:

  • Variance is high: The same prompt can produce different outputs, so you need more samples
  • Context matters: The same change might help for one task and hurt for another
  • Metrics lag: AI quality changes may take time to show up in user behavior
  • Interaction effects: A change to one part of the conversation affects all subsequent parts Design A/B tests with:
  • Sufficient sample sizes to account for output variance
  • Segmentation by task type and user experience level
  • Multiple metrics (don't optimise for one at the expense of others)
  • Guardrails to catch severe quality regressions quickly

Side-by-Side Evaluation Design

For human evaluation of AI outputs:

  • Blind evaluation: Evaluators shouldn't know which version is which
  • Consistent inputs: Compare outputs generated from the same input
  • Structured criteria: Give evaluators specific dimensions to compare on, not just "which is better"
  • Multiple evaluators: Use at least 3 evaluators per comparison for reliability
  • Diverse inputs: Test across a representative sample of real user inputs

When to Use Comparative vs. Absolute Evaluation

  • Comparative: Best for choosing between alternatives, detecting subtle quality differences, and model selection
  • Absolute: Best for measuring against a standard, tracking progress over time, and certification

Design Artefacts

  • A/B test design templates
  • Side-by-side evaluation protocols
  • Evaluator instructions and rubrics
  • Sample size calculators for AI experiments
  • Comparison result analysis frameworks

© Owl-Listener, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/evaluation/comparative-evaluation of Owl-Listener/ai-design-skills.

Open the folder on GitHubat commit f41b650

Compare with similar skills

Comparative Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Comparative Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Comparative Evaluation this skillOwl-Listener/ai-design-skills180—~614Automated safety check: PassMIT
AnalyticsNexus-JPF/note-companion8696 repos~2.2kAutomated safety check: PassMIT
Ab Test Setupfreekmurze/dotfiles1k15 repos~1.8kAutomated safety check: PassNone
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab TestingCesarjoquin/Marketing-Skills1992 repos~2.8kAutomated safety check: PassMIT
Meta Tags Optimizernowork-studio/notfair-plugin3.9k1 repos~2.7kAutomated safety check: PassMIT

Similar skills

  • Analytics

    Nexus-JPF/note-companion

    When the user wants to set up, improve, or audit analytics tracking and measurement.

    869 GitHub starsUsed in 6 repos~2.2k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Setup

    freekmurze/dotfiles

    When the user wants to plan, design, or implement an A/B test or experiment.

    1k GitHub starsUsed in 15 repos~1.8k tokens
    Marketing & SEOAuto-check passed
  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Testing

    Cesarjoquin/Marketing-Skills

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

    199 GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Meta Tags Optimizer

    nowork-studio/notfair-plugin

    Writes and improves title tags, meta descriptions, Open Graph and Twitter card tags for click-through, with character counts and A/B test variants.

    3.9k GitHub starsUsed in 1 repo~2.7k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.8k GitHub stars~1.4k tokensUpdated 14 days ago
    Marketing & SEOAuto-check passed

More from Owl-Listener/ai-design-skills

All 11 skills in this repo
  • Few Shot Patterns

    Owl-Listener/ai-design-skills

    Crafting examples that steer AI behavior effectively. An agent skill from Owl-Listener/ai-design-skills.

    180 GitHub stars~723 tokensUpdated 4 mo ago
    Auto-check passed
  • State Management

    Owl-Listener/ai-design-skills

    Managing shared context, memory, and state across multiple agents.

    180 GitHub stars~1.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Agent Role Design

    Owl-Listener/ai-design-skills

    Defining what each agent does, knows, and owns in a multi-agent system.

    180 GitHub stars~658 tokensUpdated 4 mo ago
    Auto-check passed
  • Behavioral Consistency

    Owl-Listener/ai-design-skills

    Ensuring the AI behaves predictably across sessions, edge cases, and modalities.

    180 GitHub stars~568 tokensUpdated 4 mo ago
    Auto-check passed
  • Bias Detection Design

    Owl-Listener/ai-design-skills

    Designing review workflows to surface and mitigate bias in AI outputs.

    180 GitHub stars~650 tokensUpdated 4 mo ago
    Auto-check passed
  • Chain Of Thought Design

    Owl-Listener/ai-design-skills

    Designing reasoning chains that produce better outputs. An agent skill from Owl-Listener/ai-design-skills.

    180 GitHub stars~687 tokensUpdated 4 mo ago
    Auto-check passed

Categories

Questions about Comparative Evaluation

What does Comparative Evaluation do?

A/B testing, side-by-side comparison, and preference ranking for AI outputs. Comparative Evaluation is an agent skill from Owl-Listener/ai-design-skills. A/B testing, side-by-side comparison, and preference ranking for AI outputs.

When should I use Comparative Evaluation?

Comparative Evaluation fits situations like: tasks that involve A/B testing.

How do I install Comparative Evaluation in Claude Code?

Run `npx skills add Owl-Listener/ai-design-skills --skill comparative-evaluation -a claude-code`. Or copy the skill folder (skills/evaluation/comparative-evaluation in Owl-Listener/ai-design-skills) into .claude/skills/comparative-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Comparative Evaluation in Codex?

Run `npx skills add Owl-Listener/ai-design-skills --skill comparative-evaluation -a codex`. Or copy the skill folder (skills/evaluation/comparative-evaluation in Owl-Listener/ai-design-skills) into .agents/skills/comparative-evaluation in your project. Codex loads it when a task matches its description.

Can I use Comparative Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Owl-Listener/ai-design-skills --skill comparative-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/comparative-evaluation, .gemini/skills/comparative-evaluation, .github/skills/comparative-evaluation and .opencode/skills/comparative-evaluation in your project.

What does Comparative Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Comparative Evaluation is instructions for the agent only.

Does Comparative Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Comparative Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Comparative Evaluation use?

Comparative Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Comparative Evaluation use?

About 614 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Comparative Evaluation?

Skills that share tags, products or a category with Comparative Evaluation: Analytics (Nexus-JPF/note-companion, 869 stars), Ab Test Setup (freekmurze/dotfiles, 1k stars), Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars) and Ab Testing (Cesarjoquin/Marketing-Skills, 199 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Comparative Evaluation?

Owl-Listener (a GitHub user) maintains it in Owl-Listener/ai-design-skills, which has 180 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on June 9, 2026.

Source: Owl-Listener/ai-design-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.