Agent skill

LLM-as-Judge Prompt Writer

by ai-evals-course in ai-evals-course/evals-skills

Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

Apache-2.0Auto-check passedAI & LLM Engineering

Install LLM-as-Judge Prompt Writer

skills CLI
$ npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-evals-course/evals-skills write-judge-prompt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/write-judge-prompt .claude/skills/write-judge-prompt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
write-judge-prompt
GitHub stars
1.5k
Token cost
~1.9k tokens
SKILL.md length
683 words
Files
2
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

  • Works in 4 steps: Task and Evaluation Criterion → Pass/Fail Definitions → Few-Shot Examples → …
  • Writing an evaluator for tone, faithfulness, relevance or completeness
  • SKILL.md covers Prerequisites, The Four Components, Choosing What to Pass to the… and Model Selection, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Use this after error analysis has named a failure mode and code-based checks have been ruled out. The skill asks you to have human-labeled traces first, at least 20 Pass and 20 Fail examples, and to exhaust keyword, regex and API checks before reaching for a judge. Its example is an interview coach that asks general questions, which looks like it needs semantic understanding but can often be caught by a keyword check for words like usually or typically.

Each judge prompt has four components. The task names one failure mode and avoids vague asks like rating quality from 1-5. The definitions state Pass and Fail strictly, with no scales or partial credit. The few-shot examples come from the training split, the 10-20% of labeled data set aside for this, and include one clear Pass, one clear Fail and one borderline case, kept out of the dev and test sets to prevent leakage. Two to four examples is typical. The fourth component is a structured output format, which the excerpt does not show. Two sibling skills handle code evals and validating a judge.

When your agent uses it

  • Writing an evaluator for tone, faithfulness, relevance or completeness
  • Turning a failure mode from error analysis into a judge prompt
  • Choosing few-shot examples for a judge without leaking test data

Example prompts

  • “Write a judge prompt that checks whether our support bot's replies stay faithful to the retrieved article.”
  • “Draft a Pass/Fail evaluator for the tone failure mode in our sales emails.”
  • “Pick few-shot examples for the judge from my labeled training split.”

Requirements

  • Human-labeled traces for the failure mode
  • A completed error analysis

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Task and Evaluation Criterion
  2. Pass/Fail Definitions
  3. Few-Shot Examples
  4. Structured Output Format

What it can do on your machine

Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM-as-Judge Prompt Writer loads about 1.9k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 683 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 683 words, ~1,921 tokens.

Download SKILL.mdSave it as .claude/skills/write-judge-prompt/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
write-judge-prompt
description
Design LLM-as-Judge evaluators for subjective criteria that code-based checks cannot handle. Use when a failure mode requires interpretation (tone, faithfulness, relevance, completeness). Do NOT use when the failure mode can be checked with code (regex, schema validation, execution tests); use `write-code-eval`. To validate an existing judge, use `validate-evaluator`.

Write LLM-as-Judge Prompt

Design a binary Pass/Fail LLM-as-Judge evaluator for one specific failure mode. Each judge checks exactly one thing.

Prerequisites

  • Error analysis is complete. The failure mode is identified.
  • You have human-labeled traces for this failure mode (at least 20 Pass and 20 Fail examples).
  • A code-based evaluator cannot check this failure mode. Exhaust code-based options before reaching for a judge — many failure modes that seem subjective reduce to keyword checks, regex, or API calls when you understand the domain. Example: detecting whether an AI interviewing coach suggests "general" questions (asking about typical behavior instead of a specific past event) seems to require semantic understanding, but in practice a keyword check for words like "usually," "typical," and "normally" could work quite well.

The Four Components

Every judge prompt requires exactly four components:

1. Task and Evaluation Criterion

State what the judge evaluates. One failure mode per judge.

You are an evaluator assessing whether a real estate assistant's email
uses the appropriate tone for the client's persona.

Not: "Evaluate whether the email is good" or "Rate the email quality from 1-5."

2. Pass/Fail Definitions

Outcomes are strictly binary: Pass or Fail. No Likert scales, no letter grades, no partial credit. Define exactly what constitutes Pass and Fail. These definitions come from your error analysis failure mode descriptions.

## Definitions

PASS: The email matches the expected communication style for the client persona:
- Luxury Buyers: formal language, emphasis on exclusive features, premium
  market positioning, no casual slang
- First-Time Homebuyers: warm and encouraging tone, educational explanations,
  avoids jargon, patient and supportive
- Investors: data-driven language, ROI-focused, market analytics, concise
  and professional

FAIL: The email uses a tone mismatched to the client persona. Examples:
- Using casual slang ("hey, check out this pad!") for a luxury buyer
- Using heavy financial jargon for a first-time homebuyer
- Using overly emotional language for an investor
3. Few-Shot Examples

Include labeled Pass and Fail examples from your human-labeled data.

## Examples

### Example 1: PASS
Client Persona: Luxury Buyer
Email: "Dear Mr. Harrington, I am pleased to present an exclusive listing
at 1200 Pacific Heights Drive. This distinguished property features..."
Critique: The email opens with a formal salutation and uses language
consistent with luxury positioning — "exclusive listing," "distinguished
property." No casual slang or informal phrasing. The tone matches the
luxury buyer persona throughout.
Result: Pass

### Example 2: FAIL
Client Persona: Luxury Buyer
Email: "Hey! Just found this awesome place you might like. It's got a
pool and stuff, super cool neighborhood..."
Critique: The greeting "Hey!" is informal. Phrases like "awesome place,"
"got a pool and stuff," and "super cool" are casual slang inappropriate
for a luxury buyer. The email reads like a text message, not a
professional communication for a high-end client.
Result: Fail

### Example 3: PASS (borderline)
Client Persona: First-Time Homebuyer
Email: "Hi Sarah, I found a property that might be a great fit for your
first home. The neighborhood has good schools nearby, and the monthly
payment would be similar to what you're currently paying in rent..."
Critique: The greeting is warm but not overly casual. The email explains
the property in relatable terms — comparing mortgage to rent, mentioning
schools — which is educational without being condescending. It avoids
jargon like "amortization" or "LTV ratio." While not deeply technical,
this matches the supportive tone expected for a first-time buyer.
Result: Pass

Rules for selecting examples:

  • Include at least one clear Pass, one clear Fail, and one borderline case. Borderline examples are the most valuable — they teach nuance.
  • Draw examples from the training split (10-20% of labeled data set aside for this purpose).
  • Any example used in the judge prompt must be excluded from dev and test sets. Using dev/test examples is data leakage.
  • 2-4 examples is typical. Performance plateaus after 4-8.
4. Structured Output Format

Enforce structured output using your LLM provider's schema enforcement (e.g., response_format in OpenAI, tool definitions in Anthropic) or a library like Instructor or Outlines. If the provider doesn't support schema enforcement, specify the JSON schema in the prompt.

The output must include a critique before the verdict. Placing the critique first forces the judge to articulate its assessment before committing to a decision.

json
{
  "critique": "string — detailed assessment of the output against the criterion",
  "result": "Pass or Fail"
}

Critiques must be detailed, not terse. A good critique explains what specifically was correct or incorrect and references concrete evidence from the output. The critiques in your few-shot examples set the bar for the level of detail the judge will produce.

Show full SKILL.md (294 more words)Show less

Choosing What to Pass to the Judge

Feed only what the judge needs for an accurate decision:

Failure ModeWhat the Judge Needs
Tone mismatchClient persona + generated email
Answer faithfulnessRetrieved context + generated answer
SQL correctnessUser query + generated SQL + schema
Instruction followingSystem prompt rules + generated response
Tool call justificationConversation history + tool call + tool result

For long documents, feed only the relevant snippet, not the entire document.

Model Selection

Start with the most capable model available. The same model used for the main task works as judge (the judge performs a different, narrower task). Optimize for cost later once alignment is confirmed.

Anti-Patterns

  • Vague criteria like "is this helpful?" Target a specific, observable failure mode from error analysis.
  • Holistic judge for the entire trace. A single judge covering multiple dimensions produces unactionable verdicts.
  • No few-shot examples. Without examples, the model won't know what counts as a failure in your application.
  • Dev/test examples used as few-shot. This is data leakage. Use only the training split.
  • Likert scales (1-5, letter grades, etc.). Binary pass/fail only. Likert scales produce scores that sound precise but can't be calibrated: annotators disagree on the difference between a 3 and a 4, and the judge inherits that noise. Binary forces you to define a clear decision boundary upfront, which makes inter-annotator agreement measurable and the judge's errors actionable. If you need to capture severity, use multiple binary judges (e.g., "factually wrong" and "dangerously wrong") rather than one ordinal scale.
  • Skipping validation. Measure alignment with human labels using validate-evaluator before trusting the judge.
  • Judges for specification failures without fixing the prompt first. If the prompt never asked for the behavior, add the instruction before building an evaluator. For critical requirements, a judge can still serve as a regression guard.

© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/write-judge-prompt of ai-evals-course/evals-skills.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 80d5f7b

Compare with similar skills

LLM-as-Judge Prompt Writer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM-as-Judge Prompt Writer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM-as-Judge Prompt Writer this skillai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
Context Auditundefined-ui/second-brain-os1k—~802Automated safety check: PassMIT
AI Product Managementandreaskelm/pm-brain234—~1.8kAutomated safety check: PassCustom licence
Building Agent Systemstelagod/code-abyss244—~691Automated safety check: PassMIT
LLM Eval Harnessdaymade/claude-code-skills1.4k—~4.7kAutomated safety check: PassMIT
Prompt Engineeringericrisco/rsc-harness180—~2.4kAutomated safety check: PassMIT

Similar skills

  • Context Audit

    undefined-ui/second-brain-os

    Audit an agent's context layout against the four places: system prompt, tools, history, tail.

    1k GitHub stars~802 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • AI Product Management

    andreaskelm/pm-brain

    Ship and spec AI features, LLM products, agents, copilots, and generative UX — including when to use a model vs.

    234 GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Building Agent Systems

    telagod/code-abyss

    AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt…

    244 GitHub stars~691 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Eval Harness

    daymade/claude-code-skills

    Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.

    1.4k GitHub stars~4.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineering

    ericrisco/rsc-harness

    A skill your agent uses when one prompt must give the same right answer across reruns, models, and pasted-in hostile input: forcing a fixed schema, picking the few-shot set, ordering the prompt…

    180 GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Wjs Evaling Voicedrop Prompts

    jianshuo/claude-skills

    A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×…

    131 GitHub stars~475 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from ai-evals-course/evals-skills

All 9 skills in this repo
  • LLM Trace Review Interface

    ai-evals-course/evals-skills

    Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

    1.5k GitHub stars~1.4k tokensUpdated 16 days ago
    Auto-check passed
  • LLM Eval Pipeline Audit

    ai-evals-course/evals-skills

    Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

    1.5k GitHub stars~2.5k tokensUpdated 16 days ago
    Auto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 16 days ago
    Auto-check passed
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 16 days ago
    Auto-check passed
  • LLM Judge Validation

    ai-evals-course/evals-skills

    Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

    1.5k GitHub stars~2.2k tokensUpdated 16 days ago
    Auto-check passed
  • Write Code Eval

    ai-evals-course/evals-skills

    Write code evaluators for known failure modes with objective rules.

    1.5k GitHub stars~385 tokensUpdated 16 days ago
    Auto-check passed

Questions about LLM-as-Judge Prompt Writer

What does LLM-as-Judge Prompt Writer do?

Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format. Use this after error analysis has named a failure mode and code-based checks have been ruled out. The skill asks you to have human-labeled traces first, at least 20 Pass and 20 Fail examples, and to exhaust keyword, regex and API checks before reaching for a judge.

When should I use LLM-as-Judge Prompt Writer?

LLM-as-Judge Prompt Writer fits situations like: writing an evaluator for tone, faithfulness, relevance or completeness; turning a failure mode from error analysis into a judge prompt; choosing few-shot examples for a judge without leaking test data.

How do I install LLM-as-Judge Prompt Writer in Claude Code?

Run `npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a claude-code`. Or copy the skill folder (skills/write-judge-prompt in ai-evals-course/evals-skills) into .claude/skills/write-judge-prompt in your project. Claude Code loads it when a task matches its description.

How do I install LLM-as-Judge Prompt Writer in Codex?

Run `npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a codex`. Or copy the skill folder (skills/write-judge-prompt in ai-evals-course/evals-skills) into .agents/skills/write-judge-prompt in your project. Codex loads it when a task matches its description.

Can I use LLM-as-Judge Prompt Writer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/write-judge-prompt, .gemini/skills/write-judge-prompt, .github/skills/write-judge-prompt and .opencode/skills/write-judge-prompt in your project.

What does LLM-as-Judge Prompt Writer need to run?

SKILL.md names no scripts, command-line tools or credentials: LLM-as-Judge Prompt Writer is instructions for the agent only. Our summary lists: Human-labeled traces for the failure mode; A completed error analysis.

Does LLM-as-Judge Prompt Writer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is LLM-as-Judge Prompt Writer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM-as-Judge Prompt Writer use?

LLM-as-Judge Prompt Writer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM-as-Judge Prompt Writer use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM-as-Judge Prompt Writer?

Skills that share tags, products or a category with LLM-as-Judge Prompt Writer: Context Audit (undefined-ui/second-brain-os, 1k stars), AI Product Management (andreaskelm/pm-brain, 234 stars), Building Agent Systems (telagod/code-abyss, 244 stars) and LLM Eval Harness (daymade/claude-code-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM-as-Judge Prompt Writer?

ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,479 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.

Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.