Context Audit
undefined-ui/second-brain-os
Audit an agent's context layout against the four places: system prompt, tools, history, tail.
Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.
$ npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ai-evals-course/evals-skills write-judge-prompt --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/write-judge-prompt .claude/skills/write-judge-prompt && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "write-judge-prompt" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/write-judge-prompt into .claude/skills/write-judge-prompt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "write-judge-prompt", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ai-evals-course/evals-skills/tree/main/skills/write-judge-promptType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ai-evals-course/evals-skills write-judge-prompt --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/write-judge-prompt .agents/skills/write-judge-prompt && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "write-judge-prompt" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/write-judge-prompt into .agents/skills/write-judge-prompt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "write-judge-prompt", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ai-evals-course/evals-skills write-judge-prompt --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/write-judge-prompt .cursor/skills/write-judge-prompt && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "write-judge-prompt" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/write-judge-prompt into .cursor/skills/write-judge-prompt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "write-judge-prompt", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ai-evals-course/evals-skills.git --path skills/write-judge-prompt--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ai-evals-course/evals-skills write-judge-prompt --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/write-judge-prompt .gemini/skills/write-judge-prompt && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "write-judge-prompt" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/write-judge-prompt into .gemini/skills/write-judge-prompt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "write-judge-prompt", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ai-evals-course/evals-skills write-judge-promptInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/write-judge-prompt .github/skills/write-judge-prompt && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "write-judge-prompt" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/write-judge-prompt into .github/skills/write-judge-prompt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "write-judge-prompt", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ai-evals-course/evals-skills write-judge-prompt --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/write-judge-prompt .opencode/skills/write-judge-prompt && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "write-judge-prompt" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/write-judge-prompt into .opencode/skills/write-judge-prompt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "write-judge-prompt", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
write-judge-promptDesigns a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.
Use this after error analysis has named a failure mode and code-based checks have been ruled out. The skill asks you to have human-labeled traces first, at least 20 Pass and 20 Fail examples, and to exhaust keyword, regex and API checks before reaching for a judge. Its example is an interview coach that asks general questions, which looks like it needs semantic understanding but can often be caught by a keyword check for words like usually or typically.
Each judge prompt has four components. The task names one failure mode and avoids vague asks like rating quality from 1-5. The definitions state Pass and Fail strictly, with no scales or partial credit. The few-shot examples come from the training split, the 10-20% of labeled data set aside for this, and include one clear Pass, one clear Fail and one borderline case, kept out of the dev and test sets to prevent leakage. Two to four examples is typical. The fourth component is a structured output format, which the excerpt does not show. Two sibling skills handle code evals and validating a judge.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are json).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
LLM-as-Judge Prompt Writer loads about 1.9k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 683 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 683 words, ~1,921 tokens.
.claude/skills/write-judge-prompt/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Design a binary Pass/Fail LLM-as-Judge evaluator for one specific failure mode. Each judge checks exactly one thing.
Every judge prompt requires exactly four components:
State what the judge evaluates. One failure mode per judge.
You are an evaluator assessing whether a real estate assistant's email
uses the appropriate tone for the client's persona.Not: "Evaluate whether the email is good" or "Rate the email quality from 1-5."
Outcomes are strictly binary: Pass or Fail. No Likert scales, no letter grades, no partial credit. Define exactly what constitutes Pass and Fail. These definitions come from your error analysis failure mode descriptions.
## Definitions
PASS: The email matches the expected communication style for the client persona:
- Luxury Buyers: formal language, emphasis on exclusive features, premium
market positioning, no casual slang
- First-Time Homebuyers: warm and encouraging tone, educational explanations,
avoids jargon, patient and supportive
- Investors: data-driven language, ROI-focused, market analytics, concise
and professional
FAIL: The email uses a tone mismatched to the client persona. Examples:
- Using casual slang ("hey, check out this pad!") for a luxury buyer
- Using heavy financial jargon for a first-time homebuyer
- Using overly emotional language for an investorInclude labeled Pass and Fail examples from your human-labeled data.
## Examples
### Example 1: PASS
Client Persona: Luxury Buyer
Email: "Dear Mr. Harrington, I am pleased to present an exclusive listing
at 1200 Pacific Heights Drive. This distinguished property features..."
Critique: The email opens with a formal salutation and uses language
consistent with luxury positioning — "exclusive listing," "distinguished
property." No casual slang or informal phrasing. The tone matches the
luxury buyer persona throughout.
Result: Pass
### Example 2: FAIL
Client Persona: Luxury Buyer
Email: "Hey! Just found this awesome place you might like. It's got a
pool and stuff, super cool neighborhood..."
Critique: The greeting "Hey!" is informal. Phrases like "awesome place,"
"got a pool and stuff," and "super cool" are casual slang inappropriate
for a luxury buyer. The email reads like a text message, not a
professional communication for a high-end client.
Result: Fail
### Example 3: PASS (borderline)
Client Persona: First-Time Homebuyer
Email: "Hi Sarah, I found a property that might be a great fit for your
first home. The neighborhood has good schools nearby, and the monthly
payment would be similar to what you're currently paying in rent..."
Critique: The greeting is warm but not overly casual. The email explains
the property in relatable terms — comparing mortgage to rent, mentioning
schools — which is educational without being condescending. It avoids
jargon like "amortization" or "LTV ratio." While not deeply technical,
this matches the supportive tone expected for a first-time buyer.
Result: PassRules for selecting examples:
Enforce structured output using your LLM provider's schema enforcement (e.g., response_format in OpenAI, tool definitions in Anthropic) or a library like Instructor or Outlines. If the provider doesn't support schema enforcement, specify the JSON schema in the prompt.
The output must include a critique before the verdict. Placing the critique first forces the judge to articulate its assessment before committing to a decision.
{
"critique": "string — detailed assessment of the output against the criterion",
"result": "Pass or Fail"
}Critiques must be detailed, not terse. A good critique explains what specifically was correct or incorrect and references concrete evidence from the output. The critiques in your few-shot examples set the bar for the level of detail the judge will produce.
Feed only what the judge needs for an accurate decision:
| Failure Mode | What the Judge Needs |
|---|---|
| Tone mismatch | Client persona + generated email |
| Answer faithfulness | Retrieved context + generated answer |
| SQL correctness | User query + generated SQL + schema |
| Instruction following | System prompt rules + generated response |
| Tool call justification | Conversation history + tool call + tool result |
For long documents, feed only the relevant snippet, not the entire document.
Start with the most capable model available. The same model used for the main task works as judge (the judge performs a different, narrower task). Optimize for cost later once alignment is confirmed.
© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/write-judge-prompt of ai-evals-course/evals-skills.
Open the folder on GitHubat commit 80d5f7b
LLM-as-Judge Prompt Writer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| LLM-as-Judge Prompt Writer this skillai-evals-course/evals-skills | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Context Auditundefined-ui/second-brain-os | 1k | — | ~802 | Automated safety check: Pass | MIT | |
| AI Product Managementandreaskelm/pm-brain | 234 | — | ~1.8k | Automated safety check: Pass | Custom licence | |
| Building Agent Systemstelagod/code-abyss | 244 | — | ~691 | Automated safety check: Pass | MIT | |
| LLM Eval Harnessdaymade/claude-code-skills | 1.4k | — | ~4.7k | Automated safety check: Pass | MIT | |
| Prompt Engineeringericrisco/rsc-harness | 180 | — | ~2.4k | Automated safety check: Pass | MIT |
undefined-ui/second-brain-os
Audit an agent's context layout against the four places: system prompt, tools, history, tail.
andreaskelm/pm-brain
Ship and spec AI features, LLM products, agents, copilots, and generative UX — including when to use a model vs.
telagod/code-abyss
AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt…
daymade/claude-code-skills
Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.
ericrisco/rsc-harness
A skill your agent uses when one prompt must give the same right answer across reruns, models, and pasted-in hostile input: forcing a fixed schema, picking the few-shot set, ordering the prompt…
jianshuo/claude-skills
A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×…
ai-evals-course/evals-skills
Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.
ai-evals-course/evals-skills
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
ai-evals-course/evals-skills
Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.
ai-evals-course/evals-skills
Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.
ai-evals-course/evals-skills
Write code evaluators for known failure modes with objective rules.
Categories
Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format. Use this after error analysis has named a failure mode and code-based checks have been ruled out. The skill asks you to have human-labeled traces first, at least 20 Pass and 20 Fail examples, and to exhaust keyword, regex and API checks before reaching for a judge.
LLM-as-Judge Prompt Writer fits situations like: writing an evaluator for tone, faithfulness, relevance or completeness; turning a failure mode from error analysis into a judge prompt; choosing few-shot examples for a judge without leaking test data.
Run `npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a claude-code`. Or copy the skill folder (skills/write-judge-prompt in ai-evals-course/evals-skills) into .claude/skills/write-judge-prompt in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a codex`. Or copy the skill folder (skills/write-judge-prompt in ai-evals-course/evals-skills) into .agents/skills/write-judge-prompt in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill write-judge-prompt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/write-judge-prompt, .gemini/skills/write-judge-prompt, .github/skills/write-judge-prompt and .opencode/skills/write-judge-prompt in your project.
SKILL.md names no scripts, command-line tools or credentials: LLM-as-Judge Prompt Writer is instructions for the agent only. Our summary lists: Human-labeled traces for the failure mode; A completed error analysis.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
LLM-as-Judge Prompt Writer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with LLM-as-Judge Prompt Writer: Context Audit (undefined-ui/second-brain-os, 1k stars), AI Product Management (andreaskelm/pm-brain, 234 stars), Building Agent Systems (telagod/code-abyss, 244 stars) and LLM Eval Harness (daymade/claude-code-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,479 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.
Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.