LLM Judge Validation
ai-evals-course/evals-skills
Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.
Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.
$ npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install timothywarner-org/claude-code genai-prompt-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/timothywarner-org/claude-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/genai-prompt-eval .claude/skills/genai-prompt-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "genai-prompt-eval" agent skill from https://github.com/timothywarner-org/claude-code/tree/main/.claude/skills/genai-prompt-eval into .claude/skills/genai-prompt-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genai-prompt-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/timothywarner-org/claude-code/tree/main/.claude/skills/genai-prompt-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install timothywarner-org/claude-code genai-prompt-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/timothywarner-org/claude-code.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/genai-prompt-eval .agents/skills/genai-prompt-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "genai-prompt-eval" agent skill from https://github.com/timothywarner-org/claude-code/tree/main/.claude/skills/genai-prompt-eval into .agents/skills/genai-prompt-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genai-prompt-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install timothywarner-org/claude-code genai-prompt-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/timothywarner-org/claude-code.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/genai-prompt-eval .cursor/skills/genai-prompt-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "genai-prompt-eval" agent skill from https://github.com/timothywarner-org/claude-code/tree/main/.claude/skills/genai-prompt-eval into .cursor/skills/genai-prompt-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genai-prompt-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/timothywarner-org/claude-code.git --path .claude/skills/genai-prompt-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install timothywarner-org/claude-code genai-prompt-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/timothywarner-org/claude-code.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/genai-prompt-eval .gemini/skills/genai-prompt-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "genai-prompt-eval" agent skill from https://github.com/timothywarner-org/claude-code/tree/main/.claude/skills/genai-prompt-eval into .gemini/skills/genai-prompt-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genai-prompt-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install timothywarner-org/claude-code genai-prompt-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/timothywarner-org/claude-code.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/genai-prompt-eval .github/skills/genai-prompt-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "genai-prompt-eval" agent skill from https://github.com/timothywarner-org/claude-code/tree/main/.claude/skills/genai-prompt-eval into .github/skills/genai-prompt-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genai-prompt-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install timothywarner-org/claude-code genai-prompt-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/timothywarner-org/claude-code.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/genai-prompt-eval .opencode/skills/genai-prompt-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "genai-prompt-eval" agent skill from https://github.com/timothywarner-org/claude-code/tree/main/.claude/skills/genai-prompt-eval into .opencode/skills/genai-prompt-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genai-prompt-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
genai-prompt-evalScore a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.
Genai Prompt Eval is an agent skill from timothywarner-org/claude-code. Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. Use when building a prompt-eval harness, writing eval cases for a GenAI feature, gating a deploy on quality thresholds, or measuring whether model answers stay grounded and on-topic. Triggers on "evaluate my prompts", "run evals", "groundedness score", "eval cases", "quality gate for the model", "is the answer grounded".
Its SKILL.md is about 700 tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `resources/references/EVAL-DIMENSIONS.md` and `resources/scripts/run_eval.py`).
It sits in AI & LLM Engineering, covering LLM evaluation and Quality gates. It works with Python. The repository describes itself as: Claude Code and Large-Context Reasoning (O'Reilly Live Learning). The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cb80eae. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGlobGrepBashEditWriteFrom allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Genai Prompt Eval loads about 696 tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 308 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Glob, Grep, Bash, Edit, WriteAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from timothywarner-org/claude-code at commit cb80eae, republished under its MIT licence (© timothywarner-org). 308 words, ~696 tokens.
.claude/skills/genai-prompt-eval/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.This skill measures whether a generative-AI feature produces answers that are grounded, relevant, coherent, and safe. It runs a set of eval cases through the model, scores each output on those four dimensions, and reports pass or fail against thresholds. Pair it with the azure-ai-deploy skill: evals are Gate 1 of that deploy checklist.
Read resources/references/EVAL-DIMENSIONS.md. It defines groundedness, relevance, coherence, and safety, states what each one measures, and gives a pass signal for each.
Start from resources/templates/eval_cases.jsonl. Each line is one case: an input prompt, optional context the answer must stay grounded to, and expected_criteria describing a passing answer. Add cases that mirror the questions real users send.
uv run python ${CLAUDE_SKILL_DIR}/resources/scripts/run_eval.py \
--cases ${CLAUDE_SKILL_DIR}/resources/templates/eval_cases.jsonl \
--threshold 0.8The script loads the cases, calls the model for each, scores the output on the four dimensions, prints a per-case and aggregate report, and exits non-zero when the aggregate score falls below the threshold. That non-zero exit fails a CI or deploy step.
uv run.© timothywarner-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in .claude/skills/genai-prompt-eval of timothywarner-org/claude-code.
Open the folder on GitHubat commit cb80eae
Genai Prompt Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Genai Prompt Eval this skilltimothywarner-org/claude-code | 224 | — | ~696 | Automated safety check: Notes | MIT | |
| LLM Judge Validationai-evals-course/evals-skills | 1.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | — | ~2.8k | Automated safety check: Pass | MIT | |
| Clawpathy AutoresearchClawBio/ClawBio | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Failproof AI SDK IntegrationFailproofAI/failproofai | 5.3k | — | ~6k | Automated safety check: Pass | Custom licence |
ai-evals-course/evals-skills
Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
ClawBio/ClawBio
Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
FailproofAI/failproofai
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
wso2/agent-manager
Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.
timothywarner-org/claude-code
Scaffold production-ready Python MCP servers using FastMCP. An agent skill from timothywarner-org/claude-code.
timothywarner-org/claude-code
A skill your agent uses when authoring, reviewing, or refactoring Azure Bicep code.
timothywarner-org/claude-code
Audit the CLAUDE.md hierarchy in a repo for drift between what each CLAUDE.md claims and what's actually on disk.
timothywarner-org/claude-code
Ship a Python generative-AI app to Azure the keyless way, using DefaultAzureCredential and azd.
timothywarner-org/claude-code
Kubernetes workload patterns, resource management, RBAC, probes, autoscaling, ConfigMap/Secret handling, and kubectl debugging for production-grade deployments.
timothywarner-org/claude-code
Review uncommitted local changes in the current git working tree for bugs, smells, missing tests, and CLAUDE.md voice violations.
Works with
Categories
Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. Genai Prompt Eval is an agent skill from timothywarner-org/claude-code. Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.
Genai Prompt Eval fits situations like: building a prompt-eval harness; writing eval cases for a GenAI feature; gating a deploy on quality thresholds; measuring whether model answers stay grounded and on-topic.
Run `npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a claude-code`. Or copy the skill folder (.claude/skills/genai-prompt-eval in timothywarner-org/claude-code) into .claude/skills/genai-prompt-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a codex`. Or copy the skill folder (.claude/skills/genai-prompt-eval in timothywarner-org/claude-code) into .agents/skills/genai-prompt-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/genai-prompt-eval, .gemini/skills/genai-prompt-eval, .github/skills/genai-prompt-eval and .opencode/skills/genai-prompt-eval in your project.
Going by SKILL.md and its folder, Genai Prompt Eval needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, Bash, Edit, Write.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Genai Prompt Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 696 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Genai Prompt Eval: LLM Judge Validation (ai-evals-course/evals-skills, 1.5k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Clawpathy Autoresearch (ClawBio/ClawBio, 1.2k stars) and LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
timothywarner-org (a GitHub organization) maintains it in timothywarner-org/claude-code, which has 224 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on July 20, 2026.
Source: timothywarner-org/claude-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.