Agent skill

Genai Prompt Eval

by timothywarner-org in timothywarner-org/claude-code

Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.

MITAuto-check: notesAI & LLM Engineering

Install Genai Prompt Eval

skills CLI
$ npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install timothywarner-org/claude-code genai-prompt-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/timothywarner-org/claude-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/genai-prompt-eval .claude/skills/genai-prompt-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
genai-prompt-eval
GitHub stars
224
Token cost
~696 tokens
SKILL.md length
308 words
Files
4
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.

  • Works in 4 steps: Read the dimensions → Build the eval corpus → Run the harness → …
  • Building a prompt-eval harness
  • SKILL.md covers When to reach for this, Workflow and Conventions
  • Runs Python scripts from its folder; calls uv

What it does

Genai Prompt Eval is an agent skill from timothywarner-org/claude-code. Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. Use when building a prompt-eval harness, writing eval cases for a GenAI feature, gating a deploy on quality thresholds, or measuring whether model answers stay grounded and on-topic. Triggers on "evaluate my prompts", "run evals", "groundedness score", "eval cases", "quality gate for the model", "is the answer grounded".

Its SKILL.md is about 700 tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `resources/references/EVAL-DIMENSIONS.md` and `resources/scripts/run_eval.py`).

It sits in AI & LLM Engineering, covering LLM evaluation and Quality gates. It works with Python. The repository describes itself as: Claude Code and Large-Context Reasoning (O'Reilly Live Learning). The licence is MIT.

When your agent uses it

  • Building a prompt-eval harness
  • Writing eval cases for a GenAI feature
  • Gating a deploy on quality thresholds
  • Measuring whether model answers stay grounded and on-topic

Example prompts

  • “evaluate my prompts”
  • “run evals”
  • “groundedness score”
  • “/genai-prompt-eval”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, Bash, Edit, Write

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Read the dimensions
  2. Build the eval corpus
  3. Run the harness
  4. Read the report and act

What it can do on your machine

Read from SKILL.md and the folder at commit cb80eae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • Bash
    • Edit
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Genai Prompt Eval loads about 696 tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 308 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~696

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Glob, Grep, Bash, Edit, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from timothywarner-org/claude-code at commit cb80eae, republished under its MIT licence (© timothywarner-org). 308 words, ~696 tokens.

Download SKILL.mdSave it as .claude/skills/genai-prompt-eval/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
genai-prompt-eval
description
Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. Use when building a prompt-eval harness, writing eval cases for a GenAI feature, gating a deploy on quality thresholds, or measuring whether model answers stay grounded and on-topic. Triggers on "evaluate my prompts", "run evals", "groundedness score", "eval cases", "quality gate for the model", "is the answer grounded".
allowed-tools
Read, Glob, Grep, Bash, Edit, Write

Score GenAI outputs before shipping

This skill measures whether a generative-AI feature produces answers that are grounded, relevant, coherent, and safe. It runs a set of eval cases through the model, scores each output on those four dimensions, and reports pass or fail against thresholds. Pair it with the azure-ai-deploy skill: evals are Gate 1 of that deploy checklist.

When to reach for this

  • A GenAI feature is changing and you need a regression signal on answer quality.
  • A deploy gate requires proof that outputs meet a quality bar.
  • You want a repeatable eval corpus that reflects real enterprise questions, not toy prompts.

Workflow

1. Read the dimensions

Read resources/references/EVAL-DIMENSIONS.md. It defines groundedness, relevance, coherence, and safety, states what each one measures, and gives a pass signal for each.

2. Build the eval corpus

Start from resources/templates/eval_cases.jsonl. Each line is one case: an input prompt, optional context the answer must stay grounded to, and expected_criteria describing a passing answer. Add cases that mirror the questions real users send.

3. Run the harness
bash
uv run python ${CLAUDE_SKILL_DIR}/resources/scripts/run_eval.py \
  --cases ${CLAUDE_SKILL_DIR}/resources/templates/eval_cases.jsonl \
  --threshold 0.8

The script loads the cases, calls the model for each, scores the output on the four dimensions, prints a per-case and aggregate report, and exits non-zero when the aggregate score falls below the threshold. That non-zero exit fails a CI or deploy step.

4. Read the report and act
  • Cases below threshold name the failing dimension. Fix the prompt, the retrieval context, or the guardrail, then re-run.
  • Record the aggregate score as the new baseline so the next run detects regressions.

Conventions

  • uv manages Python, not pip. Run scripts with uv run.
  • No hardcoded secrets. The scoring model client reads its endpoint and deployment from env vars.
  • Realistic cases only. Eval inputs are enterprise scenarios, never placeholder prompts.
  • Deterministic scoring where possible. Prefer a low temperature on any model-graded dimension so scores are stable across runs.

© timothywarner-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in .claude/skills/genai-prompt-eval of timothywarner-org/claude-code.

  • SKILL.md
  • resources/references/EVAL-DIMENSIONS.md
  • resources/scripts/run_eval.py
  • resources/templates/eval_cases.jsonl

Open the folder on GitHubat commit cb80eae

Compare with similar skills

Genai Prompt Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Genai Prompt Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Genai Prompt Eval this skilltimothywarner-org/claude-code224—~696Automated safety check: NotesMIT
LLM Judge Validationai-evals-course/evals-skills1.5k—~2.2kAutomated safety check: PassApache-2.0
Azure AI Projects Python SDKmicrosoft/skills3.1k—~2.8kAutomated safety check: PassMIT
Clawpathy AutoresearchClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence

Similar skills

  • LLM Judge Validation

    ai-evals-course/evals-skills

    Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

    1.5k GitHub stars~2.2k tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub stars~2.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Add Evaluator

    wso2/agent-manager

    Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.

    108 GitHub stars~710 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from timothywarner-org/claude-code

  • MCP Scaffold

    timothywarner-org/claude-code

    Scaffold production-ready Python MCP servers using FastMCP. An agent skill from timothywarner-org/claude-code.

    224 GitHub stars~940 tokensUpdated 2 mo ago
    Auto-check passed
  • Azure Bicep Skill

    timothywarner-org/claude-code

    A skill your agent uses when authoring, reviewing, or refactoring Azure Bicep code.

    224 GitHub stars~2.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Claude Md Audit

    timothywarner-org/claude-code

    Audit the CLAUDE.md hierarchy in a repo for drift between what each CLAUDE.md claims and what's actually on disk.

    224 GitHub stars~649 tokensUpdated 2 mo ago
    Auto-check passed
  • Azure AI Deploy

    timothywarner-org/claude-code

    Ship a Python generative-AI app to Azure the keyless way, using DefaultAzureCredential and azd.

    224 GitHub stars~731 tokensUpdated 2 mo ago
    Auto-check: notes
  • Kubernetes Patterns

    timothywarner-org/claude-code

    Kubernetes workload patterns, resource management, RBAC, probes, autoscaling, ConfigMap/Secret handling, and kubectl debugging for production-grade deployments.

    224 GitHub stars~4.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Review Changes

    timothywarner-org/claude-code

    Review uncommitted local changes in the current git working tree for bugs, smells, missing tests, and CLAUDE.md voice violations.

    224 GitHub stars~506 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Genai Prompt Eval

What does Genai Prompt Eval do?

Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. Genai Prompt Eval is an agent skill from timothywarner-org/claude-code. Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.

When should I use Genai Prompt Eval?

Genai Prompt Eval fits situations like: building a prompt-eval harness; writing eval cases for a GenAI feature; gating a deploy on quality thresholds; measuring whether model answers stay grounded and on-topic.

How do I install Genai Prompt Eval in Claude Code?

Run `npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a claude-code`. Or copy the skill folder (.claude/skills/genai-prompt-eval in timothywarner-org/claude-code) into .claude/skills/genai-prompt-eval in your project. Claude Code loads it when a task matches its description.

How do I install Genai Prompt Eval in Codex?

Run `npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a codex`. Or copy the skill folder (.claude/skills/genai-prompt-eval in timothywarner-org/claude-code) into .agents/skills/genai-prompt-eval in your project. Codex loads it when a task matches its description.

Can I use Genai Prompt Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add timothywarner-org/claude-code --skill genai-prompt-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/genai-prompt-eval, .gemini/skills/genai-prompt-eval, .github/skills/genai-prompt-eval and .opencode/skills/genai-prompt-eval in your project.

What does Genai Prompt Eval need to run?

Going by SKILL.md and its folder, Genai Prompt Eval needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, Bash, Edit, Write.

Does Genai Prompt Eval access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Genai Prompt Eval safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Genai Prompt Eval use?

Genai Prompt Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Genai Prompt Eval use?

About 696 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Genai Prompt Eval?

Skills that share tags, products or a category with Genai Prompt Eval: LLM Judge Validation (ai-evals-course/evals-skills, 1.5k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Clawpathy Autoresearch (ClawBio/ClawBio, 1.2k stars) and LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Genai Prompt Eval?

timothywarner-org (a GitHub organization) maintains it in timothywarner-org/claude-code, which has 224 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on July 20, 2026.

Source: timothywarner-org/claude-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.