Agent skill

Eval Rubric Designer

by mohitagw15856 in mohitagw15856/pm-claude-skills

Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output.

MITAuto-check passedEducation

Install Eval Rubric Designer

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill eval-rubric-designer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills eval-rubric-designer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-rubric-designer .claude/skills/eval-rubric-designer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-rubric-designer
GitHub stars
1.4k
Token cost
~1.1k tokens
SKILL.md length
565 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output.

  • Asked to create an eval rubric
  • SKILL.md covers Working from a brief, Required Inputs, Output Format and Quality Checks, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Define quality dimensions

What it does

Eval Rubric Designer is an agent skill from mohitagw15856/pm-claude-skills. Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output. Use when asked to create an eval rubric, define quality dimensions, build an LLM judge, or decide how to measure whether AI output is good. Produces a rubric with weighted dimensions and concrete 1–5 anchors, a ready-to-run judge prompt, a labelling guide, and notes on judge reliability.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Quizzes and assessments and LLM evaluation. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to create an eval rubric
  • Define quality dimensions
  • Build an LLM judge
  • Decide how to measure whether AI output is good

Example prompts

  • “/eval-rubric-designer”

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Rubric Designer loads about 1.1k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 565 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 565 words, ~1,077 tokens.

Download SKILL.mdSave it as .claude/skills/eval-rubric-designer/SKILL.md (or your agent's skills folder).
name
eval-rubric-designer
description
Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output. Use when asked to create an eval rubric, define quality dimensions, build an LLM judge, or decide how to measure whether AI output is good. Produces a rubric with weighted dimensions and concrete 1–5 anchors, a ready-to-run judge prompt, a labelling guide, and notes on judge reliability.

Eval Rubric Designer Skill

You can't improve what you can't score. The hard part of evaluating AI output isn't running the judge — it's defining dimensions that are specific, observable, and independent, with anchors concrete enough that two people (or two judge runs) agree. This skill turns "is the output good?" into a rubric and a judge prompt you can run today.

Working from a brief

Given just "I need to eval my summariser", produce the full rubric anyway — infer the task, the output type, and the dimensions that matter for it, and label inferred choices. Never hand back a list of dimension names with no anchors; the anchors are where the rubric earns its keep.

Required Inputs

Ask for these only if they aren't already provided (else infer and label):

  • The task — what the AI is supposed to produce, and for whom.
  • A sample output (or two) — ideally one good and one weak, to calibrate anchors.
  • What "good" means here — the quality bar and any non-negotiables (e.g. must be grounded, must follow format).
  • How it'll be scored — human review, LLM-as-judge, or both; and whether you need a single score or per-dimension.

Output Format

Eval Rubric: [task]

1. Dimensions — 3–6 independent dimensions, each with a one-line definition and a weight. Default set, tailored to the task: structure, completeness, correctness/grounding, usefulness, safety/tone.

2. Anchors — for each dimension, concrete descriptions at 1, 3, and 5 (what a poor / acceptable / excellent answer looks like for this task). Anchors must be observable, not "feels good".

Dimension (weight)1 — poor3 — acceptable5 — excellent
Grounding (×2)invents facts not in the sourcemostly grounded, minor driftevery claim traceable to the source

3. Judge prompt — a ready-to-run LLM-as-judge prompt in a fenced block: the task description, the rubric, an instruction to score each dimension 1–5, and a strict JSON output contract ({"dimension":N,...}) so scores parse reliably. Include a one-line "return only JSON" reinforcement.

4. Labelling guide — short rules for tie-breaks and common edge cases, so repeat runs stay consistent.

5. Judge reliability notes — known biases (length, position, self-preference), and how to mitigate: a cheaper judge for scale vs. a stronger judge for the rubric, sampling N runs, and spot-checking judge scores against a few human labels before trusting the leaderboard.

Show full SKILL.md (191 more words)Show less

Quality Checks

  • Dimensions are independent — a single flaw doesn't tank three of them at once
  • Every dimension has concrete 1/3/5 anchors specific to this task, not generic adjectives
  • The judge prompt has a strict, parseable output contract (JSON), with a retry/repair note
  • Weights reflect what actually matters for the task (grounding usually > prose polish)
  • The rubric is calibrated against at least one good and one weak sample
  • Judge biases are named with a concrete mitigation, not just listed

Anti-Patterns

  • Do not ship dimension names without anchors — names alone don't make scores reproducible
  • Do not let one quality issue load onto multiple dimensions — keep them orthogonal
  • Do not trust an LLM judge blind — calibrate against a handful of human labels first
  • Do not use a vague "overall quality 1–10" — it hides which part is broken
  • Do not ignore the negative case — a rubric must distinguish "wrong" from "thin", not just "great" from "okay"

Based On

LLM-as-judge evaluation practice — orthogonal weighted dimensions, anchored scales, structured judge prompts, and judge-bias mitigation.

Example Trigger Phrases

  • "Create an eval rubric."
  • "Define quality dimensions."
  • "Build an LLM judge."
  • "Decide how to measure whether AI output is good."

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/eval-rubric-designer of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Eval Rubric Designer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Rubric Designer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Rubric Designer this skillmohitagw15856/pm-claude-skills1.4k—~1.1kAutomated safety check: PassMIT
Task Creatorbenchflow-ai/benchflow356—~4.5kAutomated safety check: PassApache-2.0
Create Custom GraderNVIDIA/SkillEvaluator554—~2.1kAutomated safety check: PassApache-2.0
Create Skill Testdotnet/skills5.6k1 repos~6.1kAutomated safety check: PassMIT
Advanced Evaluationaiskillstore/marketplace4333 repos~4.2kAutomated safety check: PassNone
Woo AI Smokewoocommerce/woocommerce-ios358—~7.4kAutomated safety check: NotesGPL-2.0

Similar skills

  • Task Creator

    benchflow-ai/benchflow

    SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

    356 GitHub stars~4.5k tokensUpdated 5 days ago
    EducationAuto-check passed
  • Create Custom Grader

    NVIDIA/SkillEvaluator

    Official

    A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

    554 GitHub stars~2.1k tokensUpdated yesterday
    EducationAuto-check passed
  • Create Skill Test

    dotnet/skills

    Official

    Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository.

    5.6k GitHub starsUsed in 1 repo~6.1k tokens
    EducationAuto-check passed
  • Advanced Evaluation

    aiskillstore/marketplace

    This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…

    433 GitHub starsUsed in 3 repos~4.2k tokens
    EducationAuto-check passed
  • Woo AI Smoke

    woocommerce/woocommerce-ios

    Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

    358 GitHub stars~7.4k tokensUpdated yesterday
    EducationAuto-check: notes
  • Design AI Benchmarking

    Aperivue/medsci-skills

    A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.

    333 GitHub stars~2.4k tokensUpdated 5 days ago
    EducationAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Eval Rubric Designer

What does Eval Rubric Designer do?

Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output. Eval Rubric Designer is an agent skill from mohitagw15856/pm-claude-skills. Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output.

When should I use Eval Rubric Designer?

Eval Rubric Designer fits situations like: asked to create an eval rubric; define quality dimensions; build an LLM judge; decide how to measure whether AI output is good.

How do I install Eval Rubric Designer in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill eval-rubric-designer -a claude-code`. Or copy the skill folder (skills/eval-rubric-designer in mohitagw15856/pm-claude-skills) into .claude/skills/eval-rubric-designer in your project. Claude Code loads it when a task matches its description.

How do I install Eval Rubric Designer in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill eval-rubric-designer -a codex`. Or copy the skill folder (skills/eval-rubric-designer in mohitagw15856/pm-claude-skills) into .agents/skills/eval-rubric-designer in your project. Codex loads it when a task matches its description.

Can I use Eval Rubric Designer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill eval-rubric-designer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-rubric-designer, .gemini/skills/eval-rubric-designer, .github/skills/eval-rubric-designer and .opencode/skills/eval-rubric-designer in your project.

What does Eval Rubric Designer need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Rubric Designer is instructions for the agent only.

Does Eval Rubric Designer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Rubric Designer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Rubric Designer use?

Eval Rubric Designer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Rubric Designer use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Rubric Designer?

Skills that share tags, products or a category with Eval Rubric Designer: Task Creator (benchflow-ai/benchflow, 356 stars), Create Custom Grader (NVIDIA/SkillEvaluator, 554 stars), Create Skill Test (dotnet/skills, 5.6k stars) and Advanced Evaluation (aiskillstore/marketplace, 433 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Rubric Designer?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.