Agent skill

Eval Plan

by menkesu in menkesu/awesome-pm-skills

Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a…

Custom licenceAuto-check passedAI & LLM Engineering

Install Eval Plan

skills CLI
$ npx skills add menkesu/awesome-pm-skills --skill eval-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install menkesu/awesome-pm-skills eval-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/menkesu/awesome-pm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-plan .claude/skills/eval-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-plan
GitHub stars
433
Token cost
~4.5k tokens
SKILL.md length
2,544 words
Files
4 (incl. references, assets)
Skills in repo
25
Repo updated
First seen
Licence
Custom licence

At a glance

Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a…

  • Works in 3 steps: Diagnose (ask before answering) → Pick the play → Run it
  • You ask how do we eval this
  • SKILL.md covers When to use, Step 1 - Diagnose (ask before…, Step 2 - Pick the play and Step 3 - Run it, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Plan is an agent skill from menkesu/awesome-pm-skills. Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a launch bar. Use when you ask how do we eval this, how do I test an AI feature, our LLM outputs are inconsistent, are the evals good enough, LLM-as-judge, hallucinations, or what accuracy is good enough to launch. Draws on 41 Lenny's Podcast guests incl. Hamel Husain, Shreya Shankar and Sander Schulhoff.

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files and assets (for example `references/frameworks.md` and `references/quotes.md`).

It sits in AI & LLM Engineering, covering LLM evaluation and Podcasting. The repository describes itself as: Product management workflows grounded in Lenny's Podcast, for Claude Code, Codex, and Cursor.

When your agent uses it

  • You ask how do we eval this
  • How do I test an AI feature
  • Our LLM outputs are inconsistent
  • Are the evals good enough

Example prompts

  • “Use the eval-plan skill to build an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM…”
  • “/eval-plan”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Diagnose (ask before answering)
  2. Pick the play
  3. Run it

What it can do on your machine

Read from SKILL.md and the folder at commit c1b6e2d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Plan loads about 4.5k tokens when it runs, and up to ~9.3k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 2,544 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 2,544 words (~4,500 tokens).

“Turn a vague 'the AI feels off' into a ranked list of failure modes, a small set of trustworthy evaluators, and a CI + monitoring plan you can paste into a doc. Built from 169 insights from 41 Lenny's Podcast…”

— opening of SKILL.md by menkesu, Custom licence
name
eval-plan

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (references, assets) in skills/eval-plan of menkesu/awesome-pm-skills.

  • SKILL.md
  • assets/card.png
  • references/frameworks.md
  • references/quotes.md

Open the folder on GitHubat commit c1b6e2d

Compare with similar skills

Eval Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Plan this skillmenkesu/awesome-pm-skills433—~4.5kAutomated safety check: PassCustom licence
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from menkesu/awesome-pm-skills

All 25 skills in this repo
  • Craft Review

    menkesu/awesome-pm-skills

    Runs a product quality review on a journey, feature, landing page, onboarding flow, or AI/agent surface and returns a friction log, a scored critical-journey review, and a taste bar your team can…

    429 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed
  • Customer Interviews

    menkesu/awesome-pm-skills

    Designs customer interviews and turns the transcripts into decisions.

    429 GitHub stars~4.4k tokensUpdated 2 days ago
    Auto-check passed
  • Decide

    menkesu/awesome-pm-skills

    Turns a hard product or career decision into a one-page decision memo with the eigenquestion, options, pre-mortem with kill criteria, reversibility rating, a named decider and a review date.

    429 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed
  • Exec Comms

    menkesu/awesome-pm-skills

    Rewrite an exec update, memo, strategy doc, deck outline, Slack post or talk so it lands with senior people, and get a graded score plus the top 3 fixes.

    429 GitHub stars~4.4k tokensUpdated 2 days ago
    Auto-check passed
  • Hard Conversations

    menkesu/awesome-pm-skills

    Writes you a ready-to-use script for a feedback or hard conversation, then rehearses the pushback.

    429 GitHub stars~4.7k tokensUpdated 2 days ago
    Auto-check passed
  • Launch

    menkesu/awesome-pm-skills

    Builds a launch plan, launch story and press/distribution plan, then drafts the pitch emails and announcement copy.

    429 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed

Questions about Eval Plan

What does Eval Plan do?

Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a…. Eval Plan is an agent skill from menkesu/awesome-pm-skills. Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a launch bar.

When should I use Eval Plan?

Eval Plan fits situations like: you ask how do we eval this; how do I test an AI feature; our LLM outputs are inconsistent; are the evals good enough.

How do I install Eval Plan in Claude Code?

Run `npx skills add menkesu/awesome-pm-skills --skill eval-plan -a claude-code`. Or copy the skill folder (skills/eval-plan in menkesu/awesome-pm-skills) into .claude/skills/eval-plan in your project. Claude Code loads it when a task matches its description.

How do I install Eval Plan in Codex?

Run `npx skills add menkesu/awesome-pm-skills --skill eval-plan -a codex`. Or copy the skill folder (skills/eval-plan in menkesu/awesome-pm-skills) into .agents/skills/eval-plan in your project. Codex loads it when a task matches its description.

Can I use Eval Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add menkesu/awesome-pm-skills --skill eval-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-plan, .gemini/skills/eval-plan, .github/skills/eval-plan and .opencode/skills/eval-plan in your project.

What does Eval Plan need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Plan is instructions for the agent only.

Does Eval Plan access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Plan use?

Eval Plan has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Eval Plan use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.8k tokens, read only when the agent opens those files.

What are the alternatives to Eval Plan?

Skills that share tags, products or a category with Eval Plan: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Plan?

menkesu (a GitHub user) maintains it in menkesu/awesome-pm-skills, which has 433 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 6, 2026.

Source: menkesu/awesome-pm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.