Agent skill

AI Evals

by RefoundAI in RefoundAI/lenny-skills

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

MITAuto-check passedAI & LLM Engineering

Install AI Evals

skills CLI
$ npx skills add RefoundAI/lenny-skills --skill ai-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RefoundAI/lenny-skills ai-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RefoundAI/lenny-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-evals .claude/skills/ai-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-evals
GitHub stars
1.4k
Token cost
~1.7k tokens
SKILL.md length
947 words
Files
3 (incl. references)
Skills in repo
76
Repo updated
First seen
Licence
MIT

At a glance

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

  • Works in 4 steps: Identify Failure Modes - Help the user… → Select Eval Methods - Recommend the… → Build Gold Sets - Assist in curating a… → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers How to Help, Core Principles, Templates & Frameworks and Questions to Help Users, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

AI Evals is an agent skill from RefoundAI/lenny-skills. Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/artifacts.md` and `references/guest-insights.md`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: 86 product management skills from Lenny's Podcast for Claude Code and AI agents. Hiring, user research, strategy, shipping, and more. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/ai-evals”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Identify Failure Modes - Help the user conduct error analysis on real traces to find where the system specifically breaks.
  2. Select Eval Methods - Recommend the right mix of human, code, and LLM judges based on the specific technical use case.
  3. Build Gold Sets - Assist in curating a reference dataset of high-quality examples to act as the ground truth for your application.
  4. Operationalize - Guide the user in integrating these evaluations into a CI/CD pipeline for continuous quality improvement.

What it can do on your machine

Read from SKILL.md and the folder at commit 13598cc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Evals loads about 1.7k tokens when it runs, and up to ~21k if it reads all its reference files. Until then it costs about 43 tokens; SKILL.md has 947 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~21k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from RefoundAI/lenny-skills at commit 13598cc, republished under its MIT licence (© RefoundAI). 947 words, ~1,657 tokens.

Download SKILL.mdSave it as .claude/skills/ai-evals/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
ai-evals
description
Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

AI Evaluation Strategy

Move beyond vibe checks to systematic, empirical measurement of AI product quality and reliability.

Help the user with ai evaluation strategy using insights from 11 guests and posts across Lenny's Podcast and Newsletter.

How to Help

  1. Identify Failure Modes - Help the user conduct error analysis on real traces to find where the system specifically breaks.
  2. Select Eval Methods - Recommend the right mix of human, code, and LLM judges based on the specific technical use case.
  3. Build Gold Sets - Assist in curating a reference dataset of high-quality examples to act as the ground truth for your application.
  4. Operationalize - Guide the user in integrating these evaluations into a CI/CD pipeline for continuous quality improvement.

Core Principles

Automate the Value Chain

Brendan Foody: "I think that for enterprises especially, the core way to think about it is how can they build a test or systematic way to measure how well AI automates their core value chain? So if it's an architecture firm that's producing these architecture diagrams of what they provide to their end customer, how can they effectively measure that? And each company has its own value chain or maybe a handful of them if it's a multi-product company."

Identify the core deliverables unique to your business and develop systematic tests to measure how accurately AI can replicate those specific tasks.

Prioritize Subjective Excellence

Edwin Chen: "We are looking for a Nobel Prize-winning poetry. Is this poetry unique? Is it full of subtle imagery? Does it surprise you and target your heart? Does it teach you something about the nature of moonlight?"

True data quality is defined by deep, subjective human excellence, such as emotional resonance and uniqueness, rather than superficial binary checks.

Eliminate Vibe Checks

Hamel Husain & Shreya Shankar: "Evals help you create metrics that you can use to measure how your application is doing and kind of give you a way to improve your application with confidence. That you have a feedback signal in which to iterate against."

Create systematic metrics to track application quality over time, allowing teams to iterate on prompts or models with the same confidence as traditional software.

Structured Judge Logic

From "Beyond vibe checks: A PM’s complete guide to evals": "Clearly articulating what you want your judge-LLM to measure isn’t just a step in the process; it’s the difference between a mediocre AI and one that consistently delights users. Building these writing skills requires practice and attention."

Write effective automated evaluations by using a structured prompt that defines the role, data, success criteria, and specific labels for the judge.

Show full SKILL.md (516 more words)Show less

Templates & Frameworks

  • LLM-as-a-Judge Playbook (Building eval systems that improve your AI product) - A systematic three-step process for building, validating, and measuring an LLM judge that provides trusted binary pass/fail metrics for subjective AI quality as
  • Three Eval Approaches (Human, Code-based, LLM-based) (Beyond vibe checks: A PM’s complete guide to evals) - A decision framework for choosing the right eval approach based on your use case, with pros and cons for each.
  • The Eval Formula (Four-Part Structure) (Beyond vibe checks: A PM’s complete guide to evals) - A four-part formula for writing effective LLM-based eval prompts that any PM can use to construct judge-LLM prompts.
  • Open Coding and Axial Coding for AI Error Analysis (Building eval systems that improve your AI product) - A qualitative research methodology adapted for AI product evaluation, used to discover and categorize failure modes from user interaction data.
  • RAG Evaluation Framework (Retriever + Generator) (Building eval systems that improve your AI product) - A two-part evaluation approach for RAG systems that separately assesses the retriever and generator components with specific metrics for each.
  • AI Eval Improvement Flywheel (Building eval systems that improve your AI product) - The closed-loop process that uses CI safety nets and production discovery engines together to create continuous AI product improvement.
  • Reference Dataset Structure (Why your AI product needs a different development lifecycle) - A template for building the initial reference dataset (20-100 examples) to break the cold start and provide a baseline for AI system evaluation.
  • Transition Failure Matrix for Agentic Workflows (Building eval systems that improve your AI product) - A diagnostic tool for pinpointing exactly which step in an agent's multi-step workflow breaks down, enabling data-driven debugging.

See references/artifacts.md for the full list with details.

Questions to Help Users

  • "What are the top 3 to 5 failure modes your users are currently experiencing in production?"
  • "Do you have a single domain expert or benevolent dictator who defines what quality looks like for this feature?"
  • "What percentage of your current evaluation process is manual versus automated?"
  • "Is your system non-determinism primarily occurring in the retrieval or the generation phase?"
  • "Have you established a golden dataset of at least 20 to 50 human-labeled examples yet?"
  • "How do you currently measure the performance delta when you switch models or change a system prompt?"

Common Mistakes to Flag

  • Relying on vibe checks - Manual and anecdotal testing leads to inconsistent quality and hidden regressions that damage user trust over time.
  • Obsessing over prompt engineering - Focusing solely on prompts while neglecting the underlying evaluation system prevents teams from scaling or hill-climbing systematically.
  • Using generic metrics for product reporting - Off-the-shelf scores are useful for filtering but often fail to capture the specific value or failures unique to your business logic.
  • Ignoring component isolation - Failing to evaluate the retriever separately from the generator in RAG systems makes it impossible to know which part of the stack is failing.
  • Neglecting non-determinism - Failing to account for the stochastic nature of LLMs leads to false confidence in results that may not repeat in production.

Deep Dive

For all 33 sourced insights from 11 guests, see references/guest-insights.md

  • Ai Product Strategy
  • Ai Native Ux

© RefoundAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/ai-evals of RefoundAI/lenny-skills.

  • SKILL.md
  • references/artifacts.md
  • references/guest-insights.md

Open the folder on GitHubat commit 13598cc

Compare with similar skills

AI Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Evals this skillRefoundAI/lenny-skills1.4k—~1.7kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from RefoundAI/lenny-skills

All 76 skills in this repo
  • Customer Interviews

    RefoundAI/lenny-skills

    Help users conduct high-impact customer interviews that move beyond surface-level feature requests to identify root emotional frustrations and specific causal triggers.

    1.4k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Time Energy Management

    RefoundAI/lenny-skills

    Help users master their personal output by shifting from reactive scheduling to intentional energy management, internal trigger mastery, and proactive boundary setting.

    1.4k GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • User Onboarding Activation

    RefoundAI/lenny-skills

    Help users reach their first moment of core value by optimizing the first-run experience, removing friction, and aligning product design with psychological triggers.

    1.4k GitHub stars~2.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Acquisition Channels

    RefoundAI/lenny-skills

    Help users identify unique distribution advantages and master the lifecycle of acquisition channels to build a sustainable engine for growth and retention.

    1.4k GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • AI Assisted Prototyping

    RefoundAI/lenny-skills

    Help users build functional product prototypes from natural language or visual mocks using AI coding tools.

    1.4k GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • AI Native UX

    RefoundAI/lenny-skills

    Help users interact with probabilistic models by designing interfaces that manage fluidity, intent, and agency while maintaining trust and control.

    1.4k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed

Questions about AI Evals

What does AI Evals do?

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies. AI Evals is an agent skill from RefoundAI/lenny-skills. Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

When should I use AI Evals?

AI Evals fits situations like: tasks that involve LLM evaluation.

How do I install AI Evals in Claude Code?

Run `npx skills add RefoundAI/lenny-skills --skill ai-evals -a claude-code`. Or copy the skill folder (skills/ai-evals in RefoundAI/lenny-skills) into .claude/skills/ai-evals in your project. Claude Code loads it when a task matches its description.

How do I install AI Evals in Codex?

Run `npx skills add RefoundAI/lenny-skills --skill ai-evals -a codex`. Or copy the skill folder (skills/ai-evals in RefoundAI/lenny-skills) into .agents/skills/ai-evals in your project. Codex loads it when a task matches its description.

Can I use AI Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RefoundAI/lenny-skills --skill ai-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-evals, .gemini/skills/ai-evals, .github/skills/ai-evals and .opencode/skills/ai-evals in your project.

What does AI Evals need to run?

SKILL.md names no scripts, command-line tools or credentials: AI Evals is instructions for the agent only.

Does AI Evals access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Evals use?

AI Evals is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Evals use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to AI Evals?

Skills that share tags, products or a category with AI Evals: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Evals?

RefoundAI (a GitHub organization) maintains it in RefoundAI/lenny-skills, which has 1,381 GitHub stars. The repository holds 76 skills in this directory. The repository was last updated on July 16, 2026.

Source: RefoundAI/lenny-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.