Official agent skill

Eval Suite Planner

by microsoft in microsoft/eval-guide

Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.

OfficialMITAuto-check passedDocuments & Office

Install Eval Suite Planner

skills CLI
$ npx skills add microsoft/eval-guide --skill eval-suite-planner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/eval-guide eval-suite-planner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-suite-planner .claude/skills/eval-suite-planner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-suite-planner
GitHub stars
138
Token cost
~2.3k tokens
SKILL.md length
1,133 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.

  • Works in 5 steps: Eval owner / named approver. → Lifecycle stage and target deployment… → Whether the agent is prompt-only,… → …
  • Tasks that involve Excel spreadsheets
  • SKILL.md covers Purpose, Core rule, Question policy and Planning method, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Suite Planner is an agent skill from microsoft/eval-guide, published by the product's own GitHub organization. Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. Grounded in Practical Guidance on Agent Evaluation v5: Step 1 planning, Steps 2-3 eval-set decomposition, Step 4 gates/improvement targets, Step 5 human inputs, Step 6 grader-validation planning, Step 7 baseline placeholders, Step 8 regression partitioning, and Step 10 reusable-asset candidates. Output is a template-preserving .xlsx workbook plus an interactive HTML review page. Use…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Excel spreadsheets, HTML artifacts and Test generation. It works with Microsoft Excel and Microsoft Word. The repository describes itself as: A plugin for AI agent evaluation. Plan evals, generate test cases, interpret results for Copilot Studio agents. Grounded in Microsoft's Eval Scenario Library & Triage Playbook. The licence is MIT.

When your agent uses it

  • Tasks that involve Excel spreadsheets
  • Tasks that involve HTML artifacts
  • Tasks that involve Test generation

Example prompts

  • “/eval-suite-planner”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Eval owner / named approver.
  2. Lifecycle stage and target deployment decision.
  3. Whether the agent is prompt-only, RAG/knowledge-grounded, or agentic with tools/connectors.
  4. Regulated/compliance obligations.
  5. Authoritative sources and source owners.

What it can do on your machine

Read from SKILL.md and the folder at commit 7a22a89. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Suite Planner loads about 2.3k tokens when it runs. Until then it costs about 145 tokens; SKILL.md has 1,133 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~145
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/eval-guide at commit 7a22a89, republished under its MIT licence (© microsoft). 1,133 words, ~2,330 tokens.

Download SKILL.mdSave it as .claude/skills/eval-suite-planner/SKILL.md (or your agent's skills folder).
name
eval-suite-planner
description
Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. Grounded in Practical Guidance on Agent Evaluation v5: Step 1 planning, Steps 2-3 eval-set decomposition, Step 4 gates/improvement targets, Step 5 human inputs, Step 6 grader-validation planning, Step 7 baseline placeholders, Step 8 regression partitioning, and Step 10 reusable-asset candidates. Output is a template-preserving `.xlsx` workbook plus an interactive HTML review page. Use before generating test cases or running evals.

Purpose

This skill produces the Plan artifact of the /eval-guide lifecycle: a populated copy of the customer's Eval Suite Planning & Logging Template plus an interactive HTML review page. The workbook is the source-of-truth artifact; do not replace it with a scenario table, quality-signal table, generic spreadsheet, default .docx report, or HTML-only plan.

The skill aligns to skills/eval-guide/playbook.md and skills/eval-guide/eval-suite-template.md. Use the 10-step playbook as the methodology spine and the XLSX template as the output shape.

Core rule

Copy the blank XLSX template and populate existing cells/rows only. Do not modify the template.

Do not rename sheets, add sheets, delete sheets, add columns, change headers, rewrite README text, edit Dropdown Lists, change styles, change data validation, or convert the template into a different spreadsheet.

If a blank template workbook is available in the session, use it. If not, ask the user to provide the template; do not silently invent a new workbook.

Question policy

Ask targeted questions only when a workbook field materially affects the plan and cannot be inferred safely:

  1. Eval owner / named approver.
  2. Lifecycle stage and target deployment decision.
  3. Whether the agent is prompt-only, RAG/knowledge-grounded, or agentic with tools/connectors.
  4. Regulated/compliance obligations.
  5. Authoritative sources and source owners.

If the user wants speed or cannot answer, populate TBD - confirm before baseline.

Planning method

When invoked as /eval-suite-planner <agent description>:

  1. Extract or infer the agent's purpose, users, knowledge sources, capabilities, boundaries, architecture, lifecycle stage, and known risks.
  2. Populate Step 1 — Plan the Eval Effort:
    • one-sentence eval objective;
    • five-factor risk tier: reach, criticality of error, autonomy/blast radius, regulatory/compliance exposure, data sensitivity;
    • one accountable owner.
  3. Define eval sets, not scenarios:
    • Capability eval sets: one row per capability dimension that must be diagnostic, e.g. accuracy/correctness, faithfulness/groundedness, relevancy, style/tone, reasoning/tool use.
    • Trust & Safety eval sets: one row per refusal, boundary, or safety category, e.g. guardrails, out-of-scope handling, sensitive-data handling, prompt injection/jailbreak, compliance-specific behavior.
  4. Apply Step 4 v5 gates/improvement-target logic:
    • T&S sets use absolute pass-rate hard gates, usually near 100%.
    • Capability sets usually use a launch floor for first deployment plus regression/direction after baseline, not a standing absolute pass-rate target.
    • High-risk capabilities that function like guardrails keep explicit hard floors.
    • Use the template's existing Target pass rate, Target rationale, Gate type, Intended use, Run cadence, and Notes columns to express this; do not add a new column.
  5. Specify Step 5 human inputs:
    • grading rubric, ground truth, golden answer, or rubric + ground truth;
    • author/owner;
    • grounding source dependency;
    • whether source changes require review.
  6. Plan Step 6 grader validation without changing the template:
    • record grader type and validation expectation in the registry row's Notes;
    • for LLM-as-judge / Custom rubrics, note that human-labeled hard and borderline cases must validate the judge before baseline scores are trusted;
    • for programmatic checks, note the deterministic check to confirm;
    • for human grading, note reviewer agreement expectations where relevant.
  7. Seed Step 7 baseline placeholders in 3 . Run Log only when useful:
    • one placeholder row per eval set;
    • Run type = Baseline;
    • result fields blank;
    • Actionable next step = Validate grader, then run baseline;
    • Status = Open.
  8. Apply Step 8 regression partitioning in existing registry fields:
    • capability sets usually Intended use = Both or Regression;
    • most T&S sets are Gate; the slim subset likely affected by model/tool/policy changes can be Both or Regression;
    • set Run cadence using existing dropdown values such as Per-change, Nightly, Weekly, or Milestone-only.
  9. Flag Step 10 reusable assets in 4 . Reusable Library:
    • reusable T&S sets;
    • grading rubrics;
    • failure-pattern templates;
    • production-derived edge-case categories when applicable.

Workbook population rules

Use skills/eval-guide/eval-suite-template.md as the exact tab/column map.

README

Do not edit.

1 . Planning

Populate only existing input cells:

  • Agent identity.
  • Risk classification (5 factors).
  • Owners & roles.
  • Deployment gates / sign-off criteria.

For the template's Min pass rate - Capability row, reflect v5 Step 4 accurately: use launch floor / high-risk capability floor / regression-governance language, not a generic scenario pass-rate target.

Show full SKILL.md (493 more words)Show less
2 . Eval Suite Registry

Populate one row per eval set. Do not populate one row per test case or legacy planning artifact.

Required row semantics:

  • Category: Capability or Trust & Safety.
  • Dimension tested: capability dimension or T&S category from the template dropdowns.
  • Purpose / diagnostic signal: what failure in this set diagnoses.
  • Target pass rate: absolute gate for T&S; launch floor or Regression / direction after baseline for most capability sets.
  • Target rationale: v5 Step 4 rationale.
  • Gate type: closest existing dropdown value.
  • Intended use: Gate, Regression, or Both.
  • Run cadence: cadence for Step 8.
  • Human input type, Human input author, Grounding source dependency, Source change -> review?: Step 5.
  • Reusable asset?, Reuse tier, Set status: Step 10 and lifecycle status.
  • Notes: assumptions, open questions, Step 4 nuance, and Step 6 grader-validation plan.
3 . Run Log

Use this for Step 7 baseline/iteration logging. During planning, add placeholder baseline rows only if useful; keep result fields blank.

4 . Reusable Library

Populate candidate reusable assets only. Do not duplicate every eval set; promote assets that could help other agents.

Dropdown Lists

Do not edit.

Output

Create eval-suite-<agent-name>-<YYYY-MM-DD>.xlsx as a populated copy of the template.

Then create eval-suite-<agent-name>-<YYYY-MM-DD>-review.html next to the workbook using skills/eval-guide/plan-review-page.md.

Do not paste the summary, eval-set table, or checklist into chat. The HTML page carries that content. The final chat response should be only the workbook path, the HTML review page path, and any blocker/manual action.

Human review checkpoints

Include these in the HTML review page checklist instead of displaying them in chat:

#CheckpointWhat to verify
1Objective, risk tier, ownerThe objective is decision-oriented, the five-factor risk tier is right, and a named owner can sign off.
2Eval-set decompositionCapability sets isolate one diagnostic capability each; T&S sets remain separate from capability.
3Step 4 barsT&S has absolute hard gates; capability uses launch floors / regression-direction unless high-risk.
4Human inputsRubrics, ground truths, golden answers, and source dependencies have owners.
5Grader validationEach set has a plausible grader type and validation plan before baseline.
6Regression partitionCapability and slim T&S regression sets have cadence; gate-only T&S sets run at milestones.
7Template integrityNo sheets, columns, headers, dropdowns, README text, or formatting were changed.

Behavior rules

  • Do not generate scenario-plan tables as the Plan artifact.
  • Do not generate quality-signal sheets or quality-signal grouping as the Plan artifact.
  • Do not add columns to support missing concepts; use existing fields, especially Notes.
  • Do not create a .docx unless the user explicitly asks for a narrative report.
  • Do not produce long narrative chat output after artifact generation; use the HTML review page for the interactive summary and checkpoints.
  • Be specific to the described agent, but at eval-set granularity.

Companion skills

  • /eval-generator — Generate test cases from the populated workbook registry.
  • /eval-result-interpreter — Interpret baseline / iteration results using Step 6-7 and gate status.
  • /eval-triage-and-improvement — Diagnose failures and feed the Step 9 optimization loop.
  • /eval-library-promoter — Promote Step 10 reusable assets.
  • /eval-guide — Orchestrated workflow with dashboard review checkpoints.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/eval-suite-planner of microsoft/eval-guide.

Open the folder on GitHubat commit 7a22a89

Compare with similar skills

Eval Suite Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Suite Planner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Suite Planner this skillmicrosoft/eval-guide138—~2.3kAutomated safety check: PassMIT
Caselyaiskillstore/marketplace430—~2.5kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Docx4jplutext/docx4j2.4k—~2.5kAutomated safety check: PassNone
Cyber Pptcrazyykhllc-bit/CyberPPT1.8k—~10kAutomated safety check: PassMIT
Markitshift-labs-ai/markit1.3k—~299Automated safety check: PassMIT

Similar skills

  • Casely

    aiskillstore/marketplace

    Intelligent QA assistant that automates writing test cases from project documentation.

    430 GitHub stars~2.5k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Docx4j

    plutext/docx4j

    A skill your agent uses when writing Java code that creates, reads or edits Word (.docx), PowerPoint (.pptx) or Excel (.xlsx) files with docx4j — including generating documents, editing existing…

    2.4k GitHub stars~2.5k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Cyber Ppt

    crazyykhllc-bit/CyberPPT

    当用户需要把 DOCX、PDF、TXT、XLSX、研究报告、业务材料或原始数据转成高密度、可编辑、咨询风格 PPTX 时使用;也适用于需要 SCR 论证、视觉风格探索、详细图表和渲染质检的 PPT。

    1.8k GitHub stars~10k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Markit

    shift-labs-ai/markit

    Convert files and URLs to Markdown. An agent skill from shift-labs-ai/markit.

    1.3k GitHub stars~299 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed

More from microsoft/eval-guide

  • Eval Triage And Improvement

    microsoft/eval-guide

    Official

    A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…

    138 GitHub stars~5.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Eval Guide

    microsoft/eval-guide

    Official

    Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

    138 GitHub stars~22k tokensUpdated 3 mo ago
    Auto-check: warnings
  • Eval Faq

    microsoft/eval-guide

    Official

    Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &…

    138 GitHub stars~10k tokensUpdated 3 mo ago
    Auto-check passed
  • Eval Generator

    microsoft/eval-guide

    Official

    Generate standalone — turns the populated Eval Suite Planning workbook (output of /eval-suite-planner) into concrete capability eval sets and trust & safety eval sets.

    138 GitHub stars~7.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Eval Result Interpreter

    microsoft/eval-guide

    Official

    Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics.

    138 GitHub stars~9.9k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Eval Suite Planner

What does Eval Suite Planner do?

Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. Eval Suite Planner is an agent skill from microsoft/eval-guide, published by the product's own GitHub organization. Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.

When should I use Eval Suite Planner?

Eval Suite Planner fits situations like: tasks that involve Excel spreadsheets; tasks that involve HTML artifacts; tasks that involve Test generation.

How do I install Eval Suite Planner in Claude Code?

Run `npx skills add microsoft/eval-guide --skill eval-suite-planner -a claude-code`. Or copy the skill folder (skills/eval-suite-planner in microsoft/eval-guide) into .claude/skills/eval-suite-planner in your project. Claude Code loads it when a task matches its description.

How do I install Eval Suite Planner in Codex?

Run `npx skills add microsoft/eval-guide --skill eval-suite-planner -a codex`. Or copy the skill folder (skills/eval-suite-planner in microsoft/eval-guide) into .agents/skills/eval-suite-planner in your project. Codex loads it when a task matches its description.

Can I use Eval Suite Planner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/eval-guide --skill eval-suite-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-suite-planner, .gemini/skills/eval-suite-planner, .github/skills/eval-suite-planner and .opencode/skills/eval-suite-planner in your project.

What does Eval Suite Planner need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Suite Planner is instructions for the agent only.

Does Eval Suite Planner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Suite Planner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Suite Planner use?

Eval Suite Planner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Suite Planner use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Suite Planner?

Skills that share tags, products or a category with Eval Suite Planner: Casely (aiskillstore/marketplace, 430 stars), Markitdown (ImCa0/just-laws, 781 stars), Docx4j (plutext/docx4j, 2.4k stars) and Cyber Ppt (crazyykhllc-bit/CyberPPT, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Suite Planner?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/eval-guide, which has 138 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 24, 2026.

Source: microsoft/eval-guide on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.