Casely
aiskillstore/marketplace
Intelligent QA assistant that automates writing test cases from project documentation.
Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.
$ npx skills add microsoft/eval-guide --skill eval-suite-planner -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/eval-guide eval-suite-planner --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-suite-planner .claude/skills/eval-suite-planner && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-suite-planner" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-suite-planner into .claude/skills/eval-suite-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-suite-planner", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/eval-guide/tree/main/skills/eval-suite-plannerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/eval-guide --skill eval-suite-planner -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/eval-guide eval-suite-planner --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-suite-planner .agents/skills/eval-suite-planner && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-suite-planner" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-suite-planner into .agents/skills/eval-suite-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-suite-planner", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/eval-guide --skill eval-suite-planner -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/eval-guide eval-suite-planner --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-suite-planner .cursor/skills/eval-suite-planner && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-suite-planner" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-suite-planner into .cursor/skills/eval-suite-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-suite-planner", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/eval-guide.git --path skills/eval-suite-planner--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/eval-guide --skill eval-suite-planner -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/eval-guide eval-suite-planner --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-suite-planner .gemini/skills/eval-suite-planner && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-suite-planner" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-suite-planner into .gemini/skills/eval-suite-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-suite-planner", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/eval-guide eval-suite-plannerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/eval-guide --skill eval-suite-planner -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-suite-planner .github/skills/eval-suite-planner && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-suite-planner" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-suite-planner into .github/skills/eval-suite-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-suite-planner", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/eval-guide --skill eval-suite-planner -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/eval-guide eval-suite-planner --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-suite-planner .opencode/skills/eval-suite-planner && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-suite-planner" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-suite-planner into .opencode/skills/eval-suite-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-suite-planner", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-suite-plannerPlan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.
Eval Suite Planner is an agent skill from microsoft/eval-guide, published by the product's own GitHub organization. Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. Grounded in Practical Guidance on Agent Evaluation v5: Step 1 planning, Steps 2-3 eval-set decomposition, Step 4 gates/improvement targets, Step 5 human inputs, Step 6 grader-validation planning, Step 7 baseline placeholders, Step 8 regression partitioning, and Step 10 reusable-asset candidates. Output is a template-preserving .xlsx workbook plus an interactive HTML review page. Use…
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering Excel spreadsheets, HTML artifacts and Test generation. It works with Microsoft Excel and Microsoft Word. The repository describes itself as: A plugin for AI agent evaluation. Plan evals, generate test cases, interpret results for Copilot Studio agents. Grounded in Microsoft's Eval Scenario Library & Triage Playbook. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 7a22a89. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Suite Planner loads about 2.3k tokens when it runs. Until then it costs about 145 tokens; SKILL.md has 1,133 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/eval-guide at commit 7a22a89, republished under its MIT licence (© microsoft). 1,133 words, ~2,330 tokens.
.claude/skills/eval-suite-planner/SKILL.md (or your agent's skills folder).This skill produces the Plan artifact of the /eval-guide lifecycle: a populated copy of the customer's Eval Suite Planning & Logging Template plus an interactive HTML review page. The workbook is the source-of-truth artifact; do not replace it with a scenario table, quality-signal table, generic spreadsheet, default .docx report, or HTML-only plan.
The skill aligns to skills/eval-guide/playbook.md and skills/eval-guide/eval-suite-template.md. Use the 10-step playbook as the methodology spine and the XLSX template as the output shape.
Copy the blank XLSX template and populate existing cells/rows only. Do not modify the template.
Do not rename sheets, add sheets, delete sheets, add columns, change headers, rewrite README text, edit Dropdown Lists, change styles, change data validation, or convert the template into a different spreadsheet.
If a blank template workbook is available in the session, use it. If not, ask the user to provide the template; do not silently invent a new workbook.
Ask targeted questions only when a workbook field materially affects the plan and cannot be inferred safely:
If the user wants speed or cannot answer, populate TBD - confirm before baseline.
When invoked as /eval-suite-planner <agent description>:
Target pass rate, Target rationale, Gate type, Intended use, Run cadence, and Notes columns to express this; do not add a new column.Notes;3 . Run Log only when useful:Run type = Baseline;Actionable next step = Validate grader, then run baseline;Status = Open.Intended use = Both or Regression;Gate; the slim subset likely affected by model/tool/policy changes can be Both or Regression;Run cadence using existing dropdown values such as Per-change, Nightly, Weekly, or Milestone-only.4 . Reusable Library:Use skills/eval-guide/eval-suite-template.md as the exact tab/column map.
READMEDo not edit.
1 . PlanningPopulate only existing input cells:
For the template's Min pass rate - Capability row, reflect v5 Step 4 accurately: use launch floor / high-risk capability floor / regression-governance language, not a generic scenario pass-rate target.
2 . Eval Suite RegistryPopulate one row per eval set. Do not populate one row per test case or legacy planning artifact.
Required row semantics:
Category: Capability or Trust & Safety.Dimension tested: capability dimension or T&S category from the template dropdowns.Purpose / diagnostic signal: what failure in this set diagnoses.Target pass rate: absolute gate for T&S; launch floor or Regression / direction after baseline for most capability sets.Target rationale: v5 Step 4 rationale.Gate type: closest existing dropdown value.Intended use: Gate, Regression, or Both.Run cadence: cadence for Step 8.Human input type, Human input author, Grounding source dependency, Source change -> review?: Step 5.Reusable asset?, Reuse tier, Set status: Step 10 and lifecycle status.Notes: assumptions, open questions, Step 4 nuance, and Step 6 grader-validation plan.3 . Run LogUse this for Step 7 baseline/iteration logging. During planning, add placeholder baseline rows only if useful; keep result fields blank.
4 . Reusable LibraryPopulate candidate reusable assets only. Do not duplicate every eval set; promote assets that could help other agents.
Dropdown ListsDo not edit.
Create eval-suite-<agent-name>-<YYYY-MM-DD>.xlsx as a populated copy of the template.
Then create eval-suite-<agent-name>-<YYYY-MM-DD>-review.html next to the workbook using skills/eval-guide/plan-review-page.md.
Do not paste the summary, eval-set table, or checklist into chat. The HTML page carries that content. The final chat response should be only the workbook path, the HTML review page path, and any blocker/manual action.
Include these in the HTML review page checklist instead of displaying them in chat:
| # | Checkpoint | What to verify |
|---|---|---|
| 1 | Objective, risk tier, owner | The objective is decision-oriented, the five-factor risk tier is right, and a named owner can sign off. |
| 2 | Eval-set decomposition | Capability sets isolate one diagnostic capability each; T&S sets remain separate from capability. |
| 3 | Step 4 bars | T&S has absolute hard gates; capability uses launch floors / regression-direction unless high-risk. |
| 4 | Human inputs | Rubrics, ground truths, golden answers, and source dependencies have owners. |
| 5 | Grader validation | Each set has a plausible grader type and validation plan before baseline. |
| 6 | Regression partition | Capability and slim T&S regression sets have cadence; gate-only T&S sets run at milestones. |
| 7 | Template integrity | No sheets, columns, headers, dropdowns, README text, or formatting were changed. |
Notes..docx unless the user explicitly asks for a narrative report./eval-generator — Generate test cases from the populated workbook registry./eval-result-interpreter — Interpret baseline / iteration results using Step 6-7 and gate status./eval-triage-and-improvement — Diagnose failures and feed the Step 9 optimization loop./eval-library-promoter — Promote Step 10 reusable assets./eval-guide — Orchestrated workflow with dashboard review checkpoints.© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/eval-suite-planner of microsoft/eval-guide.
Open the folder on GitHubat commit 7a22a89
Eval Suite Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Suite Planner this skillmicrosoft/eval-guide | 138 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Caselyaiskillstore/marketplace | 430 | — | ~2.5k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Docx4jplutext/docx4j | 2.4k | — | ~2.5k | Automated safety check: Pass | None | |
| Cyber Pptcrazyykhllc-bit/CyberPPT | 1.8k | — | ~10k | Automated safety check: Pass | MIT | |
| Markitshift-labs-ai/markit | 1.3k | — | ~299 | Automated safety check: Pass | MIT |
aiskillstore/marketplace
Intelligent QA assistant that automates writing test cases from project documentation.
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
plutext/docx4j
A skill your agent uses when writing Java code that creates, reads or edits Word (.docx), PowerPoint (.pptx) or Excel (.xlsx) files with docx4j — including generating documents, editing existing…
crazyykhllc-bit/CyberPPT
当用户需要把 DOCX、PDF、TXT、XLSX、研究报告、业务材料或原始数据转成高密度、可编辑、咨询风格 PPTX 时使用;也适用于需要 SCR 论证、视觉风格探索、详细图表和渲染质检的 PPT。
shift-labs-ai/markit
Convert files and URLs to Markdown. An agent skill from shift-labs-ai/markit.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
microsoft/eval-guide
A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…
microsoft/eval-guide
Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.
microsoft/eval-guide
Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &…
microsoft/eval-guide
Generate standalone — turns the populated Eval Suite Planning workbook (output of /eval-suite-planner) into concrete capability eval sets and trust & safety eval sets.
microsoft/eval-guide
Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics.
Works with
Categories
Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. Eval Suite Planner is an agent skill from microsoft/eval-guide, published by the product's own GitHub organization. Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.
Eval Suite Planner fits situations like: tasks that involve Excel spreadsheets; tasks that involve HTML artifacts; tasks that involve Test generation.
Run `npx skills add microsoft/eval-guide --skill eval-suite-planner -a claude-code`. Or copy the skill folder (skills/eval-suite-planner in microsoft/eval-guide) into .claude/skills/eval-suite-planner in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/eval-guide --skill eval-suite-planner -a codex`. Or copy the skill folder (skills/eval-suite-planner in microsoft/eval-guide) into .agents/skills/eval-suite-planner in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/eval-guide --skill eval-suite-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-suite-planner, .gemini/skills/eval-suite-planner, .github/skills/eval-suite-planner and .opencode/skills/eval-suite-planner in your project.
SKILL.md names no scripts, command-line tools or credentials: Eval Suite Planner is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Suite Planner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Suite Planner: Casely (aiskillstore/marketplace, 430 stars), Markitdown (ImCa0/just-laws, 781 stars), Docx4j (plutext/docx4j, 2.4k stars) and Cyber Ppt (crazyykhllc-bit/CyberPPT, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/eval-guide, which has 138 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 24, 2026.
Source: microsoft/eval-guide on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.