Repository
microsoft/eval-guide agent skills
- skills
- 6
- official
- 6
- GitHub stars
- 138
GitHub description: “A plugin for AI agent evaluation. Plan evals, generate test cases, interpret results for Copilot Studio agents. Grounded in Microsoft's Eval Scenario Library & Triage Playbook.”
- Stars
- 138 (22 forks)
- Licence
- MIT
- Last push
- Jun 2026
- Created
- Mar 2026
Install all skills
npx skills add microsoft/eval-guideAdd --skill <name> for a single skill and -a <agent> to choose the agent (see the agent guides).
Skills in microsoft/eval-guide, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps… | microsoft/ | 138 | — | ~5.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately. | microsoft/ | 138 | — | ~22k | Automated safety check: Warn | MIT | 3 mo ago |
| 3 | Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &… | microsoft/ | 138 | — | ~10k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | Generate standalone — turns the populated Eval Suite Planning workbook (output of /eval-suite-planner) into concrete capability eval sets and trust & safety eval sets. | microsoft/ | 138 | — | ~7.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 5 | Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics. | microsoft/ | 138 | — | ~9.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. | microsoft/ | 138 | — | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
Questions, answered from the data.
What is the best skill in microsoft/eval-guide?
Eval Triage And Improvement (official) from microsoft/eval-guide ranks first of the 6 skills in microsoft/eval-guide listed here, with the highest score: its repository has 138 GitHub stars, its SKILL.md loads about 5.9k tokens and it passes the automated safety check with no findings. Next come Eval Guide and Eval Faq.
Are the skills in microsoft/eval-guide official?
6 of the 6 skills in microsoft/eval-guide are official, published by the vendor's own GitHub organization: Eval Triage And Improvement, Eval Guide, Eval Faq, Eval Generator, Eval Result Interpreter and 1 more.
How do I install all skills from microsoft/eval-guide?
Run npx skills add microsoft/eval-guide in your project: the open-source skills CLI installs the repository's skills into your coding agent's skills folder. To install a single skill, open its page here for the exact command.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.